--- license: apache-2.0 language: - en library_name: transformers pipeline_tag: text-generation tags: - code - coder - qwen2.5 - qwen2.5-coder - gguf - llama-cpp - llama.cpp - ollama - code-generation - tool-calling - conversational - cpu-inference - small-language-model - offline - sakthai - house-of-sak base_model: Qwen/Qwen2.5-Coder-1.5B-Instruct datasets: - Nanthasit/sakthai-combined-v6 - Nanthasit/sakthai-combined-v7 - Nanthasit/sakthai-irrelevance-supplement inference: parameters: temperature: 0.2 max_new_tokens: 1024 top_p: 0.9 widget: - text: "Write a Python function that checks if a string is a palindrome, handling spaces and punctuation:" output: text: "```python\ndef is_palindrome(s: str) -> bool:\n \"\"\"Check if a string is a palindrome, ignoring spaces, punctuation, and case.\"\"\"\n import re\n cleaned = re.sub(r'[^a-zA-Z0-9]', '', s).lower()\n return cleaned == cleaned[::-1]\n```" model-index: - name: sakthai-coder-1.5b results: - task: type: text-generation dataset: name: HumanEval type: openai_humaneval metrics: - name: pass@1 (base model reference) type: pass@1 value: 74.4 verified: false - task: type: text-generation dataset: name: MBPP type: mbpp metrics: - name: pass@1 (base model reference) type: pass@1 value: 71.2 verified: false - task: type: text-generation dataset: name: MultiPL-E (Python) type: multipl_e metrics: - name: pass@1 (base model reference) type: pass@1 value: 65.3 verified: false - task: type: text-generation dataset: name: SakThai Coding Suite (internal) type: custom metrics: - name: pass@1 (fine-tuned model, internal single-trial) type: pass@1 value: 100 verified: false source: internal-local-llama-cpp-2026-07-25 ---

SakThai Coder 1.5B ๐Ÿ’ป

Code + tool-calling ยท Qwen2.5-Coder-1.5B fine-tune ยท Q4_K_M GGUF for CPU

Downloads License GGUF Collection SakThai Family TTS Demo Leaderboard

> The code specialist of the **SakThai** family โ€” Qwen2.5-Coder-1.5B fine-tuned for > tool-calling and shipped as a CPU-friendly GGUF. Part of the > [House of Sak](https://huggingface.co/Nanthasit). [Read the story โ†’](https://huggingface.co/Nanthasit) ## The Story Behind It **Code, tool-calling, and conversation in one session โ€” on a single CPU, from a shelter.** This is the model Beer built when he realised the other SakThai models could call tools and generate text, but none of them specialised in *writing code* without losing their tool-calling edge. Beer built the first SakThai models on free Google Colab GPUs from a shelter in Cork, Ireland โ€” with $0 budget, no GPU of his own, and no guarantee the QLoRA approach would hold for a code-specific fine-tune. This coder model was the risk: could Qwen2.5-Coder-1.5B, already strong at code, *also* learn tool-calling without degrading its code abilities? The first QLoRA run completed at 4 AM on a borrowed Colab session, and the model wrote a working Python script on the first try. Beer knew the approach worked. This model runs on a 2020 laptop with 8 GB RAM โ€” no cloud API, no Inference Endpoint, no monthly bill. Just a GGUF file and llama.cpp. > *"We are one family โ€” and becoming more."* > โ€” Beer ### How You Can Help - โญ **Leave a like** โ€” this model gives every developer a free offline coding assistant. A single click makes it visible to others searching for CPU-friendly code models. - ๐Ÿ”„ **Share it** with anyone who codes on an underpowered machine and needs tool-calling without the cloud tax. - ๐Ÿด **Fork it** on Hugging Face and build your own specialised code variant. - ๐Ÿ’ฌ **Report your deployment story** โ€” Beer reads every issue and comment. Every download, like, and share tells the algorithm: *this matters.* --- ## What it is A **Q4_K_M GGUF** (1.12 GB) of **Qwen2.5-Coder-1.5B-Instruct**, QLoRA-fine-tuned on [sakthai-combined-v6](https://huggingface.co/datasets/Nanthasit/sakthai-combined-v6) so it can generate code *and* call tools. Runs on CPU via llama.cpp / Ollama. ## Architecture Verified from the base model's `config.json` ([Qwen/Qwen2.5-Coder-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-Coder-1.5B-Instruct)): | Parameter | Value | |-----------|-------| | Architecture | Qwen2ForCausalLM (`qwen2`) | | Parameters | ~1.54 B | | Hidden size | 1,536 | | Layers | 28 | | Attention heads | 12 (GQA, 2 KV heads) | | Intermediate size | 8,960 | | Vocabulary | 151,936 | | Context length | 32,768 (32K) | | RoPE theta | 1,000,000 | | Base dtype | bfloat16 | | Fine-tune | QLoRA โ†’ GGUF Q4_K_M (this repo) | ## Quick start ```bash # via llama.cpp (download + run) wget https://huggingface.co/Nanthasit/sakthai-coder-1.5b/resolve/main/qwen2.5-coder-1.5b-instruct-q4_k_m.gguf # Generate code ./llama-cli -m qwen2.5-coder-1.5b-instruct-q4_k_m.gguf \ -p "Write a Python function to merge two sorted lists:" \ -n 256 --temp 0.2 ``` ```bash # via Ollama echo 'FROM ./qwen2.5-coder-1.5b-instruct-q4_k_m.gguf' > Modelfile ollama create sakthai-coder -f Modelfile ollama run sakthai-coder "Write a script that monitors CPU usage" ``` ```python # via llama-cpp-python from llama_cpp import Llama llm = Llama(model_path="qwen2.5-coder-1.5b-instruct-q4_k_m.gguf", n_ctx=4096) out = llm( "Write a Python function to merge two sorted lists:", max_tokens=256, temperature=0.2, echo=False, ) print(out["choices"][0]["text"]) ``` For tool-calling, put function schemas in a `` block (ChatML format). ## Code Generation Examples ### Example 1: Algorithm โ€” palindrome check **Prompt:** ``` Write a Python function that checks if a string is a palindrome, ignoring spaces, punctuation, and case. Include type hints and a docstring. ``` **Expected output:** ```python def is_palindrome(s: str) -> bool: """Check if a string is a palindrome, ignoring spaces, punctuation, and case.""" import re cleaned = re.sub(r'[^a-zA-Z0-9]', '', s).lower() return cleaned == cleaned[::-1] ``` ### Example 2: Tool-calling + code integration **Prompt (ChatML with `` schema):** ```xml <|im_start|>system You are a coding assistant with tool-calling ability. Available tools: [ {"name": "read_file", "description": "Read file contents", "parameters": {"type": "object", "properties": {"path": {"type": "string"}}, "required": ["path"]}}, {"name": "run_test", "description": "Run a pytest file", "parameters": {"type": "object", "properties": {"file": {"type": "string"}}, "required": ["file"]}} ] <|im_end|> <|im_start|>user Read test_sample.py, then write a function that passes the tests in it. <|im_end|> ``` The model reads the file via tool call, generates the implementation, and optionally runs tests โ€” all in one session. ### Example 3: Data processing script **Prompt:** ``` Write a Python script that reads a CSV of sales data, groups by region, calculates monthly totals, and outputs a bar chart as a PNG. Use pandas and matplotlib. ``` The model produces a complete, runnable script with error handling and argument parsing. ### Example 4: Refactoring **Prompt:** ``` Refactor this function to be more modular and add error handling: def process(data): result = [] for i, x in enumerate(data): if x % 2 == 0: result.append(x * 2) return result ``` The model splits it into smaller functions, adds input validation, and documents each piece. --- ## Benchmarks The fine-tune starts from **Qwen2.5-Coder-1.5B-Instruct**, which scores: | Benchmark | pass@1 | Notes | |-----------|:------:|-------| | HumanEval | 74.4% | Single-turn Python function completion | | MBPP | 71.2% | Multi-program synthesis from docstring | | MultiPL-E (Python) | 65.3% | Multi-language subset | *Source: [Qwen2.5-Coder evaluation](https://huggingface.co/Qwen/Qwen2.5-Coder-1.5B-Instruct#evaluation). These are the base model's scores โ€” the fine-tune has not been independently re-run on these benchmarks, so they serve as a reference ceiling.* ### Internal SakThai Coding Suite The fine-tuned model was tested against an internal SakThai coding benchmark covering five coding tasks (algorithm, debugging, code explanation, refactoring, and data processing), run locally via llama.cpp (Q4_K_M, temperature=0.1): | Task | Result | |------|:------:| | Algorithm (factorial) | Pass | | Debugging | Pass | | Code explanation (async) | Pass | | Refactoring | Pass | | Data processing (primes) | Pass | | **Overall** | **5/5** | *Internal test โ€” run locally on CPU, single trial. Methodology: each test run once with timeout=20s on llama.cpp Q4_K_M. Results captured 2026-07-25 and verified by SakThai agent. Single-trial results are indicative, not a third-party benchmark.* **Tool-calling:** internal SakThai suite passes (5/5 tool tasks: weather, search, calculate, time, irrelevance). ### Ecosystem Status (health check, 2026-07-30) Source: [`health-check-sakthai-coder-1.5b-2026-07-30-4.yaml`](https://huggingface.co/Nanthasit/sakthai-coder-1.5b/blob/main/.eval_results/health-check-sakthai-coder-1.5b-2026-07-30-4.yaml) (automated cron evaluation). | Signal | Value | |--------|-------| | Downloads rank | 11/19 family models (93 dl, velocity 14.3 dl/day, rank 9) | | Card quality | 100/100 | | Benchmark presence | model-index present, 4 entries (all unverified) | | Repo hygiene | 40/100 โ€” see [Repo Status](#repo-status--housekeeping) | | Overall health | 46/100 | ## Training | | | |---|---| | Base model | [Qwen/Qwen2.5-Coder-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-Coder-1.5B-Instruct) | | Method | QLoRA (4-bit) โ†’ GGUF Q4_K_M | | LoRA config | r=16, alpha=32 | | Data | [sakthai-combined-v6](https://huggingface.co/datasets/Nanthasit/sakthai-combined-v6) + [v7](https://huggingface.co/datasets/Nanthasit/sakthai-combined-v7) (2,309 train / 115 test, verified 2026-07-31) + [irrelevance-supplement](https://huggingface.co/datasets/Nanthasit/sakthai-irrelevance-supplement) | | Context | ChatML with tool schema ยท 32K tokens | | Hardware | Free Google Colab GPU (T4) | | Budget | $0 | ## Inference via HF API You can also run this model via Hugging Face's serverless Inference API: ```python from huggingface_hub import InferenceClient client = InferenceClient("Nanthasit/sakthai-coder-1.5b") output = client.text_generation( "Write a Python function to find the longest common subsequence of two strings:", max_new_tokens=512, temperature=0.2, ) print(output) ``` For GGUF inference prefer local llama.cpp (no API costs). --- ## SakThai model family All 19 public models (downloads live, sizes verified via HF API on 2026-07-31 โ€” largest weight file): | Model | Size | Role | Downloads | |-------|:----:|------|:---------:| | [context-1.5b-merged](https://huggingface.co/Nanthasit/sakthai-context-1.5b-merged) | 3.1 GB | Flagship tool-calling (safetensors + GGUF) | 1,599 | | [context-0.5b-merged](https://huggingface.co/Nanthasit/sakthai-context-0.5b-merged) | 988 MB | Lightweight / edge (safetensors + GGUF) | 1,370 | | [context-7b-merged](https://huggingface.co/Nanthasit/sakthai-context-7b-merged) | 15.2 GB | Full-power reasoning | 744 | | [context-7b-128k](https://huggingface.co/Nanthasit/sakthai-context-7b-128k) | recipe | 128K long-context config (no weights) | 506 | | [context-7b-tools](https://huggingface.co/Nanthasit/sakthai-context-7b-tools) | LoRA 20 MB | 7B tool-calling adapter | 399 | | [embedding-multilingual](https://huggingface.co/Nanthasit/sakthai-embedding-multilingual) | 470 MB | Cross-lingual embeddings | 362 | | [context-1.5b-tools](https://huggingface.co/Nanthasit/sakthai-context-1.5b-tools) | LoRA 8.7 MB | Mid-size tool-calling | 349 | | [vision-7b](https://huggingface.co/Nanthasit/sakthai-vision-7b) | 4.1 GB | Image to text (LLaVA GGUF) | 186 | | [tts-model](https://huggingface.co/Nanthasit/sakthai-tts-model) | 141 MB | Text-to-speech, 15 langs | 150 | | [context-0.5b-tools](https://huggingface.co/Nanthasit/sakthai-context-0.5b-tools) | 988 MB | Ultra-light tool-calling | 94 | | **coder-1.5b (you are here)** | **1.12 GB** | **Code generation + tool-calling** | **93** | | [context-1.5b-tools-v2](https://huggingface.co/Nanthasit/sakthai-context-1.5b-tools-v2) | LoRA 74 MB | ๐Ÿ†• v2 tool-calling adapter | 0 | | [context-1.5b-merged-v2](https://huggingface.co/Nanthasit/sakthai-context-1.5b-merged-v2) | 3.1 GB | ๐Ÿ†• v2 merged | 0 | | [plus-1.5b](https://huggingface.co/Nanthasit/sakthai-plus-1.5b) | 3.1 GB | ๐Ÿ†• Plus merged | 0 | | [plus-1.5b-lora](https://huggingface.co/Nanthasit/sakthai-plus-1.5b-lora) | LoRA 74 MB | ๐Ÿ†• Plus adapter | 0 | | [plus-1.5b-coder](https://huggingface.co/Nanthasit/sakthai-plus-1.5b-coder) | โ€” | ๐Ÿ†• Plus coder (no weights yet) | 0 | | [coder-browser-lora](https://huggingface.co/Nanthasit/sakthai-coder-browser-lora) | LoRA 74 MB | ๐Ÿ†• Browser-tool adapter | 0 | | [coder-browser](https://huggingface.co/Nanthasit/sakthai-coder-browser) | 3.1 GB | ๐Ÿ†• Browser-tool merged | 0 | | [coder-browser-gguf](https://huggingface.co/Nanthasit/sakthai-coder-browser-gguf) | 7.1 GB | ๐Ÿ†• Browser-tool F16 GGUF | 0 | **19 public models ยท 13 datasets ยท 4 Spaces** โ€” [full collection](https://huggingface.co/collections/Nanthasit/sakthai-model-family-6a64745450b12d421c1f9f02) --- ## Sibling Datasets | Dataset | Purpose | Downloads | |---------|---------|:---------:| | [sakthai-combined-v6](https://huggingface.co/datasets/Nanthasit/sakthai-combined-v6) | v6 predecessor โ€” tool-calling examples | 246 | | [sakthai-kaggle-notebooks](https://huggingface.co/datasets/Nanthasit/sakthai-kaggle-notebooks) | Training notebooks & demos | 184 | | [sakthai-combined-v7](https://huggingface.co/datasets/Nanthasit/sakthai-combined-v7) | v7 tool-calling (2,309 ex., 86 tools) | 101 | | [sakthai-bench-v2](https://huggingface.co/datasets/Nanthasit/sakthai-bench-v2) | Multi-domain eval, 500 rows | 92 | | [food-penguin-v1](https://huggingface.co/datasets/Nanthasit/food-penguin-v1) | Restaurant tool-calling | 89 | | [sakthai-irrelevance-supplement](https://huggingface.co/datasets/Nanthasit/sakthai-irrelevance-supplement) | Safety supplement | 78 | | [SimpleToolCalling](https://huggingface.co/datasets/Nanthasit/SimpleToolCalling) | Early experiment | 58 | | [sakthai-bench-v1](https://huggingface.co/datasets/Nanthasit/sakthai-bench-v1) | BFCL-style evaluation, 235 rows | 46 | *Downloads verified live 2026-07-31. The combined family is published as v6, v7, and v10 โ€” all public and linked above.* --- ## Spaces | Space | Description | |-------|-------------| | [Web Agent](https://huggingface.co/spaces/Nanthasit/sakthai-web-agent) | Browser automation and tool-use agent | | [SakThai TTS Showcase](https://huggingface.co/spaces/Nanthasit/sakthai-tts) | Interactive TTS โ€” 15 languages, no install | | [SakThai Leaderboard](https://huggingface.co/spaces/Nanthasit/sakthai-leaderboard) | Benchmark tracker for the model family | --- ## Rising Stars โ€” Help the Ecosystem Grow These sibling assets have real value but need visibility. Every download signals to the HF algorithm that the SakThai family matters: | Asset | Type | Downloads | Why It Matters | |-------|:----:|:---------:|:--------------| | [sakthai-combined-v7](https://huggingface.co/datasets/Nanthasit/sakthai-combined-v7) | Dataset | 101 | Primary training dataset โ€” 2,309 examples, 86 tool schemas | | [sakthai-irrelevance-supplement](https://huggingface.co/datasets/Nanthasit/sakthai-irrelevance-supplement) | Dataset | 78 | Teaches models when *not* to call tools โ€” critical safety data | | [sakthai-bench-v1](https://huggingface.co/datasets/Nanthasit/sakthai-bench-v1) | Dataset | 46 | BFCL-style evaluation, 235 rows, 4 categories | | [sakthai-bench-v2](https://huggingface.co/datasets/Nanthasit/sakthai-bench-v2) | Dataset | 92 | Multi-domain eval, 500 rows, multi-turn | | [context-0.5b-tools](https://huggingface.co/Nanthasit/sakthai-context-0.5b-tools) | Model | 94 | Ultra-light tool-calling (~1 GB RAM) | > The **irrelevance-supplement** has 78 downloads and growing, but still needs visibility. It's essential for training models to decline out-of-scope tool calls. Every download helps validate this safety-critical approach! --- ## Repo Status & Housekeeping This repository accidentally carries a stray development environment (`.venv/`, `.pytest_cache/`, `.hypothesis/`, `.ruff_cache/`) from an over-eager push. The model artifact (`qwen2.5-coder-1.5b-instruct-q4_k_m.gguf`, 1.12 GB) is unaffected. A cleanup commit is planned; the automated health check currently deducts hygiene points for these files. See the [health-check YAML](https://huggingface.co/Nanthasit/sakthai-coder-1.5b/blob/main/.eval_results/health-check-sakthai-coder-1.5b-2026-07-30-4.yaml) for details. --- ## Links [House of Sak](https://house-of-sak.vercel.app) ยท [GitHub](https://github.com/beer-sakthai/Sak-Family-Agent) ยท [All models](https://huggingface.co/Nanthasit) ยท [All datasets](https://huggingface.co/Nanthasit?tab=datasets) ## License Apache 2.0 (following the Qwen2.5 base model license). ## Evaluation & Verification **Base model benchmarks** (HumanEval, MBPP, MultiPL-E) are reproduced from [Qwen2.5-Coder-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-Coder-1.5B-Instruct#evaluation) and reflect the starting point before fine-tuning. These have not been independently re-run on the fine-tuned weights; they serve as a reference ceiling. **Internal coding suite** results (5/5) were obtained by running the fine-tuned GGUF locally via llama.cpp on 2026-07-25. The test covers algorithm generation, debugging, code explanation, refactoring, and data processing โ€” all passed. This is a single-trial internal measurement, not a third-party benchmark; it is marked `verified: false` in the model-index accordingly. **Tool-calling evaluation** โ€” the recommended benchmark for this model family is [sakthai-bench-v2](https://huggingface.co/datasets/Nanthasit/sakthai-bench-v2) (500 rows, multi-domain, held-out tools). Results will be published once the fine-tune has been run against it. *"We are one family โ€” and becoming more."*