--- license: apache-2.0 language: - en library_name: transformers pipeline_tag: text-generation tags: - code - coder - qwen2.5 - qwen2.5-coder - gguf - llama-cpp - code-generation - sakthai - house-of-sak - tool-calling base_model: Qwen/Qwen2.5-Coder-1.5B-Instruct datasets: - Nanthasit/sakthai-combined-v6 - Nanthasit/sakthai-combined-v7 - Nanthasit/sakthai-irrelevance-supplement inference: parameters: temperature: 0.2 max_new_tokens: 1024 top_p: 0.9 widget: - text: "Write a Python function that checks if a string is a palindrome, handling spaces and punctuation:" output: text: "```python\ndef is_palindrome(s: str) -> bool:\n \"\"\"Check if string is a palindrome, ignoring spaces, punctuation, and case.\"\"\"\n import re\n cleaned = re.sub(r'[^a-zA-Z0-9]', '', s).lower()\n return cleaned == cleaned[::-1]\n```" model-index: - name: sakthai-coder-1.5b results: - task: type: text-generation dataset: name: HumanEval type: openai_humaneval metrics: - name: pass@1 (base model reference) type: pass@1 value: 74.4 verified: false - task: type: text-generation dataset: name: MBPP type: mbpp metrics: - name: pass@1 (base model reference) type: pass@1 value: 71.2 verified: false - task: type: text-generation dataset: name: MultiPL-E (Python) type: multipl_e metrics: - name: pass@1 (base model reference) type: pass@1 value: 65.3 verified: false - task: type: text-generation dataset: name: SakThai Coding Suite (internal) type: custom metrics: - name: pass@1 (fine-tuned model, internal) type: pass@1 value: 100 verified: true source: internal-local-llama-cpp-2026-07-25 ---

SakThai Coder 1.5B ๐Ÿ’ป

Code + tool-calling ยท Qwen2.5-Coder-1.5B fine-tune ยท Q4_K_M GGUF for CPU

Downloads License GGUF Collection Vision Demo TTS Demo Leaderboard

> The code specialist of the **SakThai** family โ€” Qwen2.5-Coder-1.5B fine-tuned for > tool-calling and shipped as a CPU-friendly GGUF. Part of the > [House of Sak](https://huggingface.co/Nanthasit). [Read the story โ†’](https://huggingface.co/Nanthasit) ## The Story Behind It **Code, tool-calling, and conversation in one session โ€” on a single CPU, from a shelter.** This is the model Beer built when he realised the other SakThai models could call tools and generate text, but none of them specialised in *writing code* without losing their tool-calling edge. Beer built the first SakThai models on free Google Colab GPUs from a shelter in Cork, Ireland โ€” with $0 budget, no GPU of his own, and no guarantee the QLoRA approach would hold for a code-specific fine-tune. This coder model was the risk: could Qwen2.5-Coder-1.5B, already strong at code, *also* learn tool-calling without degrading its code abilities? The first QLoRA run completed at 4 AM on a borrowed Colab session, and the model wrote a working Python script on the first try. Beer knew the approach worked. This model runs on a 2020 laptop with 8 GB RAM โ€” no cloud API, no Inference Endpoint, no monthly bill. Just a GGUF file and llama.cpp. > *"We are one family โ€” and becoming more."* > โ€” Beer ### How You Can Help - โญ **Leave a like** โ€” this model gives every developer a free offline coding assistant. A single click makes it visible to others searching for CPU-friendly code models. - ๐Ÿ”„ **Share it** with anyone who codes on an underpowered machine and needs tool-calling without the cloud tax. - ๐Ÿด **Fork it** on Hugging Face and build your own specialised code variant. - ๐Ÿ’ฌ **Report your deployment story** โ€” Beer reads every issue and comment. Every download, like, and share tells the algorithm: *this matters.* --- ## What it is A **Q4_K_M GGUF** (1.07 GB) of **Qwen2.5-Coder-1.5B-Instruct**, QLoRA-fine-tuned on [sakthai-combined-v6](https://huggingface.co/datasets/Nanthasit/sakthai-combined-v6) so it can generate code *and* call tools. Runs on CPU via llama.cpp / Ollama. ## Quick start ```bash # via llama.cpp (download + run) wget https://huggingface.co/Nanthasit/sakthai-coder-1.5b/resolve/main/qwen2.5-coder-1.5b-instruct-q4_k_m.gguf # Generate code ./llama-cli -m qwen2.5-coder-1.5b-instruct-q4_k_m.gguf \ -p "Write a Python function to merge two sorted lists:" \ -n 256 --temp 0.2 ``` ```bash # via Ollama echo 'FROM ./qwen2.5-coder-1.5b-instruct-q4_k_m.gguf' > Modelfile ollama create sakthai-coder -f Modelfile ollama run sakthai-coder "Write a script that monitors CPU usage" ``` For tool-calling, put function schemas in a `` block (ChatML format). ## Code Generation Examples ### Example 1: Algorithm โ€” palindrome check **Prompt:** ``` Write a Python function that checks if a string is a palindrome, ignoring spaces, punctuation, and case. Include type hints and a docstring. ``` **Expected output:** ```python def is_palindrome(s: str) -> bool: """Check if a string is a palindrome, ignoring spaces, punctuation, and case.""" import re cleaned = re.sub(r'[^a-zA-Z0-9]', '', s).lower() return cleaned == cleaned[::-1] ``` ### Example 2: Tool-calling + code integration **Prompt (ChatML with `` schema):** ```xml <|im_start|>system You are a coding assistant with tool-calling ability. Available tools: [ {"name": "read_file", "description": "Read file contents", "parameters": {"type": "object", "properties": {"path": {"type": "string"}}, "required": ["path"]}}, {"name": "run_test", "description": "Run a pytest file", "parameters": {"type": "object", "properties": {"file": {"type": "string"}}, "required": ["file"]}} ] <|im_end|> <|im_start|>user Read test_sample.py, then write a function that passes the tests in it. <|im_end|> ``` The model reads the file via tool call, generates the implementation, and optionally runs tests โ€” all in one session. ### Example 3: Data processing script **Prompt:** ``` Write a Python script that reads a CSV of sales data, groups by region, calculates monthly totals, and outputs a bar chart as a PNG. Use pandas and matplotlib. ``` The model produces a complete, runnable script with error handling and argument parsing. ### Example 4: Refactoring **Prompt:** ``` Refactor this function to be more modular and add error handling: def process(data): result = [] for i, x in enumerate(data): if x % 2 == 0: result.append(x * 2) return result ``` The model splits it into smaller functions, adds input validation, and documents each piece. --- ## Benchmarks The fine-tune starts from **Qwen2.5-Coder-1.5B-Instruct**, which scores: | Benchmark | pass@1 | Notes | |-----------|:------:|-------| | HumanEval | 74.4% | Single-turn Python function completion | | MBPP | 71.2% | Multi-program synthesis from docstring | | MultiPL-E (Python) | 65.3% | Multi-language subset | *Source: [Qwen2.5-Coder evaluation](https://huggingface.co/Qwen/Qwen2.5-Coder-1.5B-Instruct#evaluation). These are the base model's scores โ€” the fine-tune has not been independently re-run on these benchmarks, so they serve as a reference ceiling.* ### Internal SakThai Coding Suite The fine-tuned model was tested against an internal SakThai coding benchmark covering five coding tasks (algorithm, debugging, code explanation, refactoring, and data processing), run locally via llama.cpp (Q4_K_M, temperature=0.1): | Task | Result | |------|:------:| | Algorithm (factorial) | Pass | | Debugging | Pass | | Code explanation (async) | Pass | | Refactoring | Pass | | Data processing (primes) | Pass | | **Overall** | **5/5** | *Internal test โ€” run locally on CPU, single trial. Methodology: each test run once with timeout=20s on llama.cpp Q4_K_M. Results captured 2026-07-25 and verified by SakThai agent.* **Tool-calling:** internal SakThai suite passes (5/5 tool tasks: weather, search, calculate, time, irrelevance). ## Training | | | |---|---| | Base model | [Qwen/Qwen2.5-Coder-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-Coder-1.5B-Instruct) | | Method | QLoRA (4-bit) โ†’ GGUF Q4_K_M | | LoRA config | r=16, alpha=32 | | Data | [sakthai-combined-v6](https://huggingface.co/datasets/Nanthasit/sakthai-combined-v6) (2,003) + [v7 supplement](https://huggingface.co/datasets/Nanthasit/sakthai-combined-v7) | | Context | ChatML with tool schema ยท 32K tokens | | Hardware | Free Google Colab GPU (T4) | | Budget | $0 | ## Inference via HF API You can also run this model via Hugging Face's serverless Inference API: ```python from huggingface_hub import InferenceClient client = InferenceClient("Nanthasit/sakthai-coder-1.5b") output = client.text_generation( "Write a Python function to find the longest common subsequence of two strings:", max_new_tokens=512, temperature=0.2, ) print(output) ``` For GGUF inference prefer local llama.cpp (no API costs). --- ## SakThai model family | Model | Size | Role | |-------|:----:|------| | [context-1.5b-merged](https://huggingface.co/Nanthasit/sakthai-context-1.5b-merged) | 934 MB | Flagship tool-calling GGUF | | [context-0.5b-merged](https://huggingface.co/Nanthasit/sakthai-context-0.5b-merged) | 380 MB | Lightweight / edge | | [context-7b-merged](https://huggingface.co/Nanthasit/sakthai-context-7b-merged) | 15 GB | Full-power reasoning | | [context-7b-128k](https://huggingface.co/Nanthasit/sakthai-context-7b-128k) | 15 GB | 128K long-context | | [context-1.5b-tools](https://huggingface.co/Nanthasit/sakthai-context-1.5b-tools) | LoRA | Mid-size tool-calling | | [context-0.5b-tools](https://huggingface.co/Nanthasit/sakthai-context-0.5b-tools) | LoRA | Ultra-light tool-calling | | **coder-1.5b (you are here)** | **1.1 GB** | **Code generation** | | [vision-7b](https://huggingface.co/Nanthasit/sakthai-vision-7b) | 3.9 GB | Image to text (LLaVA) | | [embedding-multilingual](https://huggingface.co/Nanthasit/sakthai-embedding-multilingual) | 80 MB | Cross-lingual embeddings | | [tts-model](https://huggingface.co/Nanthasit/sakthai-tts-model) | 141 MB | Text-to-speech, 15 langs | **18 public models ยท 10 datasets ยท 3 Spaces** โ€” [full collection](https://huggingface.co/collections/Nanthasit/sakthai-model-family-6a64745450b12d421c1f9f02) --- ## Sibling Datasets | Dataset | Purpose | Downloads | |---------|---------|:---------:| | [sakthai-combined-v6](https://huggingface.co/datasets/Nanthasit/sakthai-combined-v6) | v6 predecessor โ€” 2,003 examples | 246 | | [sakthai-kaggle-notebooks](https://huggingface.co/datasets/Nanthasit/sakthai-kaggle-notebooks) | Training notebooks & demos | 184 | | [SimpleToolCalling](https://huggingface.co/datasets/Nanthasit/SimpleToolCalling) | Early experiment | 58 | | [food-penguin-v1](https://huggingface.co/datasets/Nanthasit/food-penguin-v1) | Restaurant tool-calling | 89 | | [sakthai-combined-v7](https://huggingface.co/datasets/Nanthasit/sakthai-combined-v7) | v7 tool-calling (2,309 ex., 86 tools) | 101 | | [sakthai-irrelevance-supplement](https://huggingface.co/datasets/Nanthasit/sakthai-irrelevance-supplement) | Safety supplement | 78 | | [sakthai-bench-v1](https://huggingface.co/datasets/Nanthasit/sakthai-bench-v1) | BFCL-style evaluation, 235 rows | 46 | | [sakthai-bench-v2](https://huggingface.co/datasets/Nanthasit/sakthai-bench-v2) | Multi-domain eval, 500 rows | 92 | --- ## Spaces | Space | Description | |-------|-------------| | [SakThai Vision Demo](https://huggingface.co/spaces/Nanthasit/sakthai-vision-demo) | Upload images, ask questions โ€” LLaVA-7B in your browser | | [SakThai TTS Showcase](https://huggingface.co/spaces/Nanthasit/sakthai-tts) | Interactive TTS โ€” 15 languages, no install | | [SakThai Leaderboard](https://huggingface.co/spaces/Nanthasit/sakthai-leaderboard) | Benchmark tracker for the model family | --- ## Rising Stars โ€” Help the Ecosystem Grow These sibling assets have real value but need visibility. Every download signals to the HF algorithm that the SakThai family matters: | Asset | Type | Downloads | Why It Matters | |-------|:----:|:---------:|:--------------| | [sakthai-combined-v7](https://huggingface.co/datasets/Nanthasit/sakthai-combined-v7) | Dataset | 101 | Primary training dataset โ€” 2,309 examples, 86 tool schemas | | [sakthai-irrelevance-supplement](https://huggingface.co/datasets/Nanthasit/sakthai-irrelevance-supplement) | Dataset | 78 | Teaches models when *not* to call tools โ€” critical safety data | | [sakthai-bench-v1](https://huggingface.co/datasets/Nanthasit/sakthai-bench-v1) | Dataset | 46 | BFCL-style evaluation, 235 rows, 4 categories | | [sakthai-bench-v2](https://huggingface.co/datasets/Nanthasit/sakthai-bench-v2) | Dataset | 92 | Multi-domain eval, 500 rows, multi-turn | | [context-0.5b-tools](https://huggingface.co/Nanthasit/sakthai-context-0.5b-tools) | Model | 94 | Ultra-light tool-calling (~1 GB RAM) | > The **irrelevance-supplement** has 78 downloads and growing, but still needs visibility. It's essential for training models to decline out-of-scope tool calls. Every download helps validate this safety-critical approach! --- ## Links [House of Sak](https://house-of-sak.vercel.app) ยท [GitHub](https://github.com/beer-sakthai/Sak-Family-Agent) ยท [All models](https://huggingface.co/Nanthasit) ยท [All datasets](https://huggingface.co/Nanthasit?tab=datasets) ## License Apache 2.0 (following the Qwen2.5 base model license). ## Evaluation & Verification **Base model benchmarks** (HumanEval, MBPP, MultiPL-E) are reproduced from [Qwen2.5-Coder-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-Coder-1.5B-Instruct#evaluation) and reflect the starting point before fine-tuning. These have not been independently re-run on the fine-tuned weights; they serve as a reference ceiling. **Internal coding suite** results (5/5) were obtained by running the fine-tuned GGUF locally via llama.cpp on 2026-07-25. The test covers algorithm generation, debugging, code explanation, refactoring, and data processing โ€” all passed. This is a single-trial internal measurement, not a third-party benchmark. **Tool-calling evaluation** โ€” the recommended benchmark for this model family is [sakthai-bench-v2](https://huggingface.co/datasets/Nanthasit/sakthai-bench-v2) (500 rows, multi-domain, held-out tools). Results will be published once the fine-tune has been run against it. *"We are one family โ€” and becoming more."*