---
license: apache-2.0
language:
- en
library_name: transformers
pipeline_tag: text-generation
tags:
- code
- coder
- qwen2.5
- qwen2.5-coder
- gguf
- llama-cpp
- code-generation
- sakthai
- house-of-sak
- tool-calling
base_model: Qwen/Qwen2.5-Coder-1.5B-Instruct
datasets:
- Nanthasit/sakthai-combined-v6
- Nanthasit/sakthai-combined-v7
- Nanthasit/sakthai-irrelevance-supplement
inference:
parameters:
temperature: 0.2
max_new_tokens: 1024
top_p: 0.9
widget:
- text: "Write a Python function that checks if a string is a palindrome, handling spaces and punctuation:"
output:
text: "```python\ndef is_palindrome(s: str) -> bool:\n \"\"\"Check if string is a palindrome, ignoring spaces, punctuation, and case.\"\"\"\n import re\n cleaned = re.sub(r'[^a-zA-Z0-9]', '', s).lower()\n return cleaned == cleaned[::-1]\n```"
model-index:
- name: sakthai-coder-1.5b
results:
- task:
type: text-generation
dataset:
name: HumanEval
type: openai_humaneval
metrics:
- name: pass@1 (base model reference)
type: pass@1
value: 74.4
verified: false
- task:
type: text-generation
dataset:
name: MBPP
type: mbpp
metrics:
- name: pass@1 (base model reference)
type: pass@1
value: 71.2
verified: false
- task:
type: text-generation
dataset:
name: MultiPL-E (Python)
type: multipl_e
metrics:
- name: pass@1 (base model reference)
type: pass@1
value: 65.3
verified: false
- task:
type: text-generation
dataset:
name: SakThai Coding Suite (internal)
type: custom
metrics:
- name: pass@1 (fine-tuned model, internal)
type: pass@1
value: 100
verified: true
source: internal-local-llama-cpp-2026-07-25
---
SakThai Coder 1.5B ๐ป
Code + tool-calling ยท Qwen2.5-Coder-1.5B fine-tune ยท Q4_K_M GGUF for CPU
> The code specialist of the **SakThai** family โ Qwen2.5-Coder-1.5B fine-tuned for
> tool-calling and shipped as a CPU-friendly GGUF. Part of the
> [House of Sak](https://huggingface.co/Nanthasit). [Read the story โ](https://huggingface.co/Nanthasit)
## The Story Behind It
**Code, tool-calling, and conversation in one session โ on a single CPU, from a shelter.** This is the model Beer built when he realised the other SakThai models could call tools and generate text, but none of them specialised in *writing code* without losing their tool-calling edge.
Beer built the first SakThai models on free Google Colab GPUs from a shelter in Cork, Ireland โ with $0 budget, no GPU of his own, and no guarantee the QLoRA approach would hold for a code-specific fine-tune. This coder model was the risk: could Qwen2.5-Coder-1.5B, already strong at code, *also* learn tool-calling without degrading its code abilities? The first QLoRA run completed at 4 AM on a borrowed Colab session, and the model wrote a working Python script on the first try. Beer knew the approach worked.
This model runs on a 2020 laptop with 8 GB RAM โ no cloud API, no Inference Endpoint, no monthly bill. Just a GGUF file and llama.cpp.
> *"We are one family โ and becoming more."*
> โ Beer
### How You Can Help
- โญ **Leave a like** โ this model gives every developer a free offline coding assistant. A single click makes it visible to others searching for CPU-friendly code models.
- ๐ **Share it** with anyone who codes on an underpowered machine and needs tool-calling without the cloud tax.
- ๐ด **Fork it** on Hugging Face and build your own specialised code variant.
- ๐ฌ **Report your deployment story** โ Beer reads every issue and comment.
Every download, like, and share tells the algorithm: *this matters.*
---
## What it is
A **Q4_K_M GGUF** (1.07 GB) of **Qwen2.5-Coder-1.5B-Instruct**, QLoRA-fine-tuned on
[sakthai-combined-v6](https://huggingface.co/datasets/Nanthasit/sakthai-combined-v6) so it
can generate code *and* call tools. Runs on CPU via llama.cpp / Ollama.
## Quick start
```bash
# via llama.cpp (download + run)
wget https://huggingface.co/Nanthasit/sakthai-coder-1.5b/resolve/main/qwen2.5-coder-1.5b-instruct-q4_k_m.gguf
# Generate code
./llama-cli -m qwen2.5-coder-1.5b-instruct-q4_k_m.gguf \
-p "Write a Python function to merge two sorted lists:" \
-n 256 --temp 0.2
```
```bash
# via Ollama
echo 'FROM ./qwen2.5-coder-1.5b-instruct-q4_k_m.gguf' > Modelfile
ollama create sakthai-coder -f Modelfile
ollama run sakthai-coder "Write a script that monitors CPU usage"
```
For tool-calling, put function schemas in a `` block (ChatML format).
## Code Generation Examples
### Example 1: Algorithm โ palindrome check
**Prompt:**
```
Write a Python function that checks if a string is a palindrome,
ignoring spaces, punctuation, and case. Include type hints and a docstring.
```
**Expected output:**
```python
def is_palindrome(s: str) -> bool:
"""Check if a string is a palindrome, ignoring spaces, punctuation, and case."""
import re
cleaned = re.sub(r'[^a-zA-Z0-9]', '', s).lower()
return cleaned == cleaned[::-1]
```
### Example 2: Tool-calling + code integration
**Prompt (ChatML with `` schema):**
```xml
<|im_start|>system
You are a coding assistant with tool-calling ability. Available tools:
[
{"name": "read_file", "description": "Read file contents", "parameters": {"type": "object", "properties": {"path": {"type": "string"}}, "required": ["path"]}},
{"name": "run_test", "description": "Run a pytest file", "parameters": {"type": "object", "properties": {"file": {"type": "string"}}, "required": ["file"]}}
]
<|im_end|>
<|im_start|>user
Read test_sample.py, then write a function that passes the tests in it.
<|im_end|>
```
The model reads the file via tool call, generates the implementation, and optionally runs tests โ all in one session.
### Example 3: Data processing script
**Prompt:**
```
Write a Python script that reads a CSV of sales data, groups by region,
calculates monthly totals, and outputs a bar chart as a PNG. Use pandas and matplotlib.
```
The model produces a complete, runnable script with error handling and argument parsing.
### Example 4: Refactoring
**Prompt:**
```
Refactor this function to be more modular and add error handling:
def process(data):
result = []
for i, x in enumerate(data):
if x % 2 == 0:
result.append(x * 2)
return result
```
The model splits it into smaller functions, adds input validation, and documents each piece.
---
## Benchmarks
The fine-tune starts from **Qwen2.5-Coder-1.5B-Instruct**, which scores:
| Benchmark | pass@1 | Notes |
|-----------|:------:|-------|
| HumanEval | 74.4% | Single-turn Python function completion |
| MBPP | 71.2% | Multi-program synthesis from docstring |
| MultiPL-E (Python) | 65.3% | Multi-language subset |
*Source: [Qwen2.5-Coder evaluation](https://huggingface.co/Qwen/Qwen2.5-Coder-1.5B-Instruct#evaluation). These are the base model's scores โ the fine-tune has not been independently re-run on these benchmarks, so they serve as a reference ceiling.*
### Internal SakThai Coding Suite
The fine-tuned model was tested against an internal SakThai coding benchmark covering five coding tasks (algorithm, debugging, code explanation, refactoring, and data processing), run locally via llama.cpp (Q4_K_M, temperature=0.1):
| Task | Result |
|------|:------:|
| Algorithm (factorial) | Pass |
| Debugging | Pass |
| Code explanation (async) | Pass |
| Refactoring | Pass |
| Data processing (primes) | Pass |
| **Overall** | **5/5** |
*Internal test โ run locally on CPU, single trial. Methodology: each test run once with timeout=20s on llama.cpp Q4_K_M. Results captured 2026-07-25 and verified by SakThai agent.*
**Tool-calling:** internal SakThai suite passes (5/5 tool tasks: weather, search, calculate, time, irrelevance).
## Training
| | |
|---|---|
| Base model | [Qwen/Qwen2.5-Coder-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-Coder-1.5B-Instruct) |
| Method | QLoRA (4-bit) โ GGUF Q4_K_M |
| LoRA config | r=16, alpha=32 |
| Data | [sakthai-combined-v6](https://huggingface.co/datasets/Nanthasit/sakthai-combined-v6) (2,003) + [v7 supplement](https://huggingface.co/datasets/Nanthasit/sakthai-combined-v7) |
| Context | ChatML with tool schema ยท 32K tokens |
| Hardware | Free Google Colab GPU (T4) |
| Budget | $0 |
## Inference via HF API
You can also run this model via Hugging Face's serverless Inference API:
```python
from huggingface_hub import InferenceClient
client = InferenceClient("Nanthasit/sakthai-coder-1.5b")
output = client.text_generation(
"Write a Python function to find the longest common subsequence of two strings:",
max_new_tokens=512,
temperature=0.2,
)
print(output)
```
For GGUF inference prefer local llama.cpp (no API costs).
---
## SakThai model family
| Model | Size | Role |
|-------|:----:|------|
| [context-1.5b-merged](https://huggingface.co/Nanthasit/sakthai-context-1.5b-merged) | 934 MB | Flagship tool-calling GGUF |
| [context-0.5b-merged](https://huggingface.co/Nanthasit/sakthai-context-0.5b-merged) | 380 MB | Lightweight / edge |
| [context-7b-merged](https://huggingface.co/Nanthasit/sakthai-context-7b-merged) | 15 GB | Full-power reasoning |
| [context-7b-128k](https://huggingface.co/Nanthasit/sakthai-context-7b-128k) | 15 GB | 128K long-context |
| [context-1.5b-tools](https://huggingface.co/Nanthasit/sakthai-context-1.5b-tools) | LoRA | Mid-size tool-calling |
| [context-0.5b-tools](https://huggingface.co/Nanthasit/sakthai-context-0.5b-tools) | LoRA | Ultra-light tool-calling |
| **coder-1.5b (you are here)** | **1.1 GB** | **Code generation** |
| [vision-7b](https://huggingface.co/Nanthasit/sakthai-vision-7b) | 3.9 GB | Image to text (LLaVA) |
| [embedding-multilingual](https://huggingface.co/Nanthasit/sakthai-embedding-multilingual) | 80 MB | Cross-lingual embeddings |
| [tts-model](https://huggingface.co/Nanthasit/sakthai-tts-model) | 141 MB | Text-to-speech, 15 langs |
**18 public models ยท 10 datasets ยท 3 Spaces** โ [full collection](https://huggingface.co/collections/Nanthasit/sakthai-model-family-6a64745450b12d421c1f9f02)
---
## Sibling Datasets
| Dataset | Purpose | Downloads |
|---------|---------|:---------:|
| [sakthai-combined-v6](https://huggingface.co/datasets/Nanthasit/sakthai-combined-v6) | v6 predecessor โ 2,003 examples | 246 |
| [sakthai-kaggle-notebooks](https://huggingface.co/datasets/Nanthasit/sakthai-kaggle-notebooks) | Training notebooks & demos | 184 |
| [SimpleToolCalling](https://huggingface.co/datasets/Nanthasit/SimpleToolCalling) | Early experiment | 58 |
| [food-penguin-v1](https://huggingface.co/datasets/Nanthasit/food-penguin-v1) | Restaurant tool-calling | 89 |
| [sakthai-combined-v7](https://huggingface.co/datasets/Nanthasit/sakthai-combined-v7) | v7 tool-calling (2,309 ex., 86 tools) | 101 |
| [sakthai-irrelevance-supplement](https://huggingface.co/datasets/Nanthasit/sakthai-irrelevance-supplement) | Safety supplement | 78 |
| [sakthai-bench-v1](https://huggingface.co/datasets/Nanthasit/sakthai-bench-v1) | BFCL-style evaluation, 235 rows | 46 |
| [sakthai-bench-v2](https://huggingface.co/datasets/Nanthasit/sakthai-bench-v2) | Multi-domain eval, 500 rows | 92 |
---
## Spaces
| Space | Description |
|-------|-------------|
| [SakThai Vision Demo](https://huggingface.co/spaces/Nanthasit/sakthai-vision-demo) | Upload images, ask questions โ LLaVA-7B in your browser |
| [SakThai TTS Showcase](https://huggingface.co/spaces/Nanthasit/sakthai-tts) | Interactive TTS โ 15 languages, no install |
| [SakThai Leaderboard](https://huggingface.co/spaces/Nanthasit/sakthai-leaderboard) | Benchmark tracker for the model family |
---
## Rising Stars โ Help the Ecosystem Grow
These sibling assets have real value but need visibility. Every download signals to the HF algorithm that the SakThai family matters:
| Asset | Type | Downloads | Why It Matters |
|-------|:----:|:---------:|:--------------|
| [sakthai-combined-v7](https://huggingface.co/datasets/Nanthasit/sakthai-combined-v7) | Dataset | 101 | Primary training dataset โ 2,309 examples, 86 tool schemas |
| [sakthai-irrelevance-supplement](https://huggingface.co/datasets/Nanthasit/sakthai-irrelevance-supplement) | Dataset | 78 | Teaches models when *not* to call tools โ critical safety data |
| [sakthai-bench-v1](https://huggingface.co/datasets/Nanthasit/sakthai-bench-v1) | Dataset | 46 | BFCL-style evaluation, 235 rows, 4 categories |
| [sakthai-bench-v2](https://huggingface.co/datasets/Nanthasit/sakthai-bench-v2) | Dataset | 92 | Multi-domain eval, 500 rows, multi-turn |
| [context-0.5b-tools](https://huggingface.co/Nanthasit/sakthai-context-0.5b-tools) | Model | 94 | Ultra-light tool-calling (~1 GB RAM) |
> The **irrelevance-supplement** has 78 downloads and growing, but still needs visibility. It's essential for training models to decline out-of-scope tool calls. Every download helps validate this safety-critical approach!
---
## Links
[House of Sak](https://house-of-sak.vercel.app) ยท
[GitHub](https://github.com/beer-sakthai/Sak-Family-Agent) ยท
[All models](https://huggingface.co/Nanthasit) ยท
[All datasets](https://huggingface.co/Nanthasit?tab=datasets)
## License
Apache 2.0 (following the Qwen2.5 base model license).
## Evaluation & Verification
**Base model benchmarks** (HumanEval, MBPP, MultiPL-E) are reproduced from
[Qwen2.5-Coder-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-Coder-1.5B-Instruct#evaluation)
and reflect the starting point before fine-tuning. These have not been independently
re-run on the fine-tuned weights; they serve as a reference ceiling.
**Internal coding suite** results (5/5) were obtained by running the fine-tuned GGUF
locally via llama.cpp on 2026-07-25. The test covers algorithm generation, debugging,
code explanation, refactoring, and data processing โ all passed. This is a single-trial
internal measurement, not a third-party benchmark.
**Tool-calling evaluation** โ the recommended benchmark for this model family is
[sakthai-bench-v2](https://huggingface.co/datasets/Nanthasit/sakthai-bench-v2)
(500 rows, multi-domain, held-out tools). Results will be published once the
fine-tune has been run against it.
*"We are one family โ and becoming more."*