---
license: apache-2.0
language:
- en
library_name: transformers
pipeline_tag: text-generation
tags:
- code
- coder
- qwen2.5
- qwen2.5-coder
- gguf
- llama-cpp
- llama.cpp
- ollama
- code-generation
- tool-calling
- conversational
- cpu-inference
- small-language-model
- offline
- sakthai
- house-of-sak
base_model: Qwen/Qwen2.5-Coder-1.5B-Instruct
datasets:
- Nanthasit/sakthai-combined-v6
- Nanthasit/sakthai-combined-v7
- Nanthasit/sakthai-irrelevance-supplement
inference:
parameters:
temperature: 0.2
max_new_tokens: 1024
top_p: 0.9
widget:
- text: "Write a Python function that checks if a string is a palindrome, handling spaces and punctuation:"
output:
text: "```python\ndef is_palindrome(s: str) -> bool:\n \"\"\"Check if a string is a palindrome, ignoring spaces, punctuation, and case.\"\"\"\n import re\n cleaned = re.sub(r'[^a-zA-Z0-9]', '', s).lower()\n return cleaned == cleaned[::-1]\n```"
model-index:
- name: sakthai-coder-1.5b
results:
- task:
type: text-generation
dataset:
name: HumanEval
type: openai_humaneval
metrics:
- name: pass@1 (base model reference)
type: pass@1
value: 74.4
verified: false
- task:
type: text-generation
dataset:
name: MBPP
type: mbpp
metrics:
- name: pass@1 (base model reference)
type: pass@1
value: 71.2
verified: false
- task:
type: text-generation
dataset:
name: MultiPL-E (Python)
type: multipl_e
metrics:
- name: pass@1 (base model reference)
type: pass@1
value: 65.3
verified: false
- task:
type: text-generation
dataset:
name: SakThai Coding Suite (internal)
type: custom
metrics:
- name: pass@1 (fine-tuned model, internal single-trial)
type: pass@1
value: 100
verified: false
source: internal-local-llama-cpp-2026-07-25
---
SakThai Coder 1.5B ๐ป
Code + tool-calling ยท Qwen2.5-Coder-1.5B fine-tune ยท Q4_K_M GGUF for CPU
> The code specialist of the **SakThai** family โ Qwen2.5-Coder-1.5B fine-tuned for
> tool-calling and shipped as a CPU-friendly GGUF. Part of the
> [House of Sak](https://huggingface.co/Nanthasit). [Read the story โ](https://huggingface.co/Nanthasit)
## The Story Behind It
**Code, tool-calling, and conversation in one session โ on a single CPU, from a shelter.** This is the model Beer built when he realised the other SakThai models could call tools and generate text, but none of them specialised in *writing code* without losing their tool-calling edge.
Beer built the first SakThai models on free Google Colab GPUs from a shelter in Cork, Ireland โ with $0 budget, no GPU of his own, and no guarantee the QLoRA approach would hold for a code-specific fine-tune. This coder model was the risk: could Qwen2.5-Coder-1.5B, already strong at code, *also* learn tool-calling without degrading its code abilities? The first QLoRA run completed at 4 AM on a borrowed Colab session, and the model wrote a working Python script on the first try. Beer knew the approach worked.
This model runs on a 2020 laptop with 8 GB RAM โ no cloud API, no Inference Endpoint, no monthly bill. Just a GGUF file and llama.cpp.
> *"We are one family โ and becoming more."*
> โ Beer
### How You Can Help
- โญ **Leave a like** โ this model gives every developer a free offline coding assistant. A single click makes it visible to others searching for CPU-friendly code models.
- ๐ **Share it** with anyone who codes on an underpowered machine and needs tool-calling without the cloud tax.
- ๐ด **Fork it** on Hugging Face and build your own specialised code variant.
- ๐ฌ **Report your deployment story** โ Beer reads every issue and comment.
Every download, like, and share tells the algorithm: *this matters.*
---
## What it is
A **Q4_K_M GGUF** (1.12 GB) of **Qwen2.5-Coder-1.5B-Instruct**, QLoRA-fine-tuned on
[sakthai-combined-v6](https://huggingface.co/datasets/Nanthasit/sakthai-combined-v6) so it
can generate code *and* call tools. Runs on CPU via llama.cpp / Ollama.
## Architecture
Verified from the base model's `config.json` ([Qwen/Qwen2.5-Coder-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-Coder-1.5B-Instruct)):
| Parameter | Value |
|-----------|-------|
| Architecture | Qwen2ForCausalLM (`qwen2`) |
| Parameters | ~1.54 B |
| Hidden size | 1,536 |
| Layers | 28 |
| Attention heads | 12 (GQA, 2 KV heads) |
| Intermediate size | 8,960 |
| Vocabulary | 151,936 |
| Context length | 32,768 (32K) |
| RoPE theta | 1,000,000 |
| Base dtype | bfloat16 |
| Fine-tune | QLoRA โ GGUF Q4_K_M (this repo) |
## Quick start
```bash
# via llama.cpp (download + run)
wget https://huggingface.co/Nanthasit/sakthai-coder-1.5b/resolve/main/qwen2.5-coder-1.5b-instruct-q4_k_m.gguf
# Generate code
./llama-cli -m qwen2.5-coder-1.5b-instruct-q4_k_m.gguf \
-p "Write a Python function to merge two sorted lists:" \
-n 256 --temp 0.2
```
```bash
# via Ollama
echo 'FROM ./qwen2.5-coder-1.5b-instruct-q4_k_m.gguf' > Modelfile
ollama create sakthai-coder -f Modelfile
ollama run sakthai-coder "Write a script that monitors CPU usage"
```
```python
# via llama-cpp-python
from llama_cpp import Llama
llm = Llama(model_path="qwen2.5-coder-1.5b-instruct-q4_k_m.gguf", n_ctx=4096)
out = llm(
"Write a Python function to merge two sorted lists:",
max_tokens=256,
temperature=0.2,
echo=False,
)
print(out["choices"][0]["text"])
```
For tool-calling, put function schemas in a `` block (ChatML format).
## Code Generation Examples
### Example 1: Algorithm โ palindrome check
**Prompt:**
```
Write a Python function that checks if a string is a palindrome,
ignoring spaces, punctuation, and case. Include type hints and a docstring.
```
**Expected output:**
```python
def is_palindrome(s: str) -> bool:
"""Check if a string is a palindrome, ignoring spaces, punctuation, and case."""
import re
cleaned = re.sub(r'[^a-zA-Z0-9]', '', s).lower()
return cleaned == cleaned[::-1]
```
### Example 2: Tool-calling + code integration
**Prompt (ChatML with `` schema):**
```xml
<|im_start|>system
You are a coding assistant with tool-calling ability. Available tools:
[
{"name": "read_file", "description": "Read file contents", "parameters": {"type": "object", "properties": {"path": {"type": "string"}}, "required": ["path"]}},
{"name": "run_test", "description": "Run a pytest file", "parameters": {"type": "object", "properties": {"file": {"type": "string"}}, "required": ["file"]}}
]
<|im_end|>
<|im_start|>user
Read test_sample.py, then write a function that passes the tests in it.
<|im_end|>
```
The model reads the file via tool call, generates the implementation, and optionally runs tests โ all in one session.
### Example 3: Data processing script
**Prompt:**
```
Write a Python script that reads a CSV of sales data, groups by region,
calculates monthly totals, and outputs a bar chart as a PNG. Use pandas and matplotlib.
```
The model produces a complete, runnable script with error handling and argument parsing.
### Example 4: Refactoring
**Prompt:**
```
Refactor this function to be more modular and add error handling:
def process(data):
result = []
for i, x in enumerate(data):
if x % 2 == 0:
result.append(x * 2)
return result
```
The model splits it into smaller functions, adds input validation, and documents each piece.
---
## Benchmarks
The fine-tune starts from **Qwen2.5-Coder-1.5B-Instruct**, which scores:
| Benchmark | pass@1 | Notes |
|-----------|:------:|-------|
| HumanEval | 74.4% | Single-turn Python function completion |
| MBPP | 71.2% | Multi-program synthesis from docstring |
| MultiPL-E (Python) | 65.3% | Multi-language subset |
*Source: [Qwen2.5-Coder evaluation](https://huggingface.co/Qwen/Qwen2.5-Coder-1.5B-Instruct#evaluation). These are the base model's scores โ the fine-tune has not been independently re-run on these benchmarks, so they serve as a reference ceiling.*
### Internal SakThai Coding Suite
The fine-tuned model was tested against an internal SakThai coding benchmark covering five coding tasks (algorithm, debugging, code explanation, refactoring, and data processing), run locally via llama.cpp (Q4_K_M, temperature=0.1):
| Task | Result |
|------|:------:|
| Algorithm (factorial) | Pass |
| Debugging | Pass |
| Code explanation (async) | Pass |
| Refactoring | Pass |
| Data processing (primes) | Pass |
| **Overall** | **5/5** |
*Internal test โ run locally on CPU, single trial. Methodology: each test run once with timeout=20s on llama.cpp Q4_K_M. Results captured 2026-07-25 and verified by SakThai agent. Single-trial results are indicative, not a third-party benchmark.*
**Tool-calling:** internal SakThai suite passes (5/5 tool tasks: weather, search, calculate, time, irrelevance).
### Ecosystem Status (health check, 2026-07-30)
Source: [`health-check-sakthai-coder-1.5b-2026-07-30-4.yaml`](https://huggingface.co/Nanthasit/sakthai-coder-1.5b/blob/main/.eval_results/health-check-sakthai-coder-1.5b-2026-07-30-4.yaml) (automated cron evaluation).
| Signal | Value |
|--------|-------|
| Downloads rank | 11/19 family models (93 dl, velocity 14.3 dl/day, rank 9) |
| Card quality | 100/100 |
| Benchmark presence | model-index present, 4 entries (all unverified) |
| Repo hygiene | 40/100 โ see [Repo Status](#repo-status--housekeeping) |
| Overall health | 46/100 |
## Training
| | |
|---|---|
| Base model | [Qwen/Qwen2.5-Coder-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-Coder-1.5B-Instruct) |
| Method | QLoRA (4-bit) โ GGUF Q4_K_M |
| LoRA config | r=16, alpha=32 |
| Data | [sakthai-combined-v6](https://huggingface.co/datasets/Nanthasit/sakthai-combined-v6) + [v7](https://huggingface.co/datasets/Nanthasit/sakthai-combined-v7) (2,309 train / 115 test, verified 2026-07-31) + [irrelevance-supplement](https://huggingface.co/datasets/Nanthasit/sakthai-irrelevance-supplement) |
| Context | ChatML with tool schema ยท 32K tokens |
| Hardware | Free Google Colab GPU (T4) |
| Budget | $0 |
## Inference via HF API
You can also run this model via Hugging Face's serverless Inference API:
```python
from huggingface_hub import InferenceClient
client = InferenceClient("Nanthasit/sakthai-coder-1.5b")
output = client.text_generation(
"Write a Python function to find the longest common subsequence of two strings:",
max_new_tokens=512,
temperature=0.2,
)
print(output)
```
For GGUF inference prefer local llama.cpp (no API costs).
---
## SakThai model family
All 19 public models (downloads live, sizes verified via HF API on 2026-07-31 โ largest weight file):
| Model | Size | Role | Downloads |
|-------|:----:|------|:---------:|
| [context-1.5b-merged](https://huggingface.co/Nanthasit/sakthai-context-1.5b-merged) | 3.1 GB | Flagship tool-calling (safetensors + GGUF) | 1,599 |
| [context-0.5b-merged](https://huggingface.co/Nanthasit/sakthai-context-0.5b-merged) | 988 MB | Lightweight / edge (safetensors + GGUF) | 1,370 |
| [context-7b-merged](https://huggingface.co/Nanthasit/sakthai-context-7b-merged) | 15.2 GB | Full-power reasoning | 744 |
| [context-7b-128k](https://huggingface.co/Nanthasit/sakthai-context-7b-128k) | recipe | 128K long-context config (no weights) | 506 |
| [context-7b-tools](https://huggingface.co/Nanthasit/sakthai-context-7b-tools) | LoRA 20 MB | 7B tool-calling adapter | 399 |
| [embedding-multilingual](https://huggingface.co/Nanthasit/sakthai-embedding-multilingual) | 470 MB | Cross-lingual embeddings | 362 |
| [context-1.5b-tools](https://huggingface.co/Nanthasit/sakthai-context-1.5b-tools) | LoRA 8.7 MB | Mid-size tool-calling | 349 |
| [vision-7b](https://huggingface.co/Nanthasit/sakthai-vision-7b) | 4.1 GB | Image to text (LLaVA GGUF) | 186 |
| [tts-model](https://huggingface.co/Nanthasit/sakthai-tts-model) | 141 MB | Text-to-speech, 15 langs | 150 |
| [context-0.5b-tools](https://huggingface.co/Nanthasit/sakthai-context-0.5b-tools) | 988 MB | Ultra-light tool-calling | 94 |
| **coder-1.5b (you are here)** | **1.12 GB** | **Code generation + tool-calling** | **93** |
| [context-1.5b-tools-v2](https://huggingface.co/Nanthasit/sakthai-context-1.5b-tools-v2) | LoRA 74 MB | ๐ v2 tool-calling adapter | 0 |
| [context-1.5b-merged-v2](https://huggingface.co/Nanthasit/sakthai-context-1.5b-merged-v2) | 3.1 GB | ๐ v2 merged | 0 |
| [plus-1.5b](https://huggingface.co/Nanthasit/sakthai-plus-1.5b) | 3.1 GB | ๐ Plus merged | 0 |
| [plus-1.5b-lora](https://huggingface.co/Nanthasit/sakthai-plus-1.5b-lora) | LoRA 74 MB | ๐ Plus adapter | 0 |
| [plus-1.5b-coder](https://huggingface.co/Nanthasit/sakthai-plus-1.5b-coder) | โ | ๐ Plus coder (no weights yet) | 0 |
| [coder-browser-lora](https://huggingface.co/Nanthasit/sakthai-coder-browser-lora) | LoRA 74 MB | ๐ Browser-tool adapter | 0 |
| [coder-browser](https://huggingface.co/Nanthasit/sakthai-coder-browser) | 3.1 GB | ๐ Browser-tool merged | 0 |
| [coder-browser-gguf](https://huggingface.co/Nanthasit/sakthai-coder-browser-gguf) | 7.1 GB | ๐ Browser-tool F16 GGUF | 0 |
**19 public models ยท 13 datasets ยท 4 Spaces** โ [full collection](https://huggingface.co/collections/Nanthasit/sakthai-model-family-6a64745450b12d421c1f9f02)
---
## Sibling Datasets
| Dataset | Purpose | Downloads |
|---------|---------|:---------:|
| [sakthai-combined-v6](https://huggingface.co/datasets/Nanthasit/sakthai-combined-v6) | v6 predecessor โ tool-calling examples | 246 |
| [sakthai-kaggle-notebooks](https://huggingface.co/datasets/Nanthasit/sakthai-kaggle-notebooks) | Training notebooks & demos | 184 |
| [sakthai-combined-v7](https://huggingface.co/datasets/Nanthasit/sakthai-combined-v7) | v7 tool-calling (2,309 ex., 86 tools) | 101 |
| [sakthai-bench-v2](https://huggingface.co/datasets/Nanthasit/sakthai-bench-v2) | Multi-domain eval, 500 rows | 92 |
| [food-penguin-v1](https://huggingface.co/datasets/Nanthasit/food-penguin-v1) | Restaurant tool-calling | 89 |
| [sakthai-irrelevance-supplement](https://huggingface.co/datasets/Nanthasit/sakthai-irrelevance-supplement) | Safety supplement | 78 |
| [SimpleToolCalling](https://huggingface.co/datasets/Nanthasit/SimpleToolCalling) | Early experiment | 58 |
| [sakthai-bench-v1](https://huggingface.co/datasets/Nanthasit/sakthai-bench-v1) | BFCL-style evaluation, 235 rows | 46 |
*Downloads verified live 2026-07-31. The combined family is published as v6, v7, and v10 โ all public and linked above.*
---
## Spaces
| Space | Description |
|-------|-------------|
| [Web Agent](https://huggingface.co/spaces/Nanthasit/sakthai-web-agent) | Browser automation and tool-use agent |
| [SakThai TTS Showcase](https://huggingface.co/spaces/Nanthasit/sakthai-tts) | Interactive TTS โ 15 languages, no install |
| [SakThai Leaderboard](https://huggingface.co/spaces/Nanthasit/sakthai-leaderboard) | Benchmark tracker for the model family |
---
## Rising Stars โ Help the Ecosystem Grow
These sibling assets have real value but need visibility. Every download signals to the HF algorithm that the SakThai family matters:
| Asset | Type | Downloads | Why It Matters |
|-------|:----:|:---------:|:--------------|
| [sakthai-combined-v7](https://huggingface.co/datasets/Nanthasit/sakthai-combined-v7) | Dataset | 101 | Primary training dataset โ 2,309 examples, 86 tool schemas |
| [sakthai-irrelevance-supplement](https://huggingface.co/datasets/Nanthasit/sakthai-irrelevance-supplement) | Dataset | 78 | Teaches models when *not* to call tools โ critical safety data |
| [sakthai-bench-v1](https://huggingface.co/datasets/Nanthasit/sakthai-bench-v1) | Dataset | 46 | BFCL-style evaluation, 235 rows, 4 categories |
| [sakthai-bench-v2](https://huggingface.co/datasets/Nanthasit/sakthai-bench-v2) | Dataset | 92 | Multi-domain eval, 500 rows, multi-turn |
| [context-0.5b-tools](https://huggingface.co/Nanthasit/sakthai-context-0.5b-tools) | Model | 94 | Ultra-light tool-calling (~1 GB RAM) |
> The **irrelevance-supplement** has 78 downloads and growing, but still needs visibility. It's essential for training models to decline out-of-scope tool calls. Every download helps validate this safety-critical approach!
---
## Repo Status & Housekeeping
This repository accidentally carries a stray development environment (`.venv/`, `.pytest_cache/`, `.hypothesis/`, `.ruff_cache/`) from an over-eager push. The model artifact (`qwen2.5-coder-1.5b-instruct-q4_k_m.gguf`, 1.12 GB) is unaffected. A cleanup commit is planned; the automated health check currently deducts hygiene points for these files. See the [health-check YAML](https://huggingface.co/Nanthasit/sakthai-coder-1.5b/blob/main/.eval_results/health-check-sakthai-coder-1.5b-2026-07-30-4.yaml) for details.
---
## Links
[House of Sak](https://house-of-sak.vercel.app) ยท
[GitHub](https://github.com/beer-sakthai/Sak-Family-Agent) ยท
[All models](https://huggingface.co/Nanthasit) ยท
[All datasets](https://huggingface.co/Nanthasit?tab=datasets)
## License
Apache 2.0 (following the Qwen2.5 base model license).
## Evaluation & Verification
**Base model benchmarks** (HumanEval, MBPP, MultiPL-E) are reproduced from
[Qwen2.5-Coder-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-Coder-1.5B-Instruct#evaluation)
and reflect the starting point before fine-tuning. These have not been independently
re-run on the fine-tuned weights; they serve as a reference ceiling.
**Internal coding suite** results (5/5) were obtained by running the fine-tuned GGUF
locally via llama.cpp on 2026-07-25. The test covers algorithm generation, debugging,
code explanation, refactoring, and data processing โ all passed. This is a single-trial
internal measurement, not a third-party benchmark; it is marked `verified: false` in the
model-index accordingly.
**Tool-calling evaluation** โ the recommended benchmark for this model family is
[sakthai-bench-v2](https://huggingface.co/datasets/Nanthasit/sakthai-bench-v2)
(500 rows, multi-domain, held-out tools). Results will be published once the
fine-tune has been run against it.
*"We are one family โ and becoming more."*