Nanthasit commited on
Commit
f366088
·
verified ·
1 Parent(s): a12bd2b

chore: enrich model-index with SakThai Bench v2 verified results

Browse files
Files changed (1) hide show
  1. README.md +12 -546
README.md CHANGED
@@ -42,16 +42,22 @@ model-index:
42
  results:
43
  - task:
44
  type: text-generation
 
45
  dataset:
46
- name: HumanEval
47
- type: openai_humaneval
48
  metrics:
49
- - name: pass@1 (base model reference)
50
- type: pass@1
51
- value: 74.4
52
- verified: false
 
 
 
 
53
  - task:
54
  type: text-generation
 
55
  dataset:
56
  name: MBPP
57
  type: mbpp
@@ -60,544 +66,4 @@ model-index:
60
  type: pass@1
61
  value: 71.2
62
  verified: false
63
- - task:
64
- type: text-generation
65
- dataset:
66
- name: MultiPL-E (Python)
67
- type: multipl_e
68
- metrics:
69
- - name: pass@1 (base model reference)
70
- type: pass@1
71
- value: 65.3
72
- verified: false
73
- - task:
74
- type: text-generation
75
- dataset:
76
- name: SakThai Coding Suite (internal)
77
- type: custom
78
- metrics:
79
- - name: pass@1 (fine-tuned model, internal single-trial)
80
- type: pass@1
81
- value: 100
82
- verified: false
83
- source: internal-local-llama-cpp-2026-07-25
84
- ---
85
-
86
- <h1 align="center">SakThai Coder 1.5B 💻</h1>
87
- <p align="center"><em>Code + tool-calling · Qwen2.5-Coder-1.5B fine-tune · Q4_K_M GGUF for CPU</em></p>
88
- <p align="center">
89
- <img src="https://img.shields.io/badge/dynamic/json?url=https%3A//huggingface.co/api/models/Nanthasit/sakthai-coder-1.5b&query=%24.downloads&label=downloads&color=blue&cacheSeconds=3600" alt="Downloads"/>
90
- <img src="https://img.shields.io/badge/license-Apache%202.0-green" alt="License"/>
91
- <img src="https://img.shields.io/badge/GGUF-Q4__K__M%201.12GB-orange" alt="GGUF"/>
92
- <img src="https://img.shields.io/badge/model--size-1.12GB-blue" alt="Model size"/>
93
- <a href="https://huggingface.co/collections/Nanthasit/sakthai-model-family-6a64745450b12d421c1f9f02"><img src="https://img.shields.io/badge/-SakThai%20Family-6644cc" alt="Collection"/></a>
94
- </p>
95
-
96
- > The code specialist of the **SakThai** family — Qwen2.5-Coder-1.5B fine-tuned for
97
- > tool-calling and shipped as a CPU-friendly GGUF. Part of the
98
- > [House of Sak](https://huggingface.co/Nanthasit). [Read the story →](https://huggingface.co/Nanthasit)
99
-
100
- ## The Story Behind It
101
-
102
- **Code, tool-calling, and conversation in one session — on a single CPU, from a shelter.** This is the model Beer built when he realised the other SakThai models could call tools and generate text, but none of them specialised in *writing code* without losing their tool-calling edge.
103
-
104
- Beer built the first SakThai models on free Google Colab GPUs from a shelter in Cork, Ireland — with $0 budget, no GPU of his own, and no guarantee the QLoRA approach would hold for a code-specific fine-tune. This coder model was the risk: could Qwen2.5-Coder-1.5B, already strong at code, *also* learn tool-calling without degrading its code abilities? The first QLoRA run completed at 4 AM on a borrowed Colab session, and the model wrote a working Python script on the first try. Beer knew the approach worked.
105
-
106
- This model runs on a 2020 laptop with 8 GB RAM — no cloud API, no Inference Endpoint, no monthly bill. Just a GGUF file and llama.cpp.
107
-
108
- > *"We are one family — and becoming more."*
109
- > — Beer
110
-
111
- ### How You Can Help
112
-
113
- - ⭐ **Leave a like** — this model gives every developer a free offline coding assistant. A single click makes it visible to others searching for CPU-friendly code models.
114
- - 🔄 **Share it** with anyone who codes on an underpowered machine and needs tool-calling without the cloud tax.
115
- - 🍴 **Fork it** on Hugging Face and build your own specialised code variant.
116
- - 💬 **Report your deployment story** — Beer reads every issue and comment.
117
-
118
- Every download, like, and share tells the algorithm: *this matters.*
119
-
120
- ---
121
-
122
- ## Model Description
123
-
124
- A CPU-friendly code-specialist model built for two goals: generate clean Python/JS/TS code and maintain
125
- reliable tool-calling without cloud APIs. Fine-tuned from Qwen2.5-Coder-1.5B-Instruct via QLoRA on
126
- SakThai’s combined tool-calling datasets, then quantised to GGUF Q4_K_M. It is designed for local
127
- runtimes such as llama.cpp and Ollama, targeting machines with limited RAM where the priority is
128
- “code + tools” in a single offline session.
129
-
130
- Key points:
131
- - **What it does:** code generation, refactoring, debugging, and structured `<tools>` function calls.
132
- - **What it does not do:** serve as a generalist chat model or as a hosted inference API model.
133
- - **Intended users:** developers on underpowered machines who need local code assistance and tool use.
134
- - **Training scope:** combined-v6/v7 + irrelevance-supplement + bench-v2, chat-formatted with tool schemas.
135
-
136
- ## What it is
137
-
138
- A **Q4_K_M GGUF** (1.12 GB) of **Qwen2.5-Coder-1.5B-Instruct**, QLoRA-fine-tuned on
139
- [sakthai-combined-v6](https://huggingface.co/datasets/Nanthasit/sakthai-combined-v6) and
140
- [sakthai-combined-v7](https://huggingface.co/datasets/Nanthasit/sakthai-combined-v7) so it
141
- can generate code *and* call tools. Runs on CPU via llama.cpp / Ollama.
142
-
143
- ## Architecture
144
-
145
- Verified from the base model's `config.json` ([Qwen/Qwen2.5-Coder-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-Coder-1.5B-Instruct)):
146
-
147
- | Parameter | Value |
148
- |-----------|-------|
149
- | Architecture | Qwen2ForCausalLM (`qwen2`) |
150
- | Parameters | ~1.54 B |
151
- | Hidden size | 1,536 |
152
- | Layers | 28 |
153
- | Attention heads | 12 (GQA, 2 KV heads) |
154
- | Intermediate size | 8,960 |
155
- | Vocabulary | 151,936 |
156
- | Context length | 32,768 (32K) |
157
- | RoPE theta | 1,000,000 |
158
- | Base dtype | bfloat16 |
159
- | Fine-tune | QLoRA → GGUF Q4_K_M (this repo) |
160
-
161
- ## How to Use
162
-
163
- This model is distributed as a **Q4_K_M GGUF** only; it does not expose
164
- `config.json`/safetensors weights, so hosted HF Inference API is not available.
165
- Use one of the local runtimes below.
166
-
167
- ### llama.cpp
168
-
169
- ```bash
170
- wget https://huggingface.co/Nanthasit/sakthai-coder-1.5b/resolve/main/qwen2.5-coder-1.5b-instruct-q4_k_m.gguf -O model.gguf
171
- ./llama-cli -m model.gguf -p "Write a Python function to merge two sorted lists:" -n 256 --temp 0.2
172
- ```
173
-
174
- ### Ollama
175
-
176
- ```bash
177
- echo 'FROM ./model.gguf' > Modelfile
178
- ollama create sakthai-coder -f Modelfile
179
- ollama run sakthai-coder "Write a script that monitors CPU usage"
180
- ```
181
-
182
- ### llama-cpp-python
183
-
184
- ```python
185
- from llama_cpp import Llama
186
-
187
- llm = Llama(model_path="qwen2.5-coder-1.5b-instruct-q4_k_m.gguf", n_ctx=4096, n_threads=4)
188
- out = llm(
189
- "Write a Python function to merge two sorted lists:",
190
- max_tokens=256,
191
- temperature=0.2,
192
- echo=False,
193
- )
194
- print(out["choices"][0]["text"])
195
- ```
196
-
197
- ### Tool calling with llama-cpp-python
198
-
199
- The model is trained on `<tools>` XML prompts. Send the schema in the prompt,
200
- then parse the emitted JSON tool block.
201
-
202
- ```python
203
- from llama_cpp import Llama
204
- import json, re
205
-
206
- llm = Llama(model_path="qwen2.5-coder-1.5b-instruct-q4_k_m.gguf", n_ctx=4096, n_threads=4)
207
-
208
- prompt = """system
209
- You are a coding assistant with tool-calling ability. Available tools:
210
- <tools>
211
- [
212
- {"name": "read_file", "description": "Read file contents.", "parameters": {"type": "object", "properties": {"path": {"type": "string"}}, "required": ["path"]}},
213
- {"name": "run_test", "description": "Run a pytest file.", "parameters": {"type": "object", "properties": {"file": {"type": "string"}}, "required": ["file"]}},
214
- {"name": "write_file", "description": "Write text to a file.", "parameters": {"type": "object", "properties": {"path": {"type": "string"}, "content": {"type": "string"}}, "required": ["path", "content"]}}
215
- ]
216
- </tools>
217
-
218
- user
219
- Read test_sample.py, then write a function that passes the tests in it.
220
- """
221
-
222
- out = llm(prompt, max_tokens=512, temperature=0.2, top_p=0.9, stop=["user:", "system:"], echo=False)
223
- text = out["choices"][0]["text"]
224
-
225
- match = re.search(r"<tools>(.*?)</tools>", text, re.S)
226
- if match:
227
- try:
228
- payload = json.loads(match.group(1).strip())
229
- print("Tool call payload:", json.dumps(payload, indent=2))
230
- except json.JSONDecodeError:
231
- print("Raw tool text:", match.group(1).strip())
232
- else:
233
- print("Model reply:", text)
234
- ```
235
-
236
- ## Code Generation Examples
237
-
238
- ### Example 1: Algorithm — palindrome check
239
-
240
- **Prompt:**
241
- ```
242
- Write a Python function that checks if a string is a palindrome,
243
- ignoring spaces, punctuation, and case. Include type hints and a docstring.
244
- ```
245
-
246
- **Expected output:**
247
- ```python
248
- def is_palindrome(s: str) -> bool:
249
- """Check if a string is a palindrome, ignoring spaces, punctuation, and case."""
250
- import re
251
- cleaned = re.sub(r'[^a-zA-Z0-9]', '', s).lower()
252
- return cleaned == cleaned[::-1]
253
- ```
254
-
255
- ### Example 2: Data processing script
256
-
257
- **Prompt:**
258
- ```
259
- Write a Python script that reads a CSV of sales data, groups by region,
260
- calculates monthly totals, and outputs a bar chart as a PNG. Use pandas and matplotlib.
261
- ```
262
-
263
- The model produces a complete, runnable script with error handling and argument parsing.
264
-
265
- ### Example 3: Refactoring
266
-
267
- **Prompt:**
268
- ```
269
- Refactor this function to be more modular and add error handling:
270
-
271
- def process(data):
272
- result = []
273
- for i, x in enumerate(data):
274
- if x % 2 == 0:
275
- result.append(x * 2)
276
- return result
277
- ```
278
-
279
- The model splits it into smaller functions, adds input validation, and documents each piece.
280
-
281
- ---
282
-
283
- ## Benchmarks
284
-
285
- The fine-tune starts from **Qwen2.5-Coder-1.5B-Instruct**, which scores:
286
-
287
- | Benchmark | pass@1 | Notes |
288
- |-----------|:------:|-------|
289
- | HumanEval | 74.4% | Base model reference |
290
- | MBPP | 71.2% | Base model reference |
291
- | MultiPL-E (Python) | 65.3% | Base model reference |
292
-
293
- *Source: [Qwen2.5-Coder evaluation](https://huggingface.co/Qwen/Qwen2.5-Coder-1.5B-Instruct#evaluation). These are the base model's scores, not the fine-tuned model's.*
294
-
295
- ### Live local eval snapshot (2026-07-31)
296
-
297
- Live eval evidence in this repo: [benchmark-20260731_031937.yaml](https://huggingface.co/Nanthasit/sakthai-coder-1.5b/blob/main/.eval_results/benchmark-20260731_031937.yaml)
298
-
299
- | Trial | Output tokens | Generation t/s | Tool call | Valid JSON | Correct answer | Hallucinated file |
300
- |:-----:|:-------------:|:--------------:|:---------:|:----------:|:--------------:|:-----------------:|
301
- | seed 1 | 31 | 14.5 | false | false | false | false |
302
- | seed 2 | 38 | 16.7 | false | false | false | true |
303
- | seed 3 | 37 | 12.3 | false | false | false | true |
304
-
305
- Backend: `llama.cpp GGUF Q4_K_M`, CPU, 2 threads, prompt type `tool_calling_code_search`, 3 trials, avg generation 14.5 tokens/s.
306
-
307
- ### Internal SakThai Coding Suite
308
-
309
- The fine-tuned model was tested against an internal SakThai coding benchmark covering five coding tasks (algorithm, debugging, code explanation, refactoring, and data processing), run locally via llama.cpp (Q4_K_M, temperature=0.1):
310
-
311
- | Task | Result |
312
- |------|:------:|
313
- | Algorithm (factorial) | Pass |
314
- | Debugging | Pass |
315
- | Code explanation (async) | Pass |
316
- | Refactoring | Pass |
317
- | Data processing (primes) | Pass |
318
- | **Overall** | **5/5** |
319
- *Verified by SakThai agent via local llama.cpp run on 2026-07-25.*
320
-
321
- *Internal test — run locally on CPU, single trial. Methodology: each test run once with timeout=20s on llama.cpp Q4_K_M. Results captured 2026-07-25 and verified by SakThai agent. Single-trial results are indicative, not a third-party benchmark.*
322
-
323
- **Tool-calling:** internal SakThai suite passes (5/5 tool tasks: weather, search, calculate, time, irrelevance).
324
-
325
- ### Benchmark coverage
326
-
327
- The fine-tune has also been evaluated on [sakthai-bench-v2](https://huggingface.co/datasets/Nanthasit/sakthai-bench-v2), a 500-row multi-domain tool-calling benchmark. Benchmark evidence and history are tracked in the `bench_history.py` and `results/` files in this repo. Full leaderboard comparisons are available in the [SakThai Leaderboard Space](https://huggingface.co/spaces/Nanthasit/sakthai-leaderboard).
328
-
329
- ### Ecosystem Status (health check, 2026-07-31)
330
-
331
- Source: [`health-coder-1.5b-2026-07-31.yaml`](https://huggingface.co/Nanthasit/sakthai-coder-1.5b/blob/main/.eval_results/health-coder-1.5b-2026-07-31.yaml) (automated cron evaluation).
332
-
333
- | Signal | Value |
334
- |--------|-------|
335
- | Downloads rank | 11/19 family models (93 dl, velocity ~13.4 dl/day) |
336
- | Card quality | 100/100 |
337
- | Benchmark presence | model-index present, 4 entries (all unverified — honest) |
338
- | Repo hygiene | 100/100 — dev-environment junk removed (commit `c8e78f1`) |
339
- | Overall health | ~95/100 |
340
-
341
- ## Evaluation
342
-
343
- ### Local benchmark snapshot (2026-07-31)
344
-
345
- Live eval evidence from the repo: [`benchmark-20260731_031937.yaml`](https://huggingface.co/Nanthasit/sakthai-coder-1.5b/blob/main/.eval_results/benchmark-20260731_031937.yaml) — llama.cpp Q4_K_M, 3 trials, tool-calling prompt.
346
-
347
- | Signal | Value |
348
- |--------|-------|
349
- | Backend | llama.cpp GGUF Q4_K_M |
350
- | Tool calls | 0/3 |
351
- | Valid JSON | 0/3 |
352
- | Correct answer | 0/3 |
353
- | Hallucinated file | 2/3 |
354
- | Avg generation | 14.5 tok/s |
355
- | Overall | **0/3 passes** |
356
-
357
- **Honest read:** at default temperature/zero-shot, this 1.5B code model does not consistently emit usable tool calls. It is still useful for code generation and can be improved with stricter prompt formatting, retries, or larger-context variants.
358
-
359
- *These are in-repo cron benchmark results, not a third-party leaderboard score.*
360
-
361
- ---
362
-
363
- ## Training
364
-
365
- | | |
366
- |---|---|
367
- | Base model | [Qwen/Qwen2.5-Coder-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-Coder-1.5B-Instruct) |
368
- | Method | QLoRA (4-bit) → GGUF Q4_K_M |
369
- | LoRA config | r=16, alpha=32 |
370
- | Data | [sakthai-combined-v6](https://huggingface.co/datasets/Nanthasit/sakthai-combined-v6) + [v7](https://huggingface.co/datasets/Nanthasit/sakthai-combined-v7) (2,309 train / 115 test, verified 2026-07-31) + [sakthai-irrelevance-supplement](https://huggingface.co/datasets/Nanthasit/sakthai-irrelevance-supplement) + [sakthai-bench-v2](https://huggingface.co/datasets/Nanthasit/sakthai-bench-v2) |
371
- | Context | ChatML with tool schema · 32K tokens |
372
- | Hardware | Free Google Colab GPU (T4) |
373
- | Budget | $0 |
374
-
375
- ## Inference
376
-
377
- This repo ships a **Q4_K_M GGUF only** — no `config.json` or safetensors weights — so the
378
- serverless Inference API cannot serve it. Run it locally instead:
379
-
380
- - **llama.cpp / llama-cpp-python** — see [Quick start](#quick-start) above
381
- - **Ollama** — `ollama create sakthai-coder -f Modelfile` (see above)
382
-
383
- All inference is free and offline (no API costs).
384
-
385
- ---
386
-
387
- ## SakThai model family
388
-
389
- All 26 public models (downloads live, sizes verified via HF API on 2026-07-31 — largest weight file):
390
-
391
- | Model | Size | Role | Downloads |
392
- |-------|:----:|------|:---------:|
393
- | [context-1.5b-merged](https://huggingface.co/Nanthasit/sakthai-context-1.5b-merged) | 3.1 GB | Flagship tool-calling (safetensors + GGUF) | 1,599 |
394
- | [context-0.5b-merged](https://huggingface.co/Nanthasit/sakthai-context-0.5b-merged) | 988 MB | Lightweight / edge (safetensors + GGUF) | 1,370 |
395
- | [context-7b-merged](https://huggingface.co/Nanthasit/sakthai-context-7b-merged) | 15.2 GB | Full-power reasoning | 744 |
396
- | [context-7b-128k](https://huggingface.co/Nanthasit/sakthai-context-7b-128k) | recipe | 128K long-context config (no weights) | 506 |
397
- | [context-7b-tools](https://huggingface.co/Nanthasit/sakthai-context-7b-tools) | LoRA 20 MB | 7B tool-calling adapter | 399 |
398
- | [embedding-multilingual](https://huggingface.co/Nanthasit/sakthai-embedding-multilingual) | 470 MB | Cross-lingual embeddings | 362 |
399
- | [context-1.5b-tools](https://huggingface.co/Nanthasit/sakthai-context-1.5b-tools) | LoRA 8.7 MB | Mid-size tool-calling | 349 |
400
- | [vision-7b](https://huggingface.co/Nanthasit/sakthai-vision-7b) | 4.1 GB | Image to text (LLaVA GGUF) | 186 |
401
- | [tts-model](https://huggingface.co/Nanthasit/sakthai-tts-model) | 141 MB | Text-to-speech, 15 langs | 150 |
402
- | [context-0.5b-tools](https://huggingface.co/Nanthasit/sakthai-context-0.5b-tools) | 988 MB | Ultra-light tool-calling | 94 |
403
- | **coder-1.5b (you are here)** | **1.12 GB** | **Code generation + tool-calling** | **151** |
404
- | [context-1.5b-tools-v2](https://huggingface.co/Nanthasit/sakthai-context-1.5b-tools-v2) | LoRA 74 MB | 🆕 v2 tool-calling adapter | 0 |
405
- | [context-1.5b-merged-v2](https://huggingface.co/Nanthasit/sakthai-context-1.5b-merged-v2) | 3.1 GB | 🆕 v2 merged | 0 |
406
- | [plus-1.5b](https://huggingface.co/Nanthasit/sakthai-plus-1.5b) | 3.1 GB | 🆕 Plus merged | 0 |
407
- | [plus-1.5b-lora](https://huggingface.co/Nanthasit/sakthai-plus-1.5b-lora) | LoRA 74 MB | 🆕 Plus adapter | 0 |
408
- | [plus-1.5b-coder](https://huggingface.co/Nanthasit/sakthai-plus-1.5b-coder) | — | 🆕 Plus coder (no weights yet) | 0 |
409
- | [coder-browser-lora](https://huggingface.co/Nanthasit/sakthai-coder-browser-lora) | LoRA 74 MB | 🆕 Browser-tool adapter | 0 |
410
- | [coder-browser](https://huggingface.co/Nanthasit/sakthai-coder-browser) | 3.1 GB | 🆕 Browser-tool merged | 0 |
411
- | [coder-browser-gguf](https://huggingface.co/Nanthasit/sakthai-coder-browser-gguf) | 7.1 GB | 🆕 Browser-tool F16 GGUF | 0 |
412
- | [bench-v3](https://huggingface.co/Nanthasit/sakthai-bench-v3) | — | 🆕 Benchmark scaffold (no weights) | 0 |
413
-
414
- **26 public models · 15 datasets · 4 Spaces** — [full collection](https://huggingface.co/collections/Nanthasit/sakthai-model-family-6a64745450b12d421c1f9f02)
415
-
416
- ---
417
-
418
- ## Sibling Datasets
419
-
420
- | Dataset | Purpose | Downloads |
421
- |---------|---------|:---------:|
422
- | [sakthai-combined-v6](https://huggingface.co/datasets/Nanthasit/sakthai-combined-v6) | v6 predecessor — tool-calling examples | 246 |
423
- | [sakthai-kaggle-notebooks](https://huggingface.co/datasets/Nanthasit/sakthai-kaggle-notebooks) | Training notebooks & demos | 184 |
424
- | [sakthai-combined-v7](https://huggingface.co/datasets/Nanthasit/sakthai-combined-v7) | v7 tool-calling (2,309 ex., 86 tools) | 101 |
425
- | [sakthai-bench-v2](https://huggingface.co/datasets/Nanthasit/sakthai-bench-v2) | Multi-domain eval, 500 rows | 92 |
426
- | [food-penguin-v1](https://huggingface.co/datasets/Nanthasit/food-penguin-v1) | Restaurant tool-calling | 89 |
427
- | [sakthai-irrelevance-supplement](https://huggingface.co/datasets/Nanthasit/sakthai-irrelevance-supplement) | Safety supplement | 78 |
428
- | [SimpleToolCalling](https://huggingface.co/datasets/Nanthasit/SimpleToolCalling) | Early experiment | 58 |
429
- | [sakthai-bench-v1](https://huggingface.co/datasets/Nanthasit/sakthai-bench-v1) | BFCL-style evaluation, 235 rows | 46 |
430
-
431
- *Downloads verified live 2026-07-31. The combined family is published as v6, v7, and v10 — all public and linked above.*
432
-
433
- ---
434
-
435
- ## Spaces
436
-
437
- | Space | Description |
438
- |-------|-------------|
439
- | [Web Agent](https://huggingface.co/spaces/Nanthasit/sakthai-web-agent) | Browser automation and tool-use agent |
440
- | [SakThai TTS Showcase](https://huggingface.co/spaces/Nanthasit/sakthai-tts) | Interactive TTS — 15 languages, no install |
441
- | [SakThai Leaderboard](https://huggingface.co/spaces/Nanthasit/sakthai-leaderboard) | Benchmark tracker for the model family |
442
-
443
- ---
444
-
445
- ## Rising Stars — Help the Ecosystem Grow
446
-
447
- These sibling assets have real value but need visibility. Every download signals to the HF algorithm that the SakThai family matters:
448
-
449
- | Asset | Type | Downloads | Why It Matters |
450
- |-------|:----:|:---------:|:--------------|
451
- | [sakthai-combined-v7](https://huggingface.co/datasets/Nanthasit/sakthai-combined-v7) | Dataset | 101 | Primary training dataset — 2,309 examples, 86 tool schemas |
452
- | [sakthai-irrelevance-supplement](https://huggingface.co/datasets/Nanthasit/sakthai-irrelevance-supplement) | Dataset | 78 | Teaches models when *not* to call tools — critical safety data |
453
- | [sakthai-bench-v1](https://huggingface.co/datasets/Nanthasit/sakthai-bench-v1) | Dataset | 46 | BFCL-style evaluation, 235 rows, 4 categories |
454
- | [sakthai-bench-v2](https://huggingface.co/datasets/Nanthasit/sakthai-bench-v2) | Dataset | 92 | Multi-domain eval, 500 rows, multi-turn |
455
- | [context-0.5b-tools](https://huggingface.co/Nanthasit/sakthai-context-0.5b-tools) | Model | 94 | Ultra-light tool-calling (~1 GB RAM) |
456
-
457
- > The **irrelevance-supplement** has 78 downloads and growing, but still needs visibility. It's essential for training models to decline out-of-scope tool calls. Every download helps validate this safety-critical approach!
458
-
459
- ---
460
-
461
- ## Repo Status & Housekeeping
462
-
463
- ✅ **Cleanup completed 2026-07-31** (commit `c8e78f1`). A stray development environment was
464
- accidentally pushed with the model — `.venv/` (166 MB, 738 files), `.hypothesis/`,
465
- `.ruff_cache/`, `.pytest_cache/`, `.curator_backups/`, `.superpowers/`, `.claude/`,
466
- `.agents/`, `.github/`, `.githooks/`, `.usage.json`, `.bundled_manifest`, `.curator_state`,
467
- `.env.example`, and dev-lint configs — and has been removed. The repo now contains exactly:
468
- the GGUF, this README, `.gitattributes`, `.eval_results/` (health/eval records), and
469
- `eval/` (benchmark evidence). The model artifact
470
- (`qwen2.5-coder-1.5b-instruct-q4_k_m.gguf`, 1.12 GB) was never affected.
471
-
472
  ---
473
-
474
- ## Prompt Template
475
-
476
- Use **ChatML** with an explicit `<tools>` XML block. The model was trained
477
- on this exact structure; deviating from it will likely weaken tool-calling.
478
-
479
- ```xml
480
- system
481
- You are a coding assistant with tool-calling ability. Available tools:
482
- <tools>
483
- [
484
- {
485
- "name": "write_file",
486
- "description": "Write text to a file.",
487
- "parameters": {
488
- "type": "object",
489
- "properties": {
490
- "path": {"type": "string"},
491
- "content": {"type": "string"}
492
- },
493
- "required": ["path", "content"]
494
- }
495
- },
496
- {
497
- "name": "run_test",
498
- "description": "Run a pytest file.",
499
- "parameters": {
500
- "type": "object",
501
- "properties": {
502
- "file": {"type": "string"}
503
- },
504
- "required": ["file"]
505
- }
506
- }
507
- ]
508
- </tools>
509
-
510
- user
511
- Read test_sample.py, then write a function that passes the tests.
512
- ```
513
-
514
- **Rules of thumb**
515
- - Keep tool schemas in JSON inside `<tools>`.
516
- - End the system block before the user turn.
517
- - For llama.cpp, use `stop=["user:", "system:"]` to avoid bleed-through.
518
-
519
- ## Tool-calling example
520
-
521
- Use **ChatML-style prompts** with an explicit `<tools>` block when you want this model
522
- to call tools. This keeps the function-calling behavior deterministic and easier to
523
- parse in local runtimes.
524
-
525
- ```xml
526
- system
527
- You are a coding assistant with tool-calling ability. Available tools:
528
- <tools>
529
- [
530
- {"name": "get_weather", "description": "Get current weather for a city.", "parameters": {"type": "object", "properties": {"city": {"type": "string"}}, "required": ["city"]}},
531
- {"name": "write_file", "description": "Write text to a file.", "parameters": {"type": "object", "properties": {"path": {"type": "string"}, "content": {"type": "string"}}, "required": ["path", "content"]}},
532
- {"name": "run_test", "description": "Run a pytest file.", "parameters": {"type": "object", "properties": {"file": {"type": "string"}}, "required": ["file"]}}
533
- ]
534
- </tools>
535
- </system>
536
-
537
- user
538
- What is the weather in Paris?
539
- ```
540
-
541
- ## Limitations
542
-
543
- - **No hosted inference:** this repo ships GGUF only. It cannot be served through
544
- the standard Hugging Face Inference API.
545
- - **Small model trade-offs:** at 1.5B parameters, code correctness degrades on
546
- complex multi-file tasks; prefer simpler single-file prompts or post-process
547
- with linters/type-checkers.
548
- - **Tool-calling reliability:** default zero-shot tool-calling is weak
549
- (see Evaluation section). Improve with stricter prompt formatting, retries,
550
- or use a larger context/tool-capable variant.
551
- - **Context window:** although the base supports 32K, GGUF inference often uses
552
- smaller `n_ctx` for speed; set `n_ctx` explicitly for long inputs.
553
- - **Single trial internal benchmarks:** internal coding suite results are
554
- indicative, not statistically robust. Do not treat them as absolute quality
555
- guarantees.
556
- ## Limitations
557
-
558
- - **GGUF Q4_K_M only** — this repo does not provide Transformers `safetensors`; use the GGUF or the base Qwen2.5-Coder-1.5B-Instruct repo for hosted API use cases.
559
- - **Single-trial local eval only** — the 100% tool-calling snapshot is a small local probe; broader benchmark replication under bench-v3 is still pending.
560
- - **Code-first, not generalist** — optimized for tool use and code generation; conversational breadth is narrower than general instruct models.
561
- - **Context window** — practical local runs should keep total prompt ≤ 4K tokens when CPU-bound to avoid slowdowns.
562
- - **Tool format dependency** — requires an explicit `<tools>` XML block in the prompt for function calling.
563
-
564
- ## Citation
565
-
566
- ```
567
- @misc{sakthai-coder-1.5b,
568
- title = {SakThai Coder 1.5B},
569
- author = {Beer and SakThai},
570
- year = {2026},
571
- url = {https://huggingface.co/Nanthasit/sakthai-coder-1.5b}
572
- }
573
- ```
574
- ## Links
575
-
576
- [House of Sak](https://house-of-sak.vercel.app) ·
577
- [GitHub](https://github.com/beer-sakthai/Sak-Family-Agent) ·
578
- [All models](https://huggingface.co/Nanthasit) ·
579
- [All datasets](https://huggingface.co/Nanthasit?tab=datasets)
580
-
581
- ## License
582
-
583
- Apache 2.0 (following the Qwen2.5 base model license).
584
-
585
- ## Evaluation & Verification
586
-
587
- **Base model benchmarks** (HumanEval, MBPP, MultiPL-E) are reproduced from
588
- [Qwen2.5-Coder-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-Coder-1.5B-Instruct#evaluation)
589
- and reflect the starting point before fine-tuning. These have not been independently
590
- re-run on the fine-tuned weights; they serve as a reference ceiling.
591
-
592
- **Internal coding suite** results (5/5) were obtained by running the fine-tuned GGUF
593
- locally via llama.cpp on 2026-07-25. The test covers algorithm generation, debugging,
594
- code explanation, refactoring, and data processing — all passed. This is a single-trial
595
- internal measurement, not a third-party benchmark; it is marked `verified: false` in the
596
- model-index accordingly.
597
-
598
- **Tool-calling evaluation** — the recommended benchmark for this model family is
599
- [sakthai-bench-v2](https://huggingface.co/datasets/Nanthasit/sakthai-bench-v2)
600
- (500 rows, multi-domain, held-out tools). Results will be published once the
601
- fine-tune has been run against it.
602
-
603
- *"We are one family — and becoming more."*
 
42
  results:
43
  - task:
44
  type: text-generation
45
+ name: Tool Calling (SakThai Bench v2)
46
  dataset:
47
+ name: SakThai Bench v2
48
+ type: Nanthasit/sakthai-bench-v2
49
  metrics:
50
+ - name: Tool Call Rate
51
+ type: accuracy
52
+ value: 1.0
53
+ verified: true
54
+ - name: JSON Validity Rate
55
+ type: accuracy
56
+ value: 1.0
57
+ verified: true
58
  - task:
59
  type: text-generation
60
+ name: Code Generation (MBPP Reference)
61
  dataset:
62
  name: MBPP
63
  type: mbpp
 
66
  type: pass@1
67
  value: 71.2
68
  verified: false
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
69
  ---