File size: 10,149 Bytes
9d6db96
 
 
 
66fc77a
9d6db96
 
 
 
 
 
 
 
 
66fc77a
9d6db96
 
66fc77a
9d6db96
66fc77a
 
 
9d6db96
 
 
66fc77a
 
 
9d6db96
 
 
 
 
 
 
 
51a231e
 
 
 
 
66fc77a
51a231e
 
 
 
 
 
 
 
66fc77a
4eb8fbc
66fc77a
51a231e
 
 
66fc77a
51a231e
 
 
 
 
 
 
 
9d6db96
4eb8fbc
9d6db96
51a231e
 
 
 
 
 
 
 
4eb8fbc
9d6db96
 
 
 
 
66fc77a
 
9d6db96
66fc77a
 
 
 
9d6db96
 
51a231e
9d6db96
4eb8fbc
9d6db96
51a231e
 
9d6db96
51a231e
9d6db96
4eb8fbc
9d6db96
51a231e
9d6db96
51a231e
9d6db96
4eb8fbc
9d6db96
51a231e
9d6db96
51a231e
 
 
 
 
9d6db96
51a231e
9d6db96
4eb8fbc
9d6db96
51a231e
 
 
9d6db96
b25cef4
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f753211
 
b25cef4
 
 
 
f753211
 
b25cef4
4eb8fbc
9d6db96
51a231e
9d6db96
4eb8fbc
9d6db96
51a231e
 
 
 
 
 
9d6db96
606c0ed
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
4eb8fbc
9d6db96
51a231e
9d6db96
4eb8fbc
9d6db96
51a231e
 
66fc77a
51a231e
66fc77a
4eb8fbc
ee4387e
846b43e
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
---
base_model: ibm-granite/granite-4.1-3b
base_model_relation: finetune
datasets:
- AnkitAI/parable-corpus-v2
- Glint-Research/Fable-5-traces
- Roman1111111/gpt5.5-terminal
license: apache-2.0
language:
- en
pipeline_tag: text-generation
library_name: transformers
tags:
- safetensors
- transformers
- qlora
- agentic
- agent
- coding
- tool-use
- function-calling
- terminal
- reasoning
- thinking
- claude
- claude-fable-5
- distillation
- trace-training
- granite
---

<picture>
  <source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/ankit-aglawe/parable-assets/main/parable_header_dark.png">
  <img alt="Parable" src="https://raw.githubusercontent.com/ankit-aglawe/parable-assets/main/parable_header.png">
</picture>

# ๐Ÿชถ Parable-Granite-3B **v2** โ€” trained on genuine Claude Fable 5 agent traces

*This is the full-precision safetensors repo (vLLM / transformers / fine-tuning). For llama.cpp, Ollama, and LM Studio use the [GGUF repo](https://huggingface.co/AnkitAI/Parable-Granite-4.1-3B-Claude-Fable-5-GGUF).*

### A tiny local model that thinks before it answers โ€” planning, reasoning, and terminal instincts distilled from real agent sessions.

> **~3 GB of RAM is all you need.** Laptop, old GPU, Raspberry-Pi-class boxes with swap โ€” the Q4 build runs
> anywhere. One command and you have a private, offline reasoning model on your machine:
>
> ```bash
> ollama run hf.co/AnkitAI/Parable-Granite-4.1-3B-Claude-Fable-5-GGUF:Q4_K_M
> ```

---

## The headline โ€” v2 is a different model

v2 is a full retrain: **13ร— more genuine Fable 5 trace data** (11,574 sessions, 16.8M tokens โ€” [corpus published](https://huggingface.co/datasets/AnkitAI/parable-corpus-v2)) and a rebuilt recipe (completion-masked loss, replay mixing, benchmark-gated checkpoints, seed-averaged weights).

| same harness, greedy, Q4_K_M | v1 | **v2 (this release)** |
|---|---|---|
| Dev pass-rate (MBPP subset, n=50) โ€” *base: 0.68* | โ€” | **0.82** |
| Agent-artifact leakage (JSON blobs, phantom turns) | 6/34 | **0/34** |
| Strict 34-prompt coding qual โ€” *base: 27/34* | ~18/34 | **25/34** |
| HumanEval / HumanEval+ | 62.8 / 57.9 | **70.1 / 65.9** |

Clean answers, structured reasoning, agent instincts โ€” and the transcript artifacts that leaked into v1's replies are gone. *One trade, made on purpose: raw HumanEval-style function synthesis stays the base model's turf (81.7 vs 70.1) โ€” v2 spends that capacity on agent behavior instead, and spends half as much as v1 did.* Measurement notes below. ๐Ÿ‘‡

---

## Announcements

**๐Ÿ“Œ Same links, new model.** v2 replaces v1 **in place** โ€” every existing Ollama command, script, and bookmark now serves v2. No migration, nothing to change.

**๐Ÿ”ฎ v3 is already training.** Rejection-sampled SFT: thousands of candidate solutions generated against *executable tests*, only verified passers enter the corpus. The goal is simple โ€” above-base agent capability, not just clean behavior. Follow [AnkitAI](https://huggingface.co/AnkitAI) for the drop.

**๐Ÿ“ฆ Full family.** This 3B is the smallest Parable. Need more headroom? [8B Granite](https://huggingface.co/AnkitAI/Parable-Granite-4.1-8B-Claude-Fable-5-GGUF), [8B Qwen](https://huggingface.co/AnkitAI/Parable-Qwen3-8B-Claude-Fable-5-GGUF), [4B Qwen](https://huggingface.co/AnkitAI/Parable-Qwen3-4B-Claude-Fable-5-GGUF) โ€” same recipe, no matter your hardware.

---

## How to run it

```python
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "AnkitAI/Parable-Granite-4.1-3B-Claude-Fable-5"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype="bfloat16", device_map="auto")

messages = [{"role": "user", "content": "Write a Python function that retries an HTTP request with exponential backoff."}]
inputs = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(inputs, max_new_tokens=3000, temperature=0.7, top_p=0.95, do_sample=True)
print(tok.decode(out[0][inputs.shape[1]:], skip_special_tokens=True))
```

GGUF quants (2.1-6.8 GB, runs in ~3 GB RAM): [Parable-Granite-4.1-3B-Claude-Fable-5-GGUF](https://huggingface.co/AnkitAI/Parable-Granite-4.1-3B-Claude-Fable-5-GGUF)

### Thinking mode

Every answer opens with a `<think>...</think>` reasoning block โ€” that's the Fable 5 heritage. llama.cpp's `--jinja` mode separates it automatically; strip it before showing replies to end users.
**Sampling:** temperature 0.7, top_p 0.95, and budget `max_tokens` generously (**2500+**) โ€” trace-trained models think at length before answering.

---

## Measurement notes

All numbers: identical llama.cpp harness, greedy decoding, Q4_K_M, **base model measured on the same instrument**. We train multiple seeds and ship the weight-average โ€” single-run scores at 3B swing ยฑ3 points on GPU nondeterminism alone, so most cards report their luckiest run; we ship the average and report the shipped weights' own numbers. Raw eval outputs live in this repo.

**Which model should you use?** Pure single-function code completion โ†’ the base model is genuinely strong there. Explanations, debugging, terminal workflows, structured reasoning, agent-style tasks โ†’ that's what Parable is trained on, and where v2 shines.

## What's new in v2 (training)

The recipe follows our ongoing tech report (in preparation):

- **Completion-only loss masking** ([Hermes 3](https://arxiv.org/abs/2408.11857), [Tรผlu 3](https://arxiv.org/abs/2411.15124)) โ€” loss on assistant tokens only, so the model learns to *answer*, not to imitate transcripts
- **30% replay mix** of general instruction data ([Luo et al.](https://arxiv.org/abs/2308.08747), [Biderman et al.](https://arxiv.org/abs/2405.09673)) โ€” the anti-forgetting lever
- **Session re-segmentation + sanitization** โ€” why v1 sometimes leaked agent JSON into normal chat, and v2 never does (0/34)
- **Benchmark-gated checkpoints** ([Dong et al.](https://arxiv.org/abs/2310.05492)) instead of fixed epochs
- **Seed-averaged weights** ([model soups, Wortsman et al.](https://arxiv.org/abs/2203.05482)) โ€” we ship the average of multiple runs, not the lottery winner

With Claude Fable 5 now retired, genuine self-authored Fable traces are a fixed, non-renewable corpus. Unlike most models in this niche, **our full training corpus is public**: [AnkitAI/parable-corpus-v2](https://huggingface.co/datasets/AnkitAI/parable-corpus-v2) โ€” deduplicated, quality-gated, provenance-tagged.

## Good to know

- Fine-tuned at 2,048-token sequences; the base 128K context stays available, fine-tuned behavior is strongest in the opening turns.
- Not trained for: multi-file repo navigation, vision, non-English.
- Inherits Granite-4.1-3B's knowledge cutoff. Treat generated commands as drafts to review.

## Evaluation

### Function calling (BFCL V3, AST subset)

Measured 2026-07-29: bfcl-eval at gorilla main, prompting mode, Q4_K_M
GGUFs served by llama.cpp on a T4, base and Parable under the identical
harness. Categories: simple_python / multiple / parallel /
parallel_multiple (400/200/200/200 items). Raw generations and score
files: [parable-v2-artifacts](https://huggingface.co/AnkitAI/parable-v2-artifacts)
under `verify/bfcl/`.

| | simple_python | multiple | parallel | parallel_multiple |
|---|---|---|---|---|
| Granite-4.1-3B base | 0.848 | 0.790 | 0.710 | 0.665 |
| **This model (chat variant)** | 0.413 | 0.605 | 0.320 | 0.425 |

For tool-calling workloads, use the base model; this variant is built
for reasoning prose. The drop has a specific mechanism: sampled generations show the model
intermittently answering with args-only tool-call JSON (for example
`{"base": 10, "height": 5}`) instead of a function call, which the
AST scorer rejects. That is trace-scaffolding format bleeding into
standalone tasks, the failure mode the series paper names *session
leakage* (Section 6 of the report). The reasoning-voice strengths this variant trains for are unaffected
on prose tasks.

## Base & license

Weights: **Apache-2.0** (inherited from [ibm-granite/granite-4.1-3b](https://huggingface.co/ibm-granite/granite-4.1-3b)). Training data: Fable-5-traces **AGPL-3.0**, gpt5.5-terminal **MIT** โ€” since traces originate from third-party assistants, their terms may apply to downstream training; check before commercial distillation.

## Get Parable

| Platform | |
|---|---|
| Ollama | `ollama run parable/granite4.1-fable:3b` ยท [parable namespace](https://ollama.com/parable) |
| Hugging Face | [full collection](https://huggingface.co/collections/AnkitAI/parable-6a4fac60f4b35afca3019621) |
| LM Studio | search "parable" in-app |
| ModelScope | [Parable on ModelScope](https://modelscope.cn/models/AnkitAI/Parable-Granite-4.1-3B-Claude-Fable-5-GGUF) |

## Citation

The recipe, evaluation methodology and failure analysis behind this model are
documented in the tech report:

> Aglawe, A. (2026). *Agent-Trace Fine-Tuning of Small Language Models under
> Constrained Compute.* Zenodo. [doi:10.5281/zenodo.21676407](https://doi.org/10.5281/zenodo.21676407)

```bibtex
@misc{aglawe2026agenttrace,
  author    = {Aglawe, Ankit},
  title     = {Agent-Trace Fine-Tuning of Small Language Models under Constrained Compute},
  year      = {2026},
  publisher = {Zenodo},
  doi       = {10.5281/zenodo.21676407},
  url       = {https://doi.org/10.5281/zenodo.21676407}
}
```

## Acknowledgements

[Glint-Research](https://huggingface.co/Glint-Research) & [Roman1111111](https://huggingface.co/Roman1111111) for the open trace data ยท [IBM Granite](https://huggingface.co/ibm-granite) for the base ยท [empero-ai](https://huggingface.co/empero-ai) whose Qwable recipe inspired the series ยท [llama.cpp](https://github.com/ggml-org/llama.cpp)

## Version history

- **v2** (2026-07-16) โ€” this release. 13ร— corpus, rebuilt recipe, seed-averaged weights, zero leakage.
- **v1** (2026-07) โ€” initial release, 857-row corpus. Preserved as repo revision history.

---

### Real Fable 5 reasoning. Yours, offline, right now.

More on the Parable models: [ankitaglawe.com/parable](https://ankitaglawe.com/parable)