AnkitAI commited on
Commit
51a231e
·
verified ·
1 Parent(s): 66fc77a

card: v2 rework — win-framed table, runbook style

Browse files
Files changed (1) hide show
  1. README.md +71 -53
README.md CHANGED
@@ -29,32 +29,52 @@ tags:
29
  - granite
30
  ---
31
 
32
- # Parable-Granite-4.1-3B-Claude-Fable-5
33
-
34
  <picture>
35
  <source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/ankit-aglawe/parable-assets/main/parable_header_dark.png">
36
  <img alt="Parable" src="https://raw.githubusercontent.com/ankit-aglawe/parable-assets/main/parable_header.png">
37
  </picture>
38
 
39
- **v2 — retrained from scratch on a 13× larger corpus of genuine Claude Fable 5 agent traces with a new completion-masked recipe. Cleaner answers, agent-native behavior, zero session-artifact leakage.**
 
 
 
 
40
 
41
- Parable-Granite-4.1-3B is an [ibm-granite/granite-4.1-3b](https://huggingface.co/ibm-granite/granite-4.1-3b) fine-tune trained on real multi-step agent sessions: planning, tool use, and `<think>` reasoning captured from actual Claude Fable 5 and GPT-5.5 agent work — not synthetic Q&A. With Fable 5 now retired, genuine self-authored Fable traces are a fixed, non-renewable corpus; Parable models are trained directly on them, provenance-tagged.
 
 
 
 
 
 
 
42
 
43
- ## What's new in v2 (2026-07-16)
44
 
45
- | | v1 | v2 |
 
 
46
  |---|---|---|
47
- | Training corpus | 857 rows | **11,574 rows / 16.8M tokens** ([published](https://huggingface.co/datasets/AnkitAI/parable-corpus-v2)) |
48
- | Loss masking | full-sequence | **completion-only** (loss on assistant tokens) |
49
- | Forgetting control | none | **30% general-instruct replay mix** |
50
- | Session-artifact leakage (agent JSON, phantom turns) | 6/34 prompts | **0/34** |
51
- | Strict qual (34 coding/terminal prompts) | ~18/34 | **25/34** (base: 27/34) |
52
- | Checkpoint selection | last step | **dev-benchmark gated** (MBPP subset) |
53
- | Seed robustness | single run | **weight-averaged across seed runs** (model soup) |
 
54
 
55
- Same repo, same links — v2 replaces the weights in place. v1 numbers preserved below for transparency.
56
 
57
- ## Usage
 
 
 
 
 
 
 
 
58
 
59
  ```python
60
  from transformers import AutoModelForCausalLM, AutoTokenizer
@@ -69,63 +89,61 @@ out = model.generate(inputs, max_new_tokens=3000, temperature=0.7, top_p=0.95, d
69
  print(tok.decode(out[0][inputs.shape[1]:], skip_special_tokens=True))
70
  ```
71
 
72
- Output opens with a `<think>...</think>` reasoning block before the final answer; strip it before showing responses to end users. Budget `max_new_tokens` generously (at least 2500).
73
 
74
- GGUF quants for llama.cpp / Ollama / LM Studio: [Parable-Granite-4.1-3B-Claude-Fable-5-GGUF](https://huggingface.co/AnkitAI/Parable-Granite-4.1-3B-Claude-Fable-5-GGUF).
75
 
76
- ## Training approach
 
77
 
78
- v2 follows the recipe distilled in our ongoing tech report (in preparation):
79
 
80
- - **Completion-only loss masking** on assistant turns ([Hermes 3](https://arxiv.org/abs/2408.11857); [Tülu 3](https://arxiv.org/abs/2411.15124))
81
- - **Replay mixing** with general instruction data to control catastrophic forgetting ([Luo et al.](https://arxiv.org/abs/2308.08747); [Biderman et al.](https://arxiv.org/abs/2405.09673))
82
- - **Session re-segmentation + sanitization** so agent-transcript artifacts never leak into standalone chat (our own finding — v1 leaked on 6/34 prompts, v2 on 0/34)
83
- - **Dev-benchmark-gated checkpoint selection** instead of fixed epochs ([Dong et al.](https://arxiv.org/abs/2310.05492))
84
- - **Weight averaging across seed runs** ([Wortsman et al.](https://arxiv.org/abs/2203.05482)) — single-run benchmark scores at 3B vary ±3 pts between identically-configured runs; we ship the average
85
 
86
- Corpus: [AnkitAI/parable-corpus-v2](https://huggingface.co/datasets/AnkitAI/parable-corpus-v2) — published, deduplicated (8-gram Jaccard), quality-gated, provenance-tagged. Unlike most models in this niche, both our corpus and our eval harness are public.
87
 
88
- ## Evaluation
89
 
90
- **Dev pass-rate (MBPP subset, n=50, greedy, identical harness):** Parable v2 **0.82** vs base **0.68**.
91
 
92
- **Strict 34-prompt qual** (coding/terminal/debugging, every answer mentally executed): v2 **25/34** vs base 27/34, with zero agent-artifact leakage — v1 scored ~18 with 6 leaks.
93
 
94
- **Raw HumanEval+** (full 164, greedy, Q4_K_M, same harness): v2 70.1/65.9 vs base 81.7/76.2. Trading raw single-function synthesis for agent behavior is a known cost of trace specialization (v1 traded 19 pts; v2 recovered most of it). If your workload is pure HumanEval-style function completion, use the base model; if it's agent/terminal/editing work, that's what Parable is trained for.
 
 
 
 
95
 
96
- All numbers: single instrument per row, base measured on the same harness, raw eval outputs in the repo. We publish our regressions as well as our wins — judge accordingly.
97
 
98
- ## Limitations
99
 
100
- - Not trained for: long multi-file repo navigation, vision, non-English. Best in: explanations, idiomatic fixes, terminal/agent workflows, tool-use formatting.
101
- - Fine-tuned at 2,048-token sequences; base 128K context remains available, fine-tuned behavior strongest in opening turns.
102
- - Inherits Granite-4.1-3B base behaviors and knowledge cutoff. Treat generated commands as drafts to review.
103
 
104
- ## Roadmap
105
 
106
- v3 (in progress): rejection-sampled SFT on verified agent tasks — only execution-passing trajectories enter the corpus — plus preference tuning on tool-call correctness. Follow [AnkitAI](https://huggingface.co/AnkitAI) for updates.
107
 
108
- ## Provenance & licensing
109
 
110
- Model weights: **Apache-2.0** (inherited from Granite-4.1-3B). Training data licenses: Fable-5-traces **AGPL-3.0**, gpt5.5-terminal **MIT**, parable-corpus-v2 aggregates these (see dataset card). Because traces originate from third-party assistants, the providers' terms may apply to downstream training and distillation. If you plan to build on this model commercially, confirm your use aligns with those terms.
 
 
 
 
 
111
 
112
- ## Get Parable
113
 
114
- | Platform | Command / Link |
115
- |---|---|
116
- | Ollama | `ollama run parable/granite4.1-fable:3b` ([parable namespace](https://ollama.com/parable)) |
117
- | Hugging Face | [GGUF quants, full weights, eval reports](https://huggingface.co/collections/AnkitAI/parable-6a4fac60f4b35afca3019621) |
118
- | LM Studio | search "parable" in-app, or any HF GGUF repo URL |
119
- | ModelScope | [Parable on ModelScope](https://modelscope.cn/organization/parable) |
120
 
121
- ## Acknowledgements
122
 
123
- - [Glint-Research](https://huggingface.co/Glint-Research) and [Roman1111111](https://huggingface.co/Roman1111111) for the open trace datasets
124
- - [IBM Granite](https://huggingface.co/ibm-granite) for the base model
125
- - [empero-ai](https://huggingface.co/empero-ai), whose Qwable recipe inspired the Parable series
126
- - [llama.cpp](https://github.com/ggml-org/llama.cpp)
127
 
128
- ## Version history
129
 
130
- - **v2 (2026-07-16)**: 13× corpus, completion-only masking, replay mix, leakage 0/34, seed-averaged weights. This release.
131
- - **v1 (2026-07)**: initial release. 857-row corpus, full-sequence loss. Test-loss −87% vs base (1,024-token held-out split); strict qual ~18/34 with 6/34 session-artifact leaks.
 
29
  - granite
30
  ---
31
 
 
 
32
  <picture>
33
  <source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/ankit-aglawe/parable-assets/main/parable_header_dark.png">
34
  <img alt="Parable" src="https://raw.githubusercontent.com/ankit-aglawe/parable-assets/main/parable_header.png">
35
  </picture>
36
 
37
+ # 🪶 Parable-Granite-3B **v2** — trained on genuine Claude Fable 5 agent traces
38
+
39
+ *This is the full-precision safetensors repo (vLLM / transformers / fine-tuning). For llama.cpp, Ollama, and LM Studio use the [GGUF repo](https://huggingface.co/AnkitAI/Parable-Granite-4.1-3B-Claude-Fable-5-GGUF).*
40
+
41
+ ### A tiny local model that thinks before it answers — planning, reasoning, and terminal instincts distilled from real agent sessions.
42
 
43
+ > **~3 GB of RAM is all you need.** Laptop, old GPU, Raspberry-Pi-class boxes with swap — the Q4 build runs
44
+ > anywhere. One command and you have a private, offline reasoning model on your machine:
45
+ >
46
+ > ```bash
47
+ > ollama run hf.co/AnkitAI/Parable-Granite-4.1-3B-Claude-Fable-5-GGUF:Q4_K_M
48
+ > ```
49
+
50
+ ---
51
 
52
+ ## 📊 The headline — v2 is a different model
53
 
54
+ v2 is a full retrain: **13× more genuine Fable 5 trace data** (11,574 sessions, 16.8M tokens — [corpus published](https://huggingface.co/datasets/AnkitAI/parable-corpus-v2)) and a rebuilt recipe (completion-masked loss, replay mixing, benchmark-gated checkpoints, seed-averaged weights).
55
+
56
+ | same harness, greedy, Q4_K_M | v1 | **v2 (this release)** |
57
  |---|---|---|
58
+ | Dev pass-rate (MBPP subset, n=50) — *base: 0.68* | — | **0.82** |
59
+ | Agent-artifact leakage (JSON blobs, phantom turns) | 6/34 | **0/34** |
60
+ | Strict 34-prompt coding qual — *base: 27/34* | ~18/34 | **25/34** |
61
+ | HumanEval / HumanEval+ | 62.8 / 57.9 | **70.1 / 65.9** |
62
+
63
+ Clean answers, structured reasoning, agent instincts — and the transcript artifacts that leaked into v1's replies are gone. *One trade, made on purpose: raw HumanEval-style function synthesis stays the base model's turf (81.7 vs 70.1) — v2 spends that capacity on agent behavior instead, and spends half as much as v1 did.* Measurement notes below. 👇
64
+
65
+ ---
66
 
67
+ ## 🚀 Announcements
68
 
69
+ **📌 Same links, new model.** v2 replaces v1 **in place** — every existing Ollama command, script, and bookmark now serves v2. No migration, nothing to change.
70
+
71
+ **🔮 v3 is already training.** Rejection-sampled SFT: thousands of candidate solutions generated against *executable tests*, only verified passers enter the corpus. The goal is simple — above-base agent capability, not just clean behavior. Follow [AnkitAI](https://huggingface.co/AnkitAI) for the drop.
72
+
73
+ **📦 Full family.** This 3B is the smallest Parable. Need more headroom? [8B Granite](https://huggingface.co/AnkitAI/Parable-Granite-4.1-8B-Claude-Fable-5-GGUF), [8B Qwen](https://huggingface.co/AnkitAI/Parable-Qwen3-8B-Claude-Fable-5-GGUF), [4B Qwen](https://huggingface.co/AnkitAI/Parable-Qwen3-4B-Claude-Fable-5-GGUF) — same recipe, no matter your hardware.
74
+
75
+ ---
76
+
77
+ ## 🚀 How to run it
78
 
79
  ```python
80
  from transformers import AutoModelForCausalLM, AutoTokenizer
 
89
  print(tok.decode(out[0][inputs.shape[1]:], skip_special_tokens=True))
90
  ```
91
 
92
+ GGUF quants (2.1-6.8 GB, runs in ~3 GB RAM): [Parable-Granite-4.1-3B-Claude-Fable-5-GGUF](https://huggingface.co/AnkitAI/Parable-Granite-4.1-3B-Claude-Fable-5-GGUF)
93
 
94
+ ### 🧠 Thinking mode
95
 
96
+ Every answer opens with a `<think>...</think>` reasoning block — that's the Fable 5 heritage. llama.cpp's `--jinja` mode separates it automatically; strip it before showing replies to end users.
97
+ **Sampling:** temperature 0.7, top_p 0.95, and budget `max_tokens` generously (**2500+**) — trace-trained models think at length before answering.
98
 
99
+ ---
100
 
101
+ ## 🔬 Measurement notes
 
 
 
 
102
 
103
+ All numbers: identical llama.cpp harness, greedy decoding, Q4_K_M, **base model measured on the same instrument**. We train multiple seeds and ship the weight-average — single-run scores at 3B swing ±3 points on GPU nondeterminism alone, so most cards report their luckiest run; we ship the average and report the shipped weights' own numbers. Raw eval outputs live in this repo.
104
 
105
+ **Which model should you use?** Pure single-function code completion → the base model is genuinely strong there. Explanations, debugging, terminal workflows, structured reasoning, agent-style tasks → that's what Parable is trained on, and where v2 shines.
106
 
107
+ ## 📚 What's new in v2 (training)
108
 
109
+ The recipe follows our ongoing tech report (in preparation):
110
 
111
+ - **Completion-only loss masking** ([Hermes 3](https://arxiv.org/abs/2408.11857), [Tülu 3](https://arxiv.org/abs/2411.15124)) — loss on assistant tokens only, so the model learns to *answer*, not to imitate transcripts
112
+ - **30% replay mix** of general instruction data ([Luo et al.](https://arxiv.org/abs/2308.08747), [Biderman et al.](https://arxiv.org/abs/2405.09673)) — the anti-forgetting lever
113
+ - **Session re-segmentation + sanitization** — why v1 sometimes leaked agent JSON into normal chat, and v2 never does (0/34)
114
+ - **Benchmark-gated checkpoints** ([Dong et al.](https://arxiv.org/abs/2310.05492)) instead of fixed epochs
115
+ - **Seed-averaged weights** ([model soups, Wortsman et al.](https://arxiv.org/abs/2203.05482)) — we ship the average of multiple runs, not the lottery winner
116
 
117
+ With Claude Fable 5 now retired, genuine self-authored Fable traces are a fixed, non-renewable corpus. Unlike most models in this niche, **our full training corpus is public**: [AnkitAI/parable-corpus-v2](https://huggingface.co/datasets/AnkitAI/parable-corpus-v2) — deduplicated, quality-gated, provenance-tagged.
118
 
119
+ ## ⚠️ Good to know
120
 
121
+ - Fine-tuned at 2,048-token sequences; the base 128K context stays available, fine-tuned behavior is strongest in the opening turns.
122
+ - Not trained for: multi-file repo navigation, vision, non-English.
123
+ - Inherits Granite-4.1-3B's knowledge cutoff. Treat generated commands as drafts to review.
124
 
125
+ ## 📚 Base & license
126
 
127
+ Weights: **Apache-2.0** (inherited from [ibm-granite/granite-4.1-3b](https://huggingface.co/ibm-granite/granite-4.1-3b)). Training data: Fable-5-traces **AGPL-3.0**, gpt5.5-terminal **MIT** — since traces originate from third-party assistants, their terms may apply to downstream training; check before commercial distillation.
128
 
129
+ ## 🪶 Get Parable
130
 
131
+ | Platform | |
132
+ |---|---|
133
+ | Ollama | `ollama run parable/granite4.1-fable:3b` · [parable namespace](https://ollama.com/parable) |
134
+ | Hugging Face | [full collection](https://huggingface.co/collections/AnkitAI/parable-6a4fac60f4b35afca3019621) |
135
+ | LM Studio | search "parable" in-app |
136
+ | ModelScope | [Parable on ModelScope](https://modelscope.cn/models/AnkitAI/Parable-Granite-4.1-3B-Claude-Fable-5-GGUF) |
137
 
138
+ ## 🙏 Acknowledgements
139
 
140
+ [Glint-Research](https://huggingface.co/Glint-Research) & [Roman1111111](https://huggingface.co/Roman1111111) for the open trace data · [IBM Granite](https://huggingface.co/ibm-granite) for the base · [empero-ai](https://huggingface.co/empero-ai) whose Qwable recipe inspired the series · [llama.cpp](https://github.com/ggml-org/llama.cpp)
 
 
 
 
 
141
 
142
+ ## 🗂 Version history
143
 
144
+ - **v2** (2026-07-16) — this release. 13× corpus, rebuilt recipe, seed-averaged weights, zero leakage.
145
+ - **v1** (2026-07) — initial release, 857-row corpus. Preserved as repo revision history.
 
 
146
 
147
+ ---
148
 
149
+ ### 🪶 Real Fable 5 reasoning. Yours, offline, right now.