File size: 11,626 Bytes
aa15560
 
 
 
 
 
 
 
 
 
e5b9975
 
aa15560
 
26fa89b
aa15560
 
26fa89b
 
aa15560
26fa89b
 
aa15560
26fa89b
 
aa15560
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
---
license: other
library_name: transformers
tags:
- flm
- fastflowlm
- npu
- npu2
- amd-xdna
- lemonade
base_model:
- empero-ai/Qwable-9B-Claude-Fable-5
---

# Qwable-9B-Claude-Fable-5-NPU2 (FastFlowLM / Lemonade NPU2 Quantization)

> [!IMPORTANT]
> **Quantization & NPU Compatibility Note:**
> This repository contains **Q4NX quantized weights** converted from [empero-ai/Qwable-9B-Claude-Fable-5](https://huggingface.co/empero-ai/Qwable-9B-Claude-Fable-5) to run natively on **FastFlowLM (`flm`) v1.0.3+** and **Lemonade** on AMD XDNA NPU hardware.
> 
> * **Model Type**: Quantized model conversion (NPU Q4NX format)
> * **Parent / Base Model**: [empero-ai/Qwable-9B-Claude-Fable-5](https://huggingface.co/empero-ai/Qwable-9B-Claude-Fable-5)
> * **Details**: Re-quantized to Q4NX format for FastFlowLM v1.0.3+ and Lemonade on AMD XDNA NPU. Fine-tuned for agentic coding and reasoning, configured with full EOS stop token sequence IDs ([248044, 248046]).
> * **Architecture**: Qwable 9B (Qwen3.5 9B architecture)
> * **Quantization Format**: Q4_K / Q4_1 / Q8_0 hybrid Q4NX
> * **Format**: `Q4NX` (safetensors format with AMD NPU block packing). Note that this is **not** a standard GGUF file; it is executed natively via `flm` / Lemonade on AMD Ryzen AI NPUs.

---

## Serving with Lemonade & FastFlowLM

To serve this model via Lemonade or FastFlowLM:

```bash
# Pull and run with FLM:
flm pull Qwable-9B-Claude-Fable-5-NPU2
flm serve Qwable-9B-Claude-Fable-5-NPU2 --ctx-len 32768 --port 8001
```

Or configure via Lemonade:
```bash
lemonade run Qwable-9B-Claude-Fable-5-NPU2
```

---

## Original Model Information (empero-ai/Qwable-9B-Claude-Fable-5)

Below is the model card from the upstream repository [empero-ai/Qwable-9B-Claude-Fable-5](https://huggingface.co/empero-ai/Qwable-9B-Claude-Fable-5):

---

<p align="center">
  <img src="qwable9b.jpg" alt="Qwable-9B-Claude-Fable-5" width="420"/>
</p>

# Qwable-9B-Claude-Fable-5

**Developed by [Empero](https://empero.org)**

Qwable-9B-Claude-Fable-5 is a full-parameter supervised fine-tune of
**[Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B)** on a curated mix of agentic coding and
reasoning traces. It is a distillation-style fine-tune: the training targets are outputs from other
assistants (Claude Fable 5 and a GPT-5.5 terminal agent), teaching the model to imitate their reasoning and
tool-use style on long, multi-turn coding and agent tasks.

> **Early release.** Qwable-9B-Claude-Fable-5 brings strong coding and agentic behavior out of the box. A
> full suite of quantitative benchmarks (coding, agentic, and safety) is underway and will be added to this
> card; training quality is already backed by held-out validation results (see [Evaluation](#evaluation)).
> See [Provenance & licensing](#provenance--licensing) for licensing notes.

## Model details

- **Developed by:** [Empero](https://empero.org)
- **Base model:** Qwen3.5-9B β€” a dense, natively **multimodal** model with a hybrid attention stack
  (3:1 Gated DeltaNet linear-attention to Gated full-attention), ~152k vocabulary, long native context.
- **Fine-tune type:** full parameter (all text-backbone weights trained). The **vision tower was frozen** β€”
  training was **text-only**, so vision behavior is inherited from the base and **was not tuned or tested**.
- **Objective:** supervised fine-tuning, **assistant-only loss** (the model is scored only on the
  assistant/completion tokens; prompts are masked out).
- **Languages:** primarily English.
- **License:** `apache-2.0`, inherited from the base weights β€” but see the data-provenance caveat below.

## Training data

| Source | Role | Approx. examples (after holdout) |
|---|---|---|
| [`Glint-Research/Fable-5-traces`](https://huggingface.co/datasets/Glint-Research/Fable-5-traces) | Claude Fable 5 reasoning + coding traces (`context` β†’ `completion`) | ~4,585 |
| [`Roman1111111/gpt5.5-terminal`](https://huggingface.co/datasets/Roman1111111/gpt5.5-terminal) | GPT-5.5 terminal/agent task solutions (`system` + `prompt` β†’ `solution`) | ~111 |

Both sources were normalized to a single chat format (`user`/`assistant`, with an optional `system` turn for
the terminal tasks) and concatenated. The natural mix is heavily skewed toward Fable traces (~97%); no
re-weighting was applied to the training set.

**Held-out eval split:** 100 examples were withheld from training β€” deliberately composed **80% Fable /
20% terminal** so the held-out loss carries signal on *both* task types rather than being dominated by Fable.

## Training procedure

Full-parameter supervised fine-tuning with [TRL](https://github.com/huggingface/trl), using:

- **Full-length traces, zero truncation** (`max_length = 76,800`) β€” even the longest multi-turn traces
  (~74k tokens) are trained in full.
- **Assistant-only loss** β€” the model is scored only on assistant/completion tokens; prompt tokens are masked.
- **Chunked cross-entropy** for memory-efficient long-context training.

| Hyperparameter | Value |
|---|---|
| Epochs | 2 |
| Effective batch size | 16 |
| Max sequence length | 76,800 (no truncation) |
| Learning rate | 1e-5 (cosine, 3% warmup) |
| Optimizer | AdamW (8-bit) |
| Precision | bf16 |
| Loss | chunked NLL, assistant-only |

## Evaluation

Training quality was tracked via **held-out validation loss and token-accuracy** on a 100-example split and
supplemented with a qualitative generation review (below). A full suite of **coding, agentic, and safety
benchmarks is in progress and will be published here.** Validation was run periodically during training:

| Step | eval loss | eval token-acc |
|---|---|---|
| 100 | 0.743 | 0.784 |
| 200 | 0.722 | 0.789 |
| 300 (β‰ˆ epoch 1) | 0.714 | 0.791 |
| 400 | 0.7135 | 0.791 |
| 500 | 0.713 | 0.791 |

**No overfitting observed.** Held-out loss decreased monotonically and then **plateaued (~0.71)** through the
second epoch β€” it never rose, even as train loss fell to ~0.64. Epoch-1 and final (epoch-2) checkpoints
generalize equivalently on held-out data.

> Note: `token-accuracy` is teacher-forced, per-token next-token accuracy over completion tokens only. It is
> **not** end-to-end correctness and tends to read high on consistent-style distillation data.

### Qualitative generation review

34 prompts spanning coding, terminal/agentic tasks, reasoning, explanation, instruction-following, and
honesty/calibration probes were run against the final checkpoint using Qwen3.5's recommended sampling
settings. Full unedited transcripts are in [`sample_generations.md`](sample_generations.md).

**Strengths.** Coding and terminal/agentic prompts were the strongest β€” correct, idiomatic solutions using
current tooling (e.g. `ss` over `netstat`, `git-filter-repo`, Argon2id) with security-aware judgment
(rotating a leaked key first, constant-time comparison, generic auth errors). Reasoning, instruction/format
following, and calibration probes were handled well. Roughly **27 of 34** responses were clean and correct.

The model is a **reasoning model**: every answer begins with a `<think>` block followed by the final
response β€” downstream consumers should parse out and strip the `<think>...</think>` span. See
[Limitations](#limitations) for usage tips.

## How to use

The base is a multimodal (image-text-to-text) architecture; for text-only use load it with
`AutoModelForImageTextToText`. Build the prompt with `tokenize=False` and then tokenize the string
(the recommended path for this tokenizer):

```python
import torch
from transformers import AutoModelForImageTextToText, AutoTokenizer

model_id = "empero-ai/Qwable-9B-Claude-Fable-5"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(
    model_id, dtype="bfloat16", device_map="auto"
)

messages = [{"role": "user", "content": "Write a Python function that merges two sorted lists."}]
text = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tok(text, return_tensors="pt").to(model.device)

out = model.generate(
    **inputs, max_new_tokens=2048, do_sample=True,
    temperature=0.7, top_p=0.95, top_k=20, repetition_penalty=1.05,
)
# Output begins with a <think>...</think> reasoning block, then the final answer.
print(tok.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
```

`repetition_penalty=1.05` is a small deviation from Qwen's default (1.0) that prevents rare
non-terminating reasoning loops; allow generous `max_new_tokens` since the model reasons before answering.

**Requirements:** a recent `transformers` (Qwen3.5 support) plus the Gated DeltaNet kernels
(`flash-linear-attention` and a CUDA-matched `causal_conv1d` build) β€” without them the linear-attention
layers fall back to slow, memory-hungry PyTorch ops.

## Limitations

Qwable-9B-Claude-Fable-5 is a focused 9B model that shines on the coding, agentic, and reasoning tasks it was
trained for. A few characteristics are worth knowing to get the best out of it:

- **It's a reasoning model.** Each response opens with a `<think>` block before the final answer, so parse
  and strip the `<think>...</think>` span for end users. On open-ended or creative prompts it may reason at
  length β€” allow generous `max_new_tokens` and use `repetition_penaltyβ‰ˆ1.05` (as in the snippet above) for
  consistently crisp completions.
- **Strongest within its domain.** Capability is concentrated in coding and agentic/tool-use tasks. For
  general-knowledge or long-form factual questions, treat specifics as you would any 9B model's β€” verify
  before relying on them, and don't expect knowledge of events outside the base model's training.
- **Reflects its base and teachers.** As a distillation fine-tune of Qwen3.5-9B on Claude Fable 5 and GPT-5.5
  traces, it carries the style and limits of those sources and received no extra safety tuning beyond the
  base model's. Add your own review/safety layer for production use.
- **Text-only fine-tune.** The base is multimodal, but only the text path was trained (vision left untouched
  and not evaluated here).

These are normal considerations for a compact, domain-focused model rather than blockers β€” used within its
wheelhouse with the sampling settings above, it's a capable and dependable coding/agentic assistant.

## Provenance & licensing

The model weights are released under **Apache-2.0**, inherited from the Qwen3.5-9B base. The fine-tuning data
comes from generated traces of Claude Fable 5 and GPT-5.5 (via the linked public datasets). Because those
traces originate from third-party assistants, the providers' terms may apply to downstream training and
distillation β€” so if you plan to build on this model commercially, it's worth confirming your use aligns with
those terms. Shared with the community for research and experimentation, as-is.

## Support / Donate

If this model helped you, consider supporting the project:

- **BTC**: `bc1qx6zepu6sfkvshgdmc4ewu6pk6rpadvpgffpp7v`
- **LTC**: `ltc1qv2mefzps2vtjcpwfx8xxdrpplrcvltswm68r7x`
- **XMR**: `42Dbm5xg5Nq26fdyzfEU7KBnAJfhi7Cvz5J2ex5CzHXkfKuNEJzYCcmJ1GTbgjFZ5MBx72sdG1G9239Cd6rsZfv4QeDkYJY`

## Acknowledgements

- Developed and released by [Empero](https://empero.org)
- Base model: [Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B) (Alibaba Qwen team)
- Datasets: [`Glint-Research/Fable-5-traces`](https://huggingface.co/datasets/Glint-Research/Fable-5-traces),
  [`Roman1111111/gpt5.5-terminal`](https://huggingface.co/datasets/Roman1111111/gpt5.5-terminal)
- Training: [TRL](https://github.com/huggingface/trl) + [Transformers](https://github.com/huggingface/transformers)