benthecarman commited on
Commit
c0087c1
·
verified ·
1 Parent(s): b773825

Rewrite the model card

Browse files
Files changed (1) hide show
  1. README.md +89 -701
README.md CHANGED
@@ -14,634 +14,128 @@ tags:
14
  - speculative-decoding
15
  ---
16
 
17
- # MiMo-V2.6-Flash-RL — EXL3 2.27 bpw (with the DFlash drafter)
18
 
19
  [XiaomiMiMo/MiMo-V2.6-Flash-RL](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL)
20
- quantized to **EXL3 ~2.27 bpw** so that a 309B-parameter MoE runs on **one 128 GB
21
- machine**. Built and measured end to end on an **NVIDIA DGX Spark (GB10, 121.6 GiB
22
- unified memory, aarch64)**.
23
-
24
- The repo also ships Xiaomi's **DFlash** drafter in `dflash/`, repaired so it actually
25
- loads and quantized to 4 bpw, so speculative decoding works out of the box (**1.58x on coding, 1.70x on
26
- reasoning — and up to 2.08x at 250K context** — see [Speed](#speed-on-a-dgx-spark)).
27
-
28
- > **Updated 2026-09-23.** Layer 47 was re-quantized with a proper fix for its fp16 overflow
29
- > (`interm_div`, see [below](#layer-47-and-interm_div)) instead of a 6 bpw override plus a raised
30
- > bad-row limit. This build needs the current
31
- > [`mimo-v2.6-flash` exllamav3 branch](https://github.com/benthecarman/exllamav3/tree/mimo-v2.6-flash);
32
- > the previous weights do not load correctly on it, and these weights do not load correctly on the
33
- > old branch. Update both together. Size, bitrate and perplexity below are for the new build. The
34
- > benchmark, speed, memory and long-context sections were measured on the previous 2.34 bpw build
35
- > (layers 0-46 are byte-identical to it) and have not been re-run.
36
 
37
  | | |
38
  |---|---|
39
- | Weights | 12 shards, **89,605,407,713 B = 89.61 GB = 83.45 GiB**; 89.68 GB / 83.52 GiB for the whole model directory |
40
- | Bitrate | **2.27 bpw** excluding head (converter's own figure); head 6.00 |
41
- | Drafter | `dflash/`, 4 bpw EXL3, 0.74 GB (the repaired BF16 original is in `dflash-bf16/`) |
42
- | Perplexity | **5.4003** wikitext-2 test, 64 x 2048 tokens (previous build: 5.3615) |
43
- | Decode | **31.5 tok/s** @2K ctx, **28.5 tok/s** @32K, batch 1, no drafter; ~42 tok/s sustained with DFlash (previous build; the new one measured 30.0 vs 30.3 tok/s for the old in the same short-prompt harness) |
44
- | Peak host memory | **95.06 GiB** of 121.63 at 32K context |
45
- | Max context | **262,144 served with the drafter attached** (393,216 measured and working, with less headroom); ~980K tokens at 1 slot within a 25.4 GiB KV budget, model max 1,048,576 |
46
- | Modalities | **text only** — no vision tower, no audio tower, no MTP head |
47
- | Runs on | [this exllamav3 branch](https://github.com/benthecarman/exllamav3/tree/mimo-v2.6-flash) + [this TabbyAPI branch](https://github.com/benthecarman/tabbyAPI/tree/mimo-v2.6-flash) |
48
-
49
- ---
50
-
51
- ## What it is
52
 
53
- MiMo-V2.6-Flash-RL is a **309B-total / 15B-active** mixture-of-experts model: 48 layers,
54
- **256 routed experts** with top-8 routing (sigmoid scoring, `noaux_tc` grouping), layer 0
55
- dense, `hidden_size` 4096, `moe_intermediate_size` 2048. Attention is hybrid: 9 global
56
- layers (64 query / 4 KV heads) and 39 sliding-window layers (window 128, 64 query / 8 KV
57
- heads, with attention sinks), fused QKV, QK head_dim 192 against V head_dim 128, and an
58
- `attention_value_scale` of 0.707.
59
 
60
- The source checkpoint is ~173 GB: routed experts are stored as OCP **MXFP4** (e2m1 nibbles
61
- with e8m0 byte scales, block 32 along K), attention and the dense layer-0 MLP as **FP8
62
- e4m3** with 128x128 `weight_scale_inv` blocks, and embeddings / `lm_head` / norms in BF16.
63
- All of that is dequantized during conversion and re-quantized as EXL3 trellis codes.
64
-
65
- **This quant is text-only.** The source repo's vision and audio towers and its three MTP
66
- heads are not converted and not present. `config.json` still carries the `vision_config`,
67
- `audio_config` and `processor_config` blocks verbatim from upstream (so the file stays a
68
- faithful description of the base model), but no corresponding tensors exist here.
69
-
70
- ### What is in this repo
71
 
72
  ```
73
- model-0000{1..12}-of-00012.safetensors EXL3 weights, 89.61 GB (83.45 GiB)
74
- model.safetensors.index.json 13.5 MB
75
- quantization_config.json 44.7 MB — per-tensor storage record, exllamav3 1.5.1
76
- config.json
77
- tokenizer.json vocab.json merges.txt tokenizer_config.json chat_template.jinja
78
- generation_config.json
79
- configuration_mimo_v2.py modeling_mimo_v2.py upstream remote code, unmodified (auto_map)
80
- dflash/ the drafter, quantized to 4 bpw EXL3 (use this one)
81
- model.safetensors 0.74 GB, includes the mask embedding
82
- config.json repaired + switches pinned (see DFlash notes)
83
- quantization_config.json
84
- dflash-bf16/ the same drafter, unquantized
85
- dflash_draft_model.safetensors 2.94 GB, byte-identical to the base repo's
86
- config.json repaired + switches pinned
87
- mask_embedding.safetensors the learned mask vector, converted from mask_embedding.pt
88
- eval/
89
- bench/RESULTS.md bench/METHOD.md the benchmark run, both sides, verbatim
90
- bench/local-*.score.json local-*.meta.json this quant's per-task scores + run metadata
91
- bench/ref-*.score.json ref-*.meta.json the unquantized FP8 reference, same harness
92
- eval-quant.json eval-overflow-64rows.json dflash-bench-*.json
93
- longctx.json the long-context sweep, every request (see Long context)
94
  ```
95
 
96
- The shards, index, `quantization_config.json`, `config.json` and the tokenizer files are the
97
- converter's output, unmodified. The files under `eval/` are from the previous build.
98
-
99
- ---
100
-
101
  ## Bitrate
102
 
103
- `-b 2.25 -hq` does not mean "2.25 bpw everywhere" — exllamav3's `-hq` plan spends bits where
104
- they matter and the per-module result is what actually landed:
105
-
106
- | module group | layers | bpw |
107
- |---|---|---|
108
- | routed experts (`mlp.experts.*.{gate,up,down}_proj`) | 24 layers (12–35) | **2.0** |
109
- | routed experts | 23 layers (1–11, 36–47) | **2.5** |
110
- | attention (`q/k/v/o_proj`) | all 48 | **4.0** |
111
- | dense MLP, layer 0 (`gate/up/down_proj`) | 1 | **3.0** |
112
- | `lm_head` | — | **6.0** |
113
- | `embed_tokens`, all norms, router weights, `e_score_correction_bias`, attention sinks | — | **16-bit** (BF16, unquantized) |
114
-
115
- Converter's summary line: `Final bitrate (excluding head): 2.27 (--hq enabled)`.
116
-
117
- ### The command
118
 
119
- Converted with [exllamav3](https://github.com/turboderp-org/exllamav3) 1.5.1 plus the
120
- MiMo-V2.6 port, on a single RTX PRO 6000 (7.3 h, ~$29 of rented GPU; layer 47 and the head
121
- were later redone from the layer-47 checkpoint in 15 min):
122
-
123
- ```sh
124
- python convert.py \
125
- -i MiMo-V2.6-Flash-RL -o mimo-2.25bpw-hq -w work \
126
- -b 2.25 -hq \
127
- -cr 250 -cc 2048
128
- ```
129
-
130
- `-cr 250 -cc 2048` is the converter default (250 calibration rows x 2048 tokens) and was
131
- kept deliberately: measuring `-cr 120` saved only 6.1% of wall time, because calibration is
132
- ~9% of an MoE layer and LDLQ + trellis search is the other ~90%.
133
-
134
- ### Layer 47 and `interm_div`
135
-
136
- On this checkpoint, layer 47's routed experts push `act(gate) * up` to about 84k on some
137
- number tokens (mostly expert 208 channel 18 and expert 70 channel 1387), past the fp16 max of
138
- 65504. `gate` and `up` on their own stay under 2.3k, and no other layer gets above 3k. Xiaomi's
139
- reference runs bf16 there and stays finite; exllamav3 runs these intermediates in fp16.
140
-
141
- The previous build worked around it at conversion time: layer 47's experts at 6 bpw, and a
142
- raised limit on non-finite calibration rows (31 of 250 were dropped). It still produced
143
- non-finite logits at inference on 33 of the converter's 250 calibration rows.
144
-
145
- This build uses `interm_div = 128` on layer 47, the same mechanism exllamav3 uses for Laguna.
146
- `up_proj` is scaled by 1/128 before quantization and `routed_scaling_factor` puts the 128 back
147
- in fp32, so the peak becomes about 660. No calibration rows went non-finite during conversion,
148
- and the finished model gives 0 non-finite rows on those same 250 rows. Layer 47's experts are
149
- back at 2.5 bpw, which accounts for the smaller file and most of the small perplexity change.
150
-
151
- Because the 1/128 lives in the quantized weights, these weights need a matching exllamav3
152
- branch. fp32 intermediates are not a substitute: the fused prefill kernel stores fp16
153
- regardless, and without it the activation kernel clamps the product to 65504.
154
-
155
- ---
156
 
157
  ## Quality
158
 
159
- ### Perplexity
160
-
161
- wikitext-2 **test** split, non-overlapping rows, via `exllamav3/eval/ppl.py`:
162
-
163
- | rows x length | perplexity | scored tokens |
164
- |---|---|---|
165
- | **64 x 2048** | **5.4003** | 131,008 |
166
-
167
- Previous build, same method: 16 x 2048 5.2410, 64 x 2048 5.3615, 146 x 2048 (the entire test
168
- split) 5.3174. Row count moves the number by ~±0.04, so quote it with the number.
169
- **0 non-finite tokens** in every run.
170
-
171
- ### Benchmarks
172
-
173
- **Measured on the previous 2.34 bpw build**, whose layer 47 differs from this one. Not re-run.
174
-
175
- **Read this first.** These are absolute scores from one specific harness, so they are **not**
176
- comparable to anyone else's published figures. They *are* comparable to each other: the
177
- "reference FP8" column is the **unquantized model**, run through the **same script, the same
178
- seeded items, the same prompts, the same greedy sampling and the same `max_tokens`** — only
179
- `--base-url` and `--model` differed. Read the caveats below the table before the numbers.
180
-
181
- | task | N | this quant (2.36 bpw EXL3) | reference FP8, same harness | delta | published by Xiaomi |
182
- |---|---|---|---|---|---|
183
- | HumanEval+ (pass@1, `plus`) | 164 (full set) | **89.6%** | **90.2%** | **−0.6 pp** | not published |
184
- | MBPP+ (pass@1, `plus`) | 378 (full set) | **77.5%** | **77.2%** | **+0.3 pp** | not published |
185
- | GSM8K (exact match) | 500 | **95.8%** (479/500) | **96.4%** (482/500) | **−0.6 pp** | not published |
186
- | MMLU-Pro (acc) | 500, stratified over all 14 categories | **72.6%** (363/500) — a **floor**, true value 72.6–74.8% | **76.4%** (382/500), floor of 76.4–78.8% | **−3.8 pp** | not published |
187
- | GPQA-Diamond (acc, 8K budget) | 64 | **46.9%** (30/64) | **56.2%** (36/64) | **−9.4 pp** | not published |
188
- | GPQA-Diamond (acc, runs that finished inside the 8K budget) | 35 / 40 | **85.7%** | **90.0%** | **−4.3 pp** | not published |
189
-
190
- Base (non-`plus`) pass@1: HumanEval **92.07** quant vs **92.68** reference (−0.6 pp); MBPP
191
- **88.10** vs **90.21** (−2.1 pp).
192
-
193
- **The short version: code and arithmetic survive 2.36 bpw almost intact (−0.6 / +0.3 / −0.6 pp),
194
- broad factual recall does not (−3.8 pp on MMLU-Pro), and long-form reasoning degrades mostly by
195
- running long rather than by being wrong** — GPQA's −9.4 pp is about half budget (the quant hit
196
- the 8,192-token cap on 45.3% of items against the reference's 37.5%) and about half quality
197
- (−4.3 pp among the runs that finished, which at n=35/40 is inside the noise). If you serve this
198
- quant for reasoning work, give it a larger `max_tokens` than you would give the FP8 model.
199
-
200
- Item-level, on exactly the same items, the disagreements on the three near-zero-delta tasks are
201
- **two-sided and balanced** — HumanEval+ 4 items the quant wins vs 5 the reference wins, MBPP+
202
- 12 vs 11, GSM8K 6 vs 9 — i.e. noise around a shared answer rather than a systematic loss. On
203
- MMLU-Pro (26 vs 45) and GPQA (1 vs 7) the asymmetry is real.
204
-
205
- **Settings, identical for every row:** greedy (`temperature 0.0`, `top_p 1.0`), zero-shot,
206
- seed 1234, batch 1; the quant column was served by TabbyAPI with the DFlash drafter on.
207
- Thinking is **off**
208
- (`chat_template_kwargs: {"enable_thinking": false}`) for HumanEval+, MBPP+, GSM8K and
209
- MMLU-Pro, and **on** for GPQA-Diamond. `max_tokens`: 1536 / 1024 / 1024 / 1536 / 8192
210
- respectively. Datasets: evalplus `HumanEvalPlus v0.1.10` and `MbppPlus v0.2.0` (scored with
211
- `evalplus.evaluate`), `openai/gsm8k`, `TIGER-Lab/MMLU-Pro`, `Idavidrein/gpqa` `gpqa_diamond`.
212
- The reference column used **every one of those settings unchanged**, and `--only-ids-from` the
213
- quant's own output so both columns cover exactly the same item ids (GPQA: the same 64).
214
-
215
- **Caveats that matter more than the numbers:**
216
-
217
- 1. **These are not comparable to Xiaomi's published numbers, because there are none.** The
218
- MiMo-V2.6 technical report's only results table is entirely *agentic* — DeepSWE v1.1
219
- 67.9, Terminal Bench 2.1 87.6, OSWorld-Verified 80.8, CyberGym 95.1, and so on. Greps for
220
- GPQA / AIME / MMLU / LiveCodeBench / HumanEval / GSM8K / IFEval / SimpleQA / HLE /
221
- MATH-500 / BBH / DROP over all 44 pages return **zero hits**. Nothing in the table above
222
- has a published baseline. If standard-benchmark numbers for the FP8 model are published
223
- later, the "published by Xiaomi" column is where they go — **it is deliberately empty
224
- rather than filled with guesses.**
225
- 2. **What the reference column actually is.** The unquantized model as
226
- **`xiaomi/mimo-v2.6-flash` on OpenRouter**, run **2026-09-22**. OpenRouter lists exactly
227
- **one provider** for it — **Xiaomi**, endpoint quantization **fp8** — so there was no routing
228
- variance and no third-party system prompt; every stored response carries
229
- `"provider": "Xiaomi"`. **The thinking flags were honoured exactly**, which was checked on all
230
- 1,606 responses rather than assumed: the 1,542 thinking-off requests
231
- (`reasoning: {"enabled": false}` plus `chat_template_kwargs: {"enable_thinking": false}`) all
232
- returned `completion_tokens_details.reasoning_tokens == 0`, an empty `reasoning` field and no
233
- `<think>` tag in `content`; the 64 thinking-on GPQA requests all returned reasoning text, at a
234
- volume closely matching the local side (mean 12,578 chars reference vs 12,668 quant).
235
- **No reference-side failures:** 0 API errors, 0 empty responses on the thinking-off tasks, no
236
- `finish_reason` other than `stop`/`length`. One honest asymmetry: OpenRouter does not list
237
- `seed` among this endpoint's supported parameters, so the harness's seed is probably ignored
238
- upstream — at `temperature 0.0 / top_p 1.0` that should not matter, but the reference side is
239
- "greedy as the provider implements it", not "greedy with seed 1234". **Total spend $0.2206**
240
- for all 1,606 requests (654,552 generated + 266,199 prompt tokens), confirmed against
241
- OpenRouter's `/api/v1/auth/key` usage field.
242
- 3. **Greedy is a deliberate deviation** from the base model card's recommended
243
- `temperature 1.0 / top_p 0.95`. One sample at temperature 1.0 has enough variance to
244
- swamp a quantization delta. The absolute scores here are therefore *not* the model's
245
- best-effort scores.
246
- 4. **GPQA's 46.9% is a budget artefact.** 29 of 64 runs (45.3%) hit the 8,192-token thinking
247
- cap, and every one scored zero because the model was still inside `<think>` and never
248
- emitted an answer line. Mean generation 4,418 tokens, median 3,018 — bimodal: settle in
249
- ~3K or run away past 8K. The honest bracket is **46.9%–85.7%**, depending on how much
250
- budget you pay for. A real GPQA number needs a 32K budget, which at ~43 tok/s batch 1 is
251
- ~21 h for 100 items. **The reference model hits the same wall** — 24 of the same 64 items
252
- (37.5%), also scoring zero on every one, bracket 56.2%–90.0%. So the cap is a property of the
253
- task-plus-budget, not of the quantization; the quant simply runs away past 8K more often
254
- (45.3% vs 37.5%), and that gap is most of the headline −9.4 pp.
255
- 5. **MMLU-Pro's 72.6% is likewise a floor.** All 10 unparseable answers are pure truncation
256
- at the 1536-token cap (11 items = 2.2% hit it, none scored). The reference has the same
257
- shape (14 items = 2.8% capped, floor 76.4% / ceiling 78.8%), so the caps do **not** explain
258
- the gap: the two floor-to-ceiling bands, 72.6–74.8% and 76.4–78.8%, do not overlap.
259
- **−3.8 pp on MMLU-Pro is the one delta that survives every caveat.**
260
-
261
- **Stability, which the deltas do not show:** across **1,606 local generations / 574,407
262
- generated tokens** (and 1,606 reference generations / 654,552 tokens) there were **0 API errors
263
- on either side, 0 non-finite artefacts on either side, 0 empty code solutions on either side**,
264
- and exactly **one** pathological response in the whole local suite (a GSM8K item that returned an
265
- immediate EOS at zero tokens; the reference answered that item normally). Nothing traced to the
266
- layer-47 fp16 overflow. A 2.36 bpw quant that was actually broken would not look like this — and
267
- now there is a same-harness reference column saying how much it gave up.
268
-
269
- Full method, prompts, seeds, per-task throughput and every failure signal:
270
- [`eval/bench/METHOD.md`](eval/bench/METHOD.md) and
271
- [`eval/bench/RESULTS.md`](eval/bench/RESULTS.md).
272
-
273
- ### Reproducing both sides
274
-
275
- The harness lives in the [exllamav3 branch's](https://github.com/benthecarman/exllamav3/tree/mimo-v2.6-flash)
276
- project notes; both sides differ only in `--base-url`, `--api-key-env` and `--model`:
277
 
278
- ```sh
279
- # local side, against a TabbyAPI server holding this quant
280
- python run_bench.py --task gsm8k --n 500 --seed 1234 \
281
- --base-url http://127.0.0.1:30001/v1 --model mimo-2.25bpw-hq \
282
- --concurrency 2 --out raw/local-gsm8k.jsonl
283
- python score_bench.py raw/local-gsm8k.jsonl
284
-
285
- # reference side: the unquantized model, same items, same prompts, same caps
286
- export OPENROUTER_API_KEY=... # model id: xiaomi/mimo-v2.6-flash
287
- ./run_reference.sh # measured $0.22 at $0.14/M prompt, $0.28/M completion
288
- ```
289
-
290
- `--only-ids-from <the local jsonl>` restricts the reference to exactly the item ids the
291
- local side generated, so the two columns are always over the same items.
292
-
293
- ---
294
 
295
- ## Speed on a DGX Spark
296
-
297
- GB10, 121.6 GiB unified memory, batch 1, greedy, FP16 KV. Measured on the previous 2.34 bpw
298
- build (86.1 GiB); the new one is 2.6 GiB smaller.
299
-
300
- ### Decode and prefill, no drafter
301
-
302
- | prompt | input tokens | TTFT | prefill | decode |
303
  |---|---|---|---|---|
304
- | 2K | 2,002 | 3.37 s | 593.5 tok/s | **31.5 tok/s** |
305
- | 32K | 32,747 | 39.63 s | 826.2 tok/s | **28.5 tok/s** |
306
-
307
- Decode falls only **9.5%** across a 16x context increase, because only 9 of 48 layers grow
308
- their KV with context — the other 39 run on a fixed-size sliding-window ring.
309
 
310
- Model load: **21–23 s** for 86.1 GiB off NVMe (~3.9 GB/s).
 
 
311
 
312
- ### With the DFlash drafter
313
 
314
- **The 4 bpw drafter now in `dflash/` is faster than the BF16 one on every prompt measured.**
315
- Same harness for both: greedy, 512 generated tokens, dynamic draft window, median of 3 runs:
316
 
317
- | prompt | no draft | BF16 drafter | **4 bpw drafter** |
318
  |---|---|---|---|
319
- | coding | 31.4 | 34.8 | **48.5** |
320
- | prose | 31.4 | 34.4 | **40.2** |
321
- | reasoning | 31.4 | 52.7 | **59.5** |
322
-
323
- Every draft step reads the drafter's weights once, so a quarter of the bytes makes drafting
324
- much cheaper. Drafted tokens are always verified by the target model, so the drafter's
325
- precision changes how many drafts are accepted, never what is generated.
326
-
327
- The rest of this section was measured earlier with the BF16 drafter, on a different harness.
328
-
329
- 512-token generations, one configuration per process. Static 7-token draft window:
330
-
331
- | prompt | no draft | DFlash | speedup | acceptance | accepted/step |
332
- |---|---|---|---|---|---|
333
- | coding | 28.16 | **44.41** | **1.58x** | **74.06%** | 6.18 |
334
- | reasoning | 31.29 | **53.18** | **1.70x** | **60.35%** | 5.22 |
335
- | prose | 31.11 | 20.72 | **0.67x** | 12.43% | 1.87 |
336
-
337
- Prose is a real **slowdown** with a static window: 1,673 draft tokens spent to win 208, and
338
- every rejected block still costs a target forward pass.
339
-
340
- `dynamic_draft` shrinks the window from observed acceptance and is the **recommended serving
341
- default**:
342
-
343
- | prompt | no draft | static | **dynamic** | dynamic speedup | acceptance static → dynamic |
344
- |---|---|---|---|---|---|
345
- | coding | 28.16 | 44.41 | 38.90 | 1.38x | 74.06% → 81.07% |
346
- | reasoning | 31.29 | 53.18 | **55.66** | **1.78x** | 60.35% → 67.54% |
347
- | prose | 31.11 | 20.72 | **27.19** | **0.87x** | 12.43% → 50.41% |
348
-
349
- Dynamic trades ~12% of the coding peak for a far better worst case (0.67x → 0.87x) and a
350
- better mean (1.35x vs 1.31x). For peak coding throughput, turn it off; for a pure prose
351
- workload, drop the drafter entirely.
352
-
353
- Across the whole 3.77 h benchmark suite above, the server sustained **~42 generated tok/s**
354
- batch 1 with dynamic drafting — i.e. ~1.35x over the 31.5 tok/s no-draft figure, on a real
355
- mixed workload.
356
-
357
- **n-gram drafting never helps** here (0.87–0.98x at 1.8–24.7% acceptance): this model does
358
- not repeat itself enough.
359
 
360
- ### Memory
361
 
362
- | | GiB |
363
- |---|---|
364
- | MemTotal (GB10) | 121.63 |
365
- | **peak in use, 32K context** | **95.06** |
366
- | minimum MemAvailable seen | 26.57 |
367
- | torch peak allocated | 86.98 |
368
- | steady state while serving at 64K + drafter | MemAvailable **21.2** |
369
-
370
- KV cost, derived from `config.json`: **27.00 KiB/token** of paged cache (9 global-attention
371
- layers) **plus a fixed 175.5 MiB ring per slot** (768-token ring x 39 sliding-window layers).
372
- With 86.20 GiB of weights and a 10 GiB reserve, the KV budget is 25.43 GiB:
373
-
374
- | | max context, 1 slot | 4 slots |
375
- |---|---|---|
376
- | ring, FP16 KV | **980,736 tok** | 240,128 |
377
- | ring, Q8 paged KV (global layers only) | 1,048,576 (model max) | 480,256 |
378
- | full (non-ring) KV, for comparison | 102,144 | 25,344 |
379
-
380
- 64K context at 1 slot costs 1.86 GiB; 32K costs 1.02 GiB. Quantized paged KV (k=v=8) loads
381
- and generates coherently on top of the previous 2.36 bpw weights, but its quality is untested — the
382
- default stays FP16. The 39 sliding-window rings are always FP16.
383
-
384
- Those are the arithmetic limits. What was actually served, loaded and measured — up to
385
- **350,091 tokens** — is in [Long context](#long-context) below; the cache is **preallocated in
386
- full at load**, so `max_seq_len` is a memory decision, not a ceiling you pay for on demand.
387
-
388
- #### A unified-memory hazard worth knowing about
389
-
390
- On GB10 there is **no separate VRAM**: host and GPU share one pool. exllamav3's shard loader
391
- uses buffered `pread`, not `mmap`, so loading 86.1 GiB leaves **86.1 GiB of clean page cache**
392
- competing with the 86.1 GiB of weights in the same 121.6 GiB pool — and
393
- `torch.cuda.mem_get_info()` reports MemFree, not MemAvailable, so nothing in-process can see
394
- it. Unified memory does not produce an OOM kill; it starves the machine, and recovery is a
395
- physical power cycle. Two mitigations are in the exllamav3 branch (upstream PR #398) and are
396
- why the numbers above are flat rather than degrading:
397
-
398
- * `Model.load_gen()` drops the shard page cache at the end of every load
399
- (`EXL3_KEEP_PAGE_CACHE=1` disables); measured **0.00 GiB of 86.08 GiB resident** while serving.
400
- * `EXL3_LOAD_DEVICE=cuda:0` takes a single-device load path that does no MemFree-based budget
401
- arithmetic. Without it, autosplit can refuse a model that actually fits.
402
-
403
- ---
404
-
405
- ## Long context
406
-
407
- Measured on the DGX Spark, batch 1, greedy, thinking off, FP16 KV, needles and haystacks built
408
- from wikitext-2 paragraphs. **Every retrieval test passed.**
409
-
410
- ### Retrieval
411
-
412
- A six-digit passcode is hidden in a wikitext haystack at 10 / 50 / 90% depth and asked for at
413
- the end. Exact match on the digits, fresh city and fresh code per request.
414
 
415
- | prompt tokens | 10% | 50% | 90% |
416
  |---|---|---|---|
417
- | ~8.2K | OK | OK | OK |
418
- | ~32.8K | OK | OK | OK |
419
- | ~65.5K | OK | OK | OK |
420
- | ~131.1K | OK | OK | OK |
421
- | ~200.1K | OK | OK | OK |
422
- | ~250.0K | OK | OK | OK |
423
- | **350,091** | **OK** | — | — |
424
-
425
- **18/18** at `max_seq_len 262144`, plus **350,091 tokens** answered correctly in a 393,216
426
- window — the largest prompt this quant has been given. Nine more cells at `max_seq_len 131072`
427
- with the drafter on also passed, and **five more at `max_seq_len 393216` with the drafter
428
- attached** (131K at 10/50/90% depth, 250K at 50% and 90%, all exact), so retrieval is
429
- unaffected by speculative decoding or by the window size.
430
-
431
- That matters more than it looks: **39 of 48 layers are sliding-window with a 128-token
432
- window** and run on a fixed 768-token ring, so a needle 25,000 tokens into a 250K prompt has
433
- to survive ~195 ring rebases and reach the question through the 9 global-attention layers
434
- alone. The ring had previously only been validated to 2,048 tokens.
435
-
436
- **Five needles at once, 128K prompt** — all five returned, in order of appearance:
437
-
438
- ```
439
- Montevideo: 604403
440
- Bratislava: 991476
441
- Kathmandu: 472495
442
- Ulaanbaatar: 379397
443
- Ljubljana: 361254
444
- ```
445
-
446
- **A real long document, ~100K tokens** — complete wikitext articles concatenated with *Plain
447
- maskray* buried in the middle, three questions answerable only from that article:
448
-
449
- ```
450
- (a) 2008
451
- (b) ~54 Ma
452
- (c) 12 and 62 m
453
- ```
454
 
455
- All three correct against the source ("In 2008, Last and William White elevated the kuhlii
456
- group…", "estimated to have occurred ~ 54 Ma", "recorded from between 12 and 62 m"), including
457
- the source's own "~" hedge.
458
 
459
- ### Speed and memory vs prompt length
460
-
461
- No drafter, 128 generated tokens, `max_seq_len 262144` (last row 393,216):
462
-
463
- | prompt | TTFT | prefill | decode | min MemAvailable |
464
- |---|---|---|---|---|
465
- | ~8.2K | 9.1–10.6 s | 768–905 tok/s | 31.0 tok/s | 19.85 GiB |
466
- | ~32.8K | 41.2–42.0 s | 780–799 tok/s | 28.4 tok/s | 19.56 GiB |
467
- | ~65.5K | 100.1–100.5 s | 652–654 tok/s | 25.3 tok/s | 19.10 GiB |
468
- | ~131.1K | 270.0–271.5 s | 483–486 tok/s | 21.7 tok/s | 18.98 GiB |
469
- | ~200.1K | 524.8–525.9 s | 380–381 tok/s | 18.3 tok/s | 18.85 GiB |
470
- | ~250.0K | 757.6–758.8 s | 329–330 tok/s | 16.7 tok/s | 17.77 GiB |
471
- | **350,091** | **1,346.8 s** | **260 tok/s** | **13.8 tok/s** | **15.91 GiB** |
472
-
473
- **Decode falls only 2.3x from 8K to 250K** — the sliding-window ring again: only 9 of 48
474
- layers grow their KV with context. Prefill falls 2.7x and TTFT is quadratic, as full
475
- attention on those 9 layers requires. The curve fits
476
-
477
- ```
478
- TTFT(seconds) ≈ 988·L + 8172·L² (L = prompt tokens in millions)
479
- ```
480
-
481
- to better than 1.5% from 64K to 350K — the 350K point was predicted at 1,347 s from a fit to
482
- the 128K and 250K points alone and came in at 1,346.8 s. Extrapolated: **400K ≈ 28 min,
483
- 500K ≈ 42 min of prefill.** Long context on one GB10 is TTFT-bound, not memory-bound.
484
-
485
- With the DFlash drafter at `max_seq_len 131072`, on the same short-answer retrieval task
486
- (real generations, ~380 tokens, unforced): 36.3 / 29.8 / 27.0 tok/s at 2K / 65K / 130K, i.e.
487
- 1.17–1.24x, live acceptance 38–85%. Those answers are three tokens long and factual, which is
488
- the worst case for a drafter.
489
-
490
- **On a real generation task the drafter is worth far more, and worth *more* the longer the
491
- context.** A repository-sized context (source files) and a task asking for a ~400-token
492
- implementation, `max_seq_len 262144`, batch 1, greedy, same prompts against a freshly loaded
493
- server:
494
-
495
- | prompt tokens | no drafter | **DFlash drafter** | **speedup** | acceptance |
496
- |---|---|---|---|---|
497
- | 64,614 | 25.08 tok/s | **43.77 tok/s** | **1.75x** | 64% |
498
- | 130,118 | 21.08 tok/s | **41.43 tok/s** | **1.97x** | 61% |
499
- | 248,993 | 16.14 tok/s | **33.63 tok/s** | **2.08x** | 58% |
500
-
501
- The speedup **grows** with context (1.75x → 2.08x) even though acceptance **falls** (64% →
502
- 58%). Verifying a block of 8 drafted tokens is a single target forward, and at 250K that
503
- forward is dominated by reading the 9 global-attention layers' K/V — a cost paid once for the
504
- whole block instead of once per token. So at long context a lower acceptance rate still buys a
505
- larger speedup. **248,993 tokens of context, decoding at 33.6 tok/s.**
506
-
507
- ### How much context fits
508
-
509
- The paged cache is **preallocated in full at load** — `cache_size` is paid up front whether or
510
- not anyone sends a long prompt. Cost per slot:
511
-
512
- ```
513
- paged KV 27.00 KiB/token (the 9 global-attention layers only)
514
- SWA ring 175.5 MiB per slot (768-token ring x 39 layers, always FP16)
515
- draft KV 35.0 MiB per slot (5 DFlash layers on a 1792-token window ring)
516
- drafter weights 2.81 GiB (BF16; about 0.7 GiB with the 4 bpw drafter)
517
- ```
518
-
519
- so **a slot costs 27.0 KiB/token with or without the drafter**, and the drafter's whole
520
- footprint is a flat 2.85 GiB at any context length (BF16 drafter; the 4 bpw one saves about
521
- 2.1 GiB more, so the MemAvailable figures in this section are conservative). That is new: the drafter used to size its
522
- K/V for the whole context at **20.00 KiB/token** — a 74% surcharge, 7.50 GiB at `max_seq_len`
523
- 393216 — although all five of its layers are sliding-window with a 1024-token window and none
524
- of them ever looks further back. They now run on a fixed per-slot ring of window + block + two
525
- pages, so draft K/V is **35.0 MiB per slot at any context**: 2,560 MiB → 35.0 MiB at 131072,
526
- 7,680 MiB → 35.0 MiB at 393216. Generations are token-identical either way.
527
-
528
- On this box the whole budget reduces to one line that held to within 0.1 GiB across every
529
- configuration:
530
-
531
- ```
532
- MemAvailable after load ≈ 29.2 GiB − (paged KV + ring + drafter weights + draft KV)
533
- ```
534
-
535
- (121.63 GiB total − 86.15 GiB of weights − ~6 GiB of runtime.)
536
-
537
- | `max_seq_len` | drafter | load | MemAvailable after load | min under load |
538
- |---|---|---|---|---|
539
- | **393,216** | **DFlash** | **23.3 s** | **16.09 GiB** | see below |
540
- | 262,144 | DFlash | 23.2 s | **19.57 GiB** | **16.70 GiB** @ 249K |
541
- | 262,144 | off | 21.3 s | 22.27 GiB | **19.59 GiB** @ 249K |
542
- | 393,216 | off | 21.3 s | 18.65 GiB | **15.91 GiB** @ 350K |
543
- | 131,072 | DFlash | 22.6 s | 20.17 GiB | **17.07 GiB** @ 130K |
544
- | 65,536 | DFlash | 22.7 s | 24.48 GiB | 22.21 GiB @ 65K |
545
-
546
- **Recommended: `max_seq_len 262144` with the drafter** — 19.57 GiB free after load and
547
- **16.70 GiB** through a 249K-token request. With the drafter's old full-length K/V cache that
548
- configuration had ~14.6 GiB after load and ~10 GiB under a full-length request, i.e. it did not
549
- fit at all; the window ring is what makes it fit.
550
-
551
- **`max_seq_len 393216` with the drafter also works** — 16.09 GiB after load, 13.45 GiB through
552
- a 348,970-token generation, retrieval 5/5 — but it is the edge of this box rather than a
553
- set-and-forget setting: over an hour of heavy serving MemAvailable drifted 16.09 → 13.0 GiB
554
- (TabbyAPI's host RSS grows several GiB under sustained load) and the longest requests bottomed
555
- at 12.66 GiB. Use it deliberately when a prompt needs it.
556
-
557
- * **524,288 does not fit at FP16.** Its 13.67 GiB of paged KV leaves ~15.3 GiB after load, and
558
- the measured 2.7–4.5 GiB transient of a full-length request would land it at ~11–12.5 GiB —
559
- at or through a 12 GiB safety floor. On unified memory that is not an OOM kill, it is a
560
- frozen machine. A 500K prefill would also take ~42 minutes.
561
- * The drafter no longer costs anything per token, so **there is no longer a window-vs-drafting
562
- trade-off** on this box. Turn it on.
563
-
564
- So on a 121 GiB box you get **both**: a 262,144-token window (393,216 at the edge) *and* the
565
- drafter, at 1.75–2.08x the decode rate on real generation. This repo's reference server runs
566
- 262,144 with the drafter.
567
-
568
- ### Recommended serving settings
569
-
570
- ```yaml
571
- model:
572
- max_seq_len: 262144 # with the drafter; the draft cache no longer scales with it.
573
- cache_size: 262144 # 393216 also fits with the drafter, with less headroom
574
- cache_mode: FP16
575
- chunk_size: 2048
576
- max_batch_size: 1 # every extra slot repeats the SWA ring, the paged span and a 35 MiB draft ring
577
- draft_model:
578
- draft_mode: model
579
- draft_model_name: mimo-dflash-draft
580
- draft_cache_mode: FP16
581
- dynamic_draft: true
582
- ```
583
-
584
- ### Caveats
585
-
586
- * **`max_seq_len` bounds the prompt, not prompt + generation.** A 131,092-token prompt against
587
- `max_seq_len 131072` returns `400 Bad Request` / `Prompt length 131092 exceeds the …` before
588
- a token is generated. Budget `max_seq_len ≥ prompt + max_tokens`, and remember the chat
589
- template adds ~20 tokens.
590
- * **Every number here is `max_batch_size: 1`.** A second slot adds another 175.5 MiB ring, its
591
- own paged span and its own draft cache; concurrency at 128K is not free.
592
- * Prefill is chunked at 2,048 tokens. Prefill rates for a prompt sharing a long prefix with the
593
- previous request are inflated by page reuse (755 vs 473 tok/s at 130K here), so cold numbers
594
- are the ones quoted above.
595
- * Q8 paged KV halves the 27 KiB/token term and would make ~786K arithmetically fit, but
596
- quantized KV quality on top of ~2.3 bpw weights is untested and the prefill time would be
597
- hours. The 39 sliding-window rings stay FP16 regardless.
598
- * Raw records: [`eval/longctx.json`](eval/longctx.json).
599
-
600
- ---
601
 
602
  ## How to run
603
 
604
- ### Requirements
 
 
605
 
606
- This quant does **not** load on released exllamav3. `MiMoV2ForCausalLM` is not upstream yet;
607
- the port is in three open PRs (turboderp-org/exllamav3
608
- [#396](https://github.com/turboderp-org/exllamav3/pull/396),
609
- [#398](https://github.com/turboderp-org/exllamav3/pull/398),
610
- [#399](https://github.com/turboderp-org/exllamav3/pull/399)) and all three sit on one branch:
611
-
612
- * **exllamav3** — https://github.com/benthecarman/exllamav3/tree/mimo-v2.6-flash
613
- * **TabbyAPI** — https://github.com/benthecarman/tabbyAPI/tree/mimo-v2.6-flash
614
- (one patch: honour `EXL3_LOAD_DEVICE` when building the split)
615
-
616
- exllamav3 1.5.1 ships no server of its own, so serving is TabbyAPI driven against an editable
617
- checkout of that exllamav3 branch. On aarch64 you also need `uvloop` installed manually —
618
- TabbyAPI marks it x86_64-only in `pyproject.toml` but imports it unconditionally on Linux.
619
 
620
  ```sh
621
  git clone -b mimo-v2.6-flash https://github.com/benthecarman/exllamav3
622
  git clone -b mimo-v2.6-flash https://github.com/benthecarman/tabbyAPI
623
- pip install -e ./exllamav3 # torch 2.12.1+cu132 / triton 3.7.1 were used here
624
-
625
- # TabbyAPI: base dependencies ONLY. Never `pip install ./tabbyAPI[cu12]` or `[cu13]` --
626
- # those extras pin torch 2.9/2.11 and a prebuilt exllamav3 wheel, which would clobber
627
- # both your torch and the editable checkout above.
628
- pip install "fastapi-slim>=0.115" "pydantic>=2.11,<3" ruamel.yaml rich "uvicorn>=0.28.1" \
629
- "jinja2>=3.0.0" loguru "sse-starlette>=2.2.0" packaging "tokenizers>=0.21.0" numpy \
630
- aiofiles aiohttp async_lru huggingface_hub psutil "httptools>=0.5.0" pillow requests setuptools
631
- pip install uvloop # aarch64 only: pyproject marks it x86_64, main.py imports it anyway
632
-
633
  hf download benthecarman/MiMo-V2.6-Flash-RL-exl3 --local-dir mimo-exl3
634
  ```
635
 
636
- ### Serving config
 
637
 
638
- TabbyAPI resolves `draft_model_name` inside `draft_model_dir`, so the drafter needs its own
639
- directory. Split the download into two:
640
 
641
  ```
642
  models/
643
- mimo-2.25bpw-hq/ everything except dflash/ and eval/
644
- mimo-dflash-draft/ the three files from dflash/
645
  ```
646
 
647
  `config.yml`:
@@ -650,18 +144,14 @@ models/
650
  model:
651
  model_dir: models
652
  model_name: mimo-2.25bpw-hq
653
- max_seq_len: 262144 # validated, with the drafter attached; 393216 also fits
654
  cache_size: 262144
655
- cache_mode: FP16 # default; Q8 paged KV works but is unvalidated for quality
656
  chunk_size: 2048
657
- max_batch_size: 1 # each extra slot costs another 175.5 MiB SWA ring + paged span + 35 MiB draft ring
658
- gpu_split_auto: true
659
- vision: false
660
  reasoning: true
661
  reasoning_start_token: "<think>"
662
  reasoning_end_token: "</think>"
663
- start_in_reasoning: auto
664
- # prompt_template unset -> TabbyAPI picks up chat_template.jinja from the model dir
665
 
666
  draft_model:
667
  draft_mode: model
@@ -671,119 +161,17 @@ draft_model:
671
  dynamic_draft: true
672
  ```
673
 
674
- TabbyAPI auto-detects the reasoning tags `<think> </think>` and tool format **`qwen3_coder`**
675
- from the template and tokenizer. Tool calling has been exercised against this quant
676
- (`finish_reason: tool_calls`, `get_weather({"city": "Paris, France"})`, `parsed 1 tool call
677
- (qwen3_coder)`). Thinking is switched off per request with
678
- `"chat_template_kwargs": {"enable_thinking": false}`.
679
-
680
- A ready-made `serve.sh` + `tabby-config.yml` live in
681
- [`examples/mimo_v2_6/`](https://github.com/benthecarman/exllamav3/tree/mimo-v2.6-flash/examples/mimo_v2_6)
682
- on the exllamav3 branch.
683
-
684
- ### Environment knobs
685
-
686
- | variable | default | what it does |
687
- |---|---|---|
688
- | `EXL3_LOAD_DEVICE` | unset | set to `cuda:0` on GB10 / any unified-memory box: single-device load, no autosplit, no MemFree arithmetic. **Recommended here.** |
689
- | `EXL3_MIMO_FP32_MLP_LAYERS` | `none` | debugging only: layers to run with fp32 MoE intermediates. Not needed with this build. |
690
- | `EXL3_KEEP_PAGE_CACHE` | unset | keep shard page cache after load (don't, on unified memory) |
691
-
692
- You should see this at load:
693
-
694
- ```
695
- -- EXL3_LOAD_DEVICE=cuda:0 (single-device load)
696
- -- Released 83.5 GiB of shard page cache
697
- ```
698
-
699
- ### The DFlash drafter
700
-
701
- `dflash-bf16/` is Xiaomi's shipped 5-layer DFlash drafter, **repaired**. The tensor file is
702
- byte-identical to the base repo's; the two small files next to it are not, and the
703
- differences are load-bearing:
704
-
705
- * **`config.json` is valid JSON here.** The upstream one ends `"use_cache": true,` with a
706
- trailing comma and cannot be parsed at all.
707
- * **`tap_shift: 0`**, not exllamav3's legacy `+1`. MiMo's `extract_context_feature` uses
708
- `hidden_states[i+1]` = the *output* of layer i = exllamav3 export index i. SGLang's
709
- `dflash_utils.py::build_target_layer_ids` documents the same convention. Taps are
710
- `[0, 11, 23, 35, 47]`, all of them global-attention layers, so the sliding-window ring is
711
- not involved.
712
- * **`attention_sink_bias: true`**, **`attention_value_scale: 0.612`**,
713
- **`bidirectional_block: true`** (from the checkpoint's `is_causal: false`), window
714
- `(1024, 7)`, `partial_rotary_factor: 0.5` → 64 of 128 rotary dims.
715
- * **`mask_embedding.safetensors`** is the learned `[4096]` vector from the upstream
716
- `mask_embedding.pt` pickle, rewritten so the loader can find it. In the quantized
717
- `dflash/` it is stored inside `model.safetensors`.
718
-
719
- **Why the shipped `dflash/dflash.py` is stale and is not included here.** It never reads the
720
- sink biases, the value scale, the sliding window, `partial_rotary_factor`, or the mask
721
- embedding. The last one is decisive: `mask_token_id` is **151675**, which is past the end of
722
- the tokenizer, and the target's `embed_tokens[151675]` is an untrained pad row (L2 norm
723
- 2e-5, against the shipped vector's 0.765). So the reference feeds the drafter **zeros** for
724
- all seven mask slots and it degenerates to one repeated token. It cannot be the code this
725
- checkpoint was trained or served with.
726
-
727
- Every one of those corrections was **cross-checked against SGLang main**, the engine
728
- Xiaomi's own README recommends, and matches: `models/dflash.py` passes
729
- `partial_rotary_factor` into `get_rope`, applies `attention_value_scale` as `v_scale`, loads
730
- per-head `attention_sink_bias`, and maps `sliding_attention` to the 1024 window;
731
- `speculative/dflash_worker_v2.py::_maybe_merge_trained_mask_embedding()` loads
732
- `mask_embedding.pt` and merges it into the target's embedding table. (This port applies the
733
- mask vector one layer earlier, inside the drafter's input layer, so it never mutates the
734
- target's weights — equivalent, since the target can never emit 151675.)
735
-
736
- Parity against a corrected HF reference, 7 drafted positions x 64 rounds: **99.78% argmax
737
- (100% including fp16 ties), KL 3.3e-5, cos 0.999992** at context 384; 98.66% / 100% at
738
- context 1536 (beyond the 1024 window). Every single disagreement is the reference's rank-1
739
- token at a top1–top2 margin ≤ 0.017, i.e. 1–2 fp16 ULP.
740
-
741
- `dflash/` is `dflash-bf16/` converted with exllamav3's `convert.py -b 4`. DFlash drafters
742
- quantize without calibration, so the conversion needs neither calibration data nor the
743
- target model. Its `config.json` is the repaired one plus `quantization_config`.
744
 
745
- ---
746
-
747
- ## Known issues
748
-
749
- * **Greedy output is not bit-stable across drafting modes.** On a prose prompt, no-draft /
750
- DFlash / n-gram / dynamic produced 369 / 447 / 359 / 312 tokens. n-gram differs from
751
- no-draft too, so this is not a DFlash bug — it is fp16 tie-break sensitivity (traced to a
752
- single divergence at a top1–top2 margin of 0.0156 = two fp16 ULP, reproduced identically
753
- with the ring disabled). Coding was token-identical across all four modes. Don't diff
754
- outputs across drafting modes and expect equality.
755
- * **Weights and code must match.** Layer 47's `interm_div` is folded into the quantized weights.
756
- This build on the old branch, or the old build on the current branch, runs without an error
757
- and gives wrong output. See [Layer 47 and `interm_div`](#layer-47-and-interm_div).
758
- * **Instruction-following on code is imperfect.** A sample `merge_intervals` completion was a
759
- correct, idiomatic, correctly-analysed algorithm with one real bug (it sorts the caller's
760
- list in place and assumes list elements, so it raises on the tuples the prompt specified)
761
- and it skipped the three asserts the prompt asked for. Consistent with a strong-but-not-
762
- perfect ~2.3 bpw quant — and the benchmark deltas above say the FP8 model is only ~0.6 pp
763
- better at HumanEval+, so most of that is the base model, not the quantization.
764
- * **Tensor parallel is unsupported.** The port has only been built and run single-device.
765
- * **Long context is now measured, and it is TTFT-bound rather than memory-bound.** 8K → 350K
766
- has been served and retrieval is perfect at every length and depth tested, but a 250K prompt
767
- costs ~12.6 minutes of prefill and a 350K prompt ~22.4 minutes on one GB10. `max_seq_len`
768
- above 393,216 (or above 131,072 with the drafter) does not fit at FP16 KV. See
769
- [Long context](#long-context).
770
- * **Batch > 1 with DFlash is untested.** Everything above is `max_batch_size: 1`.
771
- * **No vision, no audio, no MTP.** The MTP heads in the source repo are not ported, so
772
- MTP-based drafting is unavailable; DFlash is the drafting path.
773
- * The source repo's own `preprocessor_config.json` is not included, since there is no
774
- vision/audio tower to preprocess for.
775
-
776
- ---
777
 
778
- ## License and credits
 
779
 
780
- **MIT**, inherited from [XiaomiMiMo/MiMo-V2.6-Flash-RL](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL).
781
 
782
- * **Xiaomi MiMo team** — the model, the DFlash drafter and the MIT license. All the
783
- capability measured on this page is theirs.
784
- * **[turboderp](https://github.com/turboderp)** — [ExLlamaV3](https://github.com/turboderp-org/exllamav3),
785
- the EXL3 format and the quantizer, and the architecture notes in issue #124 that made the
786
- port tractable.
787
- * **[vcruz305](https://github.com/vcruz305)** — the aarch64 build guards
788
- (commit `f4993fef`) that let exllamav3's extension compile on GB10 at all.
789
- * **[theroyallab](https://github.com/theroyallab/tabbyAPI)** — TabbyAPI.
 
14
  - speculative-decoding
15
  ---
16
 
17
+ # MiMo-V2.6-Flash-RL — EXL3 2.27 bpw
18
 
19
  [XiaomiMiMo/MiMo-V2.6-Flash-RL](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL)
20
+ (309B total / 15B active MoE) quantized to EXL3 at 2.27 bpw, so it fits on a single 128 GB
21
+ machine. Built and tested on an NVIDIA DGX Spark (GB10).
22
+
23
+ The repo also includes Xiaomi's DFlash speculative-decoding drafter, fixed so it loads and
24
+ quantized to 4 bpw.
 
 
 
 
 
 
 
 
 
 
 
25
 
26
  | | |
27
  |---|---|
28
+ | Weights | 83.45 GiB, 12 shards |
29
+ | Bitrate | 2.27 bpw (excluding head), head 6 bpw |
30
+ | Perplexity | 5.40, wikitext-2 test, 64 x 2048 tokens |
31
+ | Decode, batch 1 | ~31 tok/s without drafter; 40–60 tok/s with the 4 bpw drafter (see [Speed](#speed)) |
32
+ | Context | 262,144 tokens with the drafter on a DGX Spark |
33
+ | Modalities | Text only (no vision, audio or MTP heads) |
 
 
 
 
 
 
 
34
 
35
+ **This needs a patched exllamav3.** MiMo-V2 support is not upstream yet; see
36
+ [How to run](#how-to-run).
 
 
 
 
37
 
38
+ ## Files
 
 
 
 
 
 
 
 
 
 
39
 
40
  ```
41
+ model-*.safetensors, model.safetensors.index.json EXL3 weights
42
+ quantization_config.json per-tensor storage record
43
+ config.json, tokenizer files, chat_template.jinja
44
+ dflash/ drafter, 4 bpw EXL3 (use this one)
45
+ dflash-bf16/ the same drafter, unquantized
46
+ eval/ benchmark outputs
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
47
  ```
48
 
 
 
 
 
 
49
  ## Bitrate
50
 
51
+ Converted with `convert.py -b 2.25 -hq -cr 250 -cc 2048`. The per-module result:
 
 
 
 
 
 
 
 
 
 
 
 
 
 
52
 
53
+ | module | bpw |
54
+ |---|---|
55
+ | routed experts, layers 12–35 | 2.0 |
56
+ | routed experts, layers 1–11 and 36–47 | 2.5 |
57
+ | attention | 4.0 |
58
+ | dense MLP (layer 0) | 3.0 |
59
+ | `lm_head` | 6.0 |
60
+ | embeddings, norms, router | BF16 |
61
+
62
+ **Layer 47:** some of its experts produce intermediate values past the fp16 limit.
63
+ This build scales `up_proj` down by 128 in that layer (`interm_div`) and restores the scale in
64
+ fp32. Because the scale is folded into the weights, these weights need the current
65
+ `mimo-v2.6-flash` exllamav3 branch; on an older checkout they load but produce wrong output.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
66
 
67
  ## Quality
68
 
69
+ Perplexity (wikitext-2 test, 64 x 2048): **5.4003**.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
70
 
71
+ Benchmarks compare this quant against the unquantized FP8 model (Xiaomi's endpoint via
72
+ OpenRouter) with the same harness, items, prompts and greedy sampling.
73
+ They are only comparable to each other, not to other published scores.
 
 
 
 
 
 
 
 
 
 
 
 
 
74
 
75
+ | task | N | this quant | FP8 reference | delta |
 
 
 
 
 
 
 
76
  |---|---|---|---|---|
77
+ | HumanEval+ | 164 | 89.6% | 90.2% | −0.6 |
78
+ | MBPP+ | 378 | 77.5% | 77.2% | +0.3 |
79
+ | GSM8K | 500 | 95.8% | 96.4% | −0.6 |
80
+ | MMLU-Pro | 500 | 72.6% | 76.4% | −3.8 |
81
+ | GPQA-Diamond (thinking, 8K token cap) | 64 | 46.9% | 56.2% | −9.4 |
82
 
83
+ Most of the GPQA gap comes from the token cap: the quant ran past 8K tokens without answering
84
+ on 45% of items vs 38% for the reference. Among answered items it scored 85.7% vs 90.0%.
85
+ Method and raw outputs are in [`eval/bench/`](eval/bench/).
86
 
87
+ ## Speed
88
 
89
+ DGX Spark, batch 1, greedy, 512 generated tokens, median of 3 runs:
 
90
 
91
+ | prompt | no drafter | 4 bpw drafter | speedup |
92
  |---|---|---|---|
93
+ | coding | 31.4 tok/s | 48.5 tok/s | 1.54x |
94
+ | prose | 31.4 tok/s | 40.2 tok/s | 1.28x |
95
+ | reasoning | 31.4 tok/s | 59.5 tok/s | 1.89x |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
96
 
97
+ The target model verifies every drafted token, so the drafter does not affect output quality.
98
 
99
+ Long context with the BF16 drafter, generating ~400 tokens of code against a large
100
+ repository prompt:
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
101
 
102
+ | prompt tokens | no drafter | drafter | speedup |
103
  |---|---|---|---|
104
+ | 64,614 | 25.1 tok/s | 43.8 tok/s | 1.75x |
105
+ | 130,118 | 21.1 tok/s | 41.4 tok/s | 1.97x |
106
+ | 248,993 | 16.1 tok/s | 33.6 tok/s | 2.08x |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
107
 
108
+ Decode holds up at long context because 39 of the 48 layers use a 128-token sliding window.
109
+ Prefill does not: a 250K-token prompt takes about 13 minutes on one GB10.
 
110
 
111
+ Needle retrieval passed at every length and depth tested, from 8K to 350K tokens.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
112
 
113
  ## How to run
114
 
115
+ The port is in turboderp-org/exllamav3
116
+ [#399](https://github.com/turboderp-org/exllamav3/pull/399). Until it lands, use these
117
+ branches:
118
 
119
+ * exllamav3: https://github.com/benthecarman/exllamav3/tree/mimo-v2.6-flash
120
+ * TabbyAPI: https://github.com/benthecarman/tabbyAPI/tree/mimo-v2.6-flash
 
 
 
 
 
 
 
 
 
 
 
121
 
122
  ```sh
123
  git clone -b mimo-v2.6-flash https://github.com/benthecarman/exllamav3
124
  git clone -b mimo-v2.6-flash https://github.com/benthecarman/tabbyAPI
125
+ pip install -e ./exllamav3
 
 
 
 
 
 
 
 
 
126
  hf download benthecarman/MiMo-V2.6-Flash-RL-exl3 --local-dir mimo-exl3
127
  ```
128
 
129
+ Install TabbyAPI's base dependencies only; its `cu12`/`cu13` extras replace torch and the
130
+ editable exllamav3. On aarch64, also `pip install uvloop`.
131
 
132
+ TabbyAPI looks for the drafter by name inside `draft_model_dir`, so put `dflash/` in its own
133
+ directory next to the model:
134
 
135
  ```
136
  models/
137
+ mimo-2.25bpw-hq/ everything except dflash/, dflash-bf16/ and eval/
138
+ mimo-dflash-draft/ the contents of dflash/
139
  ```
140
 
141
  `config.yml`:
 
144
  model:
145
  model_dir: models
146
  model_name: mimo-2.25bpw-hq
147
+ max_seq_len: 262144
148
  cache_size: 262144
149
+ cache_mode: FP16
150
  chunk_size: 2048
151
+ max_batch_size: 1
 
 
152
  reasoning: true
153
  reasoning_start_token: "<think>"
154
  reasoning_end_token: "</think>"
 
 
155
 
156
  draft_model:
157
  draft_mode: model
 
161
  dynamic_draft: true
162
  ```
163
 
164
+ On a DGX Spark or other unified-memory machine, start TabbyAPI with `EXL3_LOAD_DEVICE=cuda:0`.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
165
 
166
+ Thinking can be turned off per request with `"chat_template_kwargs": {"enable_thinking": false}`.
167
+ Tool calls use the `qwen3_coder` format, which TabbyAPI detects automatically.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
168
 
169
+ A ready-made `serve.sh` and config are in
170
+ [`examples/mimo_v2_6/`](https://github.com/benthecarman/exllamav3/tree/mimo-v2.6-flash/examples/mimo_v2_6).
171
 
172
+ ## Credits
173
 
174
+ * Xiaomi MiMo team: the model and the DFlash drafter (MIT).
175
+ * [turboderp](https://github.com/turboderp): ExLlamaV3 and the EXL3 format.
176
+ * [vcruz305](https://github.com/vcruz305): the aarch64 build fixes.
177
+ * [theroyallab](https://github.com/theroyallab/tabbyAPI): TabbyAPI.