Rewrite the model card
Browse files
README.md
CHANGED
|
@@ -14,634 +14,128 @@ tags:
|
|
| 14 |
- speculative-decoding
|
| 15 |
---
|
| 16 |
|
| 17 |
-
# MiMo-V2.6-Flash-RL — EXL3 2.27 bpw
|
| 18 |
|
| 19 |
[XiaomiMiMo/MiMo-V2.6-Flash-RL](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL)
|
| 20 |
-
quantized to
|
| 21 |
-
machine
|
| 22 |
-
|
| 23 |
-
|
| 24 |
-
|
| 25 |
-
loads and quantized to 4 bpw, so speculative decoding works out of the box (**1.58x on coding, 1.70x on
|
| 26 |
-
reasoning — and up to 2.08x at 250K context** — see [Speed](#speed-on-a-dgx-spark)).
|
| 27 |
-
|
| 28 |
-
> **Updated 2026-09-23.** Layer 47 was re-quantized with a proper fix for its fp16 overflow
|
| 29 |
-
> (`interm_div`, see [below](#layer-47-and-interm_div)) instead of a 6 bpw override plus a raised
|
| 30 |
-
> bad-row limit. This build needs the current
|
| 31 |
-
> [`mimo-v2.6-flash` exllamav3 branch](https://github.com/benthecarman/exllamav3/tree/mimo-v2.6-flash);
|
| 32 |
-
> the previous weights do not load correctly on it, and these weights do not load correctly on the
|
| 33 |
-
> old branch. Update both together. Size, bitrate and perplexity below are for the new build. The
|
| 34 |
-
> benchmark, speed, memory and long-context sections were measured on the previous 2.34 bpw build
|
| 35 |
-
> (layers 0-46 are byte-identical to it) and have not been re-run.
|
| 36 |
|
| 37 |
| | |
|
| 38 |
|---|---|
|
| 39 |
-
| Weights |
|
| 40 |
-
| Bitrate |
|
| 41 |
-
|
|
| 42 |
-
|
|
| 43 |
-
|
|
| 44 |
-
|
|
| 45 |
-
| Max context | **262,144 served with the drafter attached** (393,216 measured and working, with less headroom); ~980K tokens at 1 slot within a 25.4 GiB KV budget, model max 1,048,576 |
|
| 46 |
-
| Modalities | **text only** — no vision tower, no audio tower, no MTP head |
|
| 47 |
-
| Runs on | [this exllamav3 branch](https://github.com/benthecarman/exllamav3/tree/mimo-v2.6-flash) + [this TabbyAPI branch](https://github.com/benthecarman/tabbyAPI/tree/mimo-v2.6-flash) |
|
| 48 |
-
|
| 49 |
-
---
|
| 50 |
-
|
| 51 |
-
## What it is
|
| 52 |
|
| 53 |
-
|
| 54 |
-
|
| 55 |
-
dense, `hidden_size` 4096, `moe_intermediate_size` 2048. Attention is hybrid: 9 global
|
| 56 |
-
layers (64 query / 4 KV heads) and 39 sliding-window layers (window 128, 64 query / 8 KV
|
| 57 |
-
heads, with attention sinks), fused QKV, QK head_dim 192 against V head_dim 128, and an
|
| 58 |
-
`attention_value_scale` of 0.707.
|
| 59 |
|
| 60 |
-
|
| 61 |
-
with e8m0 byte scales, block 32 along K), attention and the dense layer-0 MLP as **FP8
|
| 62 |
-
e4m3** with 128x128 `weight_scale_inv` blocks, and embeddings / `lm_head` / norms in BF16.
|
| 63 |
-
All of that is dequantized during conversion and re-quantized as EXL3 trellis codes.
|
| 64 |
-
|
| 65 |
-
**This quant is text-only.** The source repo's vision and audio towers and its three MTP
|
| 66 |
-
heads are not converted and not present. `config.json` still carries the `vision_config`,
|
| 67 |
-
`audio_config` and `processor_config` blocks verbatim from upstream (so the file stays a
|
| 68 |
-
faithful description of the base model), but no corresponding tensors exist here.
|
| 69 |
-
|
| 70 |
-
### What is in this repo
|
| 71 |
|
| 72 |
```
|
| 73 |
-
model-
|
| 74 |
-
|
| 75 |
-
|
| 76 |
-
|
| 77 |
-
|
| 78 |
-
|
| 79 |
-
configuration_mimo_v2.py modeling_mimo_v2.py upstream remote code, unmodified (auto_map)
|
| 80 |
-
dflash/ the drafter, quantized to 4 bpw EXL3 (use this one)
|
| 81 |
-
model.safetensors 0.74 GB, includes the mask embedding
|
| 82 |
-
config.json repaired + switches pinned (see DFlash notes)
|
| 83 |
-
quantization_config.json
|
| 84 |
-
dflash-bf16/ the same drafter, unquantized
|
| 85 |
-
dflash_draft_model.safetensors 2.94 GB, byte-identical to the base repo's
|
| 86 |
-
config.json repaired + switches pinned
|
| 87 |
-
mask_embedding.safetensors the learned mask vector, converted from mask_embedding.pt
|
| 88 |
-
eval/
|
| 89 |
-
bench/RESULTS.md bench/METHOD.md the benchmark run, both sides, verbatim
|
| 90 |
-
bench/local-*.score.json local-*.meta.json this quant's per-task scores + run metadata
|
| 91 |
-
bench/ref-*.score.json ref-*.meta.json the unquantized FP8 reference, same harness
|
| 92 |
-
eval-quant.json eval-overflow-64rows.json dflash-bench-*.json
|
| 93 |
-
longctx.json the long-context sweep, every request (see Long context)
|
| 94 |
```
|
| 95 |
|
| 96 |
-
The shards, index, `quantization_config.json`, `config.json` and the tokenizer files are the
|
| 97 |
-
converter's output, unmodified. The files under `eval/` are from the previous build.
|
| 98 |
-
|
| 99 |
-
---
|
| 100 |
-
|
| 101 |
## Bitrate
|
| 102 |
|
| 103 |
-
`-b 2.25 -hq
|
| 104 |
-
they matter and the per-module result is what actually landed:
|
| 105 |
-
|
| 106 |
-
| module group | layers | bpw |
|
| 107 |
-
|---|---|---|
|
| 108 |
-
| routed experts (`mlp.experts.*.{gate,up,down}_proj`) | 24 layers (12–35) | **2.0** |
|
| 109 |
-
| routed experts | 23 layers (1–11, 36–47) | **2.5** |
|
| 110 |
-
| attention (`q/k/v/o_proj`) | all 48 | **4.0** |
|
| 111 |
-
| dense MLP, layer 0 (`gate/up/down_proj`) | 1 | **3.0** |
|
| 112 |
-
| `lm_head` | — | **6.0** |
|
| 113 |
-
| `embed_tokens`, all norms, router weights, `e_score_correction_bias`, attention sinks | — | **16-bit** (BF16, unquantized) |
|
| 114 |
-
|
| 115 |
-
Converter's summary line: `Final bitrate (excluding head): 2.27 (--hq enabled)`.
|
| 116 |
-
|
| 117 |
-
### The command
|
| 118 |
|
| 119 |
-
|
| 120 |
-
|
| 121 |
-
|
| 122 |
-
|
| 123 |
-
|
| 124 |
-
|
| 125 |
-
|
| 126 |
-
|
| 127 |
-
|
| 128 |
-
|
| 129 |
-
|
| 130 |
-
|
| 131 |
-
|
| 132 |
-
~9% of an MoE layer and LDLQ + trellis search is the other ~90%.
|
| 133 |
-
|
| 134 |
-
### Layer 47 and `interm_div`
|
| 135 |
-
|
| 136 |
-
On this checkpoint, layer 47's routed experts push `act(gate) * up` to about 84k on some
|
| 137 |
-
number tokens (mostly expert 208 channel 18 and expert 70 channel 1387), past the fp16 max of
|
| 138 |
-
65504. `gate` and `up` on their own stay under 2.3k, and no other layer gets above 3k. Xiaomi's
|
| 139 |
-
reference runs bf16 there and stays finite; exllamav3 runs these intermediates in fp16.
|
| 140 |
-
|
| 141 |
-
The previous build worked around it at conversion time: layer 47's experts at 6 bpw, and a
|
| 142 |
-
raised limit on non-finite calibration rows (31 of 250 were dropped). It still produced
|
| 143 |
-
non-finite logits at inference on 33 of the converter's 250 calibration rows.
|
| 144 |
-
|
| 145 |
-
This build uses `interm_div = 128` on layer 47, the same mechanism exllamav3 uses for Laguna.
|
| 146 |
-
`up_proj` is scaled by 1/128 before quantization and `routed_scaling_factor` puts the 128 back
|
| 147 |
-
in fp32, so the peak becomes about 660. No calibration rows went non-finite during conversion,
|
| 148 |
-
and the finished model gives 0 non-finite rows on those same 250 rows. Layer 47's experts are
|
| 149 |
-
back at 2.5 bpw, which accounts for the smaller file and most of the small perplexity change.
|
| 150 |
-
|
| 151 |
-
Because the 1/128 lives in the quantized weights, these weights need a matching exllamav3
|
| 152 |
-
branch. fp32 intermediates are not a substitute: the fused prefill kernel stores fp16
|
| 153 |
-
regardless, and without it the activation kernel clamps the product to 65504.
|
| 154 |
-
|
| 155 |
-
---
|
| 156 |
|
| 157 |
## Quality
|
| 158 |
|
| 159 |
-
|
| 160 |
-
|
| 161 |
-
wikitext-2 **test** split, non-overlapping rows, via `exllamav3/eval/ppl.py`:
|
| 162 |
-
|
| 163 |
-
| rows x length | perplexity | scored tokens |
|
| 164 |
-
|---|---|---|
|
| 165 |
-
| **64 x 2048** | **5.4003** | 131,008 |
|
| 166 |
-
|
| 167 |
-
Previous build, same method: 16 x 2048 5.2410, 64 x 2048 5.3615, 146 x 2048 (the entire test
|
| 168 |
-
split) 5.3174. Row count moves the number by ~±0.04, so quote it with the number.
|
| 169 |
-
**0 non-finite tokens** in every run.
|
| 170 |
-
|
| 171 |
-
### Benchmarks
|
| 172 |
-
|
| 173 |
-
**Measured on the previous 2.34 bpw build**, whose layer 47 differs from this one. Not re-run.
|
| 174 |
-
|
| 175 |
-
**Read this first.** These are absolute scores from one specific harness, so they are **not**
|
| 176 |
-
comparable to anyone else's published figures. They *are* comparable to each other: the
|
| 177 |
-
"reference FP8" column is the **unquantized model**, run through the **same script, the same
|
| 178 |
-
seeded items, the same prompts, the same greedy sampling and the same `max_tokens`** — only
|
| 179 |
-
`--base-url` and `--model` differed. Read the caveats below the table before the numbers.
|
| 180 |
-
|
| 181 |
-
| task | N | this quant (2.36 bpw EXL3) | reference FP8, same harness | delta | published by Xiaomi |
|
| 182 |
-
|---|---|---|---|---|---|
|
| 183 |
-
| HumanEval+ (pass@1, `plus`) | 164 (full set) | **89.6%** | **90.2%** | **−0.6 pp** | not published |
|
| 184 |
-
| MBPP+ (pass@1, `plus`) | 378 (full set) | **77.5%** | **77.2%** | **+0.3 pp** | not published |
|
| 185 |
-
| GSM8K (exact match) | 500 | **95.8%** (479/500) | **96.4%** (482/500) | **−0.6 pp** | not published |
|
| 186 |
-
| MMLU-Pro (acc) | 500, stratified over all 14 categories | **72.6%** (363/500) — a **floor**, true value 72.6–74.8% | **76.4%** (382/500), floor of 76.4–78.8% | **−3.8 pp** | not published |
|
| 187 |
-
| GPQA-Diamond (acc, 8K budget) | 64 | **46.9%** (30/64) | **56.2%** (36/64) | **−9.4 pp** | not published |
|
| 188 |
-
| GPQA-Diamond (acc, runs that finished inside the 8K budget) | 35 / 40 | **85.7%** | **90.0%** | **−4.3 pp** | not published |
|
| 189 |
-
|
| 190 |
-
Base (non-`plus`) pass@1: HumanEval **92.07** quant vs **92.68** reference (−0.6 pp); MBPP
|
| 191 |
-
**88.10** vs **90.21** (−2.1 pp).
|
| 192 |
-
|
| 193 |
-
**The short version: code and arithmetic survive 2.36 bpw almost intact (−0.6 / +0.3 / −0.6 pp),
|
| 194 |
-
broad factual recall does not (−3.8 pp on MMLU-Pro), and long-form reasoning degrades mostly by
|
| 195 |
-
running long rather than by being wrong** — GPQA's −9.4 pp is about half budget (the quant hit
|
| 196 |
-
the 8,192-token cap on 45.3% of items against the reference's 37.5%) and about half quality
|
| 197 |
-
(−4.3 pp among the runs that finished, which at n=35/40 is inside the noise). If you serve this
|
| 198 |
-
quant for reasoning work, give it a larger `max_tokens` than you would give the FP8 model.
|
| 199 |
-
|
| 200 |
-
Item-level, on exactly the same items, the disagreements on the three near-zero-delta tasks are
|
| 201 |
-
**two-sided and balanced** — HumanEval+ 4 items the quant wins vs 5 the reference wins, MBPP+
|
| 202 |
-
12 vs 11, GSM8K 6 vs 9 — i.e. noise around a shared answer rather than a systematic loss. On
|
| 203 |
-
MMLU-Pro (26 vs 45) and GPQA (1 vs 7) the asymmetry is real.
|
| 204 |
-
|
| 205 |
-
**Settings, identical for every row:** greedy (`temperature 0.0`, `top_p 1.0`), zero-shot,
|
| 206 |
-
seed 1234, batch 1; the quant column was served by TabbyAPI with the DFlash drafter on.
|
| 207 |
-
Thinking is **off**
|
| 208 |
-
(`chat_template_kwargs: {"enable_thinking": false}`) for HumanEval+, MBPP+, GSM8K and
|
| 209 |
-
MMLU-Pro, and **on** for GPQA-Diamond. `max_tokens`: 1536 / 1024 / 1024 / 1536 / 8192
|
| 210 |
-
respectively. Datasets: evalplus `HumanEvalPlus v0.1.10` and `MbppPlus v0.2.0` (scored with
|
| 211 |
-
`evalplus.evaluate`), `openai/gsm8k`, `TIGER-Lab/MMLU-Pro`, `Idavidrein/gpqa` `gpqa_diamond`.
|
| 212 |
-
The reference column used **every one of those settings unchanged**, and `--only-ids-from` the
|
| 213 |
-
quant's own output so both columns cover exactly the same item ids (GPQA: the same 64).
|
| 214 |
-
|
| 215 |
-
**Caveats that matter more than the numbers:**
|
| 216 |
-
|
| 217 |
-
1. **These are not comparable to Xiaomi's published numbers, because there are none.** The
|
| 218 |
-
MiMo-V2.6 technical report's only results table is entirely *agentic* — DeepSWE v1.1
|
| 219 |
-
67.9, Terminal Bench 2.1 87.6, OSWorld-Verified 80.8, CyberGym 95.1, and so on. Greps for
|
| 220 |
-
GPQA / AIME / MMLU / LiveCodeBench / HumanEval / GSM8K / IFEval / SimpleQA / HLE /
|
| 221 |
-
MATH-500 / BBH / DROP over all 44 pages return **zero hits**. Nothing in the table above
|
| 222 |
-
has a published baseline. If standard-benchmark numbers for the FP8 model are published
|
| 223 |
-
later, the "published by Xiaomi" column is where they go — **it is deliberately empty
|
| 224 |
-
rather than filled with guesses.**
|
| 225 |
-
2. **What the reference column actually is.** The unquantized model as
|
| 226 |
-
**`xiaomi/mimo-v2.6-flash` on OpenRouter**, run **2026-09-22**. OpenRouter lists exactly
|
| 227 |
-
**one provider** for it — **Xiaomi**, endpoint quantization **fp8** — so there was no routing
|
| 228 |
-
variance and no third-party system prompt; every stored response carries
|
| 229 |
-
`"provider": "Xiaomi"`. **The thinking flags were honoured exactly**, which was checked on all
|
| 230 |
-
1,606 responses rather than assumed: the 1,542 thinking-off requests
|
| 231 |
-
(`reasoning: {"enabled": false}` plus `chat_template_kwargs: {"enable_thinking": false}`) all
|
| 232 |
-
returned `completion_tokens_details.reasoning_tokens == 0`, an empty `reasoning` field and no
|
| 233 |
-
`<think>` tag in `content`; the 64 thinking-on GPQA requests all returned reasoning text, at a
|
| 234 |
-
volume closely matching the local side (mean 12,578 chars reference vs 12,668 quant).
|
| 235 |
-
**No reference-side failures:** 0 API errors, 0 empty responses on the thinking-off tasks, no
|
| 236 |
-
`finish_reason` other than `stop`/`length`. One honest asymmetry: OpenRouter does not list
|
| 237 |
-
`seed` among this endpoint's supported parameters, so the harness's seed is probably ignored
|
| 238 |
-
upstream — at `temperature 0.0 / top_p 1.0` that should not matter, but the reference side is
|
| 239 |
-
"greedy as the provider implements it", not "greedy with seed 1234". **Total spend $0.2206**
|
| 240 |
-
for all 1,606 requests (654,552 generated + 266,199 prompt tokens), confirmed against
|
| 241 |
-
OpenRouter's `/api/v1/auth/key` usage field.
|
| 242 |
-
3. **Greedy is a deliberate deviation** from the base model card's recommended
|
| 243 |
-
`temperature 1.0 / top_p 0.95`. One sample at temperature 1.0 has enough variance to
|
| 244 |
-
swamp a quantization delta. The absolute scores here are therefore *not* the model's
|
| 245 |
-
best-effort scores.
|
| 246 |
-
4. **GPQA's 46.9% is a budget artefact.** 29 of 64 runs (45.3%) hit the 8,192-token thinking
|
| 247 |
-
cap, and every one scored zero because the model was still inside `<think>` and never
|
| 248 |
-
emitted an answer line. Mean generation 4,418 tokens, median 3,018 — bimodal: settle in
|
| 249 |
-
~3K or run away past 8K. The honest bracket is **46.9%–85.7%**, depending on how much
|
| 250 |
-
budget you pay for. A real GPQA number needs a 32K budget, which at ~43 tok/s batch 1 is
|
| 251 |
-
~21 h for 100 items. **The reference model hits the same wall** — 24 of the same 64 items
|
| 252 |
-
(37.5%), also scoring zero on every one, bracket 56.2%–90.0%. So the cap is a property of the
|
| 253 |
-
task-plus-budget, not of the quantization; the quant simply runs away past 8K more often
|
| 254 |
-
(45.3% vs 37.5%), and that gap is most of the headline −9.4 pp.
|
| 255 |
-
5. **MMLU-Pro's 72.6% is likewise a floor.** All 10 unparseable answers are pure truncation
|
| 256 |
-
at the 1536-token cap (11 items = 2.2% hit it, none scored). The reference has the same
|
| 257 |
-
shape (14 items = 2.8% capped, floor 76.4% / ceiling 78.8%), so the caps do **not** explain
|
| 258 |
-
the gap: the two floor-to-ceiling bands, 72.6–74.8% and 76.4–78.8%, do not overlap.
|
| 259 |
-
**−3.8 pp on MMLU-Pro is the one delta that survives every caveat.**
|
| 260 |
-
|
| 261 |
-
**Stability, which the deltas do not show:** across **1,606 local generations / 574,407
|
| 262 |
-
generated tokens** (and 1,606 reference generations / 654,552 tokens) there were **0 API errors
|
| 263 |
-
on either side, 0 non-finite artefacts on either side, 0 empty code solutions on either side**,
|
| 264 |
-
and exactly **one** pathological response in the whole local suite (a GSM8K item that returned an
|
| 265 |
-
immediate EOS at zero tokens; the reference answered that item normally). Nothing traced to the
|
| 266 |
-
layer-47 fp16 overflow. A 2.36 bpw quant that was actually broken would not look like this — and
|
| 267 |
-
now there is a same-harness reference column saying how much it gave up.
|
| 268 |
-
|
| 269 |
-
Full method, prompts, seeds, per-task throughput and every failure signal:
|
| 270 |
-
[`eval/bench/METHOD.md`](eval/bench/METHOD.md) and
|
| 271 |
-
[`eval/bench/RESULTS.md`](eval/bench/RESULTS.md).
|
| 272 |
-
|
| 273 |
-
### Reproducing both sides
|
| 274 |
-
|
| 275 |
-
The harness lives in the [exllamav3 branch's](https://github.com/benthecarman/exllamav3/tree/mimo-v2.6-flash)
|
| 276 |
-
project notes; both sides differ only in `--base-url`, `--api-key-env` and `--model`:
|
| 277 |
|
| 278 |
-
|
| 279 |
-
|
| 280 |
-
|
| 281 |
-
--base-url http://127.0.0.1:30001/v1 --model mimo-2.25bpw-hq \
|
| 282 |
-
--concurrency 2 --out raw/local-gsm8k.jsonl
|
| 283 |
-
python score_bench.py raw/local-gsm8k.jsonl
|
| 284 |
-
|
| 285 |
-
# reference side: the unquantized model, same items, same prompts, same caps
|
| 286 |
-
export OPENROUTER_API_KEY=... # model id: xiaomi/mimo-v2.6-flash
|
| 287 |
-
./run_reference.sh # measured $0.22 at $0.14/M prompt, $0.28/M completion
|
| 288 |
-
```
|
| 289 |
-
|
| 290 |
-
`--only-ids-from <the local jsonl>` restricts the reference to exactly the item ids the
|
| 291 |
-
local side generated, so the two columns are always over the same items.
|
| 292 |
-
|
| 293 |
-
---
|
| 294 |
|
| 295 |
-
|
| 296 |
-
|
| 297 |
-
GB10, 121.6 GiB unified memory, batch 1, greedy, FP16 KV. Measured on the previous 2.34 bpw
|
| 298 |
-
build (86.1 GiB); the new one is 2.6 GiB smaller.
|
| 299 |
-
|
| 300 |
-
### Decode and prefill, no drafter
|
| 301 |
-
|
| 302 |
-
| prompt | input tokens | TTFT | prefill | decode |
|
| 303 |
|---|---|---|---|---|
|
| 304 |
-
|
|
| 305 |
-
|
|
| 306 |
-
|
| 307 |
-
|
| 308 |
-
|
| 309 |
|
| 310 |
-
|
|
|
|
|
|
|
| 311 |
|
| 312 |
-
##
|
| 313 |
|
| 314 |
-
|
| 315 |
-
Same harness for both: greedy, 512 generated tokens, dynamic draft window, median of 3 runs:
|
| 316 |
|
| 317 |
-
| prompt | no
|
| 318 |
|---|---|---|---|
|
| 319 |
-
| coding
|
| 320 |
-
| prose
|
| 321 |
-
| reasoning | 31.4 |
|
| 322 |
-
|
| 323 |
-
Every draft step reads the drafter's weights once, so a quarter of the bytes makes drafting
|
| 324 |
-
much cheaper. Drafted tokens are always verified by the target model, so the drafter's
|
| 325 |
-
precision changes how many drafts are accepted, never what is generated.
|
| 326 |
-
|
| 327 |
-
The rest of this section was measured earlier with the BF16 drafter, on a different harness.
|
| 328 |
-
|
| 329 |
-
512-token generations, one configuration per process. Static 7-token draft window:
|
| 330 |
-
|
| 331 |
-
| prompt | no draft | DFlash | speedup | acceptance | accepted/step |
|
| 332 |
-
|---|---|---|---|---|---|
|
| 333 |
-
| coding | 28.16 | **44.41** | **1.58x** | **74.06%** | 6.18 |
|
| 334 |
-
| reasoning | 31.29 | **53.18** | **1.70x** | **60.35%** | 5.22 |
|
| 335 |
-
| prose | 31.11 | 20.72 | **0.67x** | 12.43% | 1.87 |
|
| 336 |
-
|
| 337 |
-
Prose is a real **slowdown** with a static window: 1,673 draft tokens spent to win 208, and
|
| 338 |
-
every rejected block still costs a target forward pass.
|
| 339 |
-
|
| 340 |
-
`dynamic_draft` shrinks the window from observed acceptance and is the **recommended serving
|
| 341 |
-
default**:
|
| 342 |
-
|
| 343 |
-
| prompt | no draft | static | **dynamic** | dynamic speedup | acceptance static → dynamic |
|
| 344 |
-
|---|---|---|---|---|---|
|
| 345 |
-
| coding | 28.16 | 44.41 | 38.90 | 1.38x | 74.06% → 81.07% |
|
| 346 |
-
| reasoning | 31.29 | 53.18 | **55.66** | **1.78x** | 60.35% → 67.54% |
|
| 347 |
-
| prose | 31.11 | 20.72 | **27.19** | **0.87x** | 12.43% → 50.41% |
|
| 348 |
-
|
| 349 |
-
Dynamic trades ~12% of the coding peak for a far better worst case (0.67x → 0.87x) and a
|
| 350 |
-
better mean (1.35x vs 1.31x). For peak coding throughput, turn it off; for a pure prose
|
| 351 |
-
workload, drop the drafter entirely.
|
| 352 |
-
|
| 353 |
-
Across the whole 3.77 h benchmark suite above, the server sustained **~42 generated tok/s**
|
| 354 |
-
batch 1 with dynamic drafting — i.e. ~1.35x over the 31.5 tok/s no-draft figure, on a real
|
| 355 |
-
mixed workload.
|
| 356 |
-
|
| 357 |
-
**n-gram drafting never helps** here (0.87–0.98x at 1.8–24.7% acceptance): this model does
|
| 358 |
-
not repeat itself enough.
|
| 359 |
|
| 360 |
-
|
| 361 |
|
| 362 |
-
|
| 363 |
-
|
| 364 |
-
| MemTotal (GB10) | 121.63 |
|
| 365 |
-
| **peak in use, 32K context** | **95.06** |
|
| 366 |
-
| minimum MemAvailable seen | 26.57 |
|
| 367 |
-
| torch peak allocated | 86.98 |
|
| 368 |
-
| steady state while serving at 64K + drafter | MemAvailable **21.2** |
|
| 369 |
-
|
| 370 |
-
KV cost, derived from `config.json`: **27.00 KiB/token** of paged cache (9 global-attention
|
| 371 |
-
layers) **plus a fixed 175.5 MiB ring per slot** (768-token ring x 39 sliding-window layers).
|
| 372 |
-
With 86.20 GiB of weights and a 10 GiB reserve, the KV budget is 25.43 GiB:
|
| 373 |
-
|
| 374 |
-
| | max context, 1 slot | 4 slots |
|
| 375 |
-
|---|---|---|
|
| 376 |
-
| ring, FP16 KV | **980,736 tok** | 240,128 |
|
| 377 |
-
| ring, Q8 paged KV (global layers only) | 1,048,576 (model max) | 480,256 |
|
| 378 |
-
| full (non-ring) KV, for comparison | 102,144 | 25,344 |
|
| 379 |
-
|
| 380 |
-
64K context at 1 slot costs 1.86 GiB; 32K costs 1.02 GiB. Quantized paged KV (k=v=8) loads
|
| 381 |
-
and generates coherently on top of the previous 2.36 bpw weights, but its quality is untested — the
|
| 382 |
-
default stays FP16. The 39 sliding-window rings are always FP16.
|
| 383 |
-
|
| 384 |
-
Those are the arithmetic limits. What was actually served, loaded and measured — up to
|
| 385 |
-
**350,091 tokens** — is in [Long context](#long-context) below; the cache is **preallocated in
|
| 386 |
-
full at load**, so `max_seq_len` is a memory decision, not a ceiling you pay for on demand.
|
| 387 |
-
|
| 388 |
-
#### A unified-memory hazard worth knowing about
|
| 389 |
-
|
| 390 |
-
On GB10 there is **no separate VRAM**: host and GPU share one pool. exllamav3's shard loader
|
| 391 |
-
uses buffered `pread`, not `mmap`, so loading 86.1 GiB leaves **86.1 GiB of clean page cache**
|
| 392 |
-
competing with the 86.1 GiB of weights in the same 121.6 GiB pool — and
|
| 393 |
-
`torch.cuda.mem_get_info()` reports MemFree, not MemAvailable, so nothing in-process can see
|
| 394 |
-
it. Unified memory does not produce an OOM kill; it starves the machine, and recovery is a
|
| 395 |
-
physical power cycle. Two mitigations are in the exllamav3 branch (upstream PR #398) and are
|
| 396 |
-
why the numbers above are flat rather than degrading:
|
| 397 |
-
|
| 398 |
-
* `Model.load_gen()` drops the shard page cache at the end of every load
|
| 399 |
-
(`EXL3_KEEP_PAGE_CACHE=1` disables); measured **0.00 GiB of 86.08 GiB resident** while serving.
|
| 400 |
-
* `EXL3_LOAD_DEVICE=cuda:0` takes a single-device load path that does no MemFree-based budget
|
| 401 |
-
arithmetic. Without it, autosplit can refuse a model that actually fits.
|
| 402 |
-
|
| 403 |
-
---
|
| 404 |
-
|
| 405 |
-
## Long context
|
| 406 |
-
|
| 407 |
-
Measured on the DGX Spark, batch 1, greedy, thinking off, FP16 KV, needles and haystacks built
|
| 408 |
-
from wikitext-2 paragraphs. **Every retrieval test passed.**
|
| 409 |
-
|
| 410 |
-
### Retrieval
|
| 411 |
-
|
| 412 |
-
A six-digit passcode is hidden in a wikitext haystack at 10 / 50 / 90% depth and asked for at
|
| 413 |
-
the end. Exact match on the digits, fresh city and fresh code per request.
|
| 414 |
|
| 415 |
-
| prompt tokens |
|
| 416 |
|---|---|---|---|
|
| 417 |
-
|
|
| 418 |
-
|
|
| 419 |
-
|
|
| 420 |
-
| ~131.1K | OK | OK | OK |
|
| 421 |
-
| ~200.1K | OK | OK | OK |
|
| 422 |
-
| ~250.0K | OK | OK | OK |
|
| 423 |
-
| **350,091** | **OK** | — | — |
|
| 424 |
-
|
| 425 |
-
**18/18** at `max_seq_len 262144`, plus **350,091 tokens** answered correctly in a 393,216
|
| 426 |
-
window — the largest prompt this quant has been given. Nine more cells at `max_seq_len 131072`
|
| 427 |
-
with the drafter on also passed, and **five more at `max_seq_len 393216` with the drafter
|
| 428 |
-
attached** (131K at 10/50/90% depth, 250K at 50% and 90%, all exact), so retrieval is
|
| 429 |
-
unaffected by speculative decoding or by the window size.
|
| 430 |
-
|
| 431 |
-
That matters more than it looks: **39 of 48 layers are sliding-window with a 128-token
|
| 432 |
-
window** and run on a fixed 768-token ring, so a needle 25,000 tokens into a 250K prompt has
|
| 433 |
-
to survive ~195 ring rebases and reach the question through the 9 global-attention layers
|
| 434 |
-
alone. The ring had previously only been validated to 2,048 tokens.
|
| 435 |
-
|
| 436 |
-
**Five needles at once, 128K prompt** — all five returned, in order of appearance:
|
| 437 |
-
|
| 438 |
-
```
|
| 439 |
-
Montevideo: 604403
|
| 440 |
-
Bratislava: 991476
|
| 441 |
-
Kathmandu: 472495
|
| 442 |
-
Ulaanbaatar: 379397
|
| 443 |
-
Ljubljana: 361254
|
| 444 |
-
```
|
| 445 |
-
|
| 446 |
-
**A real long document, ~100K tokens** — complete wikitext articles concatenated with *Plain
|
| 447 |
-
maskray* buried in the middle, three questions answerable only from that article:
|
| 448 |
-
|
| 449 |
-
```
|
| 450 |
-
(a) 2008
|
| 451 |
-
(b) ~54 Ma
|
| 452 |
-
(c) 12 and 62 m
|
| 453 |
-
```
|
| 454 |
|
| 455 |
-
|
| 456 |
-
|
| 457 |
-
the source's own "~" hedge.
|
| 458 |
|
| 459 |
-
|
| 460 |
-
|
| 461 |
-
No drafter, 128 generated tokens, `max_seq_len 262144` (last row 393,216):
|
| 462 |
-
|
| 463 |
-
| prompt | TTFT | prefill | decode | min MemAvailable |
|
| 464 |
-
|---|---|---|---|---|
|
| 465 |
-
| ~8.2K | 9.1–10.6 s | 768–905 tok/s | 31.0 tok/s | 19.85 GiB |
|
| 466 |
-
| ~32.8K | 41.2–42.0 s | 780–799 tok/s | 28.4 tok/s | 19.56 GiB |
|
| 467 |
-
| ~65.5K | 100.1–100.5 s | 652–654 tok/s | 25.3 tok/s | 19.10 GiB |
|
| 468 |
-
| ~131.1K | 270.0–271.5 s | 483–486 tok/s | 21.7 tok/s | 18.98 GiB |
|
| 469 |
-
| ~200.1K | 524.8–525.9 s | 380–381 tok/s | 18.3 tok/s | 18.85 GiB |
|
| 470 |
-
| ~250.0K | 757.6–758.8 s | 329–330 tok/s | 16.7 tok/s | 17.77 GiB |
|
| 471 |
-
| **350,091** | **1,346.8 s** | **260 tok/s** | **13.8 tok/s** | **15.91 GiB** |
|
| 472 |
-
|
| 473 |
-
**Decode falls only 2.3x from 8K to 250K** — the sliding-window ring again: only 9 of 48
|
| 474 |
-
layers grow their KV with context. Prefill falls 2.7x and TTFT is quadratic, as full
|
| 475 |
-
attention on those 9 layers requires. The curve fits
|
| 476 |
-
|
| 477 |
-
```
|
| 478 |
-
TTFT(seconds) ≈ 988·L + 8172·L² (L = prompt tokens in millions)
|
| 479 |
-
```
|
| 480 |
-
|
| 481 |
-
to better than 1.5% from 64K to 350K — the 350K point was predicted at 1,347 s from a fit to
|
| 482 |
-
the 128K and 250K points alone and came in at 1,346.8 s. Extrapolated: **400K ≈ 28 min,
|
| 483 |
-
500K ≈ 42 min of prefill.** Long context on one GB10 is TTFT-bound, not memory-bound.
|
| 484 |
-
|
| 485 |
-
With the DFlash drafter at `max_seq_len 131072`, on the same short-answer retrieval task
|
| 486 |
-
(real generations, ~380 tokens, unforced): 36.3 / 29.8 / 27.0 tok/s at 2K / 65K / 130K, i.e.
|
| 487 |
-
1.17–1.24x, live acceptance 38–85%. Those answers are three tokens long and factual, which is
|
| 488 |
-
the worst case for a drafter.
|
| 489 |
-
|
| 490 |
-
**On a real generation task the drafter is worth far more, and worth *more* the longer the
|
| 491 |
-
context.** A repository-sized context (source files) and a task asking for a ~400-token
|
| 492 |
-
implementation, `max_seq_len 262144`, batch 1, greedy, same prompts against a freshly loaded
|
| 493 |
-
server:
|
| 494 |
-
|
| 495 |
-
| prompt tokens | no drafter | **DFlash drafter** | **speedup** | acceptance |
|
| 496 |
-
|---|---|---|---|---|
|
| 497 |
-
| 64,614 | 25.08 tok/s | **43.77 tok/s** | **1.75x** | 64% |
|
| 498 |
-
| 130,118 | 21.08 tok/s | **41.43 tok/s** | **1.97x** | 61% |
|
| 499 |
-
| 248,993 | 16.14 tok/s | **33.63 tok/s** | **2.08x** | 58% |
|
| 500 |
-
|
| 501 |
-
The speedup **grows** with context (1.75x → 2.08x) even though acceptance **falls** (64% →
|
| 502 |
-
58%). Verifying a block of 8 drafted tokens is a single target forward, and at 250K that
|
| 503 |
-
forward is dominated by reading the 9 global-attention layers' K/V — a cost paid once for the
|
| 504 |
-
whole block instead of once per token. So at long context a lower acceptance rate still buys a
|
| 505 |
-
larger speedup. **248,993 tokens of context, decoding at 33.6 tok/s.**
|
| 506 |
-
|
| 507 |
-
### How much context fits
|
| 508 |
-
|
| 509 |
-
The paged cache is **preallocated in full at load** — `cache_size` is paid up front whether or
|
| 510 |
-
not anyone sends a long prompt. Cost per slot:
|
| 511 |
-
|
| 512 |
-
```
|
| 513 |
-
paged KV 27.00 KiB/token (the 9 global-attention layers only)
|
| 514 |
-
SWA ring 175.5 MiB per slot (768-token ring x 39 layers, always FP16)
|
| 515 |
-
draft KV 35.0 MiB per slot (5 DFlash layers on a 1792-token window ring)
|
| 516 |
-
drafter weights 2.81 GiB (BF16; about 0.7 GiB with the 4 bpw drafter)
|
| 517 |
-
```
|
| 518 |
-
|
| 519 |
-
so **a slot costs 27.0 KiB/token with or without the drafter**, and the drafter's whole
|
| 520 |
-
footprint is a flat 2.85 GiB at any context length (BF16 drafter; the 4 bpw one saves about
|
| 521 |
-
2.1 GiB more, so the MemAvailable figures in this section are conservative). That is new: the drafter used to size its
|
| 522 |
-
K/V for the whole context at **20.00 KiB/token** — a 74% surcharge, 7.50 GiB at `max_seq_len`
|
| 523 |
-
393216 — although all five of its layers are sliding-window with a 1024-token window and none
|
| 524 |
-
of them ever looks further back. They now run on a fixed per-slot ring of window + block + two
|
| 525 |
-
pages, so draft K/V is **35.0 MiB per slot at any context**: 2,560 MiB → 35.0 MiB at 131072,
|
| 526 |
-
7,680 MiB → 35.0 MiB at 393216. Generations are token-identical either way.
|
| 527 |
-
|
| 528 |
-
On this box the whole budget reduces to one line that held to within 0.1 GiB across every
|
| 529 |
-
configuration:
|
| 530 |
-
|
| 531 |
-
```
|
| 532 |
-
MemAvailable after load ≈ 29.2 GiB − (paged KV + ring + drafter weights + draft KV)
|
| 533 |
-
```
|
| 534 |
-
|
| 535 |
-
(121.63 GiB total − 86.15 GiB of weights − ~6 GiB of runtime.)
|
| 536 |
-
|
| 537 |
-
| `max_seq_len` | drafter | load | MemAvailable after load | min under load |
|
| 538 |
-
|---|---|---|---|---|
|
| 539 |
-
| **393,216** | **DFlash** | **23.3 s** | **16.09 GiB** | see below |
|
| 540 |
-
| 262,144 | DFlash | 23.2 s | **19.57 GiB** | **16.70 GiB** @ 249K |
|
| 541 |
-
| 262,144 | off | 21.3 s | 22.27 GiB | **19.59 GiB** @ 249K |
|
| 542 |
-
| 393,216 | off | 21.3 s | 18.65 GiB | **15.91 GiB** @ 350K |
|
| 543 |
-
| 131,072 | DFlash | 22.6 s | 20.17 GiB | **17.07 GiB** @ 130K |
|
| 544 |
-
| 65,536 | DFlash | 22.7 s | 24.48 GiB | 22.21 GiB @ 65K |
|
| 545 |
-
|
| 546 |
-
**Recommended: `max_seq_len 262144` with the drafter** — 19.57 GiB free after load and
|
| 547 |
-
**16.70 GiB** through a 249K-token request. With the drafter's old full-length K/V cache that
|
| 548 |
-
configuration had ~14.6 GiB after load and ~10 GiB under a full-length request, i.e. it did not
|
| 549 |
-
fit at all; the window ring is what makes it fit.
|
| 550 |
-
|
| 551 |
-
**`max_seq_len 393216` with the drafter also works** — 16.09 GiB after load, 13.45 GiB through
|
| 552 |
-
a 348,970-token generation, retrieval 5/5 — but it is the edge of this box rather than a
|
| 553 |
-
set-and-forget setting: over an hour of heavy serving MemAvailable drifted 16.09 → 13.0 GiB
|
| 554 |
-
(TabbyAPI's host RSS grows several GiB under sustained load) and the longest requests bottomed
|
| 555 |
-
at 12.66 GiB. Use it deliberately when a prompt needs it.
|
| 556 |
-
|
| 557 |
-
* **524,288 does not fit at FP16.** Its 13.67 GiB of paged KV leaves ~15.3 GiB after load, and
|
| 558 |
-
the measured 2.7–4.5 GiB transient of a full-length request would land it at ~11–12.5 GiB —
|
| 559 |
-
at or through a 12 GiB safety floor. On unified memory that is not an OOM kill, it is a
|
| 560 |
-
frozen machine. A 500K prefill would also take ~42 minutes.
|
| 561 |
-
* The drafter no longer costs anything per token, so **there is no longer a window-vs-drafting
|
| 562 |
-
trade-off** on this box. Turn it on.
|
| 563 |
-
|
| 564 |
-
So on a 121 GiB box you get **both**: a 262,144-token window (393,216 at the edge) *and* the
|
| 565 |
-
drafter, at 1.75–2.08x the decode rate on real generation. This repo's reference server runs
|
| 566 |
-
262,144 with the drafter.
|
| 567 |
-
|
| 568 |
-
### Recommended serving settings
|
| 569 |
-
|
| 570 |
-
```yaml
|
| 571 |
-
model:
|
| 572 |
-
max_seq_len: 262144 # with the drafter; the draft cache no longer scales with it.
|
| 573 |
-
cache_size: 262144 # 393216 also fits with the drafter, with less headroom
|
| 574 |
-
cache_mode: FP16
|
| 575 |
-
chunk_size: 2048
|
| 576 |
-
max_batch_size: 1 # every extra slot repeats the SWA ring, the paged span and a 35 MiB draft ring
|
| 577 |
-
draft_model:
|
| 578 |
-
draft_mode: model
|
| 579 |
-
draft_model_name: mimo-dflash-draft
|
| 580 |
-
draft_cache_mode: FP16
|
| 581 |
-
dynamic_draft: true
|
| 582 |
-
```
|
| 583 |
-
|
| 584 |
-
### Caveats
|
| 585 |
-
|
| 586 |
-
* **`max_seq_len` bounds the prompt, not prompt + generation.** A 131,092-token prompt against
|
| 587 |
-
`max_seq_len 131072` returns `400 Bad Request` / `Prompt length 131092 exceeds the …` before
|
| 588 |
-
a token is generated. Budget `max_seq_len ≥ prompt + max_tokens`, and remember the chat
|
| 589 |
-
template adds ~20 tokens.
|
| 590 |
-
* **Every number here is `max_batch_size: 1`.** A second slot adds another 175.5 MiB ring, its
|
| 591 |
-
own paged span and its own draft cache; concurrency at 128K is not free.
|
| 592 |
-
* Prefill is chunked at 2,048 tokens. Prefill rates for a prompt sharing a long prefix with the
|
| 593 |
-
previous request are inflated by page reuse (755 vs 473 tok/s at 130K here), so cold numbers
|
| 594 |
-
are the ones quoted above.
|
| 595 |
-
* Q8 paged KV halves the 27 KiB/token term and would make ~786K arithmetically fit, but
|
| 596 |
-
quantized KV quality on top of ~2.3 bpw weights is untested and the prefill time would be
|
| 597 |
-
hours. The 39 sliding-window rings stay FP16 regardless.
|
| 598 |
-
* Raw records: [`eval/longctx.json`](eval/longctx.json).
|
| 599 |
-
|
| 600 |
-
---
|
| 601 |
|
| 602 |
## How to run
|
| 603 |
|
| 604 |
-
|
|
|
|
|
|
|
| 605 |
|
| 606 |
-
|
| 607 |
-
|
| 608 |
-
[#396](https://github.com/turboderp-org/exllamav3/pull/396),
|
| 609 |
-
[#398](https://github.com/turboderp-org/exllamav3/pull/398),
|
| 610 |
-
[#399](https://github.com/turboderp-org/exllamav3/pull/399)) and all three sit on one branch:
|
| 611 |
-
|
| 612 |
-
* **exllamav3** — https://github.com/benthecarman/exllamav3/tree/mimo-v2.6-flash
|
| 613 |
-
* **TabbyAPI** — https://github.com/benthecarman/tabbyAPI/tree/mimo-v2.6-flash
|
| 614 |
-
(one patch: honour `EXL3_LOAD_DEVICE` when building the split)
|
| 615 |
-
|
| 616 |
-
exllamav3 1.5.1 ships no server of its own, so serving is TabbyAPI driven against an editable
|
| 617 |
-
checkout of that exllamav3 branch. On aarch64 you also need `uvloop` installed manually —
|
| 618 |
-
TabbyAPI marks it x86_64-only in `pyproject.toml` but imports it unconditionally on Linux.
|
| 619 |
|
| 620 |
```sh
|
| 621 |
git clone -b mimo-v2.6-flash https://github.com/benthecarman/exllamav3
|
| 622 |
git clone -b mimo-v2.6-flash https://github.com/benthecarman/tabbyAPI
|
| 623 |
-
pip install -e ./exllamav3
|
| 624 |
-
|
| 625 |
-
# TabbyAPI: base dependencies ONLY. Never `pip install ./tabbyAPI[cu12]` or `[cu13]` --
|
| 626 |
-
# those extras pin torch 2.9/2.11 and a prebuilt exllamav3 wheel, which would clobber
|
| 627 |
-
# both your torch and the editable checkout above.
|
| 628 |
-
pip install "fastapi-slim>=0.115" "pydantic>=2.11,<3" ruamel.yaml rich "uvicorn>=0.28.1" \
|
| 629 |
-
"jinja2>=3.0.0" loguru "sse-starlette>=2.2.0" packaging "tokenizers>=0.21.0" numpy \
|
| 630 |
-
aiofiles aiohttp async_lru huggingface_hub psutil "httptools>=0.5.0" pillow requests setuptools
|
| 631 |
-
pip install uvloop # aarch64 only: pyproject marks it x86_64, main.py imports it anyway
|
| 632 |
-
|
| 633 |
hf download benthecarman/MiMo-V2.6-Flash-RL-exl3 --local-dir mimo-exl3
|
| 634 |
```
|
| 635 |
|
| 636 |
-
|
|
|
|
| 637 |
|
| 638 |
-
TabbyAPI
|
| 639 |
-
directory
|
| 640 |
|
| 641 |
```
|
| 642 |
models/
|
| 643 |
-
mimo-2.25bpw-hq/
|
| 644 |
-
mimo-dflash-draft/
|
| 645 |
```
|
| 646 |
|
| 647 |
`config.yml`:
|
|
@@ -650,18 +144,14 @@ models/
|
|
| 650 |
model:
|
| 651 |
model_dir: models
|
| 652 |
model_name: mimo-2.25bpw-hq
|
| 653 |
-
max_seq_len: 262144
|
| 654 |
cache_size: 262144
|
| 655 |
-
cache_mode: FP16
|
| 656 |
chunk_size: 2048
|
| 657 |
-
max_batch_size: 1
|
| 658 |
-
gpu_split_auto: true
|
| 659 |
-
vision: false
|
| 660 |
reasoning: true
|
| 661 |
reasoning_start_token: "<think>"
|
| 662 |
reasoning_end_token: "</think>"
|
| 663 |
-
start_in_reasoning: auto
|
| 664 |
-
# prompt_template unset -> TabbyAPI picks up chat_template.jinja from the model dir
|
| 665 |
|
| 666 |
draft_model:
|
| 667 |
draft_mode: model
|
|
@@ -671,119 +161,17 @@ draft_model:
|
|
| 671 |
dynamic_draft: true
|
| 672 |
```
|
| 673 |
|
| 674 |
-
|
| 675 |
-
from the template and tokenizer. Tool calling has been exercised against this quant
|
| 676 |
-
(`finish_reason: tool_calls`, `get_weather({"city": "Paris, France"})`, `parsed 1 tool call
|
| 677 |
-
(qwen3_coder)`). Thinking is switched off per request with
|
| 678 |
-
`"chat_template_kwargs": {"enable_thinking": false}`.
|
| 679 |
-
|
| 680 |
-
A ready-made `serve.sh` + `tabby-config.yml` live in
|
| 681 |
-
[`examples/mimo_v2_6/`](https://github.com/benthecarman/exllamav3/tree/mimo-v2.6-flash/examples/mimo_v2_6)
|
| 682 |
-
on the exllamav3 branch.
|
| 683 |
-
|
| 684 |
-
### Environment knobs
|
| 685 |
-
|
| 686 |
-
| variable | default | what it does |
|
| 687 |
-
|---|---|---|
|
| 688 |
-
| `EXL3_LOAD_DEVICE` | unset | set to `cuda:0` on GB10 / any unified-memory box: single-device load, no autosplit, no MemFree arithmetic. **Recommended here.** |
|
| 689 |
-
| `EXL3_MIMO_FP32_MLP_LAYERS` | `none` | debugging only: layers to run with fp32 MoE intermediates. Not needed with this build. |
|
| 690 |
-
| `EXL3_KEEP_PAGE_CACHE` | unset | keep shard page cache after load (don't, on unified memory) |
|
| 691 |
-
|
| 692 |
-
You should see this at load:
|
| 693 |
-
|
| 694 |
-
```
|
| 695 |
-
-- EXL3_LOAD_DEVICE=cuda:0 (single-device load)
|
| 696 |
-
-- Released 83.5 GiB of shard page cache
|
| 697 |
-
```
|
| 698 |
-
|
| 699 |
-
### The DFlash drafter
|
| 700 |
-
|
| 701 |
-
`dflash-bf16/` is Xiaomi's shipped 5-layer DFlash drafter, **repaired**. The tensor file is
|
| 702 |
-
byte-identical to the base repo's; the two small files next to it are not, and the
|
| 703 |
-
differences are load-bearing:
|
| 704 |
-
|
| 705 |
-
* **`config.json` is valid JSON here.** The upstream one ends `"use_cache": true,` with a
|
| 706 |
-
trailing comma and cannot be parsed at all.
|
| 707 |
-
* **`tap_shift: 0`**, not exllamav3's legacy `+1`. MiMo's `extract_context_feature` uses
|
| 708 |
-
`hidden_states[i+1]` = the *output* of layer i = exllamav3 export index i. SGLang's
|
| 709 |
-
`dflash_utils.py::build_target_layer_ids` documents the same convention. Taps are
|
| 710 |
-
`[0, 11, 23, 35, 47]`, all of them global-attention layers, so the sliding-window ring is
|
| 711 |
-
not involved.
|
| 712 |
-
* **`attention_sink_bias: true`**, **`attention_value_scale: 0.612`**,
|
| 713 |
-
**`bidirectional_block: true`** (from the checkpoint's `is_causal: false`), window
|
| 714 |
-
`(1024, 7)`, `partial_rotary_factor: 0.5` → 64 of 128 rotary dims.
|
| 715 |
-
* **`mask_embedding.safetensors`** is the learned `[4096]` vector from the upstream
|
| 716 |
-
`mask_embedding.pt` pickle, rewritten so the loader can find it. In the quantized
|
| 717 |
-
`dflash/` it is stored inside `model.safetensors`.
|
| 718 |
-
|
| 719 |
-
**Why the shipped `dflash/dflash.py` is stale and is not included here.** It never reads the
|
| 720 |
-
sink biases, the value scale, the sliding window, `partial_rotary_factor`, or the mask
|
| 721 |
-
embedding. The last one is decisive: `mask_token_id` is **151675**, which is past the end of
|
| 722 |
-
the tokenizer, and the target's `embed_tokens[151675]` is an untrained pad row (L2 norm
|
| 723 |
-
2e-5, against the shipped vector's 0.765). So the reference feeds the drafter **zeros** for
|
| 724 |
-
all seven mask slots and it degenerates to one repeated token. It cannot be the code this
|
| 725 |
-
checkpoint was trained or served with.
|
| 726 |
-
|
| 727 |
-
Every one of those corrections was **cross-checked against SGLang main**, the engine
|
| 728 |
-
Xiaomi's own README recommends, and matches: `models/dflash.py` passes
|
| 729 |
-
`partial_rotary_factor` into `get_rope`, applies `attention_value_scale` as `v_scale`, loads
|
| 730 |
-
per-head `attention_sink_bias`, and maps `sliding_attention` to the 1024 window;
|
| 731 |
-
`speculative/dflash_worker_v2.py::_maybe_merge_trained_mask_embedding()` loads
|
| 732 |
-
`mask_embedding.pt` and merges it into the target's embedding table. (This port applies the
|
| 733 |
-
mask vector one layer earlier, inside the drafter's input layer, so it never mutates the
|
| 734 |
-
target's weights — equivalent, since the target can never emit 151675.)
|
| 735 |
-
|
| 736 |
-
Parity against a corrected HF reference, 7 drafted positions x 64 rounds: **99.78% argmax
|
| 737 |
-
(100% including fp16 ties), KL 3.3e-5, cos 0.999992** at context 384; 98.66% / 100% at
|
| 738 |
-
context 1536 (beyond the 1024 window). Every single disagreement is the reference's rank-1
|
| 739 |
-
token at a top1–top2 margin ≤ 0.017, i.e. 1–2 fp16 ULP.
|
| 740 |
-
|
| 741 |
-
`dflash/` is `dflash-bf16/` converted with exllamav3's `convert.py -b 4`. DFlash drafters
|
| 742 |
-
quantize without calibration, so the conversion needs neither calibration data nor the
|
| 743 |
-
target model. Its `config.json` is the repaired one plus `quantization_config`.
|
| 744 |
|
| 745 |
-
|
| 746 |
-
|
| 747 |
-
## Known issues
|
| 748 |
-
|
| 749 |
-
* **Greedy output is not bit-stable across drafting modes.** On a prose prompt, no-draft /
|
| 750 |
-
DFlash / n-gram / dynamic produced 369 / 447 / 359 / 312 tokens. n-gram differs from
|
| 751 |
-
no-draft too, so this is not a DFlash bug — it is fp16 tie-break sensitivity (traced to a
|
| 752 |
-
single divergence at a top1–top2 margin of 0.0156 = two fp16 ULP, reproduced identically
|
| 753 |
-
with the ring disabled). Coding was token-identical across all four modes. Don't diff
|
| 754 |
-
outputs across drafting modes and expect equality.
|
| 755 |
-
* **Weights and code must match.** Layer 47's `interm_div` is folded into the quantized weights.
|
| 756 |
-
This build on the old branch, or the old build on the current branch, runs without an error
|
| 757 |
-
and gives wrong output. See [Layer 47 and `interm_div`](#layer-47-and-interm_div).
|
| 758 |
-
* **Instruction-following on code is imperfect.** A sample `merge_intervals` completion was a
|
| 759 |
-
correct, idiomatic, correctly-analysed algorithm with one real bug (it sorts the caller's
|
| 760 |
-
list in place and assumes list elements, so it raises on the tuples the prompt specified)
|
| 761 |
-
and it skipped the three asserts the prompt asked for. Consistent with a strong-but-not-
|
| 762 |
-
perfect ~2.3 bpw quant — and the benchmark deltas above say the FP8 model is only ~0.6 pp
|
| 763 |
-
better at HumanEval+, so most of that is the base model, not the quantization.
|
| 764 |
-
* **Tensor parallel is unsupported.** The port has only been built and run single-device.
|
| 765 |
-
* **Long context is now measured, and it is TTFT-bound rather than memory-bound.** 8K → 350K
|
| 766 |
-
has been served and retrieval is perfect at every length and depth tested, but a 250K prompt
|
| 767 |
-
costs ~12.6 minutes of prefill and a 350K prompt ~22.4 minutes on one GB10. `max_seq_len`
|
| 768 |
-
above 393,216 (or above 131,072 with the drafter) does not fit at FP16 KV. See
|
| 769 |
-
[Long context](#long-context).
|
| 770 |
-
* **Batch > 1 with DFlash is untested.** Everything above is `max_batch_size: 1`.
|
| 771 |
-
* **No vision, no audio, no MTP.** The MTP heads in the source repo are not ported, so
|
| 772 |
-
MTP-based drafting is unavailable; DFlash is the drafting path.
|
| 773 |
-
* The source repo's own `preprocessor_config.json` is not included, since there is no
|
| 774 |
-
vision/audio tower to preprocess for.
|
| 775 |
-
|
| 776 |
-
---
|
| 777 |
|
| 778 |
-
|
|
|
|
| 779 |
|
| 780 |
-
|
| 781 |
|
| 782 |
-
*
|
| 783 |
-
|
| 784 |
-
*
|
| 785 |
-
|
| 786 |
-
port tractable.
|
| 787 |
-
* **[vcruz305](https://github.com/vcruz305)** — the aarch64 build guards
|
| 788 |
-
(commit `f4993fef`) that let exllamav3's extension compile on GB10 at all.
|
| 789 |
-
* **[theroyallab](https://github.com/theroyallab/tabbyAPI)** — TabbyAPI.
|
|
|
|
| 14 |
- speculative-decoding
|
| 15 |
---
|
| 16 |
|
| 17 |
+
# MiMo-V2.6-Flash-RL — EXL3 2.27 bpw
|
| 18 |
|
| 19 |
[XiaomiMiMo/MiMo-V2.6-Flash-RL](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL)
|
| 20 |
+
(309B total / 15B active MoE) quantized to EXL3 at 2.27 bpw, so it fits on a single 128 GB
|
| 21 |
+
machine. Built and tested on an NVIDIA DGX Spark (GB10).
|
| 22 |
+
|
| 23 |
+
The repo also includes Xiaomi's DFlash speculative-decoding drafter, fixed so it loads and
|
| 24 |
+
quantized to 4 bpw.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 25 |
|
| 26 |
| | |
|
| 27 |
|---|---|
|
| 28 |
+
| Weights | 83.45 GiB, 12 shards |
|
| 29 |
+
| Bitrate | 2.27 bpw (excluding head), head 6 bpw |
|
| 30 |
+
| Perplexity | 5.40, wikitext-2 test, 64 x 2048 tokens |
|
| 31 |
+
| Decode, batch 1 | ~31 tok/s without drafter; 40–60 tok/s with the 4 bpw drafter (see [Speed](#speed)) |
|
| 32 |
+
| Context | 262,144 tokens with the drafter on a DGX Spark |
|
| 33 |
+
| Modalities | Text only (no vision, audio or MTP heads) |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 34 |
|
| 35 |
+
**This needs a patched exllamav3.** MiMo-V2 support is not upstream yet; see
|
| 36 |
+
[How to run](#how-to-run).
|
|
|
|
|
|
|
|
|
|
|
|
|
| 37 |
|
| 38 |
+
## Files
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 39 |
|
| 40 |
```
|
| 41 |
+
model-*.safetensors, model.safetensors.index.json EXL3 weights
|
| 42 |
+
quantization_config.json per-tensor storage record
|
| 43 |
+
config.json, tokenizer files, chat_template.jinja
|
| 44 |
+
dflash/ drafter, 4 bpw EXL3 (use this one)
|
| 45 |
+
dflash-bf16/ the same drafter, unquantized
|
| 46 |
+
eval/ benchmark outputs
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 47 |
```
|
| 48 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 49 |
## Bitrate
|
| 50 |
|
| 51 |
+
Converted with `convert.py -b 2.25 -hq -cr 250 -cc 2048`. The per-module result:
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 52 |
|
| 53 |
+
| module | bpw |
|
| 54 |
+
|---|---|
|
| 55 |
+
| routed experts, layers 12–35 | 2.0 |
|
| 56 |
+
| routed experts, layers 1–11 and 36–47 | 2.5 |
|
| 57 |
+
| attention | 4.0 |
|
| 58 |
+
| dense MLP (layer 0) | 3.0 |
|
| 59 |
+
| `lm_head` | 6.0 |
|
| 60 |
+
| embeddings, norms, router | BF16 |
|
| 61 |
+
|
| 62 |
+
**Layer 47:** some of its experts produce intermediate values past the fp16 limit.
|
| 63 |
+
This build scales `up_proj` down by 128 in that layer (`interm_div`) and restores the scale in
|
| 64 |
+
fp32. Because the scale is folded into the weights, these weights need the current
|
| 65 |
+
`mimo-v2.6-flash` exllamav3 branch; on an older checkout they load but produce wrong output.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 66 |
|
| 67 |
## Quality
|
| 68 |
|
| 69 |
+
Perplexity (wikitext-2 test, 64 x 2048): **5.4003**.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 70 |
|
| 71 |
+
Benchmarks compare this quant against the unquantized FP8 model (Xiaomi's endpoint via
|
| 72 |
+
OpenRouter) with the same harness, items, prompts and greedy sampling.
|
| 73 |
+
They are only comparable to each other, not to other published scores.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 74 |
|
| 75 |
+
| task | N | this quant | FP8 reference | delta |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 76 |
|---|---|---|---|---|
|
| 77 |
+
| HumanEval+ | 164 | 89.6% | 90.2% | −0.6 |
|
| 78 |
+
| MBPP+ | 378 | 77.5% | 77.2% | +0.3 |
|
| 79 |
+
| GSM8K | 500 | 95.8% | 96.4% | −0.6 |
|
| 80 |
+
| MMLU-Pro | 500 | 72.6% | 76.4% | −3.8 |
|
| 81 |
+
| GPQA-Diamond (thinking, 8K token cap) | 64 | 46.9% | 56.2% | −9.4 |
|
| 82 |
|
| 83 |
+
Most of the GPQA gap comes from the token cap: the quant ran past 8K tokens without answering
|
| 84 |
+
on 45% of items vs 38% for the reference. Among answered items it scored 85.7% vs 90.0%.
|
| 85 |
+
Method and raw outputs are in [`eval/bench/`](eval/bench/).
|
| 86 |
|
| 87 |
+
## Speed
|
| 88 |
|
| 89 |
+
DGX Spark, batch 1, greedy, 512 generated tokens, median of 3 runs:
|
|
|
|
| 90 |
|
| 91 |
+
| prompt | no drafter | 4 bpw drafter | speedup |
|
| 92 |
|---|---|---|---|
|
| 93 |
+
| coding | 31.4 tok/s | 48.5 tok/s | 1.54x |
|
| 94 |
+
| prose | 31.4 tok/s | 40.2 tok/s | 1.28x |
|
| 95 |
+
| reasoning | 31.4 tok/s | 59.5 tok/s | 1.89x |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 96 |
|
| 97 |
+
The target model verifies every drafted token, so the drafter does not affect output quality.
|
| 98 |
|
| 99 |
+
Long context with the BF16 drafter, generating ~400 tokens of code against a large
|
| 100 |
+
repository prompt:
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 101 |
|
| 102 |
+
| prompt tokens | no drafter | drafter | speedup |
|
| 103 |
|---|---|---|---|
|
| 104 |
+
| 64,614 | 25.1 tok/s | 43.8 tok/s | 1.75x |
|
| 105 |
+
| 130,118 | 21.1 tok/s | 41.4 tok/s | 1.97x |
|
| 106 |
+
| 248,993 | 16.1 tok/s | 33.6 tok/s | 2.08x |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 107 |
|
| 108 |
+
Decode holds up at long context because 39 of the 48 layers use a 128-token sliding window.
|
| 109 |
+
Prefill does not: a 250K-token prompt takes about 13 minutes on one GB10.
|
|
|
|
| 110 |
|
| 111 |
+
Needle retrieval passed at every length and depth tested, from 8K to 350K tokens.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 112 |
|
| 113 |
## How to run
|
| 114 |
|
| 115 |
+
The port is in turboderp-org/exllamav3
|
| 116 |
+
[#399](https://github.com/turboderp-org/exllamav3/pull/399). Until it lands, use these
|
| 117 |
+
branches:
|
| 118 |
|
| 119 |
+
* exllamav3: https://github.com/benthecarman/exllamav3/tree/mimo-v2.6-flash
|
| 120 |
+
* TabbyAPI: https://github.com/benthecarman/tabbyAPI/tree/mimo-v2.6-flash
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 121 |
|
| 122 |
```sh
|
| 123 |
git clone -b mimo-v2.6-flash https://github.com/benthecarman/exllamav3
|
| 124 |
git clone -b mimo-v2.6-flash https://github.com/benthecarman/tabbyAPI
|
| 125 |
+
pip install -e ./exllamav3
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 126 |
hf download benthecarman/MiMo-V2.6-Flash-RL-exl3 --local-dir mimo-exl3
|
| 127 |
```
|
| 128 |
|
| 129 |
+
Install TabbyAPI's base dependencies only; its `cu12`/`cu13` extras replace torch and the
|
| 130 |
+
editable exllamav3. On aarch64, also `pip install uvloop`.
|
| 131 |
|
| 132 |
+
TabbyAPI looks for the drafter by name inside `draft_model_dir`, so put `dflash/` in its own
|
| 133 |
+
directory next to the model:
|
| 134 |
|
| 135 |
```
|
| 136 |
models/
|
| 137 |
+
mimo-2.25bpw-hq/ everything except dflash/, dflash-bf16/ and eval/
|
| 138 |
+
mimo-dflash-draft/ the contents of dflash/
|
| 139 |
```
|
| 140 |
|
| 141 |
`config.yml`:
|
|
|
|
| 144 |
model:
|
| 145 |
model_dir: models
|
| 146 |
model_name: mimo-2.25bpw-hq
|
| 147 |
+
max_seq_len: 262144
|
| 148 |
cache_size: 262144
|
| 149 |
+
cache_mode: FP16
|
| 150 |
chunk_size: 2048
|
| 151 |
+
max_batch_size: 1
|
|
|
|
|
|
|
| 152 |
reasoning: true
|
| 153 |
reasoning_start_token: "<think>"
|
| 154 |
reasoning_end_token: "</think>"
|
|
|
|
|
|
|
| 155 |
|
| 156 |
draft_model:
|
| 157 |
draft_mode: model
|
|
|
|
| 161 |
dynamic_draft: true
|
| 162 |
```
|
| 163 |
|
| 164 |
+
On a DGX Spark or other unified-memory machine, start TabbyAPI with `EXL3_LOAD_DEVICE=cuda:0`.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 165 |
|
| 166 |
+
Thinking can be turned off per request with `"chat_template_kwargs": {"enable_thinking": false}`.
|
| 167 |
+
Tool calls use the `qwen3_coder` format, which TabbyAPI detects automatically.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 168 |
|
| 169 |
+
A ready-made `serve.sh` and config are in
|
| 170 |
+
[`examples/mimo_v2_6/`](https://github.com/benthecarman/exllamav3/tree/mimo-v2.6-flash/examples/mimo_v2_6).
|
| 171 |
|
| 172 |
+
## Credits
|
| 173 |
|
| 174 |
+
* Xiaomi MiMo team: the model and the DFlash drafter (MIT).
|
| 175 |
+
* [turboderp](https://github.com/turboderp): ExLlamaV3 and the EXL3 format.
|
| 176 |
+
* [vcruz305](https://github.com/vcruz305): the aarch64 build fixes.
|
| 177 |
+
* [theroyallab](https://github.com/theroyallab/tabbyAPI): TabbyAPI.
|
|
|
|
|
|
|
|
|
|
|
|