File size: 5,964 Bytes
3d2885d c0087c1 3d2885d c0087c1 5ae87b8 c1e3fc1 3d2885d 5ae87b8 c0087c1 5ae87b8 3d2885d 6cbdacb 3d2885d c0087c1 3d2885d c0087c1 5ae87b8 c0087c1 3d2885d c0087c1 3d2885d c0087c1 5ae87b8 c0087c1 6cbdacb 3d2885d c0087c1 3d2885d c0087c1 3d2885d c0087c1 3d2885d d453d5a 3d2885d c0087c1 3d2885d 5ae87b8 027d461 5ae87b8 027d461 5ae87b8 3d2885d 5ae87b8 3d2885d c0087c1 0f8d7e0 c0087c1 0f8d7e0 c0087c1 0f8d7e0 c0087c1 6cbdacb 0f8d7e0 c0087c1 0f8d7e0 3d2885d 6cbdacb 3d2885d 6cbdacb 3d2885d c0087c1 3d2885d c0087c1 3d2885d c0087c1 225cc88 c0087c1 3d2885d c0087c1 3d2885d 5ae87b8 c0087c1 3d2885d c0087c1 3d2885d c0087c1 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 | ---
license: mit
base_model: XiaomiMiMo/MiMo-V2.6-Flash-RL
base_model_relation: quantized
pipeline_tag: text-generation
language:
- en
- zh
tags:
- exl3
- exllamav3
- mimo
- dgx-spark
- speculative-decoding
---
# MiMo-V2.6-Flash-RL β EXL3 2.27 bpw
[XiaomiMiMo/MiMo-V2.6-Flash-RL](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL)
(309B total / 15B active MoE) quantized to EXL3 at 2.27 bpw, so it fits on a single 128 GB
machine. Built and tested on an NVIDIA DGX Spark (GB10).
The repo also includes Xiaomi's DFlash speculative-decoding drafter, fixed so it loads and
quantized to 4 bpw, and the model's own MTP heads and vision tower.
| | |
|---|---|
| Weights | 85.28 GiB, 12 shards, including the MTP heads and the vision tower |
| Bitrate | 2.27 bpw (excluding head), head 6 bpw |
| Perplexity | 5.40, wikitext-2 test, 64 x 2048 tokens |
| Decode, batch 1 | ~31 tok/s without a drafter; 35β80 tok/s with the DFlash drafter (see [Speed](#speed)) |
| Context | 262,144 tokens on a DGX Spark, with the drafter and vision loaded |
| Modalities | Text and images (no audio) |
Requires exllamav3 v1.5.2 or later; see [How to run](#how-to-run).
## Files
```
model-*.safetensors, model.safetensors.index.json EXL3 weights
quantization_config.json per-tensor storage record
config.json, tokenizer files, chat_template.jinja, preprocessor_config.json
dflash/ drafter, 4 bpw EXL3 (use this one)
dflash-bf16/ the same drafter, unquantized
eval/ benchmark outputs
```
## Bitrate
Converted with `convert.py -b 2.25 -hq -cr 250 -cc 2048`. The per-module result:
| module | bpw |
|---|---|
| routed experts, layers 12β35 | 2.0 |
| routed experts, layers 1β11 and 36β47 | 2.5 |
| attention | 4.0 |
| dense MLP (layer 0) | 3.0 |
| `lm_head` | 6.0 |
| MTP heads | 4.0 |
| embeddings, norms, router, vision tower | BF16 |
**Layer 47:** some of its experts produce intermediate values past the fp16 limit.
This build scales `up_proj` down by 128 in that layer (`interm_div`) and restores the scale in
fp32. The scale is folded into the weights.
## Quality
Perplexity (wikitext-2 test, 64 x 2048): **5.4003**.
Benchmarks compare this quant against the unquantized FP8 model (Xiaomi's endpoint via
OpenRouter) with the same harness, items, prompts and greedy sampling.
They are only comparable to each other, not to other published scores.
| task | N | this quant | FP8 reference | delta |
|---|---|---|---|---|
| HumanEval+ | 164 | 88.4% | 90.2% | β1.8 |
| MBPP+ | 378 | 77.2% | 77.2% | 0.0 |
| GSM8K | 500 | 95.2% | 96.4% | β1.2 |
| MMLU-Pro | 500 | 73.8% | 76.4% | β2.6 |
| GPQA-Diamond (thinking, 8K token cap) | 64 | 50.0% | 56.2% | β6.2 |
On GPQA both models ran past the 8K token cap without answering on 38β39% of items, so the
gap is in the answers themselves: 82.1% vs 90.0% among answered items (39 and 40 of 64).
Per-task scores and run settings are in [`eval/bench/`](eval/bench/).
## Speed
exllamav3 v1.5.2 on a DGX Spark, batch 1, greedy, 512 generated tokens, median of 3 runs, in
tok/s:
| prompt | no drafter | DFlash drafter | MTP heads |
|---|---|---|---|
| coding | 31.5 | 49.5 | 40.7 |
| prose | 31.3 | 35.4 | 31.8 |
| reasoning | 31.2 | 61.1 | 45.1 |
| code edit | 31.0 | 80.7 | 51.3 |
DFlash is the faster drafter on every prompt. The MTP heads draft 3 tokens per step; asking for
6 was at most 5% faster. Prose gains the least and varies the most from prompt to prompt. The target model verifies every drafted token, so neither drafter
affects output quality.
Long context with the BF16 drafter, generating ~400 tokens of code against a large
repository prompt:
| prompt tokens | no drafter | drafter | speedup |
|---|---|---|---|
| 64,614 | 25.1 tok/s | 43.8 tok/s | 1.75x |
| 130,118 | 21.1 tok/s | 41.4 tok/s | 1.97x |
| 248,993 | 16.1 tok/s | 33.6 tok/s | 2.08x |
Decode holds up at long context because 39 of the 48 layers use a 128-token sliding window.
Prefill does not: a 250K-token prompt takes about 11 minutes on one GB10.
Needle retrieval passed at every length and depth tested, from 8K to 350K tokens.
## How to run
Supported in [exllamav3](https://github.com/turboderp-org/exllamav3) v1.5.2 and later. Serve it
with [TabbyAPI](https://github.com/theroyallab/tabbyAPI).
```sh
hf download benthecarman/MiMo-V2.6-Flash-RL-exl3 --local-dir mimo-exl3
```
There are no prebuilt exllamav3 wheels for aarch64, so on a DGX Spark install it from source.
TabbyAPI also needs `pip install uvloop` there.
TabbyAPI looks for the drafter by name inside `draft_model_dir`, so put `dflash/` in its own
directory next to the model:
```
models/
mimo-2.25bpw-hq/ everything except dflash/, dflash-bf16/ and eval/
mimo-dflash-draft/ the contents of dflash/
```
`config.yml`:
```yaml
model:
model_dir: models
model_name: mimo-2.25bpw-hq
max_seq_len: 262144
cache_size: 262144
cache_mode: FP16
chunk_size: 2048
max_batch_size: 1
reasoning: true
reasoning_start_token: "<think>"
reasoning_end_token: "</think>"
draft_model:
draft_mode: model
draft_model_dir: models
draft_model_name: mimo-dflash-draft
draft_cache_mode: FP16
dynamic_draft: true
```
To use the MTP heads instead of the DFlash drafter:
```yaml
draft_model:
draft_mode: mtp
dynamic_draft: true
```
For image input, add `vision: true` under `model:`. The vision tower adds about 1.3 GiB; with it
and the DFlash drafter loaded at 262,144 context, a 250K-token prompt still left 15 GiB free.
Thinking can be turned off per request with `"chat_template_kwargs": {"enable_thinking": false}`.
Tool calls use the `qwen3_coder` format, which TabbyAPI detects automatically.
## Credits
* Xiaomi MiMo team: the model and the DFlash drafter (MIT).
* [turboderp](https://github.com/turboderp): ExLlamaV3 and the EXL3 format.
* [vcruz305](https://github.com/vcruz305): the aarch64 build fixes.
* [theroyallab](https://github.com/theroyallab/tabbyAPI): TabbyAPI.
|