File size: 5,964 Bytes
3d2885d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
c0087c1
3d2885d
 
c0087c1
 
 
 
5ae87b8
c1e3fc1
3d2885d
 
5ae87b8
c0087c1
 
5ae87b8
 
 
3d2885d
6cbdacb
3d2885d
c0087c1
3d2885d
 
c0087c1
 
5ae87b8
c0087c1
 
 
3d2885d
 
 
 
c0087c1
3d2885d
c0087c1
 
 
 
 
 
 
5ae87b8
 
c0087c1
 
 
6cbdacb
3d2885d
 
 
c0087c1
3d2885d
c0087c1
 
 
3d2885d
c0087c1
3d2885d
d453d5a
 
 
 
 
 
 
 
 
3d2885d
c0087c1
3d2885d
5ae87b8
 
027d461
5ae87b8
027d461
5ae87b8
 
 
 
3d2885d
5ae87b8
 
 
3d2885d
c0087c1
 
0f8d7e0
c0087c1
0f8d7e0
c0087c1
 
 
0f8d7e0
c0087c1
6cbdacb
0f8d7e0
c0087c1
0f8d7e0
3d2885d
 
6cbdacb
 
3d2885d
 
 
 
 
6cbdacb
 
3d2885d
c0087c1
 
3d2885d
 
 
c0087c1
 
3d2885d
 
 
 
 
 
 
 
c0087c1
225cc88
c0087c1
3d2885d
c0087c1
3d2885d
 
 
 
 
 
 
 
 
 
 
 
5ae87b8
 
 
 
 
 
 
 
 
 
 
c0087c1
 
3d2885d
c0087c1
3d2885d
c0087c1
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
---
license: mit
base_model: XiaomiMiMo/MiMo-V2.6-Flash-RL
base_model_relation: quantized
pipeline_tag: text-generation
language:
- en
- zh
tags:
- exl3
- exllamav3
- mimo
- dgx-spark
- speculative-decoding
---

# MiMo-V2.6-Flash-RL β€” EXL3 2.27 bpw

[XiaomiMiMo/MiMo-V2.6-Flash-RL](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL)
(309B total / 15B active MoE) quantized to EXL3 at 2.27 bpw, so it fits on a single 128 GB
machine. Built and tested on an NVIDIA DGX Spark (GB10).

The repo also includes Xiaomi's DFlash speculative-decoding drafter, fixed so it loads and
quantized to 4 bpw, and the model's own MTP heads and vision tower.

| | |
|---|---|
| Weights | 85.28 GiB, 12 shards, including the MTP heads and the vision tower |
| Bitrate | 2.27 bpw (excluding head), head 6 bpw |
| Perplexity | 5.40, wikitext-2 test, 64 x 2048 tokens |
| Decode, batch 1 | ~31 tok/s without a drafter; 35–80 tok/s with the DFlash drafter (see [Speed](#speed)) |
| Context | 262,144 tokens on a DGX Spark, with the drafter and vision loaded |
| Modalities | Text and images (no audio) |

Requires exllamav3 v1.5.2 or later; see [How to run](#how-to-run).

## Files

```
model-*.safetensors, model.safetensors.index.json   EXL3 weights
quantization_config.json                            per-tensor storage record
config.json, tokenizer files, chat_template.jinja, preprocessor_config.json
dflash/        drafter, 4 bpw EXL3 (use this one)
dflash-bf16/   the same drafter, unquantized
eval/          benchmark outputs
```

## Bitrate

Converted with `convert.py -b 2.25 -hq -cr 250 -cc 2048`. The per-module result:

| module | bpw |
|---|---|
| routed experts, layers 12–35 | 2.0 |
| routed experts, layers 1–11 and 36–47 | 2.5 |
| attention | 4.0 |
| dense MLP (layer 0) | 3.0 |
| `lm_head` | 6.0 |
| MTP heads | 4.0 |
| embeddings, norms, router, vision tower | BF16 |

**Layer 47:** some of its experts produce intermediate values past the fp16 limit.
This build scales `up_proj` down by 128 in that layer (`interm_div`) and restores the scale in
fp32. The scale is folded into the weights.

## Quality

Perplexity (wikitext-2 test, 64 x 2048): **5.4003**.

Benchmarks compare this quant against the unquantized FP8 model (Xiaomi's endpoint via
OpenRouter) with the same harness, items, prompts and greedy sampling.
They are only comparable to each other, not to other published scores.

| task | N | this quant | FP8 reference | delta |
|---|---|---|---|---|
| HumanEval+ | 164 | 88.4% | 90.2% | βˆ’1.8 |
| MBPP+ | 378 | 77.2% | 77.2% | 0.0 |
| GSM8K | 500 | 95.2% | 96.4% | βˆ’1.2 |
| MMLU-Pro | 500 | 73.8% | 76.4% | βˆ’2.6 |
| GPQA-Diamond (thinking, 8K token cap) | 64 | 50.0% | 56.2% | βˆ’6.2 |

On GPQA both models ran past the 8K token cap without answering on 38–39% of items, so the
gap is in the answers themselves: 82.1% vs 90.0% among answered items (39 and 40 of 64).
Per-task scores and run settings are in [`eval/bench/`](eval/bench/).

## Speed

exllamav3 v1.5.2 on a DGX Spark, batch 1, greedy, 512 generated tokens, median of 3 runs, in
tok/s:

| prompt | no drafter | DFlash drafter | MTP heads |
|---|---|---|---|
| coding | 31.5 | 49.5 | 40.7 |
| prose | 31.3 | 35.4 | 31.8 |
| reasoning | 31.2 | 61.1 | 45.1 |
| code edit | 31.0 | 80.7 | 51.3 |

DFlash is the faster drafter on every prompt. The MTP heads draft 3 tokens per step; asking for
6 was at most 5% faster. Prose gains the least and varies the most from prompt to prompt. The target model verifies every drafted token, so neither drafter
affects output quality.

Long context with the BF16 drafter, generating ~400 tokens of code against a large
repository prompt:

| prompt tokens | no drafter | drafter | speedup |
|---|---|---|---|
| 64,614 | 25.1 tok/s | 43.8 tok/s | 1.75x |
| 130,118 | 21.1 tok/s | 41.4 tok/s | 1.97x |
| 248,993 | 16.1 tok/s | 33.6 tok/s | 2.08x |

Decode holds up at long context because 39 of the 48 layers use a 128-token sliding window.
Prefill does not: a 250K-token prompt takes about 11 minutes on one GB10.

Needle retrieval passed at every length and depth tested, from 8K to 350K tokens.

## How to run

Supported in [exllamav3](https://github.com/turboderp-org/exllamav3) v1.5.2 and later. Serve it
with [TabbyAPI](https://github.com/theroyallab/tabbyAPI).

```sh
hf download benthecarman/MiMo-V2.6-Flash-RL-exl3 --local-dir mimo-exl3
```

There are no prebuilt exllamav3 wheels for aarch64, so on a DGX Spark install it from source.
TabbyAPI also needs `pip install uvloop` there.

TabbyAPI looks for the drafter by name inside `draft_model_dir`, so put `dflash/` in its own
directory next to the model:

```
models/
  mimo-2.25bpw-hq/     everything except dflash/, dflash-bf16/ and eval/
  mimo-dflash-draft/   the contents of dflash/
```

`config.yml`:

```yaml
model:
  model_dir: models
  model_name: mimo-2.25bpw-hq
  max_seq_len: 262144
  cache_size: 262144
  cache_mode: FP16
  chunk_size: 2048
  max_batch_size: 1
  reasoning: true
  reasoning_start_token: "<think>"
  reasoning_end_token: "</think>"

draft_model:
  draft_mode: model
  draft_model_dir: models
  draft_model_name: mimo-dflash-draft
  draft_cache_mode: FP16
  dynamic_draft: true
```

To use the MTP heads instead of the DFlash drafter:

```yaml
draft_model:
  draft_mode: mtp
  dynamic_draft: true
```

For image input, add `vision: true` under `model:`. The vision tower adds about 1.3 GiB; with it
and the DFlash drafter loaded at 262,144 context, a 250K-token prompt still left 15 GiB free.

Thinking can be turned off per request with `"chat_template_kwargs": {"enable_thinking": false}`.
Tool calls use the `qwen3_coder` format, which TabbyAPI detects automatically.

## Credits

* Xiaomi MiMo team: the model and the DFlash drafter (MIT).
* [turboderp](https://github.com/turboderp): ExLlamaV3 and the EXL3 format.
* [vcruz305](https://github.com/vcruz305): the aarch64 build fixes.
* [theroyallab](https://github.com/theroyallab/tabbyAPI): TabbyAPI.