MiMo-V2.6-Flash-RL — EXL3 2.27 bpw
XiaomiMiMo/MiMo-V2.6-Flash-RL (309B total / 15B active MoE) quantized to EXL3 at 2.27 bpw, so it fits on a single 128 GB machine. Built and tested on an NVIDIA DGX Spark (GB10).
The repo also includes Xiaomi's DFlash speculative-decoding drafter, fixed so it loads and quantized to 4 bpw, and the model's own MTP heads and vision tower.
| Weights | 85.28 GiB, 12 shards, including the MTP heads and the vision tower |
| Bitrate | 2.27 bpw (excluding head), head 6 bpw |
| Perplexity | 5.40, wikitext-2 test, 64 x 2048 tokens |
| Decode, batch 1 | ~31 tok/s without a drafter; 35–80 tok/s with the DFlash drafter (see Speed) |
| Context | 262,144 tokens on a DGX Spark, with the drafter and vision loaded |
| Modalities | Text and images (no audio) |
Requires exllamav3 v1.5.2 or later; see How to run.
Files
model-*.safetensors, model.safetensors.index.json EXL3 weights
quantization_config.json per-tensor storage record
config.json, tokenizer files, chat_template.jinja, preprocessor_config.json
dflash/ drafter, 4 bpw EXL3 (use this one)
dflash-bf16/ the same drafter, unquantized
eval/ benchmark outputs
Bitrate
Converted with convert.py -b 2.25 -hq -cr 250 -cc 2048. The per-module result:
| module | bpw |
|---|---|
| routed experts, layers 12–35 | 2.0 |
| routed experts, layers 1–11 and 36–47 | 2.5 |
| attention | 4.0 |
| dense MLP (layer 0) | 3.0 |
lm_head |
6.0 |
| MTP heads | 4.0 |
| embeddings, norms, router, vision tower | BF16 |
Layer 47: some of its experts produce intermediate values past the fp16 limit.
This build scales up_proj down by 128 in that layer (interm_div) and restores the scale in
fp32. The scale is folded into the weights.
Quality
Perplexity (wikitext-2 test, 64 x 2048): 5.4003.
Benchmarks compare this quant against the unquantized FP8 model (Xiaomi's endpoint via OpenRouter) with the same harness, items, prompts and greedy sampling. They are only comparable to each other, not to other published scores.
| task | N | this quant | FP8 reference | delta |
|---|---|---|---|---|
| HumanEval+ | 164 | 88.4% | 90.2% | −1.8 |
| MBPP+ | 378 | 77.2% | 77.2% | 0.0 |
| GSM8K | 500 | 95.2% | 96.4% | −1.2 |
| MMLU-Pro | 500 | 73.8% | 76.4% | −2.6 |
| GPQA-Diamond (thinking, 8K token cap) | 64 | 50.0% | 56.2% | −6.2 |
On GPQA both models ran past the 8K token cap without answering on 38–39% of items, so the
gap is in the answers themselves: 82.1% vs 90.0% among answered items (39 and 40 of 64).
Per-task scores and run settings are in eval/bench/.
Speed
exllamav3 v1.5.2 on a DGX Spark, batch 1, greedy, 512 generated tokens, median of 3 runs, in tok/s:
| prompt | no drafter | DFlash drafter | MTP heads |
|---|---|---|---|
| coding | 31.5 | 49.5 | 40.7 |
| prose | 31.3 | 35.4 | 31.8 |
| reasoning | 31.2 | 61.1 | 45.1 |
| code edit | 31.0 | 80.7 | 51.3 |
DFlash is the faster drafter on every prompt. The MTP heads draft 3 tokens per step; asking for 6 was at most 5% faster. Prose gains the least and varies the most from prompt to prompt. The target model verifies every drafted token, so neither drafter affects output quality.
Long context with the BF16 drafter, generating ~400 tokens of code against a large repository prompt:
| prompt tokens | no drafter | drafter | speedup |
|---|---|---|---|
| 64,614 | 25.1 tok/s | 43.8 tok/s | 1.75x |
| 130,118 | 21.1 tok/s | 41.4 tok/s | 1.97x |
| 248,993 | 16.1 tok/s | 33.6 tok/s | 2.08x |
Decode holds up at long context because 39 of the 48 layers use a 128-token sliding window. Prefill does not: a 250K-token prompt takes about 11 minutes on one GB10.
Needle retrieval passed at every length and depth tested, from 8K to 350K tokens.
How to run
Supported in exllamav3 v1.5.2 and later. Serve it with TabbyAPI.
hf download benthecarman/MiMo-V2.6-Flash-RL-exl3 --local-dir mimo-exl3
There are no prebuilt exllamav3 wheels for aarch64, so on a DGX Spark install it from source.
TabbyAPI also needs pip install uvloop there.
TabbyAPI looks for the drafter by name inside draft_model_dir, so put dflash/ in its own
directory next to the model:
models/
mimo-2.25bpw-hq/ everything except dflash/, dflash-bf16/ and eval/
mimo-dflash-draft/ the contents of dflash/
config.yml:
model:
model_dir: models
model_name: mimo-2.25bpw-hq
max_seq_len: 262144
cache_size: 262144
cache_mode: FP16
chunk_size: 2048
max_batch_size: 1
reasoning: true
reasoning_start_token: "<think>"
reasoning_end_token: "</think>"
draft_model:
draft_mode: model
draft_model_dir: models
draft_model_name: mimo-dflash-draft
draft_cache_mode: FP16
dynamic_draft: true
To use the MTP heads instead of the DFlash drafter:
draft_model:
draft_mode: mtp
dynamic_draft: true
For image input, add vision: true under model:. The vision tower adds about 1.3 GiB; with it
and the DFlash drafter loaded at 262,144 context, a 250K-token prompt still left 15 GiB free.
Thinking can be turned off per request with "chat_template_kwargs": {"enable_thinking": false}.
Tool calls use the qwen3_coder format, which TabbyAPI detects automatically.
Credits
- Xiaomi MiMo team: the model and the DFlash drafter (MIT).
- turboderp: ExLlamaV3 and the EXL3 format.
- vcruz305: the aarch64 build fixes.
- theroyallab: TabbyAPI.
- Downloads last month
- 1,373
Model tree for benthecarman/MiMo-V2.6-Flash-RL-exl3
Base model
XiaomiMiMo/MiMo-V2.6-Flash-RL