MiMo-V2.6-Flash-RL — EXL3 2.27 bpw

XiaomiMiMo/MiMo-V2.6-Flash-RL (309B total / 15B active MoE) quantized to EXL3 at 2.27 bpw, so it fits on a single 128 GB machine. Built and tested on an NVIDIA DGX Spark (GB10).

The repo also includes Xiaomi's DFlash speculative-decoding drafter, fixed so it loads and quantized to 4 bpw, and the model's own MTP heads and vision tower.

Weights 85.28 GiB, 12 shards, including the MTP heads and the vision tower
Bitrate 2.27 bpw (excluding head), head 6 bpw
Perplexity 5.40, wikitext-2 test, 64 x 2048 tokens
Decode, batch 1 ~31 tok/s without a drafter; 35–80 tok/s with the DFlash drafter (see Speed)
Context 262,144 tokens on a DGX Spark, with the drafter and vision loaded
Modalities Text and images (no audio)

Requires exllamav3 v1.5.2 or later; see How to run.

Files

model-*.safetensors, model.safetensors.index.json   EXL3 weights
quantization_config.json                            per-tensor storage record
config.json, tokenizer files, chat_template.jinja, preprocessor_config.json
dflash/        drafter, 4 bpw EXL3 (use this one)
dflash-bf16/   the same drafter, unquantized
eval/          benchmark outputs

Bitrate

Converted with convert.py -b 2.25 -hq -cr 250 -cc 2048. The per-module result:

module bpw
routed experts, layers 12–35 2.0
routed experts, layers 1–11 and 36–47 2.5
attention 4.0
dense MLP (layer 0) 3.0
lm_head 6.0
MTP heads 4.0
embeddings, norms, router, vision tower BF16

Layer 47: some of its experts produce intermediate values past the fp16 limit. This build scales up_proj down by 128 in that layer (interm_div) and restores the scale in fp32. The scale is folded into the weights.

Quality

Perplexity (wikitext-2 test, 64 x 2048): 5.4003.

Benchmarks compare this quant against the unquantized FP8 model (Xiaomi's endpoint via OpenRouter) with the same harness, items, prompts and greedy sampling. They are only comparable to each other, not to other published scores.

task N this quant FP8 reference delta
HumanEval+ 164 88.4% 90.2% −1.8
MBPP+ 378 77.2% 77.2% 0.0
GSM8K 500 95.2% 96.4% −1.2
MMLU-Pro 500 73.8% 76.4% −2.6
GPQA-Diamond (thinking, 8K token cap) 64 50.0% 56.2% −6.2

On GPQA both models ran past the 8K token cap without answering on 38–39% of items, so the gap is in the answers themselves: 82.1% vs 90.0% among answered items (39 and 40 of 64). Per-task scores and run settings are in eval/bench/.

Speed

exllamav3 v1.5.2 on a DGX Spark, batch 1, greedy, 512 generated tokens, median of 3 runs, in tok/s:

prompt no drafter DFlash drafter MTP heads
coding 31.5 49.5 40.7
prose 31.3 35.4 31.8
reasoning 31.2 61.1 45.1
code edit 31.0 80.7 51.3

DFlash is the faster drafter on every prompt. The MTP heads draft 3 tokens per step; asking for 6 was at most 5% faster. Prose gains the least and varies the most from prompt to prompt. The target model verifies every drafted token, so neither drafter affects output quality.

Long context with the BF16 drafter, generating ~400 tokens of code against a large repository prompt:

prompt tokens no drafter drafter speedup
64,614 25.1 tok/s 43.8 tok/s 1.75x
130,118 21.1 tok/s 41.4 tok/s 1.97x
248,993 16.1 tok/s 33.6 tok/s 2.08x

Decode holds up at long context because 39 of the 48 layers use a 128-token sliding window. Prefill does not: a 250K-token prompt takes about 11 minutes on one GB10.

Needle retrieval passed at every length and depth tested, from 8K to 350K tokens.

How to run

Supported in exllamav3 v1.5.2 and later. Serve it with TabbyAPI.

hf download benthecarman/MiMo-V2.6-Flash-RL-exl3 --local-dir mimo-exl3

There are no prebuilt exllamav3 wheels for aarch64, so on a DGX Spark install it from source. TabbyAPI also needs pip install uvloop there.

TabbyAPI looks for the drafter by name inside draft_model_dir, so put dflash/ in its own directory next to the model:

models/
  mimo-2.25bpw-hq/     everything except dflash/, dflash-bf16/ and eval/
  mimo-dflash-draft/   the contents of dflash/

config.yml:

model:
  model_dir: models
  model_name: mimo-2.25bpw-hq
  max_seq_len: 262144
  cache_size: 262144
  cache_mode: FP16
  chunk_size: 2048
  max_batch_size: 1
  reasoning: true
  reasoning_start_token: "<think>"
  reasoning_end_token: "</think>"

draft_model:
  draft_mode: model
  draft_model_dir: models
  draft_model_name: mimo-dflash-draft
  draft_cache_mode: FP16
  dynamic_draft: true

To use the MTP heads instead of the DFlash drafter:

draft_model:
  draft_mode: mtp
  dynamic_draft: true

For image input, add vision: true under model:. The vision tower adds about 1.3 GiB; with it and the DFlash drafter loaded at 262,144 context, a 250K-token prompt still left 15 GiB free.

Thinking can be turned off per request with "chat_template_kwargs": {"enable_thinking": false}. Tool calls use the qwen3_coder format, which TabbyAPI detects automatically.

Credits

  • Xiaomi MiMo team: the model and the DFlash drafter (MIT).
  • turboderp: ExLlamaV3 and the EXL3 format.
  • vcruz305: the aarch64 build fixes.
  • theroyallab: TabbyAPI.
Downloads last month
1,373
Safetensors
Model size
46B params
Tensor type
BF16
·
F16
·
I16
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for benthecarman/MiMo-V2.6-Flash-RL-exl3

Quantized
(35)
this model