--- license: mit base_model: XiaomiMiMo/MiMo-V2.6-Flash-RL base_model_relation: quantized pipeline_tag: text-generation language: - en - zh tags: - exl3 - exllamav3 - mimo - dgx-spark - speculative-decoding --- # MiMo-V2.6-Flash-RL — EXL3 2.27 bpw [XiaomiMiMo/MiMo-V2.6-Flash-RL](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL) (309B total / 15B active MoE) quantized to EXL3 at 2.27 bpw, so it fits on a single 128 GB machine. Built and tested on an NVIDIA DGX Spark (GB10). The repo also includes Xiaomi's DFlash speculative-decoding drafter, fixed so it loads and quantized to 4 bpw, and the model's own MTP heads and vision tower. | | | |---|---| | Weights | 85.28 GiB, 12 shards, including the MTP heads and the vision tower | | Bitrate | 2.27 bpw (excluding head), head 6 bpw | | Perplexity | 5.40, wikitext-2 test, 64 x 2048 tokens | | Decode, batch 1 | ~31 tok/s without a drafter; 35–80 tok/s with the DFlash drafter (see [Speed](#speed)) | | Context | 262,144 tokens on a DGX Spark, with the drafter and vision loaded | | Modalities | Text and images (no audio) | Requires exllamav3 v1.5.2 or later; see [How to run](#how-to-run). ## Files ``` model-*.safetensors, model.safetensors.index.json EXL3 weights quantization_config.json per-tensor storage record config.json, tokenizer files, chat_template.jinja, preprocessor_config.json dflash/ drafter, 4 bpw EXL3 (use this one) dflash-bf16/ the same drafter, unquantized eval/ benchmark outputs ``` ## Bitrate Converted with `convert.py -b 2.25 -hq -cr 250 -cc 2048`. The per-module result: | module | bpw | |---|---| | routed experts, layers 12–35 | 2.0 | | routed experts, layers 1–11 and 36–47 | 2.5 | | attention | 4.0 | | dense MLP (layer 0) | 3.0 | | `lm_head` | 6.0 | | MTP heads | 4.0 | | embeddings, norms, router, vision tower | BF16 | **Layer 47:** some of its experts produce intermediate values past the fp16 limit. This build scales `up_proj` down by 128 in that layer (`interm_div`) and restores the scale in fp32. The scale is folded into the weights. ## Quality Perplexity (wikitext-2 test, 64 x 2048): **5.4003**. Benchmarks compare this quant against the unquantized FP8 model (Xiaomi's endpoint via OpenRouter) with the same harness, items, prompts and greedy sampling. They are only comparable to each other, not to other published scores. | task | N | this quant | FP8 reference | delta | |---|---|---|---|---| | HumanEval+ | 164 | 88.4% | 90.2% | −1.8 | | MBPP+ | 378 | 77.2% | 77.2% | 0.0 | | GSM8K | 500 | 95.2% | 96.4% | −1.2 | | MMLU-Pro | 500 | 73.8% | 76.4% | −2.6 | | GPQA-Diamond (thinking, 8K token cap) | 64 | 50.0% | 56.2% | −6.2 | On GPQA both models ran past the 8K token cap without answering on 38–39% of items, so the gap is in the answers themselves: 82.1% vs 90.0% among answered items (39 and 40 of 64). Per-task scores and run settings are in [`eval/bench/`](eval/bench/). ## Speed exllamav3 v1.5.2 on a DGX Spark, batch 1, greedy, 512 generated tokens, median of 3 runs, in tok/s: | prompt | no drafter | DFlash drafter | MTP heads | |---|---|---|---| | coding | 31.5 | 49.5 | 40.7 | | prose | 31.3 | 35.4 | 31.8 | | reasoning | 31.2 | 61.1 | 45.1 | | code edit | 31.0 | 80.7 | 51.3 | DFlash is the faster drafter on every prompt. The MTP heads draft 3 tokens per step; asking for 6 was at most 5% faster. Prose gains the least and varies the most from prompt to prompt. The target model verifies every drafted token, so neither drafter affects output quality. Long context with the BF16 drafter, generating ~400 tokens of code against a large repository prompt: | prompt tokens | no drafter | drafter | speedup | |---|---|---|---| | 64,614 | 25.1 tok/s | 43.8 tok/s | 1.75x | | 130,118 | 21.1 tok/s | 41.4 tok/s | 1.97x | | 248,993 | 16.1 tok/s | 33.6 tok/s | 2.08x | Decode holds up at long context because 39 of the 48 layers use a 128-token sliding window. Prefill does not: a 250K-token prompt takes about 11 minutes on one GB10. Needle retrieval passed at every length and depth tested, from 8K to 350K tokens. ## How to run Supported in [exllamav3](https://github.com/turboderp-org/exllamav3) v1.5.2 and later. Serve it with [TabbyAPI](https://github.com/theroyallab/tabbyAPI). ```sh hf download benthecarman/MiMo-V2.6-Flash-RL-exl3 --local-dir mimo-exl3 ``` There are no prebuilt exllamav3 wheels for aarch64, so on a DGX Spark install it from source. TabbyAPI also needs `pip install uvloop` there. TabbyAPI looks for the drafter by name inside `draft_model_dir`, so put `dflash/` in its own directory next to the model: ``` models/ mimo-2.25bpw-hq/ everything except dflash/, dflash-bf16/ and eval/ mimo-dflash-draft/ the contents of dflash/ ``` `config.yml`: ```yaml model: model_dir: models model_name: mimo-2.25bpw-hq max_seq_len: 262144 cache_size: 262144 cache_mode: FP16 chunk_size: 2048 max_batch_size: 1 reasoning: true reasoning_start_token: "" reasoning_end_token: "" draft_model: draft_mode: model draft_model_dir: models draft_model_name: mimo-dflash-draft draft_cache_mode: FP16 dynamic_draft: true ``` To use the MTP heads instead of the DFlash drafter: ```yaml draft_model: draft_mode: mtp dynamic_draft: true ``` For image input, add `vision: true` under `model:`. The vision tower adds about 1.3 GiB; with it and the DFlash drafter loaded at 262,144 context, a 250K-token prompt still left 15 GiB free. Thinking can be turned off per request with `"chat_template_kwargs": {"enable_thinking": false}`. Tool calls use the `qwen3_coder` format, which TabbyAPI detects automatically. ## Credits * Xiaomi MiMo team: the model and the DFlash drafter (MIT). * [turboderp](https://github.com/turboderp): ExLlamaV3 and the EXL3 format. * [vcruz305](https://github.com/vcruz305): the aarch64 build fixes. * [theroyallab](https://github.com/theroyallab/tabbyAPI): TabbyAPI.