Instructions to use primitive-ai/MiMo-V2.6-Flash-RL-NVFP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use primitive-ai/MiMo-V2.6-Flash-RL-NVFP4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="primitive-ai/MiMo-V2.6-Flash-RL-NVFP4", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("primitive-ai/MiMo-V2.6-Flash-RL-NVFP4", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use primitive-ai/MiMo-V2.6-Flash-RL-NVFP4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "primitive-ai/MiMo-V2.6-Flash-RL-NVFP4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "primitive-ai/MiMo-V2.6-Flash-RL-NVFP4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/primitive-ai/MiMo-V2.6-Flash-RL-NVFP4
- SGLang
How to use primitive-ai/MiMo-V2.6-Flash-RL-NVFP4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "primitive-ai/MiMo-V2.6-Flash-RL-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "primitive-ai/MiMo-V2.6-Flash-RL-NVFP4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "primitive-ai/MiMo-V2.6-Flash-RL-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "primitive-ai/MiMo-V2.6-Flash-RL-NVFP4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use primitive-ai/MiMo-V2.6-Flash-RL-NVFP4 with Docker Model Runner:
docker model run hf.co/primitive-ai/MiMo-V2.6-Flash-RL-NVFP4
Xiaomi's own 4-bit experts, re-encoded for Blackwell's FP4 tensor cores.
NVFP4 build of XiaomiMiMo/MiMo-V2.6-Flash-RL at
182.40 GB, on two 96 GB cards.
Every routed-expert weight keeps its released value; qkv stays FP8 and everything else BF16, as released.
There is no BF16 to start from. Xiaomi ships this model's routed experts as MXFP4 and its attention as FP8, so the release is already a 4-bit checkpoint and it is the reference row in every table below. This build changes the expert format, not the expert values.
Why this quant
- 🔁 The release's weights, bit for bit. All 302,795,194,368 routed-expert values reproduce the released MXFP4 values exactly. Each E8M0 scale over 32 weights becomes two E4M3 scales over 16 against a power-of-two global, and no value needed rounding.
- ⚡ The experts run on native FP4. vLLM serves them with FlashInfer CUTLASS W4A4 on sm_120. Over the full knowledge suite at concurrency 16 that is 532 tok/s against the release's 402.
- 🎯 Tool-calling level with the release. 78.6 over five runs, the release also 78.6. Knowledge 91.1 against 90.3.
- 📏 Activation scales set per layer. FP4 activations need one global scale per layer. The last MoE layer's expert intermediate reaches about 60,000 on the prompts we measured, 22× past what a scale of 1.0 can hold, so a single constant would clip it.
- 🧮 The cost is KV cache. NVFP4 stores 4.5 bits per weight to MXFP4's 4.25, so this build is 9.48 GB larger than the release. On two 96 GB cards that leaves a 250,094-token pool, against 627,110 for the release.
Serve it
hf download primitive-ai/MiMo-V2.6-Flash-RL-NVFP4 --local-dir ./MiMo-V2.6-Flash-RL-NVFP4
SP=/usr/local/lib/python3.12/dist-packages/vllm
docker run --gpus all --ipc=host --shm-size 32g -p 8000:8000 -v $PWD:/models \
-v $PWD/MiMo-V2.6-Flash-RL-NVFP4/vllm_patch/mimo_v2.py:$SP/model_executor/models/mimo_v2.py:ro \
vllm/vllm-openai:mimo-v26-x86_64-cu130 \
--model /models/MiMo-V2.6-Flash-RL-NVFP4 \
--tensor-parallel-size 2 --trust-remote-code --generation-config vllm \
--max-model-len 131072 --gpu-memory-utilization 0.95 \
--limit-mm-per-prompt '{"image":4,"video":0,"audio":0}' \
--reasoning-parser mimo --tool-call-parser mimo --enable-auto-tool-choice
The day-0 image (vLLM 0.29.1rc1.dev449) loads ModelOpt mixed-precision checkpoints, but its MiMo
loader assumes the release's own FP8 layout: the fused qkv loader writes to weight_scale_inv,
while ModelOpt's block-FP8 layers name that parameter weight_scale and store it 4-D.
vllm_patch/mimo_v2.py is the image's file with those spots bridged; the 21-line diff is
vllm_patch/mimo_v2.diff. Without the mount, loading stops with a KeyError.
Setting audio to 0 skips loading the audio encoder. The DFlash drafter ships as released, but speculative decoding was not measured with this build.
Measured
Two RTX PRO 6000 Blackwell, 96 GB each, tensor parallel 2. The 1,170-item knowledge suite and the
200-item tool-calling suite, temperature 0.6 / top_p 0.95 / top_k 20, thinking on, a
16,384-token budget, concurrency 16, auto-scored with no LLM judge. bf16 KV, no speculative decoding,
same box, same day.
| build | size | knowledge | tool-calling | call | abstain | finished | tok/s @1 | tok/s @16 |
|---|---|---|---|---|---|---|---|---|
| release, MXFP4 experts | 172.92 GB | 90.3 | 78.6 | 82.8 | 62.0 | 99.0% | 112.2 | 402 |
| this repo, NVFP4 | 182.40 GB | 91.1 | 78.6 | 83.0 | 61.0 | 99.0% | 121.1 | 532 |
Tool-calling is the mean of five runs per build, with a within-build standard deviation of 0.8 to
0.9, so the column is one band. Knowledge is one run per quantized build and two for the
release; gaps under 1.2 are ties. tok/s @1 is 60 real prompts one at a time; tok/s @16 is the
knowledge suite's own aggregate.
Comparable with our other models
Accuracy numbers move for reasons that have nothing to do with the model: a shorter token budget, a
different temperature, or whether the model was allowed to reason at all. So every number in this
table, on this card and on our other cards, comes from the one fixed protocol described above, the
same 1,370 items, auto-scored, no LLM judge.
| model | shape | size | overall | knowledge | call | abstain | finished | out/answer |
|---|---|---|---|---|---|---|---|---|
| Laguna-XS-2.1 | 31 B MoE | 19.3 GiB | 81.7 | 83.8 | 68.4 | 73.5 | 98.9% | 1097 |
| Nemotron-3.5-Lightning-30B-A3B | 30 B MoE+Mamba | 19.2 GiB | 87.1 | 87.9 | 85.4 | 70.5 | 97.9% | 1429 |
| Ornith-1.5-35B-A3B | 35 B MoE | 22.6 GiB | 88.7 | 91.7 | 74.4 | 60.0 | 99.3% | 760 |
| Muse-Glimmer-30B | 30 B MoE | 20.4 GiB | 86.6 | 88.8 | 78.6 | 54.5 | 99.7% | 800 |
| Qwen3.8-27B | 27 B dense | 20.7 GiB | 88.8 | 90.4 | 85.5 | 54.5 | 99.7% | 651 |
| Granite-4.2-30B | 30 B dense | 18.1 GB | 85.5 | 86.2 | 85.8 | 60.8 | 98.5% | 1502 |
| Nex-N2.5-mini NVFP4 | 35 B MoE, 3 B active | 23.91 GB | 88.5 | 90.5 | 81.9 | 57.5 | 99.3% | 524 |
| Nex-N2.5-mini mixed | 35 B MoE, 3 B active | 26.04 GB | 88.5 | 90.6 | 80.9 | 58.7 | 99.2% | 504 |
| Nex-N2.5-mini FP8 | 35 B MoE, 3 B active | 38.13 GB | 88.7 | 90.9 | 79.4 | 60.0 | 99.5% | 545 |
| K2-Horizon-MoVA-36B-A4B NVFP4 | 37 B MoE+MoVA, 4 B active | 36.7 GB | 84.1 | 86.5 | 71.8 | 60.4 | 95.6% | 1234 |
| K2-Horizon-MoVA-36B-A4B mixed | 37 B MoE+MoVA, 4 B active | 44.5 GB | 84.9 | 87.3 | 73.5 | 58.9 | 96.3% | 1118 |
| Laguna-S-2.1 | 110 B MoE | 64.0 GiB | 84.3 | 87.1 | 64.6 | 81.0 | 97.3% | 995 |
| Qwen3.8-Flash-Next | 180 B MoE, 6 B active | 183.7 GB | 90.3 | 92.2 | 84.8 | 56.7 | 99.5% | 686 |
| MiMo-V2.6-Flash-RL NVFP4 (this repo) | 309 B MoE, 15 B active | 182.40 GB | 89.3 | 91.1 | 83.0 | 61.0 | 99.0% | 664 |
overall pools the two suites as 1,370 items, weighted 85.4% knowledge and 14.6% tool calling by
item count. Read it with finished: overall scores an answer that overran the token budget as
wrong, and cannot say whether the model needed the room or failed to stop. A gap under 1.0 in
overall is a tie. Sizes are as each card reports them, which mixes GB and GiB.
What's quantized to what
MiMo-V2.6-Flash-RL is 309 B parameters, 15 B active, and 98% of those parameters are routed experts.
| tensors | count | format |
|---|---|---|
| routed experts, layers 1 to 47 | 36,096 projections | NVFP4: E2M1 codes, E4M3 scale per 16, FP32 global, W4A4 |
qkv_proj ×48, layer-0 dense MLP, MTP layers |
63 modules | FP8 E4M3, 128×128 blocks, byte-identical to the release |
o_proj, embeddings, lm_head, routers, vision and audio encoders |
the rest | BF16, byte-identical to the release |
ModelOpt MIXED_PRECISION, one quantized_layers entry per module; the dense FP8 entries are tagged
FP8_PB_WO, the name vLLM's ModelOpt linear table accepts for block FP8. gate and up of each expert
share one global scale, so the fused projection vLLM builds keeps the exact values. vLLM reduces
expert input scales to one per layer, and the checkpoint stores them that way.
The expert weights are not re-quantized. The activation scales are the only new numbers in this checkpoint.
![]()
primitive ·
more models ·
inference economics for production LLM systems
- Downloads last month
- 236
Model tree for primitive-ai/MiMo-V2.6-Flash-RL-NVFP4
Base model
XiaomiMiMo/MiMo-V2.6-Flash-RL