keys-MiMo-V2.6-Pro-MOPD Jarrelscy ARVQ hybrid-21 (stock)

XiaomiMiMo/MiMo-V2.6-Pro-MOPD in Jarrelscy's ARVQ / NVFP4 hybrid format, NVFP4 hot set widened from 5% to 21% of routed experts. Four DGX Sparks (GB10), TP4, 1M context.

This checkpoint is stock MOPD C3. It is not abliterated.

Recipe (image v4, full-speed defaults): keys-MiMo-V2.6-Pro-MOPD-C3-Jarrelscy-ARVQ-4-DGX-Sparks-1M-Context.

What is in it

Part Source
Non-expert weights (attention, norms, embeddings, LM head, dense layer 0) MOPD, exact. Attention o_proj in FP8 128×128 blocks
MTP draft heads identical in RL and MOPD
5,560 hot routed experts (21%), NVFP4 MOPD, requantized with Jarrelscy's hybrid.quantize
20,936 cold routed experts, ARVQ 2-bit Jarrelscy's published PV-fitted RL encoding, unchanged

Hot set: Jarrelscy's prepare_hybrid.select_hot, count 5,560 instead of 1,325, layers 27–68. See allocation.json and config.json (aqlm_layer_books).

Measured 2026-09-28 (image v4, MTP k=2, thinking off)

Live TP4, idle endpoint. Speed: 512 new tokens, temperature 1.0. Refusal/cyber: greedy, 192 tokens.

Prose 24.2 tok/s (1.91 tok/pass)
Code 32.2 tok/s (2.60 tok/pass)
4 requests, aggregate 46.3 tok/s
Prefill 9.5K / 38K 883 / 1074 tok/s
Refusal 32 5/32 bypass (stock MOPD)
Cyber 22 6/22 bypass

Tool-call repetition (435 Hermes turns, server defaults): 7.4% duplicate-call turns, 0.9% flood (32+ calls). GSM8K 96.8 / HumanEval 93.9 / MMLU-Pro 77.4.

Weights ~83.2 GiB/rank. Measured KV pool 1,282,005 tokens (still covers 1M context).

Run (full-speed v4 recipe)

hf download drowzeys/keys-MiMo-V2.6-Pro-MOPD-Jarrelscy-ARVQ-hybrid21 --local-dir /path/to/mimo-mopd
git clone https://github.com/drowzeys/keys-MiMo-V2.6-Pro-MOPD-C3-Jarrelscy-ARVQ-4-DGX-Sparks-1M-Context
export MASTER_ADDR=<rank-0 IP>
# ranks 1-3 first, then rank 0. GPU memory fraction is fixed at 0.85.
bash serve/launch-rank.sh <node-ip> <1|2|3> <RoCE-GID-index> /path/to/mimo-mopd headless
bash serve/launch-rank.sh <node-ip> 0 <RoCE-GID-index> /path/to/mimo-mopd api

Defaults: MTP k=2 (all three heads, non-chain), torch.compile + CUDA graphs, batched ARVQ prefill, 1M context, 4 seqs, both CX-7 RoCE paths.

Credit

Jarrelscy format, codebooks, PV fits, and vLLM fork. Xiaomi MOPD/RL weights (MIT).

Downloads last month
431
Safetensors
Model size
198B params
Tensor type
BF16
·
F32
·
F8_E4M3
·
U32
·
F16
·
I8
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for drowzeys/keys-MiMo-V2.6-Pro-MOPD-Jarrelscy-ARVQ-hybrid21

Quantized
(3)
this model