MiMo-V2.6-Pro-RL ARVQ / NVFP4 hybrid

Complete uploaded experimental 5% hot checkpoint: 1325 of 26496 routed experts use NVFP4. Remaining experts use per-expert ARVQ FP4 books and FP16 block scales. Global allocation ranks routing-weighted output-error benefit measured on 65536 training-only calibration tokens. All 69 layers have weights.

The remaining-layer PV pass (layers 21–69) is complete. Layers 1–20 retain the earlier accepted candidates (layer 9 retained initialization). See pv_progress.json for the exact provenance of each uploaded layer. These mixed origins must not be interpreted as a fresh 69-layer hybrid PV campaign.

Layer 69 repair

Expert 7 had an FP16 SwiGLU overflow, including on one training token with its original fitted weights. Halving its up-projection scales and doubling its down-projection scales preserves the mathematical function while reducing the intermediate magnitude; activation quantization/rounding can differ. All 104129 training rows routed to that expert passed a finite-range check (peak product 32892.5). The rebalanced gate/up was held fixed during PV; all other eligible projections remained trainable. Layer 69 used codebook/scales learning rates 0.0048/0.0032 and completed 69 updates over 18006461 training tokens.

Relative MoE validation RMS error: 0.787677 → 0.579308. Audit RMS error: 0.870055 → 0.766571. Export replay matched validation exactly. See evaluation/layer69_recovery.json and pv_reports/. These relative errors normalize by MoE output, not the full residual block.

Independent full-model emulation

On 16368 scored tokens from documents excluded from PV training, validation and checkpoint selection:

  • Teacher-to-student KL: 0.824378 nats/token.
  • Native source teacher perplexity: 8.799630.
  • Hybrid student perplexity: 10.273446.
  • Excess NLL: 0.154853 nats/token.

See evaluation/full_model.json for the report and exact test selection. This is calibration-arithmetic emulation, not a serving-runtime parity test. Quality acceptance, SM120 serving, and 1M-context memory fit remain unqualified. The checkpoint is experimental; production_ready=false.

All backbone, MTP, vision, audio, audio-tokenizer and DFlash weights are retained. Calibration and this quality evaluation are text-only. upload_verification.json identifies the snapshot whose full weight hashes were verified.

Downloads last month
849
Safetensors
Model size
119B params
Tensor type
BF16
·
F32
·
F8_E4M3
·
U32
·
F16
·
I8
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jarrelscy/MiMo-V2.6-Pro-RL-ARVQ-hybrid

Quantized
(9)
this model
Quantizations
1 model