MiMo-V2.6-Pro-RL ARVQ / NVFP4 hybrid
Complete uploaded experimental 5% hot checkpoint: 1325 of 26496 routed experts use NVFP4. Remaining experts use per-expert ARVQ FP4 books and FP16 block scales. Global allocation ranks routing-weighted output-error benefit measured on 65536 training-only calibration tokens. All 69 layers have weights.
The remaining-layer PV pass (layers 21–69) is complete. Layers 1–20 retain the earlier accepted candidates (layer 9 retained initialization). See pv_progress.json for the exact provenance of each uploaded layer. These mixed origins must not be interpreted as a fresh 69-layer hybrid PV campaign.
Layer 69 repair
Expert 7 had an FP16 SwiGLU overflow, including on one training token with its original fitted weights. Halving its up-projection scales and doubling its down-projection scales preserves the mathematical function while reducing the intermediate magnitude; activation quantization/rounding can differ. All 104129 training rows routed to that expert passed a finite-range check (peak product 32892.5). The rebalanced gate/up was held fixed during PV; all other eligible projections remained trainable. Layer 69 used codebook/scales learning rates 0.0048/0.0032 and completed 69 updates over 18006461 training tokens.
Relative MoE validation RMS error: 0.787677 → 0.579308. Audit RMS error: 0.870055 → 0.766571. Export replay matched validation exactly. See evaluation/layer69_recovery.json and pv_reports/. These relative errors normalize by MoE output, not the full residual block.
Independent full-model emulation
On 16368 scored tokens from documents excluded from PV training, validation and checkpoint selection:
- Teacher-to-student KL: 0.824378 nats/token.
- Native source teacher perplexity: 8.799630.
- Hybrid student perplexity: 10.273446.
- Excess NLL: 0.154853 nats/token.
See evaluation/full_model.json for the report and exact test selection. This is calibration-arithmetic emulation, not a serving-runtime parity test. Quality acceptance, SM120 serving, and 1M-context memory fit remain unqualified. The checkpoint is experimental; production_ready=false.
All backbone, MTP, vision, audio, audio-tokenizer and DFlash weights are retained. Calibration and this quality evaluation are text-only. upload_verification.json identifies the snapshot whose full weight hashes were verified.
- Downloads last month
- 849