Baekpica's picture
Publish original-representation MXFP4 BF16 calibration reference
de01a8a verified
|
Raw History Blame
3.47 kB
metadata
license: mit
base_model: XiaomiMiMo/MiMo-V2.6-Flash-RL
base_model_relation: quantized
library_name: gguf
pipeline_tag: text-generation
tags:
  - gguf
  - mimo_v2
  - mxfp4
  - calibration

MiMo-V2.6-Flash-RL GGUF — calibration reference

This repository holds intermediate artifacts actually generated while building MiMo-V2.6-Flash-RL Mixed-Quant GGUF. The final compact mixed variant belongs in that separate repository.

The source is XiaomiMiMo/MiMo-V2.6-Flash-RL, pinned to 3b38d063180c3e4aed9691fdc735f3d10b266ee4.

Reference representation

The four MXFP4-BF16 shards preserve the original routed experts through an exact MXFP4 repack and expand source FP8 dense matrices to BF16. Control tensors are stored as F32. This is not a full-BF16 source checkpoint or a Q8_0 baseline. It provides original-checkpoint values for importance-matrix collection without recalibrating from the final IQ2 weights.

Shard Bytes
MiMo-V2.6-Flash-RL-MXFP4-BF16-00001-of-00004.gguf 44,499,128,352
MiMo-V2.6-Flash-RL-MXFP4-BF16-00002-of-00004.gguf 44,493,179,840
MiMo-V2.6-Flash-RL-MXFP4-BF16-00003-of-00004.gguf 44,493,179,840
MiMo-V2.6-Flash-RL-MXFP4-BF16-00004-of-00004.gguf 41,422,766,592
Total 174,908,254,624

Keep all shards together and open the first shard. The language artifact includes the checkpoint's three embedded MTP blocks; this is not evidence of validated speculative decoding. Multimodal encoders and the separate DFlash model are separate components and are not supplied by these four shards alone.

Validation status

  • All 90 downloaded source repository files passed Hub checksum verification.
  • Independent MXFP4 repacking and tensor-parallel QKV ordering checks passed.
  • All 36,096 expert matrices were covered by a deterministic sampled-row audit: 108,257 rows matched the independently repacked source. See reference-audit.json; this is not an exhaustive payload comparison.
  • The reference loaded on a B300 GPU and passed a short arithmetic decode check.
  • Original-representation text imatrix collection is running as of 2026-09-22.
  • Final mixed-model quality, multimodal end-to-end behavior, MTP/DFlash execution, and DGX Spark serving remain pending. No throughput or benchmark qualification is claimed here.

artifact-manifest.json records exact file sizes and SHA-256 digests. SHA256SUMS can be checked after download:

hf download Baekpica/MiMo-V2.6-Flash-RL-GGUF --local-dir ./MiMo-reference
cd ./MiMo-reference
sha256sum -c SHA256SUMS

Chat template and reproduction

chat_template.jinja is copied byte for byte from the pinned source. Use the original MiMo tokenizer and special-token mapping. Correct image, audio, and video processing additionally requires the corresponding native encoder and input protocol.

The conversion uses a local MXFP4 extension to llama.cpp revision 5836771. Reproduction scripts will accompany the mixed release and private Spark handoff. The reference's approximately 175 GB file size is not the compact DGX Spark target; see the separate mixed model card for that recipe and its current estimates.

License

MIT, inherited from the pinned upstream model.