Xiaomi MiMo

MiMo-V2.6-Flash-RL — unpruned MLX 4-bit with vision, audio & MTP

A high-quality MiMo for Apple Silicon that understands images, video, and audio, retains all 256 experts, and delivers 59.4 tokens/s in measured text decode on an M3 Ultra.

Original model · Multimodal runtime · Native MTP runtime · sayyidfareed

Why this quant

This is an unpruned MiMo-V2.6-Flash-RL conversion for Apple Silicon, with the input encoders and draft heads bundled alongside the text model. The 4-bit MLX model keeps all of Xiaomi's routed experts and comes with the vision encoder, audio encoder, audio tokenizer, native three-layer MTP head, and five-layer DFlash drafter. One download contains the weights needed for text, image, sampled-video, and audio understanding in a compatible runtime.

The dense projections are 4-bit affine (group size 64). MiMo's experts were already trained in MXFP4 (group size 32), so their quantized format is preserved rather than requantized. The main model averages 4.257 bits per weight. Fused attention weights are reconstructed in the source checkpoint's tensor-parallel order before conversion so the resulting model generates correctly.

Compared with mlx-community's text-only mxfp4-q8, this repository includes the vision and audio weights and both draft heads. That conversion uses 8-bit attention and dense layers; this one uses 4-bit affine dense projections. REAP50 fits a 128 GB Mac by pruning half the experts and carries over this conversion's auxiliary weights. This release is for the 256 GB Mac that can run all 256 experts per MoE layer, with the input encoders and draft heads ready in the same download.

Measured on a Mac Studio

Test Result
Chip / memory M3 Ultra / 256 GB unified memory
Serial text decode, 128 tokens 59.4 tok/s average
Text prefill, 512 tokens 477.1 tok/s
Text prefill, 2,048 tokens 562.8 tok/s
Peak memory, short / 2,048-token prompt 164.3 / 166.8 GB
Main text model on disk about 154 GiB
Complete repository on disk about 160 GiB

Measured for serial text generation with oMLX 0.7.0.dev2, MLX 0.32.2, and mlx-lm 0.31.3. The model also answered an image question in an ongoing conversation, read labels in a two-frame video, transcribed The secret phrase is silver moonlight. exactly, and understood Xiaomi's published audio sample on the M3 Ultra. Native MTP and DFlash weights are included; the figures above measure the base text decode path.

Get the model running

A 256 GB Apple Silicon Mac is recommended to leave room for the model, KV cache, and the rest of the system. MiMo is configured for up to one million tokens of context; memory needs rise with the length of the conversation.

hf download sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP \
  --local-dir ./MiMo-V2.6-Flash-RL-MLX-4bit-MTP

Text generation works with a compatible MLX-LM installation (0.31.3 was used for the original text test):

python -m mlx_lm generate \
  --model ./MiMo-V2.6-Flash-RL-MLX-4bit-MTP \
  --prompt "Write a Python function that checks whether an integer is prime." \
  --max-tokens 256 --temp 0.6

For image, video, and audio input, use an oMLX build with MiMo multimodal support, including the multi-turn image fix. For native MTP on text, add the separate MTP runtime changes. Both runtime PRs are currently open; standard MLX-LM runs the text model.

A compatible oMLX server accepts standard OpenAI-style image parts:

{
  "model": "MiMo-V2.6-Flash-RL-MLX-4bit-MTP",
  "messages": [{
    "role": "user",
    "content": [
      {"type": "text", "text": "What is in this picture?"},
      {"type": "image_url", "image_url": {"url": "data:image/png;base64,<base64 image>"}}
    ]
  }],
  "max_tokens": 128
}

For speech, use {"type":"input_audio","input_audio":{"data":"<base64 wav>","format":"wav"}} with a transcription question. Video input is sampled into ordered vision frames; its audio track is not processed by that path.

What is in the download

Component Format File
Language model 4-bit affine dense projections; native MXFP4 experts Root model-*.safetensors
Native MTP Three-layer, 4-bit affine mtp/model_mtp.safetensors
DFlash Five-layer, upstream BF16 dflash/model.safetensors
Vision encoder Upstream BF16 omnimodal/vision_encoder.safetensors
Audio encoder and bridge Upstream BF16 omnimodal/audio_encoder.safetensors
Audio tokenizer Upstream weights audio_tokenizer/model.safetensors

The sidecars are outside the root text-model index, so MLX-LM loads text without loading them. omnimodal/manifest.json records their source. With the compatible oMLX runtime, native MTP drafts for text-only requests. Image, video, and audio requests use normal decoding, keeping their external embeddings on the correct path. DFlash weights are included for compatible future runtimes.

MiMo-V2.6-Flash-RL has 48 transformer layers, 256 routed experts with eight active per token, and a configured one-million-token context. See Xiaomi's model card for its evaluations, architecture, intended uses, and limits.

Scope and credit

This runtime handles image understanding, sampled video frames, and audio input. Video frames are processed without the video's audio track; audio input currently supports one prompt batch per request. The linked oMLX changes are under review, so use those branches for the multimodal and MTP paths described above.

Xiaomi's MiMo team designed and trained the model and released it under MIT. The MLX conversion and Apple Silicon integration are by sayyidfareed; the original converted weights were published at Vontra.

@misc{mimo2026v26flash,
  title={MiMo-V2.6-Flash-RL},
  author={{Xiaomi MiMo Team}},
  year={2026},
  howpublished={\url{https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL}},
}

Follow sayyidfareed for Apple Silicon model updates.

Downloads last month
270
Safetensors
Model size
309B params
Tensor type
U32
·
BF16
·
F32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP

Quantized
(37)
this model