--- license: mit language: - en - zh library_name: mlx pipeline_tag: image-text-to-text base_model: XiaomiMiMo/MiMo-V2.6-Flash-RL base_model_relation: quantized tags: - mlx - apple-silicon - mimo-v2 - mixture-of-experts - multimodal - 4-bit - mtp ---
Xiaomi MiMo

MiMo-V2.6-Flash-RL — unpruned MLX 4-bit with vision, audio & MTP

A high-quality MiMo for Apple Silicon that understands images, video, and audio, retains all 256 experts, and delivers 59.4 tokens/s in measured text decode on an M3 Ultra.

Original model · Multimodal runtime · Native MTP runtime · sayyidfareed

## Why this quant This is an **unpruned MiMo-V2.6-Flash-RL conversion for Apple Silicon**, with the input encoders and draft heads bundled alongside the text model. The 4-bit MLX model keeps all of Xiaomi's routed experts and comes with the vision encoder, audio encoder, audio tokenizer, native three-layer MTP head, and five-layer DFlash drafter. One download contains the weights needed for text, image, sampled-video, and audio understanding in a compatible runtime. The dense projections are 4-bit affine (group size 64). MiMo's experts were already trained in MXFP4 (group size 32), so their quantized format is preserved rather than requantized. The main model averages **4.257 bits per weight**. Fused attention weights are reconstructed in the source checkpoint's tensor-parallel order before conversion so the resulting model generates correctly. Compared with [mlx-community's text-only mxfp4-q8](https://huggingface.co/mlx-community/MiMo-V2.6-Flash-RL-mxfp4-q8), this repository includes the vision and audio weights **and** both draft heads. That conversion uses 8-bit attention and dense layers; this one uses 4-bit affine dense projections. [REAP50](https://huggingface.co/tacodevs/MiMo-V2.6-Flash-RL-MLX-REAP50-mxfp4-MTP) fits a 128 GB Mac by pruning half the experts and carries over this conversion's auxiliary weights. This release is for the 256 GB Mac that can run **all 256 experts per MoE layer**, with the input encoders and draft heads ready in the same download. ## Measured on a Mac Studio | Test | Result | | --- | ---: | | Chip / memory | M3 Ultra / 256 GB unified memory | | Serial text decode, 128 tokens | **59.4 tok/s** average | | Text prefill, 512 tokens | 477.1 tok/s | | Text prefill, 2,048 tokens | 562.8 tok/s | | Peak memory, short / 2,048-token prompt | 164.3 / 166.8 GB | | Main text model on disk | about 154 GiB | | Complete repository on disk | about 160 GiB | Measured for serial text generation with oMLX 0.7.0.dev2, MLX 0.32.2, and mlx-lm 0.31.3. The model also answered an image question in an ongoing conversation, read labels in a two-frame video, transcribed `The secret phrase is silver moonlight.` exactly, and understood Xiaomi's published audio sample on the M3 Ultra. Native MTP and DFlash weights are included; the figures above measure the base text decode path. ## Get the model running A **256 GB Apple Silicon Mac** is recommended to leave room for the model, KV cache, and the rest of the system. MiMo is configured for up to one million tokens of context; memory needs rise with the length of the conversation. ```bash hf download sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP \ --local-dir ./MiMo-V2.6-Flash-RL-MLX-4bit-MTP ``` Text generation works with a compatible MLX-LM installation (0.31.3 was used for the original text test): ```bash python -m mlx_lm generate \ --model ./MiMo-V2.6-Flash-RL-MLX-4bit-MTP \ --prompt "Write a Python function that checks whether an integer is prime." \ --max-tokens 256 --temp 0.6 ``` For image, video, and audio input, use an oMLX build with [MiMo multimodal support](https://github.com/jundot/omlx/pull/3860), including the [multi-turn image fix](https://github.com/jundot/omlx/pull/3860/commits/8ae555cb60fcd0d33abaece2af720340017232e8). For native MTP on text, add [the separate MTP runtime changes](https://github.com/jundot/omlx/pull/3861). Both runtime PRs are currently open; standard MLX-LM runs the text model. A compatible oMLX server accepts standard OpenAI-style image parts: ```json { "model": "MiMo-V2.6-Flash-RL-MLX-4bit-MTP", "messages": [{ "role": "user", "content": [ {"type": "text", "text": "What is in this picture?"}, {"type": "image_url", "image_url": {"url": "data:image/png;base64,"}} ] }], "max_tokens": 128 } ``` For speech, use `{"type":"input_audio","input_audio":{"data":"","format":"wav"}}` with a transcription question. Video input is sampled into ordered vision frames; its audio track is not processed by that path. ## What is in the download | Component | Format | File | | --- | --- | --- | | Language model | 4-bit affine dense projections; native MXFP4 experts | Root `model-*.safetensors` | | Native MTP | Three-layer, 4-bit affine | `mtp/model_mtp.safetensors` | | DFlash | Five-layer, upstream BF16 | `dflash/model.safetensors` | | Vision encoder | Upstream BF16 | `omnimodal/vision_encoder.safetensors` | | Audio encoder and bridge | Upstream BF16 | `omnimodal/audio_encoder.safetensors` | | Audio tokenizer | Upstream weights | `audio_tokenizer/model.safetensors` | The sidecars are outside the root text-model index, so MLX-LM loads text without loading them. `omnimodal/manifest.json` records their source. With the compatible oMLX runtime, native MTP drafts for **text-only** requests. Image, video, and audio requests use normal decoding, keeping their external embeddings on the correct path. DFlash weights are included for compatible future runtimes. MiMo-V2.6-Flash-RL has 48 transformer layers, 256 routed experts with eight active per token, and a configured one-million-token context. See [Xiaomi's model card](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL) for its evaluations, architecture, intended uses, and limits. ## Scope and credit This runtime handles image understanding, sampled video frames, and audio input. Video frames are processed without the video's audio track; audio input currently supports one prompt batch per request. The linked oMLX changes are under review, so use those branches for the multimodal and MTP paths described above. Xiaomi's MiMo team designed and trained the model and released it under MIT. The MLX conversion and Apple Silicon integration are by [sayyidfareed](https://huggingface.co/sayyidfareed); the original converted weights were published at [Vontra](https://huggingface.co/Vontra/MiMo-V2.6-Flash-RL-MLX-4bit-MTP). ```bibtex @misc{mimo2026v26flash, title={MiMo-V2.6-Flash-RL}, author={{Xiaomi MiMo Team}}, year={2026}, howpublished={\url{https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL}}, } ``` [Follow sayyidfareed for Apple Silicon model updates.](https://huggingface.co/sayyidfareed)