--- license: mit language: - en - zh library_name: mlx pipeline_tag: image-text-to-text base_model: XiaomiMiMo/MiMo-V2.6-Flash-RL base_model_relation: quantized tags: - mlx - apple-silicon - mimo-v2 - mixture-of-experts - multimodal - 4-bit - mtp ---
A high-quality MiMo for Apple Silicon that understands images, video, and audio, retains all 256 experts, and delivers 59.4 tokens/s in measured text decode on an M3 Ultra.
Original model · Multimodal runtime · Native MTP runtime · sayyidfareed
## Why this quant This is an **unpruned MiMo-V2.6-Flash-RL conversion for Apple Silicon**, with the input encoders and draft heads bundled alongside the text model. The 4-bit MLX model keeps all of Xiaomi's routed experts and comes with the vision encoder, audio encoder, audio tokenizer, native three-layer MTP head, and five-layer DFlash drafter. One download contains the weights needed for text, image, sampled-video, and audio understanding in a compatible runtime. The dense projections are 4-bit affine (group size 64). MiMo's experts were already trained in MXFP4 (group size 32), so their quantized format is preserved rather than requantized. The main model averages **4.257 bits per weight**. Fused attention weights are reconstructed in the source checkpoint's tensor-parallel order before conversion so the resulting model generates correctly. Compared with [mlx-community's text-only mxfp4-q8](https://huggingface.co/mlx-community/MiMo-V2.6-Flash-RL-mxfp4-q8), this repository includes the vision and audio weights **and** both draft heads. That conversion uses 8-bit attention and dense layers; this one uses 4-bit affine dense projections. [REAP50](https://huggingface.co/tacodevs/MiMo-V2.6-Flash-RL-MLX-REAP50-mxfp4-MTP) fits a 128 GB Mac by pruning half the experts and carries over this conversion's auxiliary weights. This release is for the 256 GB Mac that can run **all 256 experts per MoE layer**, with the input encoders and draft heads ready in the same download. ## Measured on a Mac Studio | Test | Result | | --- | ---: | | Chip / memory | M3 Ultra / 256 GB unified memory | | Serial text decode, 128 tokens | **59.4 tok/s** average | | Text prefill, 512 tokens | 477.1 tok/s | | Text prefill, 2,048 tokens | 562.8 tok/s | | Peak memory, short / 2,048-token prompt | 164.3 / 166.8 GB | | Main text model on disk | about 154 GiB | | Complete repository on disk | about 160 GiB | Measured for serial text generation with oMLX 0.7.0.dev2, MLX 0.32.2, and mlx-lm 0.31.3. The model also answered an image question in an ongoing conversation, read labels in a two-frame video, transcribed `The secret phrase is silver moonlight.` exactly, and understood Xiaomi's published audio sample on the M3 Ultra. Native MTP and DFlash weights are included; the figures above measure the base text decode path. ## Get the model running A **256 GB Apple Silicon Mac** is recommended to leave room for the model, KV cache, and the rest of the system. MiMo is configured for up to one million tokens of context; memory needs rise with the length of the conversation. ```bash hf download sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP \ --local-dir ./MiMo-V2.6-Flash-RL-MLX-4bit-MTP ``` Text generation works with a compatible MLX-LM installation (0.31.3 was used for the original text test): ```bash python -m mlx_lm generate \ --model ./MiMo-V2.6-Flash-RL-MLX-4bit-MTP \ --prompt "Write a Python function that checks whether an integer is prime." \ --max-tokens 256 --temp 0.6 ``` For image, video, and audio input, use an oMLX build with [MiMo multimodal support](https://github.com/jundot/omlx/pull/3860), including the [multi-turn image fix](https://github.com/jundot/omlx/pull/3860/commits/8ae555cb60fcd0d33abaece2af720340017232e8). For native MTP on text, add [the separate MTP runtime changes](https://github.com/jundot/omlx/pull/3861). Both runtime PRs are currently open; standard MLX-LM runs the text model. A compatible oMLX server accepts standard OpenAI-style image parts: ```json { "model": "MiMo-V2.6-Flash-RL-MLX-4bit-MTP", "messages": [{ "role": "user", "content": [ {"type": "text", "text": "What is in this picture?"}, {"type": "image_url", "image_url": {"url": "data:image/png;base64,