Image-Text-to-Text
MLX
Safetensors
English
Chinese
mimo_v2_flash
apple-silicon
mimo-v2
mixture-of-experts
multimodal
4-bit precision
mtp
conversational
Instructions to use sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP") config = load_config("sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent
How to use sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Document full MiMo V2.6 MLX checkpoint and M3 Ultra performance
Browse files
README.md
CHANGED
|
@@ -4,7 +4,7 @@ language:
|
|
| 4 |
- en
|
| 5 |
- zh
|
| 6 |
library_name: mlx
|
| 7 |
-
pipeline_tag: text-
|
| 8 |
base_model: XiaomiMiMo/MiMo-V2.6-Flash-RL
|
| 9 |
base_model_relation: quantized
|
| 10 |
tags:
|
|
@@ -12,6 +12,7 @@ tags:
|
|
| 12 |
- apple-silicon
|
| 13 |
- mimo-v2
|
| 14 |
- mixture-of-experts
|
|
|
|
| 15 |
- 4-bit
|
| 16 |
- mtp
|
| 17 |
---
|
|
@@ -23,75 +24,97 @@ tags:
|
|
| 23 |
</picture>
|
| 24 |
</div>
|
| 25 |
|
| 26 |
-
<h1 align="center">MiMo-V2.6-Flash-RL MLX 4-bit MTP</h1>
|
|
|
|
|
|
|
| 27 |
|
| 28 |
<p align="center">
|
| 29 |
-
|
| 30 |
-
<a href="https://
|
| 31 |
-
|
|
|
|
| 32 |
</p>
|
| 33 |
|
| 34 |
-
##
|
| 35 |
-
|
| 36 |
-
This is the text backbone of MiMo-V2.6-Flash-RL converted for MLX. The dense projections use 4-bit affine quantization with group size 64, while the model's native MXFP4 MoE experts remain in their original group-size-32 format. The resulting main model averages 4.257 bits per weight.
|
| 37 |
-
|
| 38 |
-
The checkpoint includes its native three-layer MTP payload in `mtp/model_mtp.safetensors`, converted to 4-bit affine weights. It also carries the upstream five-layer DFlash drafter, vision encoder, audio encoder, and audio tokenizer so those assets do not need a second download.
|
| 39 |
|
| 40 |
-
|
| 41 |
-
| --- | --- | --- |
|
| 42 |
-
| MTP predictor | `mtp/model_mtp.safetensors` | MLX 4-bit affine |
|
| 43 |
-
| DFlash drafter | `dflash/model.safetensors` | Upstream BF16 |
|
| 44 |
-
| Vision encoder | `omnimodal/vision_encoder.safetensors` | Upstream BF16 |
|
| 45 |
-
| Audio encoder | `omnimodal/audio_encoder.safetensors` | Upstream BF16 |
|
| 46 |
-
| Audio tokenizer | `audio_tokenizer/model.safetensors` | Upstream weights |
|
| 47 |
|
| 48 |
-
|
| 49 |
|
| 50 |
-
|
| 51 |
|
| 52 |
-
|
| 53 |
|
| 54 |
| Test | Result |
|
| 55 |
| --- | ---: |
|
| 56 |
-
|
|
| 57 |
-
|
|
| 58 |
-
|
|
| 59 |
-
|
|
| 60 |
-
| Peak
|
| 61 |
-
|
|
| 62 |
-
| Complete repository
|
|
|
|
|
|
|
| 63 |
|
| 64 |
-
|
| 65 |
|
| 66 |
-
|
| 67 |
|
| 68 |
```bash
|
| 69 |
-
|
|
|
|
|
|
|
| 70 |
|
|
|
|
|
|
|
|
|
|
| 71 |
python -m mlx_lm generate \
|
| 72 |
-
--model
|
| 73 |
--prompt "Write a Python function that checks whether an integer is prime." \
|
| 74 |
-
--max-tokens 256
|
| 75 |
-
--temp 0.6
|
| 76 |
```
|
| 77 |
|
| 78 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 79 |
|
| 80 |
-
|
| 81 |
|
| 82 |
-
|
| 83 |
|
| 84 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 85 |
|
| 86 |
-
The
|
| 87 |
|
| 88 |
-
|
| 89 |
|
| 90 |
-
|
| 91 |
|
| 92 |
-
|
| 93 |
|
| 94 |
-
|
| 95 |
|
| 96 |
```bibtex
|
| 97 |
@misc{mimo2026v26flash,
|
|
@@ -102,4 +125,4 @@ The upstream model is released under the MIT license. All model architecture, tr
|
|
| 102 |
}
|
| 103 |
```
|
| 104 |
|
| 105 |
-
Follow
|
|
|
|
| 4 |
- en
|
| 5 |
- zh
|
| 6 |
library_name: mlx
|
| 7 |
+
pipeline_tag: image-text-to-text
|
| 8 |
base_model: XiaomiMiMo/MiMo-V2.6-Flash-RL
|
| 9 |
base_model_relation: quantized
|
| 10 |
tags:
|
|
|
|
| 12 |
- apple-silicon
|
| 13 |
- mimo-v2
|
| 14 |
- mixture-of-experts
|
| 15 |
+
- multimodal
|
| 16 |
- 4-bit
|
| 17 |
- mtp
|
| 18 |
---
|
|
|
|
| 24 |
</picture>
|
| 25 |
</div>
|
| 26 |
|
| 27 |
+
<h1 align="center">MiMo-V2.6-Flash-RL — unpruned MLX 4-bit with vision, audio & MTP</h1>
|
| 28 |
+
|
| 29 |
+
<p align="center">A high-quality MiMo for Apple Silicon that understands images, video, and audio, retains all 256 experts, and delivers 59.4 tokens/s in measured text decode on an M3 Ultra.</p>
|
| 30 |
|
| 31 |
<p align="center">
|
| 32 |
+
<a href="https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL">Original model</a> ·
|
| 33 |
+
<a href="https://github.com/jundot/omlx/pull/3860">Multimodal runtime</a> ·
|
| 34 |
+
<a href="https://github.com/jundot/omlx/pull/3861">Native MTP runtime</a> ·
|
| 35 |
+
<a href="https://huggingface.co/sayyidfareed">sayyidfareed</a>
|
| 36 |
</p>
|
| 37 |
|
| 38 |
+
## Why this quant
|
|
|
|
|
|
|
|
|
|
|
|
|
| 39 |
|
| 40 |
+
This is an **unpruned MiMo-V2.6-Flash-RL conversion for Apple Silicon**, with the input encoders and draft heads bundled alongside the text model. The 4-bit MLX model keeps all of Xiaomi's routed experts and comes with the vision encoder, audio encoder, audio tokenizer, native three-layer MTP head, and five-layer DFlash drafter. One download contains the weights needed for text, image, sampled-video, and audio understanding in a compatible runtime.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 41 |
|
| 42 |
+
The dense projections are 4-bit affine (group size 64). MiMo's experts were already trained in MXFP4 (group size 32), so their quantized format is preserved rather than requantized. The main model averages **4.257 bits per weight**. Fused attention weights are reconstructed in the source checkpoint's tensor-parallel order before conversion so the resulting model generates correctly.
|
| 43 |
|
| 44 |
+
Compared with [mlx-community's text-only mxfp4-q8](https://huggingface.co/mlx-community/MiMo-V2.6-Flash-RL-mxfp4-q8), this repository includes the vision and audio weights **and** both draft heads. That conversion uses 8-bit attention and dense layers; this one uses 4-bit affine dense projections. [REAP50](https://huggingface.co/tacodevs/MiMo-V2.6-Flash-RL-MLX-REAP50-mxfp4-MTP) fits a 128 GB Mac by pruning half the experts and carries over this conversion's auxiliary weights. This release is for the 256 GB Mac that can run **all 256 experts per MoE layer**, with the input encoders and draft heads ready in the same download.
|
| 45 |
|
| 46 |
+
## Measured on a Mac Studio
|
| 47 |
|
| 48 |
| Test | Result |
|
| 49 |
| --- | ---: |
|
| 50 |
+
| Chip / memory | M3 Ultra / 256 GB unified memory |
|
| 51 |
+
| Serial text decode, 128 tokens | **59.4 tok/s** average |
|
| 52 |
+
| Text prefill, 512 tokens | 477.1 tok/s |
|
| 53 |
+
| Text prefill, 2,048 tokens | 562.8 tok/s |
|
| 54 |
+
| Peak memory, short / 2,048-token prompt | 164.3 / 166.8 GB |
|
| 55 |
+
| Main text model on disk | about 154 GiB |
|
| 56 |
+
| Complete repository on disk | about 160 GiB |
|
| 57 |
+
|
| 58 |
+
Measured for serial text generation with oMLX 0.7.0.dev2, MLX 0.32.2, and mlx-lm 0.31.3. The model also answered an image question in an ongoing conversation, read labels in a two-frame video, transcribed `The secret phrase is silver moonlight.` exactly, and understood Xiaomi's published audio sample on the M3 Ultra. Native MTP and DFlash weights are included; the figures above measure the base text decode path.
|
| 59 |
|
| 60 |
+
## Get the model running
|
| 61 |
|
| 62 |
+
A **256 GB Apple Silicon Mac** is recommended to leave room for the model, KV cache, and the rest of the system. MiMo is configured for up to one million tokens of context; memory needs rise with the length of the conversation.
|
| 63 |
|
| 64 |
```bash
|
| 65 |
+
hf download sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP \
|
| 66 |
+
--local-dir ./MiMo-V2.6-Flash-RL-MLX-4bit-MTP
|
| 67 |
+
```
|
| 68 |
|
| 69 |
+
Text generation works with a compatible MLX-LM installation (0.31.3 was used for the original text test):
|
| 70 |
+
|
| 71 |
+
```bash
|
| 72 |
python -m mlx_lm generate \
|
| 73 |
+
--model ./MiMo-V2.6-Flash-RL-MLX-4bit-MTP \
|
| 74 |
--prompt "Write a Python function that checks whether an integer is prime." \
|
| 75 |
+
--max-tokens 256 --temp 0.6
|
|
|
|
| 76 |
```
|
| 77 |
|
| 78 |
+
For image, video, and audio input, use an oMLX build with [MiMo multimodal support](https://github.com/jundot/omlx/pull/3860), including the [multi-turn image fix](https://github.com/jundot/omlx/pull/3860/commits/8ae555cb60fcd0d33abaece2af720340017232e8). For native MTP on text, add [the separate MTP runtime changes](https://github.com/jundot/omlx/pull/3861). Both runtime PRs are currently open; standard MLX-LM runs the text model.
|
| 79 |
+
|
| 80 |
+
A compatible oMLX server accepts standard OpenAI-style image parts:
|
| 81 |
+
|
| 82 |
+
```json
|
| 83 |
+
{
|
| 84 |
+
"model": "MiMo-V2.6-Flash-RL-MLX-4bit-MTP",
|
| 85 |
+
"messages": [{
|
| 86 |
+
"role": "user",
|
| 87 |
+
"content": [
|
| 88 |
+
{"type": "text", "text": "What is in this picture?"},
|
| 89 |
+
{"type": "image_url", "image_url": {"url": "data:image/png;base64,<base64 image>"}}
|
| 90 |
+
]
|
| 91 |
+
}],
|
| 92 |
+
"max_tokens": 128
|
| 93 |
+
}
|
| 94 |
+
```
|
| 95 |
|
| 96 |
+
For speech, use `{"type":"input_audio","input_audio":{"data":"<base64 wav>","format":"wav"}}` with a transcription question. Video input is sampled into ordered vision frames; its audio track is not processed by that path.
|
| 97 |
|
| 98 |
+
## What is in the download
|
| 99 |
|
| 100 |
+
| Component | Format | File |
|
| 101 |
+
| --- | --- | --- |
|
| 102 |
+
| Language model | 4-bit affine dense projections; native MXFP4 experts | Root `model-*.safetensors` |
|
| 103 |
+
| Native MTP | Three-layer, 4-bit affine | `mtp/model_mtp.safetensors` |
|
| 104 |
+
| DFlash | Five-layer, upstream BF16 | `dflash/model.safetensors` |
|
| 105 |
+
| Vision encoder | Upstream BF16 | `omnimodal/vision_encoder.safetensors` |
|
| 106 |
+
| Audio encoder and bridge | Upstream BF16 | `omnimodal/audio_encoder.safetensors` |
|
| 107 |
+
| Audio tokenizer | Upstream weights | `audio_tokenizer/model.safetensors` |
|
| 108 |
|
| 109 |
+
The sidecars are outside the root text-model index, so MLX-LM loads text without loading them. `omnimodal/manifest.json` records their source. With the compatible oMLX runtime, native MTP drafts for **text-only** requests. Image, video, and audio requests use normal decoding, keeping their external embeddings on the correct path. DFlash weights are included for compatible future runtimes.
|
| 110 |
|
| 111 |
+
MiMo-V2.6-Flash-RL has 48 transformer layers, 256 routed experts with eight active per token, and a configured one-million-token context. See [Xiaomi's model card](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL) for its evaluations, architecture, intended uses, and limits.
|
| 112 |
|
| 113 |
+
## Scope and credit
|
| 114 |
|
| 115 |
+
This runtime handles image understanding, sampled video frames, and audio input. Video frames are processed without the video's audio track; audio input currently supports one prompt batch per request. The linked oMLX changes are under review, so use those branches for the multimodal and MTP paths described above.
|
| 116 |
|
| 117 |
+
Xiaomi's MiMo team designed and trained the model and released it under MIT. The MLX conversion and Apple Silicon integration are by [sayyidfareed](https://huggingface.co/sayyidfareed); the original converted weights were published at [Vontra](https://huggingface.co/Vontra/MiMo-V2.6-Flash-RL-MLX-4bit-MTP).
|
| 118 |
|
| 119 |
```bibtex
|
| 120 |
@misc{mimo2026v26flash,
|
|
|
|
| 125 |
}
|
| 126 |
```
|
| 127 |
|
| 128 |
+
[Follow sayyidfareed for Apple Silicon model updates.](https://huggingface.co/sayyidfareed)
|