Image-Text-to-Text
MLX
Safetensors
English
Chinese
mimo_v2_flash
apple-silicon
mimo-v2
mixture-of-experts
multimodal
4-bit precision
mtp
conversational
Instructions to use sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP") config = load_config("sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent
How to use sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
|
Download README.md from sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP: direct link, hf CLI and curl.
- Browser
- Download file 7.4 kB
-
https://huggingface.co/sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP/resolve/bf4a5edbac9a53b5a2dd896528806ed319ac67d4/README.md
- Command line
-
hf download hf://sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP@bf4a5edbac9a53b5a2dd896528806ed319ac67d4/README.md
-
curl -L -o README.md https://huggingface.co/sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP/resolve/bf4a5edbac9a53b5a2dd896528806ed319ac67d4/README.md
7.4 kB
| license: mit | |
| language: | |
| - en | |
| - zh | |
| library_name: mlx | |
| pipeline_tag: image-text-to-text | |
| base_model: XiaomiMiMo/MiMo-V2.6-Flash-RL | |
| base_model_relation: quantized | |
| tags: | |
| - mlx | |
| - apple-silicon | |
| - mimo-v2 | |
| - mixture-of-experts | |
| - multimodal | |
| - 4-bit | |
| - mtp | |
| <div align="center"> | |
| <picture> | |
| <source srcset="https://github.com/XiaomiMiMo/MiMo/raw/main/figures/Xiaomi_MiMo_darkmode.png?raw=true" media="(prefers-color-scheme: dark)"> | |
| <img src="https://github.com/XiaomiMiMo/MiMo/raw/main/figures/Xiaomi_MiMo.png?raw=true" width="48%" alt="Xiaomi MiMo"> | |
| </picture> | |
| </div> | |
| <h1 align="center">MiMo-V2.6-Flash-RL — unpruned MLX 4-bit with vision, audio & MTP</h1> | |
| <p align="center">A high-quality MiMo for Apple Silicon that understands images, video, and audio, retains all 256 experts, and delivers 59.4 tokens/s in measured text decode on an M3 Ultra.</p> | |
| <p align="center"> | |
| <a href="https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL">Original model</a> · | |
| <a href="https://github.com/jundot/omlx/pull/3860">Multimodal runtime</a> · | |
| <a href="https://github.com/jundot/omlx/pull/3861">Native MTP runtime</a> · | |
| <a href="https://huggingface.co/sayyidfareed">sayyidfareed</a> | |
| </p> | |
| ## Why this quant | |
| This is an **unpruned MiMo-V2.6-Flash-RL conversion for Apple Silicon**, with the input encoders and draft heads bundled alongside the text model. The 4-bit MLX model keeps all of Xiaomi's routed experts and comes with the vision encoder, audio encoder, audio tokenizer, native three-layer MTP head, and five-layer DFlash drafter. One download contains the weights needed for text, image, sampled-video, and audio understanding in a compatible runtime. | |
| The dense projections are 4-bit affine (group size 64). MiMo's experts were already trained in MXFP4 (group size 32), so their quantized format is preserved rather than requantized. The main model averages **4.257 bits per weight**. Fused attention weights are reconstructed in the source checkpoint's tensor-parallel order before conversion so the resulting model generates correctly. | |
| Compared with [mlx-community's text-only mxfp4-q8](https://huggingface.co/mlx-community/MiMo-V2.6-Flash-RL-mxfp4-q8), this repository includes the vision and audio weights **and** both draft heads. That conversion uses 8-bit attention and dense layers; this one uses 4-bit affine dense projections. [REAP50](https://huggingface.co/tacodevs/MiMo-V2.6-Flash-RL-MLX-REAP50-mxfp4-MTP) fits a 128 GB Mac by pruning half the experts and carries over this conversion's auxiliary weights. This release is for the 256 GB Mac that can run **all 256 experts per MoE layer**, with the input encoders and draft heads ready in the same download. | |
| ## Measured on a Mac Studio | |
| | Test | Result | | |
| | --- | ---: | | |
| | Chip / memory | M3 Ultra / 256 GB unified memory | | |
| | Serial text decode, 128 tokens | **59.4 tok/s** average | | |
| | Text prefill, 512 tokens | 477.1 tok/s | | |
| | Text prefill, 2,048 tokens | 562.8 tok/s | | |
| | Peak memory, short / 2,048-token prompt | 164.3 / 166.8 GB | | |
| | Main text model on disk | about 154 GiB | | |
| | Complete repository on disk | about 160 GiB | | |
| Measured for serial text generation with oMLX 0.7.0.dev2, MLX 0.32.2, and mlx-lm 0.31.3. The model also answered an image question in an ongoing conversation, read labels in a two-frame video, transcribed `The secret phrase is silver moonlight.` exactly, and understood Xiaomi's published audio sample on the M3 Ultra. Native MTP and DFlash weights are included; the figures above measure the base text decode path. | |
| ## Get the model running | |
| A **256 GB Apple Silicon Mac** is recommended to leave room for the model, KV cache, and the rest of the system. MiMo is configured for up to one million tokens of context; memory needs rise with the length of the conversation. | |
| ```bash | |
| hf download sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP \ | |
| --local-dir ./MiMo-V2.6-Flash-RL-MLX-4bit-MTP | |
| ``` | |
| Text generation works with a compatible MLX-LM installation (0.31.3 was used for the original text test): | |
| ```bash | |
| python -m mlx_lm generate \ | |
| --model ./MiMo-V2.6-Flash-RL-MLX-4bit-MTP \ | |
| --prompt "Write a Python function that checks whether an integer is prime." \ | |
| --max-tokens 256 --temp 0.6 | |
| ``` | |
| For image, video, and audio input, use an oMLX build with [MiMo multimodal support](https://github.com/jundot/omlx/pull/3860), including the [multi-turn image fix](https://github.com/jundot/omlx/pull/3860/commits/8ae555cb60fcd0d33abaece2af720340017232e8). For native MTP on text, add [the separate MTP runtime changes](https://github.com/jundot/omlx/pull/3861). Both runtime PRs are currently open; standard MLX-LM runs the text model. | |
| A compatible oMLX server accepts standard OpenAI-style image parts: | |
| ```json | |
| { | |
| "model": "MiMo-V2.6-Flash-RL-MLX-4bit-MTP", | |
| "messages": [{ | |
| "role": "user", | |
| "content": [ | |
| {"type": "text", "text": "What is in this picture?"}, | |
| {"type": "image_url", "image_url": {"url": "data:image/png;base64,<base64 image>"}} | |
| ] | |
| }], | |
| "max_tokens": 128 | |
| } | |
| ``` | |
| For speech, use `{"type":"input_audio","input_audio":{"data":"<base64 wav>","format":"wav"}}` with a transcription question. Video input is sampled into ordered vision frames; its audio track is not processed by that path. | |
| ## What is in the download | |
| | Component | Format | File | | |
| | --- | --- | --- | | |
| | Language model | 4-bit affine dense projections; native MXFP4 experts | Root `model-*.safetensors` | | |
| | Native MTP | Three-layer, 4-bit affine | `mtp/model_mtp.safetensors` | | |
| | DFlash | Five-layer, upstream BF16 | `dflash/model.safetensors` | | |
| | Vision encoder | Upstream BF16 | `omnimodal/vision_encoder.safetensors` | | |
| | Audio encoder and bridge | Upstream BF16 | `omnimodal/audio_encoder.safetensors` | | |
| | Audio tokenizer | Upstream weights | `audio_tokenizer/model.safetensors` | | |
| The sidecars are outside the root text-model index, so MLX-LM loads text without loading them. `omnimodal/manifest.json` records their source. With the compatible oMLX runtime, native MTP drafts for **text-only** requests. Image, video, and audio requests use normal decoding, keeping their external embeddings on the correct path. DFlash weights are included for compatible future runtimes. | |
| MiMo-V2.6-Flash-RL has 48 transformer layers, 256 routed experts with eight active per token, and a configured one-million-token context. See [Xiaomi's model card](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL) for its evaluations, architecture, intended uses, and limits. | |
| ## Scope and credit | |
| This runtime handles image understanding, sampled video frames, and audio input. Video frames are processed without the video's audio track; audio input currently supports one prompt batch per request. The linked oMLX changes are under review, so use those branches for the multimodal and MTP paths described above. | |
| Xiaomi's MiMo team designed and trained the model and released it under MIT. The MLX conversion and Apple Silicon integration are by [sayyidfareed](https://huggingface.co/sayyidfareed); the original converted weights were published at [Vontra](https://huggingface.co/Vontra/MiMo-V2.6-Flash-RL-MLX-4bit-MTP). | |
| ```bibtex | |
| @misc{mimo2026v26flash, | |
| title={MiMo-V2.6-Flash-RL}, | |
| author={{Xiaomi MiMo Team}}, | |
| year={2026}, | |
| howpublished={\url{https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL}}, | |
| } | |
| ``` | |
| [Follow sayyidfareed for Apple Silicon model updates.](https://huggingface.co/sayyidfareed) | |