Instructions to use sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP") config = load_config("sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent
How to use sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Configure Hermes
# Install Hermes:
curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash
hermes setup# Point Hermes at the local server:
hermes config set model.provider custom
hermes config set model.base_url http://127.0.0.1:8080/v1
hermes config set model.default sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTPRun Hermes
hermes
MiMo-V2.6-Flash-RL — unpruned MLX 4-bit with vision, audio & MTP
A high-quality MiMo for Apple Silicon that understands images, video, and audio, retains all 256 experts, and delivers 59.4 tokens/s in measured text decode on an M3 Ultra.
Original model · Multimodal runtime · Native MTP runtime · sayyidfareed
Why this quant
This is an unpruned MiMo-V2.6-Flash-RL conversion for Apple Silicon, with the input encoders and draft heads bundled alongside the text model. The 4-bit MLX model keeps all of Xiaomi's routed experts and comes with the vision encoder, audio encoder, audio tokenizer, native three-layer MTP head, and five-layer DFlash drafter. One download contains the weights needed for text, image, sampled-video, and audio understanding in a compatible runtime.
The dense projections are 4-bit affine (group size 64). MiMo's experts were already trained in MXFP4 (group size 32), so their quantized format is preserved rather than requantized. The main model averages 4.257 bits per weight. Fused attention weights are reconstructed in the source checkpoint's tensor-parallel order before conversion so the resulting model generates correctly.
Compared with mlx-community's text-only mxfp4-q8, this repository includes the vision and audio weights and both draft heads. That conversion uses 8-bit attention and dense layers; this one uses 4-bit affine dense projections. REAP50 fits a 128 GB Mac by pruning half the experts and carries over this conversion's auxiliary weights. This release is for the 256 GB Mac that can run all 256 experts per MoE layer, with the input encoders and draft heads ready in the same download.
Measured on a Mac Studio
| Test | Result |
|---|---|
| Chip / memory | M3 Ultra / 256 GB unified memory |
| Serial text decode, 128 tokens | 59.4 tok/s average |
| Text prefill, 512 tokens | 477.1 tok/s |
| Text prefill, 2,048 tokens | 562.8 tok/s |
| Peak memory, short / 2,048-token prompt | 164.3 / 166.8 GB |
| Main text model on disk | about 154 GiB |
| Complete repository on disk | about 160 GiB |
Measured for serial text generation with oMLX 0.7.0.dev2, MLX 0.32.2, and mlx-lm 0.31.3. The model also answered an image question in an ongoing conversation, read labels in a two-frame video, transcribed The secret phrase is silver moonlight. exactly, and understood Xiaomi's published audio sample on the M3 Ultra. Native MTP and DFlash weights are included; the figures above measure the base text decode path.
Get the model running
A 256 GB Apple Silicon Mac is recommended to leave room for the model, KV cache, and the rest of the system. MiMo is configured for up to one million tokens of context; memory needs rise with the length of the conversation.
hf download sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP \
--local-dir ./MiMo-V2.6-Flash-RL-MLX-4bit-MTP
Text generation works with a compatible MLX-LM installation (0.31.3 was used for the original text test):
python -m mlx_lm generate \
--model ./MiMo-V2.6-Flash-RL-MLX-4bit-MTP \
--prompt "Write a Python function that checks whether an integer is prime." \
--max-tokens 256 --temp 0.6
For image, video, and audio input, use an oMLX build with MiMo multimodal support, including the multi-turn image fix. For native MTP on text, add the separate MTP runtime changes. Both runtime PRs are currently open; standard MLX-LM runs the text model.
A compatible oMLX server accepts standard OpenAI-style image parts:
{
"model": "MiMo-V2.6-Flash-RL-MLX-4bit-MTP",
"messages": [{
"role": "user",
"content": [
{"type": "text", "text": "What is in this picture?"},
{"type": "image_url", "image_url": {"url": "data:image/png;base64,<base64 image>"}}
]
}],
"max_tokens": 128
}
For speech, use {"type":"input_audio","input_audio":{"data":"<base64 wav>","format":"wav"}} with a transcription question. Video input is sampled into ordered vision frames; its audio track is not processed by that path.
What is in the download
| Component | Format | File |
|---|---|---|
| Language model | 4-bit affine dense projections; native MXFP4 experts | Root model-*.safetensors |
| Native MTP | Three-layer, 4-bit affine | mtp/model_mtp.safetensors |
| DFlash | Five-layer, upstream BF16 | dflash/model.safetensors |
| Vision encoder | Upstream BF16 | omnimodal/vision_encoder.safetensors |
| Audio encoder and bridge | Upstream BF16 | omnimodal/audio_encoder.safetensors |
| Audio tokenizer | Upstream weights | audio_tokenizer/model.safetensors |
The sidecars are outside the root text-model index, so MLX-LM loads text without loading them. omnimodal/manifest.json records their source. With the compatible oMLX runtime, native MTP drafts for text-only requests. Image, video, and audio requests use normal decoding, keeping their external embeddings on the correct path. DFlash weights are included for compatible future runtimes.
MiMo-V2.6-Flash-RL has 48 transformer layers, 256 routed experts with eight active per token, and a configured one-million-token context. See Xiaomi's model card for its evaluations, architecture, intended uses, and limits.
Scope and credit
This runtime handles image understanding, sampled video frames, and audio input. Video frames are processed without the video's audio track; audio input currently supports one prompt batch per request. The linked oMLX changes are under review, so use those branches for the multimodal and MTP paths described above.
Xiaomi's MiMo team designed and trained the model and released it under MIT. The MLX conversion and Apple Silicon integration are by sayyidfareed; the original converted weights were published at Vontra.
@misc{mimo2026v26flash,
title={MiMo-V2.6-Flash-RL},
author={{Xiaomi MiMo Team}},
year={2026},
howpublished={\url{https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL}},
}
- Downloads last month
- 270
4-bit
Model tree for sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP
Base model
XiaomiMiMo/MiMo-V2.6-Flash-RL
Start the MLX server
# Install MLX LM: uv tool install mlx-lm# Start a local OpenAI-compatible server: mlx_lm.server --model "sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP"