sayyidfareed commited on
Commit
bf4a5ed
·
verified ·
1 Parent(s): 9ca55c6

Document full MiMo V2.6 MLX checkpoint and M3 Ultra performance

Browse files
Files changed (1) hide show
  1. README.md +66 -43
README.md CHANGED
@@ -4,7 +4,7 @@ language:
4
  - en
5
  - zh
6
  library_name: mlx
7
- pipeline_tag: text-generation
8
  base_model: XiaomiMiMo/MiMo-V2.6-Flash-RL
9
  base_model_relation: quantized
10
  tags:
@@ -12,6 +12,7 @@ tags:
12
  - apple-silicon
13
  - mimo-v2
14
  - mixture-of-experts
 
15
  - 4-bit
16
  - mtp
17
  ---
@@ -23,75 +24,97 @@ tags:
23
  </picture>
24
  </div>
25
 
26
- <h1 align="center">MiMo-V2.6-Flash-RL MLX 4-bit MTP</h1>
 
 
27
 
28
  <p align="center">
29
- A tested Apple Silicon conversion of
30
- <a href="https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL">XiaomiMiMo/MiMo-V2.6-Flash-RL</a>,
31
- published by <a href="https://huggingface.co/Vontra">Vontra</a>.
 
32
  </p>
33
 
34
- ## What is in this release
35
-
36
- This is the text backbone of MiMo-V2.6-Flash-RL converted for MLX. The dense projections use 4-bit affine quantization with group size 64, while the model's native MXFP4 MoE experts remain in their original group-size-32 format. The resulting main model averages 4.257 bits per weight.
37
-
38
- The checkpoint includes its native three-layer MTP payload in `mtp/model_mtp.safetensors`, converted to 4-bit affine weights. It also carries the upstream five-layer DFlash drafter, vision encoder, audio encoder, and audio tokenizer so those assets do not need a second download.
39
 
40
- | Native component | Path | Format |
41
- | --- | --- | --- |
42
- | MTP predictor | `mtp/model_mtp.safetensors` | MLX 4-bit affine |
43
- | DFlash drafter | `dflash/model.safetensors` | Upstream BF16 |
44
- | Vision encoder | `omnimodal/vision_encoder.safetensors` | Upstream BF16 |
45
- | Audio encoder | `omnimodal/audio_encoder.safetensors` | Upstream BF16 |
46
- | Audio tokenizer | `audio_tokenizer/model.safetensors` | Upstream weights |
47
 
48
- Current oMLX and MLX text generation run the target model correctly but do not automatically execute MiMo's MTP, DFlash, vision, or audio paths. The speed figures below are serial text decode measurements. The auxiliary tensors and configs are packaged for MiMo-aware runtimes and ongoing MLX integration, not advertised as working oMLX controls.
49
 
50
- ## Measured on Apple Silicon
51
 
52
- Tested on a 256 GB M3 Ultra Mac Studio with oMLX 0.7.0.dev2, MLX 0.32.2, and mlx-lm 0.31.3.
53
 
54
  | Test | Result |
55
  | --- | ---: |
56
- | Sustained generation, 128-token decode | 59.4 tok/s average |
57
- | Prompt processing, 512 tokens | 477.1 tok/s |
58
- | Prompt processing, 2,048 tokens | 562.8 tok/s |
59
- | Peak unified memory, short context | 164.3 GB |
60
- | Peak unified memory, 2,048-token prompt | 166.8 GB |
61
- | Quantized text model size on disk | about 154 GiB |
62
- | Complete repository size | about 160 GiB |
 
 
63
 
64
- The model produced correct arithmetic, a clear factual explanation, and coherent Python in repeated smoke tests. A 256 GB Mac is recommended so there is room for the model, KV cache, and the rest of the system.
65
 
66
- ## Run with mlx-lm
67
 
68
  ```bash
69
- pip install -U "mlx-lm>=0.31.3"
 
 
70
 
 
 
 
71
  python -m mlx_lm generate \
72
- --model Vontra/MiMo-V2.6-Flash-RL-MLX-4bit-MTP \
73
  --prompt "Write a Python function that checks whether an integer is prime." \
74
- --max-tokens 256 \
75
- --temp 0.6
76
  ```
77
 
78
- The upstream tokenizer chat template is included. In oMLX, download or select `Vontra/MiMo-V2.6-Flash-RL-MLX-4bit-MTP` as an MLX model.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
79
 
80
- ## Quantization notes
81
 
82
- MiMo-V2.6 stores fused attention tensors in checkpoint tensor-parallel order and pads the FP8 scale grid separately for each shard. This conversion reconstructs those shards before quantization. Skipping that step produces a model that loads but returns broken output.
83
 
84
- The quantized MTP payload lives in its own `mtp/` directory with a manifest describing its tensors. It is not a standalone drafter for `mlx_lm.generate --draft-model` today. The upstream DFlash payload retains its trained mask embedding and corrected JSON config; an MLX smoke test matched serial greedy output, but it did not beat serial decode in the current experimental runtime.
 
 
 
 
 
 
 
85
 
86
- The vision and audio tensors are kept outside the root text-model index so `mlx_lm` and oMLX continue to load the tested text checkpoint unchanged. `omnimodal/manifest.json` records every auxiliary path and its upstream source.
87
 
88
- ## About MiMo-V2.6-Flash-RL
89
 
90
- Xiaomi describes MiMo-V2.6-Flash-RL as a sparse 309B-parameter MoE with 15B active parameters, 48 transformer layers, 256 routed experts, and eight active experts per token. The full upstream release supports a one-million-token context and omnimodal inputs. See the [original model card](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL) for architecture details, evaluations, deployment recipes, intended use, and limitations.
91
 
92
- ## License and credit
93
 
94
- The upstream model is released under the MIT license. All model architecture, training, tokenizer work, and original branding belong to the Xiaomi MiMo team. This repository contains a community MLX conversion and measured Apple Silicon results.
95
 
96
  ```bibtex
97
  @misc{mimo2026v26flash,
@@ -102,4 +125,4 @@ The upstream model is released under the MIT license. All model architecture, tr
102
  }
103
  ```
104
 
105
- Follow [Vontra](https://huggingface.co/Vontra) for new Apple Silicon releases and fixes.
 
4
  - en
5
  - zh
6
  library_name: mlx
7
+ pipeline_tag: image-text-to-text
8
  base_model: XiaomiMiMo/MiMo-V2.6-Flash-RL
9
  base_model_relation: quantized
10
  tags:
 
12
  - apple-silicon
13
  - mimo-v2
14
  - mixture-of-experts
15
+ - multimodal
16
  - 4-bit
17
  - mtp
18
  ---
 
24
  </picture>
25
  </div>
26
 
27
+ <h1 align="center">MiMo-V2.6-Flash-RL — unpruned MLX 4-bit with vision, audio & MTP</h1>
28
+
29
+ <p align="center">A high-quality MiMo for Apple Silicon that understands images, video, and audio, retains all 256 experts, and delivers 59.4 tokens/s in measured text decode on an M3 Ultra.</p>
30
 
31
  <p align="center">
32
+ <a href="https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL">Original model</a> ·
33
+ <a href="https://github.com/jundot/omlx/pull/3860">Multimodal runtime</a> ·
34
+ <a href="https://github.com/jundot/omlx/pull/3861">Native MTP runtime</a> ·
35
+ <a href="https://huggingface.co/sayyidfareed">sayyidfareed</a>
36
  </p>
37
 
38
+ ## Why this quant
 
 
 
 
39
 
40
+ This is an **unpruned MiMo-V2.6-Flash-RL conversion for Apple Silicon**, with the input encoders and draft heads bundled alongside the text model. The 4-bit MLX model keeps all of Xiaomi's routed experts and comes with the vision encoder, audio encoder, audio tokenizer, native three-layer MTP head, and five-layer DFlash drafter. One download contains the weights needed for text, image, sampled-video, and audio understanding in a compatible runtime.
 
 
 
 
 
 
41
 
42
+ The dense projections are 4-bit affine (group size 64). MiMo's experts were already trained in MXFP4 (group size 32), so their quantized format is preserved rather than requantized. The main model averages **4.257 bits per weight**. Fused attention weights are reconstructed in the source checkpoint's tensor-parallel order before conversion so the resulting model generates correctly.
43
 
44
+ Compared with [mlx-community's text-only mxfp4-q8](https://huggingface.co/mlx-community/MiMo-V2.6-Flash-RL-mxfp4-q8), this repository includes the vision and audio weights **and** both draft heads. That conversion uses 8-bit attention and dense layers; this one uses 4-bit affine dense projections. [REAP50](https://huggingface.co/tacodevs/MiMo-V2.6-Flash-RL-MLX-REAP50-mxfp4-MTP) fits a 128 GB Mac by pruning half the experts and carries over this conversion's auxiliary weights. This release is for the 256 GB Mac that can run **all 256 experts per MoE layer**, with the input encoders and draft heads ready in the same download.
45
 
46
+ ## Measured on a Mac Studio
47
 
48
  | Test | Result |
49
  | --- | ---: |
50
+ | Chip / memory | M3 Ultra / 256 GB unified memory |
51
+ | Serial text decode, 128 tokens | **59.4 tok/s** average |
52
+ | Text prefill, 512 tokens | 477.1 tok/s |
53
+ | Text prefill, 2,048 tokens | 562.8 tok/s |
54
+ | Peak memory, short / 2,048-token prompt | 164.3 / 166.8 GB |
55
+ | Main text model on disk | about 154 GiB |
56
+ | Complete repository on disk | about 160 GiB |
57
+
58
+ Measured for serial text generation with oMLX 0.7.0.dev2, MLX 0.32.2, and mlx-lm 0.31.3. The model also answered an image question in an ongoing conversation, read labels in a two-frame video, transcribed `The secret phrase is silver moonlight.` exactly, and understood Xiaomi's published audio sample on the M3 Ultra. Native MTP and DFlash weights are included; the figures above measure the base text decode path.
59
 
60
+ ## Get the model running
61
 
62
+ A **256 GB Apple Silicon Mac** is recommended to leave room for the model, KV cache, and the rest of the system. MiMo is configured for up to one million tokens of context; memory needs rise with the length of the conversation.
63
 
64
  ```bash
65
+ hf download sayyidfareed/MiMo-V2.6-Flash-RL-MLX-4bit-MTP \
66
+ --local-dir ./MiMo-V2.6-Flash-RL-MLX-4bit-MTP
67
+ ```
68
 
69
+ Text generation works with a compatible MLX-LM installation (0.31.3 was used for the original text test):
70
+
71
+ ```bash
72
  python -m mlx_lm generate \
73
+ --model ./MiMo-V2.6-Flash-RL-MLX-4bit-MTP \
74
  --prompt "Write a Python function that checks whether an integer is prime." \
75
+ --max-tokens 256 --temp 0.6
 
76
  ```
77
 
78
+ For image, video, and audio input, use an oMLX build with [MiMo multimodal support](https://github.com/jundot/omlx/pull/3860), including the [multi-turn image fix](https://github.com/jundot/omlx/pull/3860/commits/8ae555cb60fcd0d33abaece2af720340017232e8). For native MTP on text, add [the separate MTP runtime changes](https://github.com/jundot/omlx/pull/3861). Both runtime PRs are currently open; standard MLX-LM runs the text model.
79
+
80
+ A compatible oMLX server accepts standard OpenAI-style image parts:
81
+
82
+ ```json
83
+ {
84
+ "model": "MiMo-V2.6-Flash-RL-MLX-4bit-MTP",
85
+ "messages": [{
86
+ "role": "user",
87
+ "content": [
88
+ {"type": "text", "text": "What is in this picture?"},
89
+ {"type": "image_url", "image_url": {"url": "data:image/png;base64,<base64 image>"}}
90
+ ]
91
+ }],
92
+ "max_tokens": 128
93
+ }
94
+ ```
95
 
96
+ For speech, use `{"type":"input_audio","input_audio":{"data":"<base64 wav>","format":"wav"}}` with a transcription question. Video input is sampled into ordered vision frames; its audio track is not processed by that path.
97
 
98
+ ## What is in the download
99
 
100
+ | Component | Format | File |
101
+ | --- | --- | --- |
102
+ | Language model | 4-bit affine dense projections; native MXFP4 experts | Root `model-*.safetensors` |
103
+ | Native MTP | Three-layer, 4-bit affine | `mtp/model_mtp.safetensors` |
104
+ | DFlash | Five-layer, upstream BF16 | `dflash/model.safetensors` |
105
+ | Vision encoder | Upstream BF16 | `omnimodal/vision_encoder.safetensors` |
106
+ | Audio encoder and bridge | Upstream BF16 | `omnimodal/audio_encoder.safetensors` |
107
+ | Audio tokenizer | Upstream weights | `audio_tokenizer/model.safetensors` |
108
 
109
+ The sidecars are outside the root text-model index, so MLX-LM loads text without loading them. `omnimodal/manifest.json` records their source. With the compatible oMLX runtime, native MTP drafts for **text-only** requests. Image, video, and audio requests use normal decoding, keeping their external embeddings on the correct path. DFlash weights are included for compatible future runtimes.
110
 
111
+ MiMo-V2.6-Flash-RL has 48 transformer layers, 256 routed experts with eight active per token, and a configured one-million-token context. See [Xiaomi's model card](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL) for its evaluations, architecture, intended uses, and limits.
112
 
113
+ ## Scope and credit
114
 
115
+ This runtime handles image understanding, sampled video frames, and audio input. Video frames are processed without the video's audio track; audio input currently supports one prompt batch per request. The linked oMLX changes are under review, so use those branches for the multimodal and MTP paths described above.
116
 
117
+ Xiaomi's MiMo team designed and trained the model and released it under MIT. The MLX conversion and Apple Silicon integration are by [sayyidfareed](https://huggingface.co/sayyidfareed); the original converted weights were published at [Vontra](https://huggingface.co/Vontra/MiMo-V2.6-Flash-RL-MLX-4bit-MTP).
118
 
119
  ```bibtex
120
  @misc{mimo2026v26flash,
 
125
  }
126
  ```
127
 
128
+ [Follow sayyidfareed for Apple Silicon model updates.](https://huggingface.co/sayyidfareed)