patrickbdevaney commited on
Commit
58505a9
·
verified ·
1 Parent(s): 37dc172

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +113 -0
README.md ADDED
@@ -0,0 +1,113 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: mit
3
+ base_model:
4
+ - XiaomiMiMo/MiMo-V2.6-Flash
5
+ - patrickbdevaney/MiMo-V2.6-Flash-REAP50
6
+ tags:
7
+ - gguf
8
+ - llama.cpp
9
+ - reap
10
+ - hope
11
+ - moe
12
+ - pruned
13
+ - multimodal
14
+ - vision
15
+ - audio
16
+ - mtp
17
+ ---
18
+
19
+ # Xiaomi MiMo-V2.6-Flash REAP-50 — GGUF
20
+
21
+ Official GGUF quantisations of **MiMo-V2.6-Flash-REAP50**, a 50% routed-expert pruned checkpoint of `XiaomiMiMo/MiMo-V2.6-Flash` created with [REAP](https://github.com/CerebrasResearch/reap) and **HOPE** second-order saliency pruning.
22
+
23
+ * **Base HF Checkpoint**: [patrickbdevaney/MiMo-V2.6-Flash-REAP50](https://huggingface.co/patrickbdevaney/MiMo-V2.6-Flash-REAP50)
24
+ * **Experts Retained**: **128 of 256** routed experts per layer across 47 MoE layers (1 dense layer, 47 MoE layers).
25
+ * **Base Architecture**: Native packed **MXFP4** (`U8`, block size 32) experts with unquantized pure **BF16** attention and embeddings.
26
+ * **Towers Included**: Vision & Audio multimodal projectors (`mmproj`) and Multi-Token Prediction speculative draft heads (`mtp`).
27
+
28
+ ---
29
+
30
+ ## Quantization Ladder
31
+
32
+ | Filename | Quant Type | Size | Description | Recommended VRAM / RAM |
33
+ | :--- | :--- | :--- | :--- | :--- |
34
+ | `MiMo-V2.6-Flash-REAP50-MXFP4_MOE.gguf` | **MXFP4_MOE** | Pending | Native packed MXFP4 experts (32 blk) + BF16 attention/trunk | 96 GiB+ / 1x 128GB Thor or 2x 48GB |
35
+ | `MiMo-V2.6-Flash-REAP50-Q4_K_M.gguf` | **Q4_K_M** | Pending | 4-bit medium k-quant; optimal quality/speed balance | 64 GiB+ / 2x 3090/4090 or Mac 64GB |
36
+ | `MiMo-V2.6-Flash-REAP50-Q3_K_M.gguf` | **Q3_K_M** | Pending | 3-bit medium k-quant; high compression | 48 GiB+ / 2x 24GB or Mac 48GB |
37
+ | `MiMo-V2.6-Flash-REAP50-IQ4_XS.gguf` | **IQ4_XS** | Pending | 4-bit non-linear importance-quant; small footprint | 56 GiB+ / 2x 32GB or Mac 64GB |
38
+ | `MiMo-V2.6-Flash-REAP50-IQ3_M.gguf` | **IQ3_M** | Pending | 3-bit non-linear importance-quant mixture | 48 GiB+ / 2x 24GB or Mac 48GB |
39
+
40
+ ### Supporting Towers (Vision, Audio & MTP)
41
+
42
+ | Filename | Size | Description |
43
+ | :--- | :--- | :--- |
44
+ | `mmproj-MiMo-V2.6-Flash-REAP50-BF16.gguf` | 2.56 GiB | Multimodal projector (Vision + Audio) in BF16 |
45
+ | `mmproj-MiMo-V2.6-Flash-REAP50-Q8_0.gguf` | 1.46 GiB | Multimodal projector (Vision + Audio) quantized to Q8_0 |
46
+ | `mtp-MiMo-V2.6-Flash-REAP50-BF16.gguf` | 4.17 GiB | Multi-Token Prediction (MTP) draft head (3 next-n layers) in BF16 |
47
+ | `mtp-MiMo-V2.6-Flash-REAP50-Q8_0.gguf` | 2.22 GiB | Multi-Token Prediction (MTP) draft head (3 next-n layers) in Q8_0 |
48
+
49
+ ---
50
+
51
+ ## Key Features
52
+
53
+ 1. **Native MXFP4 MoE Preservation**:
54
+ In the base model, 92.9% of weights are stored as native packed `mxfp4` (32 block size). Our GGUF converter natively repacks these blocks directly into `GGMLQuantizationType.MXFP4`, avoiding costly lossy dequantization cycles while preserving exact native numerical precision.
55
+
56
+ 2. **Multimodal Projectors (`mmproj`)**:
57
+ Xiaomi MiMo-V2.6-Flash incorporates both visual and audio processing towers:
58
+ - Vision encoder (28-layer ViT, 560px patch representation)
59
+ - Audio tokenizer / RVQ speech representations
60
+ Both are packed into standard GGUF multimodal projectors (`mmproj-*-BF16.gguf` and `mmproj-*-Q8_0.gguf`) compatible with `llama.cpp`'s multimodal pipeline.
61
+
62
+ 3. **Multi-Token Prediction (`mtp`)**:
63
+ MiMo-V2.6-Flash includes 3 trained MTP layers for speculative decoding. We ship standalone MTP draft models (`mtp-*-BF16.gguf` and `mtp-*-Q8_0.gguf`) that can be loaded alongside the trunk model with `--draft-model` to accelerate generation.
64
+
65
+ ---
66
+
67
+ ## Running with llama.cpp
68
+
69
+ ### 1. Standard Text Inference
70
+ ```bash
71
+ ./llama-cli \
72
+ -m MiMo-V2.6-Flash-REAP50-Q4_K_M.gguf \
73
+ -p "You are MiMo, an AI assistant developed by Xiaomi. Explain how MoE expert pruning works:" \
74
+ -n 512 --temp 0.6
75
+ ```
76
+
77
+ ### 2. Speculative Decoding with MTP Draft Head
78
+ ```bash
79
+ ./llama-cli \
80
+ -m MiMo-V2.6-Flash-REAP50-Q4_K_M.gguf \
81
+ --draft-model mtp-MiMo-V2.6-Flash-REAP50-Q8_0.gguf \
82
+ -p "Explain quantum teleportation in detail:" \
83
+ -n 512
84
+ ```
85
+
86
+ ### 3. Multimodal Inference (Vision & Audio)
87
+ ```bash
88
+ ./llama-cli \
89
+ -m MiMo-V2.6-Flash-REAP50-Q4_K_M.gguf \
90
+ --mmproj mmproj-MiMo-V2.6-Flash-REAP50-Q8_0.gguf \
91
+ --image input.jpg \
92
+ -p "Describe the contents of this image in detail."
93
+ ```
94
+
95
+ ### 4. OpenAI-Compatible API Server
96
+ ```bash
97
+ ./llama-server \
98
+ -m MiMo-V2.6-Flash-REAP50-Q4_K_M.gguf \
99
+ --mmproj mmproj-MiMo-V2.6-Flash-REAP50-Q8_0.gguf \
100
+ --port 8080 \
101
+ -ngl 99
102
+ ```
103
+
104
+ ---
105
+
106
+ ## Background & Pruning Method
107
+
108
+ Pruned using **HOPE** (Higher-Order Pruning of Experts) over a diverse calibration corpus spanning code, math, conversational text, and multimodal reasoning tasks. Rather than relying solely on first-order activation frequencies, HOPE accounts for inter-expert interaction terms:
109
+ $$\Delta \mathcal{L} \approx \sum_{i} g_i^T \Delta w_i + \frac{1}{2} \sum_{i,j} \Delta w_i^T H_{ij} \Delta w_j$$
110
+ By computing cross-expert Hessian blocks during the calibration pass, 128 experts per layer were optimally selected to minimize perplexity loss under 50% parameter reduction.
111
+
112
+ ---
113
+ *Created by [patrickbdevaney](https://huggingface.co/patrickbdevaney).*