patrickbdevaney commited on
Commit
8d0abe0
·
verified ·
1 Parent(s): ce5806d

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +13 -5
README.md CHANGED
@@ -63,10 +63,18 @@ Official GGUF quantisations of **MiMo-V2.6-Flash-REAP50**, a 50% routed-expert p
63
 
64
  ## Running with llama.cpp
65
 
66
- ### 1. Standard Text Inference
67
  ```bash
68
  ./llama-cli \
69
- -m MiMo-V2.6-Flash-REAP50-Q4_K_M.gguf \
 
 
 
 
 
 
 
 
70
  -p "You are MiMo, an AI assistant developed by Xiaomi. Explain how MoE expert pruning works:" \
71
  -n 512 --temp 0.6
72
  ```
@@ -74,7 +82,7 @@ Official GGUF quantisations of **MiMo-V2.6-Flash-REAP50**, a 50% routed-expert p
74
  ### 2. Speculative Decoding with MTP Draft Head
75
  ```bash
76
  ./llama-cli \
77
- -m MiMo-V2.6-Flash-REAP50-Q4_K_M.gguf \
78
  --draft-model mtp-MiMo-V2.6-Flash-REAP50-Q8_0.gguf \
79
  -p "Explain quantum teleportation in detail:" \
80
  -n 512
@@ -83,7 +91,7 @@ Official GGUF quantisations of **MiMo-V2.6-Flash-REAP50**, a 50% routed-expert p
83
  ### 3. Multimodal Inference (Vision & Audio)
84
  ```bash
85
  ./llama-cli \
86
- -m MiMo-V2.6-Flash-REAP50-Q4_K_M.gguf \
87
  --mmproj mmproj-MiMo-V2.6-Flash-REAP50-Q8_0.gguf \
88
  --image input.jpg \
89
  -p "Describe the contents of this image in detail."
@@ -92,7 +100,7 @@ Official GGUF quantisations of **MiMo-V2.6-Flash-REAP50**, a 50% routed-expert p
92
  ### 4. OpenAI-Compatible API Server
93
  ```bash
94
  ./llama-server \
95
- -m MiMo-V2.6-Flash-REAP50-Q4_K_M.gguf \
96
  --mmproj mmproj-MiMo-V2.6-Flash-REAP50-Q8_0.gguf \
97
  --port 8080 \
98
  -ngl 99
 
63
 
64
  ## Running with llama.cpp
65
 
66
+ ### 1. Standard Text Inference (Optimal Hybrid Q2_K)
67
  ```bash
68
  ./llama-cli \
69
+ -m MiMo-V2.6-Flash-REAP50-Q2_K.gguf \
70
+ -p "You are MiMo, an AI assistant developed by Xiaomi. Explain how MoE expert pruning works:" \
71
+ -n 512 --temp 0.6
72
+ ```
73
+
74
+ Or run the flagship bit-for-bit native MXFP4 checkpoint:
75
+ ```bash
76
+ ./llama-cli \
77
+ -m MiMo-V2.6-Flash-REAP50-MXFP4_MOE.gguf \
78
  -p "You are MiMo, an AI assistant developed by Xiaomi. Explain how MoE expert pruning works:" \
79
  -n 512 --temp 0.6
80
  ```
 
82
  ### 2. Speculative Decoding with MTP Draft Head
83
  ```bash
84
  ./llama-cli \
85
+ -m MiMo-V2.6-Flash-REAP50-Q2_K.gguf \
86
  --draft-model mtp-MiMo-V2.6-Flash-REAP50-Q8_0.gguf \
87
  -p "Explain quantum teleportation in detail:" \
88
  -n 512
 
91
  ### 3. Multimodal Inference (Vision & Audio)
92
  ```bash
93
  ./llama-cli \
94
+ -m MiMo-V2.6-Flash-REAP50-Q2_K.gguf \
95
  --mmproj mmproj-MiMo-V2.6-Flash-REAP50-Q8_0.gguf \
96
  --image input.jpg \
97
  -p "Describe the contents of this image in detail."
 
100
  ### 4. OpenAI-Compatible API Server
101
  ```bash
102
  ./llama-server \
103
+ -m MiMo-V2.6-Flash-REAP50-Q2_K.gguf \
104
  --mmproj mmproj-MiMo-V2.6-Flash-REAP50-Q8_0.gguf \
105
  --port 8080 \
106
  -ngl 99