patrickbdevaney commited on
Commit
ce5806d
·
verified ·
1 Parent(s): 223bb01

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +1 -1
README.md CHANGED
@@ -32,7 +32,7 @@ Official GGUF quantisations of **MiMo-V2.6-Flash-REAP50**, a 50% routed-expert p
32
  | Filename | Quant Type | Size | Description | Recommended VRAM / RAM |
33
  | :--- | :--- | :--- | :--- | :--- |
34
  | `MiMo-V2.6-Flash-REAP50-MXFP4_MOE.gguf` | **MXFP4_MOE** | 86.06 GiB | Flagship: 1-to-1 native packed MXFP4 experts (32 blk) + BF16 attention/trunk. Exact bit-level fidelity to REAP base. | 96 GiB+ / 1x 128GB Thor or 2x 48GB |
35
- | `MiMo-V2.6-Flash-REAP50-Q2_K.gguf` | **Q2_K** | Pending | Optimal Hybrid MoE: sensitive down-projections kept in native MXFP4, gate/up in Q2_K, trunk in Q8_0 (~3.36 BPW). | 64 GiB+ / 3x 24GB GPUs (72GB) or Mac 64-96GB |
36
 
37
  ### Supporting Towers (Vision, Audio & MTP)
38
 
 
32
  | Filename | Quant Type | Size | Description | Recommended VRAM / RAM |
33
  | :--- | :--- | :--- | :--- | :--- |
34
  | `MiMo-V2.6-Flash-REAP50-MXFP4_MOE.gguf` | **MXFP4_MOE** | 86.06 GiB | Flagship: 1-to-1 native packed MXFP4 experts (32 blk) + BF16 attention/trunk. Exact bit-level fidelity to REAP base. | 96 GiB+ / 1x 128GB Thor or 2x 48GB |
35
+ | `MiMo-V2.6-Flash-REAP50-Q2_K.gguf` | **Q2_K** | 61.64 GiB | Optimal Hybrid MoE: sensitive down-projections kept in native MXFP4, gate/up in Q2_K, trunk in Q8_0 (~3.36 BPW). | 64 GiB+ / 3x 24GB GPUs (72GB) or Mac 64-96GB |
36
 
37
  ### Supporting Towers (Vision, Audio & MTP)
38