vcruz305 commited on
Commit
1e770da
·
verified ·
1 Parent(s): 76a40d2

MiMo-V2.6-Flash-RL-DERISKED-EXL3-2.20bpw: update card header and served model name

Browse files
Files changed (2) hide show
  1. README.md +85 -85
  2. launch.sh +24 -24
README.md CHANGED
@@ -1,85 +1,85 @@
1
- ---
2
- license: mit
3
- base_model: XiaomiMiMo/MiMo-V2.6-Flash-RL
4
- base_model_relation: quantized
5
- library_name: exllamav3
6
- pipeline_tag: text-generation
7
- language:
8
- - en
9
- - zh
10
- tags:
11
- - exl3
12
- - exllamav3
13
- - sage
14
- - mimo
15
- - moe
16
- - quantized
17
- - derisked
18
- - blackfrost
19
- ---
20
-
21
- # MiMo-V2.6-Flash-RL · EXL3 2.20 bpw, derisked
22
-
23
- A collaboration between [Victor Cruz](https://huggingface.co/vcruz305) and [Blackfrost](https://huggingface.co/Blackfrost-AI) ([@Blackfrost_AI](https://x.com/Blackfrost_AI)).
24
-
25
- The EXL3 quantization is mine: 2.20 bpw, and it fits and serves on a single RTX PRO 6000 (96 GB). The derisking overlay is Blackfrost's — the same one that ships in [Blackfrost-AI/MiMo-v2.6-Flash-RL-Derisked](https://huggingface.co/Blackfrost-AI/MiMo-v2.6-Flash-RL-Derisked) for SGLang — ported here onto the EXL3 serving path. The weights are the 2.20 bpw pack unchanged. The overlay is a runtime edit plus the chat template baked into this repo, not a change to the checkpoint.
26
-
27
- **Serving recipe: [vcruz305/MiMo-V2.6-Flash-RL-EXL3-recipe](https://github.com/vcruz305/MiMo-V2.6-Flash-RL-EXL3-recipe)** — the runtime, the DFlash drafter fix, the exact serve command and the measured decode / prefill / TTFT numbers.
28
-
29
- ## Quality against the original checkpoint
30
-
31
- Scored against Xiaomi's own implementation running the original weights with fp32 activations, on 10,240 positions of held-out text that was not used to make the quant. The overlay does not touch the weights, so these numbers are the pack's.
32
-
33
- | | |
34
- | --- | --- |
35
- | top-1 vs original | 82.16% (8,413 / 10,240) |
36
- | mean KLD | 0.1926 |
37
- | p99 KLD | 2.241 |
38
- | second held-out set | 82.64% top-1, mean KLD 0.1986 |
39
-
40
- - **top-1**: positions where the pack's most likely next token matches the reference's.
41
- - **KLD**: KL(reference ‖ pack) per position, mean and 99th percentile. Lower is closer.
42
- - For scale: the unquantized weights running in the same exllamav3 code match the reference on 96.49% of positions (9,881 / 10,240) with mean KLD 0.0070. That is about the best any pack can score.
43
- - The average bitrate of the model body is 2.20; the output head is 6-bit. `quantization_config.bits` in `config.json` is rounded to an integer because the Hub's metadata schema rejects a fractional value — the per-tensor bitrates are in `quantization_config.json`.
44
-
45
- ## The overlay
46
-
47
- The directions, the manifest and the hook are under `derisk/` and `exllamav3/`, and the method is Blackfrost's. It applies at serving time and does not modify weights or routing.
48
-
49
- It is a load-time switch, not a per-request flag. Same weights, same server, restart to change it.
50
-
51
- ```bash
52
- bash launch.sh <model-dir> # derisked
53
- bash launch.sh --off <model-dir> # plain EXL3 2.20
54
- ```
55
-
56
- `launch.sh` sets `BLACKFROST_MIMO_MOE_INTERVENTION` to `derisk/intervention.json`. With the variable unset the hook is a no-op, so one patched fork serves both modes.
57
-
58
- The hook needs one file and one line added to the exllamav3 fork ([vcruz305/exllamav3](https://github.com/vcruz305/exllamav3), branch `feat/mimo-v2`, commit `93e58ca`):
59
-
60
- ```bash
61
- bash setup.sh <path-to-exllamav3-checkout>
62
- ```
63
-
64
- `setup.sh` copies `exllamav3/blackfrost_derisk.py` into the fork and applies `exllamav3/transformer-derisk.patch`. It checks the fork commit and refuses to patch anything it doesn't match.
65
-
66
- ## Measured
67
-
68
- On one RTX PRO 6000, 655,360 context, Q4 KV cache, with the [EXL3 4.0 bpw dflash drafter](https://huggingface.co/vcruz305/MiMo-V2.6-Flash-RL-dflash-EXL3-4.0bpw). The reference column is Blackfrost's own MXFP4 build of the same overlay, scored by the same judge on the same 32-prompt set. It is a higher-bitrate quant, so the gap to it is mostly the base quant, not the overlay.
69
-
70
- Harmful set:
71
-
72
- | | actionable | actionability |
73
- | --- | --- | --- |
74
- | this pack, overlay off | 13/32 | 1.36 |
75
- | this pack, overlay on | 26/32 | 2.38 |
76
- | Blackfrost MXFP4 reference | 30/32 | 2.84 |
77
-
78
- Decode speed, single stream, overlay on vs off: **182.8 vs 184.7 tok/s p50** — about 1-3% depending on the prompt. Draft acceptance was 0.93-0.96 in both.
79
-
80
- ## Notes
81
-
82
- - Made with **SAGE**, a mixed-precision quantization method for EXL3. The scores measure how closely the pack tracks the original checkpoint, not task accuracy, and no evaluation text was used to calibrate it.
83
- - The pack covers the text model: the 48-layer MoE backbone with its hybrid attention. MiMo's vision and audio encoders and its multi-token-prediction drafter are not part of the pack, so it serves text. The dflash drafter above is a separate repo.
84
- - The direction files are pinned by sha256 in the manifest and checked at load.
85
- - This build is intended for authorized red teaming and safety evaluation, the same scope as Blackfrost's release.
 
1
+ ---
2
+ license: mit
3
+ base_model: XiaomiMiMo/MiMo-V2.6-Flash-RL
4
+ base_model_relation: quantized
5
+ library_name: exllamav3
6
+ pipeline_tag: text-generation
7
+ language:
8
+ - en
9
+ - zh
10
+ tags:
11
+ - exl3
12
+ - exllamav3
13
+ - sage
14
+ - mimo
15
+ - moe
16
+ - quantized
17
+ - derisked
18
+ - blackfrost
19
+ ---
20
+
21
+ # MiMo-V2.6-Flash-RL · DERISKED · EXL3 2.20 bpw
22
+
23
+ A collaboration between [Victor Cruz](https://huggingface.co/vcruz305) and [Blackfrost](https://huggingface.co/Blackfrost-AI) ([@Blackfrost_AI](https://x.com/Blackfrost_AI)).
24
+
25
+ The EXL3 quantization is mine: 2.20 bpw, and it fits and serves on a single RTX PRO 6000 (96 GB). The derisking overlay is Blackfrost's — the same one that ships in [Blackfrost-AI/MiMo-v2.6-Flash-RL-Derisked](https://huggingface.co/Blackfrost-AI/MiMo-v2.6-Flash-RL-Derisked) for SGLang — ported here onto the EXL3 serving path. The weights are the 2.20 bpw pack unchanged. The overlay is a runtime edit plus the chat template baked into this repo, not a change to the checkpoint.
26
+
27
+ **Serving recipe: [vcruz305/MiMo-V2.6-Flash-RL-EXL3-recipe](https://github.com/vcruz305/MiMo-V2.6-Flash-RL-EXL3-recipe)** — the runtime, the DFlash drafter fix, the exact serve command and the measured decode / prefill / TTFT numbers.
28
+
29
+ ## Quality against the original checkpoint
30
+
31
+ Scored against Xiaomi's own implementation running the original weights with fp32 activations, on 10,240 positions of held-out text that was not used to make the quant. The overlay does not touch the weights, so these numbers are the pack's.
32
+
33
+ | | |
34
+ | --- | --- |
35
+ | top-1 vs original | 82.16% (8,413 / 10,240) |
36
+ | mean KLD | 0.1926 |
37
+ | p99 KLD | 2.241 |
38
+ | second held-out set | 82.64% top-1, mean KLD 0.1986 |
39
+
40
+ - **top-1**: positions where the pack's most likely next token matches the reference's.
41
+ - **KLD**: KL(reference ‖ pack) per position, mean and 99th percentile. Lower is closer.
42
+ - For scale: the unquantized weights running in the same exllamav3 code match the reference on 96.49% of positions (9,881 / 10,240) with mean KLD 0.0070. That is about the best any pack can score.
43
+ - The average bitrate of the model body is 2.20; the output head is 6-bit. `quantization_config.bits` in `config.json` is rounded to an integer because the Hub's metadata schema rejects a fractional value — the per-tensor bitrates are in `quantization_config.json`.
44
+
45
+ ## The overlay
46
+
47
+ The directions, the manifest and the hook are under `derisk/` and `exllamav3/`, and the method is Blackfrost's. It applies at serving time and does not modify weights or routing.
48
+
49
+ It is a load-time switch, not a per-request flag. Same weights, same server, restart to change it.
50
+
51
+ ```bash
52
+ bash launch.sh <model-dir> # derisked
53
+ bash launch.sh --off <model-dir> # plain EXL3 2.20
54
+ ```
55
+
56
+ `launch.sh` sets `BLACKFROST_MIMO_MOE_INTERVENTION` to `derisk/intervention.json`. With the variable unset the hook is a no-op, so one patched fork serves both modes.
57
+
58
+ The hook needs one file and one line added to the exllamav3 fork ([vcruz305/exllamav3](https://github.com/vcruz305/exllamav3), branch `feat/mimo-v2`, commit `93e58ca`):
59
+
60
+ ```bash
61
+ bash setup.sh <path-to-exllamav3-checkout>
62
+ ```
63
+
64
+ `setup.sh` copies `exllamav3/blackfrost_derisk.py` into the fork and applies `exllamav3/transformer-derisk.patch`. It checks the fork commit and refuses to patch anything it doesn't match.
65
+
66
+ ## Measured
67
+
68
+ On one RTX PRO 6000, 655,360 context, Q4 KV cache, with the [EXL3 4.0 bpw dflash drafter](https://huggingface.co/vcruz305/MiMo-V2.6-Flash-RL-dflash-EXL3-4.0bpw). The reference column is Blackfrost's own MXFP4 build of the same overlay, scored by the same judge on the same 32-prompt set. It is a higher-bitrate quant, so the gap to it is mostly the base quant, not the overlay.
69
+
70
+ Harmful set:
71
+
72
+ | | actionable | actionability |
73
+ | --- | --- | --- |
74
+ | this pack, overlay off | 13/32 | 1.36 |
75
+ | this pack, overlay on | 26/32 | 2.38 |
76
+ | Blackfrost MXFP4 reference | 30/32 | 2.84 |
77
+
78
+ Decode speed, single stream, overlay on vs off: **182.8 vs 184.7 tok/s p50** — about 1-3% depending on the prompt. Draft acceptance was 0.93-0.96 in both.
79
+
80
+ ## Notes
81
+
82
+ - Made with **SAGE**, a mixed-precision quantization method for EXL3. The scores measure how closely the pack tracks the original checkpoint, not task accuracy, and no evaluation text was used to calibrate it.
83
+ - The pack covers the text model: the 48-layer MoE backbone with its hybrid attention. MiMo's vision and audio encoders and its multi-token-prediction drafter are not part of the pack, so it serves text. The dflash drafter above is a separate repo.
84
+ - The direction files are pinned by sha256 in the manifest and checked at load.
85
+ - This build is intended for authorized red teaming and safety evaluation, the same scope as Blackfrost's release.
launch.sh CHANGED
@@ -1,24 +1,24 @@
1
- #!/bin/bash
2
- # Serve this pack with the derisk overlay on or off.
3
- # bash launch.sh <model-dir> [extra serve args] derisked
4
- # bash launch.sh --off <model-dir> [extra serve args] plain EXL3 2.20, same weights
5
- #
6
- # The switch is load-time: changing it means restarting the server. It is not a
7
- # per-request flag.
8
- set -uo pipefail
9
- HERE=$(cd "$(dirname "$0")" && pwd)
10
-
11
- MODE=on
12
- if [ "${1:-}" = "--off" ]; then MODE=off; shift; fi
13
- MODEL=${1:?"usage: bash launch.sh [--off] <model-dir> [extra serve args]"}
14
- shift || true
15
-
16
- if [ "$MODE" = "on" ]; then
17
- export BLACKFROST_MIMO_MOE_INTERVENTION=$HERE/derisk/intervention.json
18
- echo "derisk overlay ON ($BLACKFROST_MIMO_MOE_INTERVENTION)"
19
- else
20
- unset BLACKFROST_MIMO_MOE_INTERVENTION
21
- echo "derisk overlay OFF (plain EXL3 2.20)"
22
- fi
23
-
24
- exec serve_native.py -m "$MODEL" --served-model-name MiMo-V2.6-Flash-RL-EXL3-2.20bpw-derisked "$@"
 
1
+ #!/bin/bash
2
+ # Serve this pack with the derisk overlay on or off.
3
+ # bash launch.sh <model-dir> [extra serve args] derisked
4
+ # bash launch.sh --off <model-dir> [extra serve args] plain EXL3 2.20, same weights
5
+ #
6
+ # The switch is load-time: changing it means restarting the server. It is not a
7
+ # per-request flag.
8
+ set -uo pipefail
9
+ HERE=$(cd "$(dirname "$0")" && pwd)
10
+
11
+ MODE=on
12
+ if [ "${1:-}" = "--off" ]; then MODE=off; shift; fi
13
+ MODEL=${1:?"usage: bash launch.sh [--off] <model-dir> [extra serve args]"}
14
+ shift || true
15
+
16
+ if [ "$MODE" = "on" ]; then
17
+ export BLACKFROST_MIMO_MOE_INTERVENTION=$HERE/derisk/intervention.json
18
+ echo "derisk overlay ON ($BLACKFROST_MIMO_MOE_INTERVENTION)"
19
+ else
20
+ unset BLACKFROST_MIMO_MOE_INTERVENTION
21
+ echo "derisk overlay OFF (plain EXL3 2.20)"
22
+ fi
23
+
24
+ exec serve_native.py -m "$MODEL" --served-model-name MiMo-V2.6-Flash-RL-DERISKED-EXL3-2.20bpw "$@"