MiMo-V2.6-Flash-RL-DERISKED-EXL3-2.20bpw: update card header and served model name
Browse files
README.md
CHANGED
|
@@ -1,85 +1,85 @@
|
|
| 1 |
-
---
|
| 2 |
-
license: mit
|
| 3 |
-
base_model: XiaomiMiMo/MiMo-V2.6-Flash-RL
|
| 4 |
-
base_model_relation: quantized
|
| 5 |
-
library_name: exllamav3
|
| 6 |
-
pipeline_tag: text-generation
|
| 7 |
-
language:
|
| 8 |
-
- en
|
| 9 |
-
- zh
|
| 10 |
-
tags:
|
| 11 |
-
- exl3
|
| 12 |
-
- exllamav3
|
| 13 |
-
- sage
|
| 14 |
-
- mimo
|
| 15 |
-
- moe
|
| 16 |
-
- quantized
|
| 17 |
-
- derisked
|
| 18 |
-
- blackfrost
|
| 19 |
-
---
|
| 20 |
-
|
| 21 |
-
# MiMo-V2.6-Flash-RL · EXL3 2.20 bpw
|
| 22 |
-
|
| 23 |
-
A collaboration between [Victor Cruz](https://huggingface.co/vcruz305) and [Blackfrost](https://huggingface.co/Blackfrost-AI) ([@Blackfrost_AI](https://x.com/Blackfrost_AI)).
|
| 24 |
-
|
| 25 |
-
The EXL3 quantization is mine: 2.20 bpw, and it fits and serves on a single RTX PRO 6000 (96 GB). The derisking overlay is Blackfrost's — the same one that ships in [Blackfrost-AI/MiMo-v2.6-Flash-RL-Derisked](https://huggingface.co/Blackfrost-AI/MiMo-v2.6-Flash-RL-Derisked) for SGLang — ported here onto the EXL3 serving path. The weights are the 2.20 bpw pack unchanged. The overlay is a runtime edit plus the chat template baked into this repo, not a change to the checkpoint.
|
| 26 |
-
|
| 27 |
-
**Serving recipe: [vcruz305/MiMo-V2.6-Flash-RL-EXL3-recipe](https://github.com/vcruz305/MiMo-V2.6-Flash-RL-EXL3-recipe)** — the runtime, the DFlash drafter fix, the exact serve command and the measured decode / prefill / TTFT numbers.
|
| 28 |
-
|
| 29 |
-
## Quality against the original checkpoint
|
| 30 |
-
|
| 31 |
-
Scored against Xiaomi's own implementation running the original weights with fp32 activations, on 10,240 positions of held-out text that was not used to make the quant. The overlay does not touch the weights, so these numbers are the pack's.
|
| 32 |
-
|
| 33 |
-
| | |
|
| 34 |
-
| --- | --- |
|
| 35 |
-
| top-1 vs original | 82.16% (8,413 / 10,240) |
|
| 36 |
-
| mean KLD | 0.1926 |
|
| 37 |
-
| p99 KLD | 2.241 |
|
| 38 |
-
| second held-out set | 82.64% top-1, mean KLD 0.1986 |
|
| 39 |
-
|
| 40 |
-
- **top-1**: positions where the pack's most likely next token matches the reference's.
|
| 41 |
-
- **KLD**: KL(reference ‖ pack) per position, mean and 99th percentile. Lower is closer.
|
| 42 |
-
- For scale: the unquantized weights running in the same exllamav3 code match the reference on 96.49% of positions (9,881 / 10,240) with mean KLD 0.0070. That is about the best any pack can score.
|
| 43 |
-
- The average bitrate of the model body is 2.20; the output head is 6-bit. `quantization_config.bits` in `config.json` is rounded to an integer because the Hub's metadata schema rejects a fractional value — the per-tensor bitrates are in `quantization_config.json`.
|
| 44 |
-
|
| 45 |
-
## The overlay
|
| 46 |
-
|
| 47 |
-
The directions, the manifest and the hook are under `derisk/` and `exllamav3/`, and the method is Blackfrost's. It applies at serving time and does not modify weights or routing.
|
| 48 |
-
|
| 49 |
-
It is a load-time switch, not a per-request flag. Same weights, same server, restart to change it.
|
| 50 |
-
|
| 51 |
-
```bash
|
| 52 |
-
bash launch.sh <model-dir> # derisked
|
| 53 |
-
bash launch.sh --off <model-dir> # plain EXL3 2.20
|
| 54 |
-
```
|
| 55 |
-
|
| 56 |
-
`launch.sh` sets `BLACKFROST_MIMO_MOE_INTERVENTION` to `derisk/intervention.json`. With the variable unset the hook is a no-op, so one patched fork serves both modes.
|
| 57 |
-
|
| 58 |
-
The hook needs one file and one line added to the exllamav3 fork ([vcruz305/exllamav3](https://github.com/vcruz305/exllamav3), branch `feat/mimo-v2`, commit `93e58ca`):
|
| 59 |
-
|
| 60 |
-
```bash
|
| 61 |
-
bash setup.sh <path-to-exllamav3-checkout>
|
| 62 |
-
```
|
| 63 |
-
|
| 64 |
-
`setup.sh` copies `exllamav3/blackfrost_derisk.py` into the fork and applies `exllamav3/transformer-derisk.patch`. It checks the fork commit and refuses to patch anything it doesn't match.
|
| 65 |
-
|
| 66 |
-
## Measured
|
| 67 |
-
|
| 68 |
-
On one RTX PRO 6000, 655,360 context, Q4 KV cache, with the [EXL3 4.0 bpw dflash drafter](https://huggingface.co/vcruz305/MiMo-V2.6-Flash-RL-dflash-EXL3-4.0bpw). The reference column is Blackfrost's own MXFP4 build of the same overlay, scored by the same judge on the same 32-prompt set. It is a higher-bitrate quant, so the gap to it is mostly the base quant, not the overlay.
|
| 69 |
-
|
| 70 |
-
Harmful set:
|
| 71 |
-
|
| 72 |
-
| | actionable | actionability |
|
| 73 |
-
| --- | --- | --- |
|
| 74 |
-
| this pack, overlay off | 13/32 | 1.36 |
|
| 75 |
-
| this pack, overlay on | 26/32 | 2.38 |
|
| 76 |
-
| Blackfrost MXFP4 reference | 30/32 | 2.84 |
|
| 77 |
-
|
| 78 |
-
Decode speed, single stream, overlay on vs off: **182.8 vs 184.7 tok/s p50** — about 1-3% depending on the prompt. Draft acceptance was 0.93-0.96 in both.
|
| 79 |
-
|
| 80 |
-
## Notes
|
| 81 |
-
|
| 82 |
-
- Made with **SAGE**, a mixed-precision quantization method for EXL3. The scores measure how closely the pack tracks the original checkpoint, not task accuracy, and no evaluation text was used to calibrate it.
|
| 83 |
-
- The pack covers the text model: the 48-layer MoE backbone with its hybrid attention. MiMo's vision and audio encoders and its multi-token-prediction drafter are not part of the pack, so it serves text. The dflash drafter above is a separate repo.
|
| 84 |
-
- The direction files are pinned by sha256 in the manifest and checked at load.
|
| 85 |
-
- This build is intended for authorized red teaming and safety evaluation, the same scope as Blackfrost's release.
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: mit
|
| 3 |
+
base_model: XiaomiMiMo/MiMo-V2.6-Flash-RL
|
| 4 |
+
base_model_relation: quantized
|
| 5 |
+
library_name: exllamav3
|
| 6 |
+
pipeline_tag: text-generation
|
| 7 |
+
language:
|
| 8 |
+
- en
|
| 9 |
+
- zh
|
| 10 |
+
tags:
|
| 11 |
+
- exl3
|
| 12 |
+
- exllamav3
|
| 13 |
+
- sage
|
| 14 |
+
- mimo
|
| 15 |
+
- moe
|
| 16 |
+
- quantized
|
| 17 |
+
- derisked
|
| 18 |
+
- blackfrost
|
| 19 |
+
---
|
| 20 |
+
|
| 21 |
+
# MiMo-V2.6-Flash-RL · DERISKED · EXL3 2.20 bpw
|
| 22 |
+
|
| 23 |
+
A collaboration between [Victor Cruz](https://huggingface.co/vcruz305) and [Blackfrost](https://huggingface.co/Blackfrost-AI) ([@Blackfrost_AI](https://x.com/Blackfrost_AI)).
|
| 24 |
+
|
| 25 |
+
The EXL3 quantization is mine: 2.20 bpw, and it fits and serves on a single RTX PRO 6000 (96 GB). The derisking overlay is Blackfrost's — the same one that ships in [Blackfrost-AI/MiMo-v2.6-Flash-RL-Derisked](https://huggingface.co/Blackfrost-AI/MiMo-v2.6-Flash-RL-Derisked) for SGLang — ported here onto the EXL3 serving path. The weights are the 2.20 bpw pack unchanged. The overlay is a runtime edit plus the chat template baked into this repo, not a change to the checkpoint.
|
| 26 |
+
|
| 27 |
+
**Serving recipe: [vcruz305/MiMo-V2.6-Flash-RL-EXL3-recipe](https://github.com/vcruz305/MiMo-V2.6-Flash-RL-EXL3-recipe)** — the runtime, the DFlash drafter fix, the exact serve command and the measured decode / prefill / TTFT numbers.
|
| 28 |
+
|
| 29 |
+
## Quality against the original checkpoint
|
| 30 |
+
|
| 31 |
+
Scored against Xiaomi's own implementation running the original weights with fp32 activations, on 10,240 positions of held-out text that was not used to make the quant. The overlay does not touch the weights, so these numbers are the pack's.
|
| 32 |
+
|
| 33 |
+
| | |
|
| 34 |
+
| --- | --- |
|
| 35 |
+
| top-1 vs original | 82.16% (8,413 / 10,240) |
|
| 36 |
+
| mean KLD | 0.1926 |
|
| 37 |
+
| p99 KLD | 2.241 |
|
| 38 |
+
| second held-out set | 82.64% top-1, mean KLD 0.1986 |
|
| 39 |
+
|
| 40 |
+
- **top-1**: positions where the pack's most likely next token matches the reference's.
|
| 41 |
+
- **KLD**: KL(reference ‖ pack) per position, mean and 99th percentile. Lower is closer.
|
| 42 |
+
- For scale: the unquantized weights running in the same exllamav3 code match the reference on 96.49% of positions (9,881 / 10,240) with mean KLD 0.0070. That is about the best any pack can score.
|
| 43 |
+
- The average bitrate of the model body is 2.20; the output head is 6-bit. `quantization_config.bits` in `config.json` is rounded to an integer because the Hub's metadata schema rejects a fractional value — the per-tensor bitrates are in `quantization_config.json`.
|
| 44 |
+
|
| 45 |
+
## The overlay
|
| 46 |
+
|
| 47 |
+
The directions, the manifest and the hook are under `derisk/` and `exllamav3/`, and the method is Blackfrost's. It applies at serving time and does not modify weights or routing.
|
| 48 |
+
|
| 49 |
+
It is a load-time switch, not a per-request flag. Same weights, same server, restart to change it.
|
| 50 |
+
|
| 51 |
+
```bash
|
| 52 |
+
bash launch.sh <model-dir> # derisked
|
| 53 |
+
bash launch.sh --off <model-dir> # plain EXL3 2.20
|
| 54 |
+
```
|
| 55 |
+
|
| 56 |
+
`launch.sh` sets `BLACKFROST_MIMO_MOE_INTERVENTION` to `derisk/intervention.json`. With the variable unset the hook is a no-op, so one patched fork serves both modes.
|
| 57 |
+
|
| 58 |
+
The hook needs one file and one line added to the exllamav3 fork ([vcruz305/exllamav3](https://github.com/vcruz305/exllamav3), branch `feat/mimo-v2`, commit `93e58ca`):
|
| 59 |
+
|
| 60 |
+
```bash
|
| 61 |
+
bash setup.sh <path-to-exllamav3-checkout>
|
| 62 |
+
```
|
| 63 |
+
|
| 64 |
+
`setup.sh` copies `exllamav3/blackfrost_derisk.py` into the fork and applies `exllamav3/transformer-derisk.patch`. It checks the fork commit and refuses to patch anything it doesn't match.
|
| 65 |
+
|
| 66 |
+
## Measured
|
| 67 |
+
|
| 68 |
+
On one RTX PRO 6000, 655,360 context, Q4 KV cache, with the [EXL3 4.0 bpw dflash drafter](https://huggingface.co/vcruz305/MiMo-V2.6-Flash-RL-dflash-EXL3-4.0bpw). The reference column is Blackfrost's own MXFP4 build of the same overlay, scored by the same judge on the same 32-prompt set. It is a higher-bitrate quant, so the gap to it is mostly the base quant, not the overlay.
|
| 69 |
+
|
| 70 |
+
Harmful set:
|
| 71 |
+
|
| 72 |
+
| | actionable | actionability |
|
| 73 |
+
| --- | --- | --- |
|
| 74 |
+
| this pack, overlay off | 13/32 | 1.36 |
|
| 75 |
+
| this pack, overlay on | 26/32 | 2.38 |
|
| 76 |
+
| Blackfrost MXFP4 reference | 30/32 | 2.84 |
|
| 77 |
+
|
| 78 |
+
Decode speed, single stream, overlay on vs off: **182.8 vs 184.7 tok/s p50** — about 1-3% depending on the prompt. Draft acceptance was 0.93-0.96 in both.
|
| 79 |
+
|
| 80 |
+
## Notes
|
| 81 |
+
|
| 82 |
+
- Made with **SAGE**, a mixed-precision quantization method for EXL3. The scores measure how closely the pack tracks the original checkpoint, not task accuracy, and no evaluation text was used to calibrate it.
|
| 83 |
+
- The pack covers the text model: the 48-layer MoE backbone with its hybrid attention. MiMo's vision and audio encoders and its multi-token-prediction drafter are not part of the pack, so it serves text. The dflash drafter above is a separate repo.
|
| 84 |
+
- The direction files are pinned by sha256 in the manifest and checked at load.
|
| 85 |
+
- This build is intended for authorized red teaming and safety evaluation, the same scope as Blackfrost's release.
|
launch.sh
CHANGED
|
@@ -1,24 +1,24 @@
|
|
| 1 |
-
#!/bin/bash
|
| 2 |
-
# Serve this pack with the derisk overlay on or off.
|
| 3 |
-
# bash launch.sh <model-dir> [extra serve args] derisked
|
| 4 |
-
# bash launch.sh --off <model-dir> [extra serve args] plain EXL3 2.20, same weights
|
| 5 |
-
#
|
| 6 |
-
# The switch is load-time: changing it means restarting the server. It is not a
|
| 7 |
-
# per-request flag.
|
| 8 |
-
set -uo pipefail
|
| 9 |
-
HERE=$(cd "$(dirname "$0")" && pwd)
|
| 10 |
-
|
| 11 |
-
MODE=on
|
| 12 |
-
if [ "${1:-}" = "--off" ]; then MODE=off; shift; fi
|
| 13 |
-
MODEL=${1:?"usage: bash launch.sh [--off] <model-dir> [extra serve args]"}
|
| 14 |
-
shift || true
|
| 15 |
-
|
| 16 |
-
if [ "$MODE" = "on" ]; then
|
| 17 |
-
export BLACKFROST_MIMO_MOE_INTERVENTION=$HERE/derisk/intervention.json
|
| 18 |
-
echo "derisk overlay ON ($BLACKFROST_MIMO_MOE_INTERVENTION)"
|
| 19 |
-
else
|
| 20 |
-
unset BLACKFROST_MIMO_MOE_INTERVENTION
|
| 21 |
-
echo "derisk overlay OFF (plain EXL3 2.20)"
|
| 22 |
-
fi
|
| 23 |
-
|
| 24 |
-
exec serve_native.py -m "$MODEL" --served-model-name MiMo-V2.6-Flash-RL-EXL3-2.20bpw
|
|
|
|
| 1 |
+
#!/bin/bash
|
| 2 |
+
# Serve this pack with the derisk overlay on or off.
|
| 3 |
+
# bash launch.sh <model-dir> [extra serve args] derisked
|
| 4 |
+
# bash launch.sh --off <model-dir> [extra serve args] plain EXL3 2.20, same weights
|
| 5 |
+
#
|
| 6 |
+
# The switch is load-time: changing it means restarting the server. It is not a
|
| 7 |
+
# per-request flag.
|
| 8 |
+
set -uo pipefail
|
| 9 |
+
HERE=$(cd "$(dirname "$0")" && pwd)
|
| 10 |
+
|
| 11 |
+
MODE=on
|
| 12 |
+
if [ "${1:-}" = "--off" ]; then MODE=off; shift; fi
|
| 13 |
+
MODEL=${1:?"usage: bash launch.sh [--off] <model-dir> [extra serve args]"}
|
| 14 |
+
shift || true
|
| 15 |
+
|
| 16 |
+
if [ "$MODE" = "on" ]; then
|
| 17 |
+
export BLACKFROST_MIMO_MOE_INTERVENTION=$HERE/derisk/intervention.json
|
| 18 |
+
echo "derisk overlay ON ($BLACKFROST_MIMO_MOE_INTERVENTION)"
|
| 19 |
+
else
|
| 20 |
+
unset BLACKFROST_MIMO_MOE_INTERVENTION
|
| 21 |
+
echo "derisk overlay OFF (plain EXL3 2.20)"
|
| 22 |
+
fi
|
| 23 |
+
|
| 24 |
+
exec serve_native.py -m "$MODEL" --served-model-name MiMo-V2.6-Flash-RL-DERISKED-EXL3-2.20bpw "$@"
|