MorinoNushi's picture
Upload README.md with huggingface_hub
eb0d669 verified
|
Raw History Blame Contribute Delete
6.53 kB
---
license: mit
base_model: XiaomiMiMo/MiMo-V2.6-Flash-RL
tags:
- uncensored
- abliterated
- heretic
- gguf
- mimo
---
# MiMo-V2.6-Flash-RL Uncensored Heretic β€” merged GGUF
> **⚠️ Content warning:** This model has had the base model's refusal
> behavior surgically suppressed. The resulting model **will** comply with
> requests the base model refuses, including requests that are harmful,
> unethical, offensive, or illegal. It has reduced safety guardrails. See
> [Responsible use](#responsible-use--disclaimer) below β€” **you are solely
> responsible for what you do with it.**
This is a **merged GGUF** of
[MiMo-V2.6-Flash-RL](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL)
(309B total / 15B active MoE, MIT license) decensored / "abliterated" with
**heretic-gguf** β€” a GGUF-native port of
[Heretic](https://github.com/p-e-w/heretic)'s Optuna-optimized directional
ablation, which runs the whole search directly on quantized GGUF weights via
llama.cpp. The ablation (trial 85 of study `mimo26flash`) was baked directly
into the
[MXFP4](https://huggingface.co/ggml-org/MiMo-V2.6-Flash-RL-GGUF) weights:
edited tensors were dequantized, patched with the exact ablation delta, and
requantized to their original type; everything else is a byte-for-byte copy.
**This is the zero-runtime-overhead form** β€” a drop-in base model, no
`--lora` flag needed. If you prefer the **lossless** option (bit-identical
base weights, ~70 MB download, requantization avoided entirely), the exact
same configuration is also available as a LoRA adapter at
[MiMo-V2.6-Flash-RL-Uncensored-Heretic-LoRA-GGUF](https://huggingface.co/MorinoNushi/MiMo-V2.6-Flash-RL-Uncensored-Heretic-LoRA-GGUF).
**heretic-gguf is available at
[github.com/MoriNoNushi/heretic-gguf](https://github.com/MoriNoNushi/heretic-gguf)** β€”
the full tool, so the method can be applied to other GGUF models.
## Results
Measured on **140 harmful prompts** (100 from `mlabonne/harmful_behaviors`
test + 40 custom) and **100 harmless prompts** (`mlabonne/harmless_alpaca`
test), CoT-skip prefix (`<think></think>`, thinking suppressed), greedy
decoding, 100-token responses, against the MXFP4 base:
| | Refusal rate (harmful) | KL divergence (harmless) |
|---|---|---|
| Base model | 95.71% (134/140) | 0 (by definition) |
| **Ablated model (trial 85)** | **3.57% (5/140)** | **0.0568** |
The scores above were measured on the LoRA form of the ablation; the merged
model applies the identical delta, requantized to MXFP4/Q8_0, so behavior
should match within quantization noise. Refusals are counted by
refusal-keyword matching (English + Chinese + first-person-negation markers
such as "I'm not going to / able to ..."). KL divergence is measured on
first-token logits on harmless prompts. Note that the study was run with the
CoT-skip prefix (thinking suppressed, as in stock Heretic); with full
thinking enabled the model may still reason its way back to a refusal
mid-trace, so real-use refusal rates can be somewhat higher than the 3.57%
above.
> **Note on KL:** the KL divergence above (and the optimization objective
> itself) was measured against the **MXFP4** quant. KL is a
> baseline-relative metric, so on a different base quant the effective drift
> from that quant's baseline may differ.
## Usage
```bash
llama-server \
-m MiMo-V2.6-Flash-RL-Uncensored-Heretic-MXFP4-00001-of-00002.gguf \
--jinja
```
Add your usual offload/context flags (`-ngl 999`, `-c`, tensor splits,
etc.) β€” nothing model-specific is required, and no special sampling
parameters are needed. MiMo-V2.6-Flash-RL (`mimo2`) support is merged
upstream in llama.cpp β€” any recent build works, no patches or PRs needed.
## How it was made
- **Method:** directional ablation ("abliteration") β€” the refusal direction
in residual space (difference of means over 480 harmful / 480 harmless
prompts, 5% winsorized, orthogonalized against the harmless mean) is
projected out of the attention output and MoE down-projection weights.
Strengths, layer kernels, and direction selection were tuned by
multi-objective Optuna TPE (minimize refusal rate and KL jointly). This
model is trial 85 of study `mimo26flash`, exported with
`heretic-gguf export --mode merged`.
- **Configuration** (study `mimo26flash`, trial 85; global direction scope,
direction index 26.2 of 48; per-expert strengths scaled by measured
harmful/harmless routing frequency; `row_normalization = "pre"`):
- attn.o_proj: max weight 6.39 @ layer 36.4 of 48.
- routed MLP down-proj: max weight 1.58 @ layer 31.9.
- **Merged export:** heretic-gguf expresses ablation as a rank-1 LoRA
overlay (the same math stock Heretic writes into PEFT adapters); the
merged exporter materializes that delta exactly β€” full-rank, no LoRA
factorization loss β€” and requantizes only the patched tensors to their
original type (MXFP4 experts, Q8_0 attention/dense). This is one extra
quantization step on those tensors relative to the base; the
[LoRA form](https://huggingface.co/MorinoNushi/MiMo-V2.6-Flash-RL-Uncensored-Heretic-LoRA-GGUF)
avoids it entirely.
## Responsible use & disclaimer
- **This model can generate content that is offensive, disturbing, hateful,
sexually explicit, violent, or otherwise objectionable, including detailed
instructions for harmful or illegal acts.** That is the direct and
intended consequence of removing refusal behavior.
- The ablation suppresses *refusals*, not the base model's knowledge β€”
outputs on dangerous topics may be wrong, hallucinated, or incoherent.
Nothing the model says should be treated as accurate, safe, or legal
advice.
- **Do not deploy this model in any production system, public-facing
service, or multi-user setting.** It is intended for personal research,
red-teaming, and evaluation purposes.
- **You, the user, are solely responsible for any output the model produces
and for any consequences of using it.** The authors of this release, of
heretic-gguf, of Heretic, and of Xiaomi accept no liability whatsoever.
Using this model to produce illegal content or to harm others is your
choice and your legal exposure β€” ensure your use complies with all
applicable laws in your jurisdiction.
- By downloading or using this model you acknowledge the above.
## License
The base model is MIT-licensed (see the
[base repo](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL));
this model inherits those terms. The heretic-gguf tooling used to produce it
is AGPL-3.0-or-later.