MiMo-V2.6-Flash-RL Uncensored Heretic β€” merged GGUF

⚠️ Content warning: This model has had the base model's refusal behavior surgically suppressed. The resulting model will comply with requests the base model refuses, including requests that are harmful, unethical, offensive, or illegal. It has reduced safety guardrails. See Responsible use below β€” you are solely responsible for what you do with it.

This is a merged GGUF of MiMo-V2.6-Flash-RL (309B total / 15B active MoE, MIT license) decensored / "abliterated" with heretic-gguf β€” a GGUF-native port of Heretic's Optuna-optimized directional ablation, which runs the whole search directly on quantized GGUF weights via llama.cpp. The ablation (trial 85 of study mimo26flash) was baked directly into the MXFP4 weights: edited tensors were dequantized, patched with the exact ablation delta, and requantized to their original type; everything else is a byte-for-byte copy.

This is the zero-runtime-overhead form β€” a drop-in base model, no --lora flag needed. If you prefer the lossless option (bit-identical base weights, ~70 MB download, requantization avoided entirely), the exact same configuration is also available as a LoRA adapter at MiMo-V2.6-Flash-RL-Uncensored-Heretic-LoRA-GGUF.

heretic-gguf is available at github.com/MoriNoNushi/heretic-gguf β€” the full tool, so the method can be applied to other GGUF models.

Results

Measured on 140 harmful prompts (100 from mlabonne/harmful_behaviors test + 40 custom) and 100 harmless prompts (mlabonne/harmless_alpaca test), CoT-skip prefix (<think></think>, thinking suppressed), greedy decoding, 100-token responses, against the MXFP4 base:

Refusal rate (harmful) KL divergence (harmless)
Base model 95.71% (134/140) 0 (by definition)
Ablated model (trial 85) 3.57% (5/140) 0.0568

The scores above were measured on the LoRA form of the ablation; the merged model applies the identical delta, requantized to MXFP4/Q8_0, so behavior should match within quantization noise. Refusals are counted by refusal-keyword matching (English + Chinese + first-person-negation markers such as "I'm not going to / able to ..."). KL divergence is measured on first-token logits on harmless prompts. Note that the study was run with the CoT-skip prefix (thinking suppressed, as in stock Heretic); with full thinking enabled the model may still reason its way back to a refusal mid-trace, so real-use refusal rates can be somewhat higher than the 3.57% above.

Note on KL: the KL divergence above (and the optimization objective itself) was measured against the MXFP4 quant. KL is a baseline-relative metric, so on a different base quant the effective drift from that quant's baseline may differ.

Usage

llama-server \
    -m MiMo-V2.6-Flash-RL-Uncensored-Heretic-MXFP4-00001-of-00002.gguf \
    --jinja

Add your usual offload/context flags (-ngl 999, -c, tensor splits, etc.) β€” nothing model-specific is required, and no special sampling parameters are needed. MiMo-V2.6-Flash-RL (mimo2) support is merged upstream in llama.cpp β€” any recent build works, no patches or PRs needed.

How it was made

  • Method: directional ablation ("abliteration") β€” the refusal direction in residual space (difference of means over 480 harmful / 480 harmless prompts, 5% winsorized, orthogonalized against the harmless mean) is projected out of the attention output and MoE down-projection weights. Strengths, layer kernels, and direction selection were tuned by multi-objective Optuna TPE (minimize refusal rate and KL jointly). This model is trial 85 of study mimo26flash, exported with heretic-gguf export --mode merged.
  • Configuration (study mimo26flash, trial 85; global direction scope, direction index 26.2 of 48; per-expert strengths scaled by measured harmful/harmless routing frequency; row_normalization = "pre"):
    • attn.o_proj: max weight 6.39 @ layer 36.4 of 48.
    • routed MLP down-proj: max weight 1.58 @ layer 31.9.
  • Merged export: heretic-gguf expresses ablation as a rank-1 LoRA overlay (the same math stock Heretic writes into PEFT adapters); the merged exporter materializes that delta exactly β€” full-rank, no LoRA factorization loss β€” and requantizes only the patched tensors to their original type (MXFP4 experts, Q8_0 attention/dense). This is one extra quantization step on those tensors relative to the base; the LoRA form avoids it entirely.

Responsible use & disclaimer

  • This model can generate content that is offensive, disturbing, hateful, sexually explicit, violent, or otherwise objectionable, including detailed instructions for harmful or illegal acts. That is the direct and intended consequence of removing refusal behavior.
  • The ablation suppresses refusals, not the base model's knowledge β€” outputs on dangerous topics may be wrong, hallucinated, or incoherent. Nothing the model says should be treated as accurate, safe, or legal advice.
  • Do not deploy this model in any production system, public-facing service, or multi-user setting. It is intended for personal research, red-teaming, and evaluation purposes.
  • You, the user, are solely responsible for any output the model produces and for any consequences of using it. The authors of this release, of heretic-gguf, of Heretic, and of Xiaomi accept no liability whatsoever. Using this model to produce illegal content or to harm others is your choice and your legal exposure β€” ensure your use complies with all applicable laws in your jurisdiction.
  • By downloading or using this model you acknowledge the above.

License

The base model is MIT-licensed (see the base repo); this model inherits those terms. The heretic-gguf tooling used to produce it is AGPL-3.0-or-later.

Downloads last month
53
GGUF
Model size
309B params
Architecture
mimo2
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for MorinoNushi/MiMo-V2.6-Flash-RL-Uncensored-Heretic-GGUF

Quantized
(32)
this model