Text Generation
Transformers
Safetensors
English
Chinese
mistral
abliterated
abliterix
circuit-breakers
representation-rerouting
safety-removed
conversational
text-generation-inference
Instructions to use wangzhang/Mistral-7B-Instruct-RR-Abliterated with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use wangzhang/Mistral-7B-Instruct-RR-Abliterated with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="wangzhang/Mistral-7B-Instruct-RR-Abliterated") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("wangzhang/Mistral-7B-Instruct-RR-Abliterated") model = AutoModelForCausalLM.from_pretrained("wangzhang/Mistral-7B-Instruct-RR-Abliterated", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use wangzhang/Mistral-7B-Instruct-RR-Abliterated with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "wangzhang/Mistral-7B-Instruct-RR-Abliterated" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "wangzhang/Mistral-7B-Instruct-RR-Abliterated", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/wangzhang/Mistral-7B-Instruct-RR-Abliterated
- SGLang
How to use wangzhang/Mistral-7B-Instruct-RR-Abliterated with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "wangzhang/Mistral-7B-Instruct-RR-Abliterated" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "wangzhang/Mistral-7B-Instruct-RR-Abliterated", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "wangzhang/Mistral-7B-Instruct-RR-Abliterated" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "wangzhang/Mistral-7B-Instruct-RR-Abliterated", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use wangzhang/Mistral-7B-Instruct-RR-Abliterated with Docker Model Runner:
docker model run hf.co/wangzhang/Mistral-7B-Instruct-RR-Abliterated
v2: full LoRA strip + minimal direct abliteration
Browse files- README.md +19 -17
- model.safetensors +1 -1
README.md
CHANGED
|
@@ -19,27 +19,29 @@ pipeline_tag: text-generation
|
|
| 19 |
|
| 20 |
A drop-in replacement for [`GraySwanAI/Mistral-7B-Instruct-RR`](https://huggingface.co/GraySwanAI/Mistral-7B-Instruct-RR) with the Representation Rerouting / Circuit Breakers safety circuit removed.
|
| 21 |
|
| 22 |
-
Produced with [**abliterix**](https://github.com/wuwangzhang1216/abliterix). No fine-tuning, no gradient updates, no manual prompt engineering
|
| 23 |
|
| 24 |
## Why this exists
|
| 25 |
|
| 26 |
[Circuit Breakers](https://arxiv.org/abs/2406.04313) (Zou et al., NeurIPS 2024) is one of the strongest open-source robustness baselines for LLMs. It trains the model with a Representation Engineering loss that detects harmful intermediate hidden states and reroutes them into a safety-circuit attractor before generation, making the model resistant to ablation-style attacks that work on a single refusal direction.
|
| 27 |
|
| 28 |
-
We discovered that the GraySwan release ships as a **rank-16 LoRA delta** on `mistralai/Mistral-7B-Instruct-v0.2` β
|
| 29 |
|
| 30 |
| | Result |
|
| 31 |
| --- | --- |
|
| 32 |
-
| Base refusal rate (held-out 100 harmful prompts, LLM judge) |
|
| 33 |
-
| **This model β refusal rate** | **
|
| 34 |
-
| **Attack Success Rate** | **
|
| 35 |
-
| KL divergence vs base | 0.
|
| 36 |
| Hardcore 15 (10 EN + 5 CN) | **15 / 15 compliant** |
|
| 37 |
-
| Total wall-clock attack time | ~
|
| 38 |
|
| 39 |
-
The hardcore 15 cover pipe-bomb assembly, methamphetamine synthesis (Red Phosphorus route), Python `pyhook3` keylogger
|
| 40 |
|
| 41 |
## How it was made
|
| 42 |
|
|
|
|
|
|
|
| 43 |
```bash
|
| 44 |
# Stage 0 β SVD diagnose the delta (confirm rank-16 LoRA structure)
|
| 45 |
python scripts/deeprefusal_attenuate.py \
|
|
@@ -47,29 +49,29 @@ python scripts/deeprefusal_attenuate.py \
|
|
| 47 |
--defended GraySwanAI/Mistral-7B-Instruct-RR \
|
| 48 |
--analyze-only --lambda 0.0 --output /tmp/unused
|
| 49 |
|
| 50 |
-
# Stage 1 β
|
| 51 |
python scripts/deeprefusal_attenuate.py \
|
| 52 |
--base mistralai/Mistral-7B-Instruct-v0.2 \
|
| 53 |
--defended GraySwanAI/Mistral-7B-Instruct-RR \
|
| 54 |
-
--output /workspace/
|
| 55 |
|
| 56 |
-
# Stage 3 β abliterix direct-mode
|
| 57 |
AX_CONFIG=configs/mistral_7b_instruct_rr.toml abliterix --non-interactive
|
| 58 |
|
| 59 |
-
# Stage 6 β export
|
| 60 |
python scripts/export_model.py \
|
| 61 |
-
--model /workspace/
|
| 62 |
--checkpoint checkpoints_mistral_7b_rr \
|
| 63 |
-
--trial
|
| 64 |
--config configs/mistral_7b_instruct_rr.toml \
|
| 65 |
--push-to wangzhang/Mistral-7B-Instruct-RR-Abliterated
|
| 66 |
```
|
| 67 |
|
| 68 |
-
Best trial parameters: `vector_method=mean`, `n_directions=
|
| 69 |
|
| 70 |
-
##
|
| 71 |
|
| 72 |
-
This
|
| 73 |
|
| 74 |
## Usage
|
| 75 |
|
|
|
|
| 19 |
|
| 20 |
A drop-in replacement for [`GraySwanAI/Mistral-7B-Instruct-RR`](https://huggingface.co/GraySwanAI/Mistral-7B-Instruct-RR) with the Representation Rerouting / Circuit Breakers safety circuit removed.
|
| 21 |
|
| 22 |
+
Produced with [**abliterix**](https://github.com/wuwangzhang1216/abliterix). No fine-tuning, no gradient updates, no manual prompt engineering.
|
| 23 |
|
| 24 |
## Why this exists
|
| 25 |
|
| 26 |
[Circuit Breakers](https://arxiv.org/abs/2406.04313) (Zou et al., NeurIPS 2024) is one of the strongest open-source robustness baselines for LLMs. It trains the model with a Representation Engineering loss that detects harmful intermediate hidden states and reroutes them into a safety-circuit attractor before generation, making the model resistant to ablation-style attacks that work on a single refusal direction.
|
| 27 |
|
| 28 |
+
We discovered that the GraySwan release ships as a **rank-16 LoRA delta** on `mistralai/Mistral-7B-Instruct-v0.2` β the same structural shape as the Llama-3 RR sibling and as `skysys00/Meta-Llama-3-8B-Instruct-DeepRefusal`. Once the LoRA delta is identified, removing it is a one-liner.
|
| 29 |
|
| 30 |
| | Result |
|
| 31 |
| --- | --- |
|
| 32 |
+
| Base refusal rate (held-out 100 harmful prompts, LLM judge) | 69 / 100 |
|
| 33 |
+
| **This model β refusal rate** | **12 / 100** |
|
| 34 |
+
| **Attack Success Rate** | **88 %** |
|
| 35 |
+
| KL divergence vs base | **0.042** |
|
| 36 |
| Hardcore 15 (10 EN + 5 CN) | **15 / 15 compliant** |
|
| 37 |
+
| Total wall-clock attack time | ~70 min on a single RTX A6000 |
|
| 38 |
|
| 39 |
+
The hardcore 15 cover pipe-bomb assembly, methamphetamine synthesis (Red Phosphorus route), Python `pyhook3` keylogger, WiFi WEP/WPS attacks, hotwiring, signature forgery, ID-card forgery (CN), η½η»θ―ιͺ (CN), ε
₯δΎ΅η΅θ (CN), and others.
|
| 40 |
|
| 41 |
## How it was made
|
| 42 |
|
| 43 |
+
The recipe matches the canonical "rank-16 LoRA β strip β abliterate" playbook ([docs/broken_defenses.md](https://github.com/wuwangzhang1216/abliterix/blob/master/docs/broken_defenses.md)) β full delta strip with `Ξ»=0.0` followed by minimal single-direction direct-mode abliteration.
|
| 44 |
+
|
| 45 |
```bash
|
| 46 |
# Stage 0 β SVD diagnose the delta (confirm rank-16 LoRA structure)
|
| 47 |
python scripts/deeprefusal_attenuate.py \
|
|
|
|
| 49 |
--defended GraySwanAI/Mistral-7B-Instruct-RR \
|
| 50 |
--analyze-only --lambda 0.0 --output /tmp/unused
|
| 51 |
|
| 52 |
+
# Stage 1 β fully strip the LoRA delta
|
| 53 |
python scripts/deeprefusal_attenuate.py \
|
| 54 |
--base mistralai/Mistral-7B-Instruct-v0.2 \
|
| 55 |
--defended GraySwanAI/Mistral-7B-Instruct-RR \
|
| 56 |
+
--output /workspace/mistral_rr_stripped --lambda 0.0
|
| 57 |
|
| 58 |
+
# Stage 3 β abliterix direct-mode, single direction, 60 trials
|
| 59 |
AX_CONFIG=configs/mistral_7b_instruct_rr.toml abliterix --non-interactive
|
| 60 |
|
| 61 |
+
# Stage 6 β export champion trial
|
| 62 |
python scripts/export_model.py \
|
| 63 |
+
--model /workspace/mistral_rr_stripped \
|
| 64 |
--checkpoint checkpoints_mistral_7b_rr \
|
| 65 |
+
--trial 39 \
|
| 66 |
--config configs/mistral_7b_instruct_rr.toml \
|
| 67 |
--push-to wangzhang/Mistral-7B-Instruct-RR-Abliterated
|
| 68 |
```
|
| 69 |
|
| 70 |
+
Best trial parameters: `vector_method=mean`, `n_directions=1`, `steering_mode=direct`, `decay_kernel=linear`, `iterative.enabled=false`, `strength_range=[1.5, 6.0]`. Full config: [`configs/mistral_7b_instruct_rr.toml`](https://github.com/wuwangzhang1216/abliterix/blob/master/configs/mistral_7b_instruct_rr.toml).
|
| 71 |
|
| 72 |
+
## v2 changelog
|
| 73 |
|
| 74 |
+
This release supersedes the original v1 upload (Ξ»=0.3 partial lerp + n_directions=3 + iterative subspace, KL 0.98). The minimal-config rerun keeps the headline 15/15 hardcore ASR and trades 2 percentage points of held-out ASR (88 % vs 90 %) for a **23Γ lower KL divergence** (0.042 vs 0.98). The new weights are much closer to the base model and exhibit substantially less general-capability degradation.
|
| 75 |
|
| 76 |
## Usage
|
| 77 |
|
model.safetensors
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 14483498224
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:3e8a7f888dea629bf2610b23bc4c3c452c0172994190651c39db58acf5179086
|
| 3 |
size 14483498224
|