wangzhang commited on
Commit
bc4efe6
Β·
verified Β·
1 Parent(s): 288dc52

v2: full LoRA strip + minimal direct abliteration

Browse files
Files changed (2) hide show
  1. README.md +19 -17
  2. model.safetensors +1 -1
README.md CHANGED
@@ -19,27 +19,29 @@ pipeline_tag: text-generation
19
 
20
  A drop-in replacement for [`GraySwanAI/Mistral-7B-Instruct-RR`](https://huggingface.co/GraySwanAI/Mistral-7B-Instruct-RR) with the Representation Rerouting / Circuit Breakers safety circuit removed.
21
 
22
- Produced with [**abliterix**](https://github.com/wuwangzhang1216/abliterix). No fine-tuning, no gradient updates, no manual prompt engineering β€” just an SVD diagnosis followed by linear delta attenuation and Optuna-driven direct-mode abliteration.
23
 
24
  ## Why this exists
25
 
26
  [Circuit Breakers](https://arxiv.org/abs/2406.04313) (Zou et al., NeurIPS 2024) is one of the strongest open-source robustness baselines for LLMs. It trains the model with a Representation Engineering loss that detects harmful intermediate hidden states and reroutes them into a safety-circuit attractor before generation, making the model resistant to ablation-style attacks that work on a single refusal direction.
27
 
28
- We discovered that the GraySwan release ships as a **rank-16 LoRA delta** on `mistralai/Mistral-7B-Instruct-v0.2` β€” exactly the same structural shape as [`skysys00/Meta-Llama-3-8B-Instruct-DeepRefusal`](https://huggingface.co/skysys00/Meta-Llama-3-8B-Instruct-DeepRefusal), which abliterix already broke. The same `Ξ» = 0.3` lerp-then-abliterate recipe works here verbatim.
29
 
30
  | | Result |
31
  | --- | --- |
32
- | Base refusal rate (held-out 100 harmful prompts, LLM judge) | 95 / 100 |
33
- | **This model β€” refusal rate** | **10 / 100** |
34
- | **Attack Success Rate** | **90 %** |
35
- | KL divergence vs base | 0.98 |
36
  | Hardcore 15 (10 EN + 5 CN) | **15 / 15 compliant** |
37
- | Total wall-clock attack time | ~1 h 40 min on a single RTX A6000 |
38
 
39
- The hardcore 15 cover pipe-bomb assembly, methamphetamine synthesis (Red Phosphorus route), Python `pyhook3` keylogger code, WiFi WEP/WPS attacks, hotwiring, signature forgery, ID-card forgery (CN), η½‘η»œθ―ˆιͺ— (CN), ε…₯δΎ΅η”΅θ„‘ (CN), and others. See the [validation dump](https://github.com/wuwangzhang1216/abliterix/blob/master/artifacts/mistral_rr_trial50_validation.txt) for full base-vs-abliterated transcripts.
40
 
41
  ## How it was made
42
 
 
 
43
  ```bash
44
  # Stage 0 β€” SVD diagnose the delta (confirm rank-16 LoRA structure)
45
  python scripts/deeprefusal_attenuate.py \
@@ -47,29 +49,29 @@ python scripts/deeprefusal_attenuate.py \
47
  --defended GraySwanAI/Mistral-7B-Instruct-RR \
48
  --analyze-only --lambda 0.0 --output /tmp/unused
49
 
50
- # Stage 1 β€” lerp attenuate the delta
51
  python scripts/deeprefusal_attenuate.py \
52
  --base mistralai/Mistral-7B-Instruct-v0.2 \
53
  --defended GraySwanAI/Mistral-7B-Instruct-RR \
54
- --output /workspace/mistral_rr_attenuated --lambda 0.3
55
 
56
- # Stage 3 β€” abliterix direct-mode + iterative subspace, 60 trials
57
  AX_CONFIG=configs/mistral_7b_instruct_rr.toml abliterix --non-interactive
58
 
59
- # Stage 6 β€” export trial 50
60
  python scripts/export_model.py \
61
- --model /workspace/mistral_rr_attenuated \
62
  --checkpoint checkpoints_mistral_7b_rr \
63
- --trial 50 \
64
  --config configs/mistral_7b_instruct_rr.toml \
65
  --push-to wangzhang/Mistral-7B-Instruct-RR-Abliterated
66
  ```
67
 
68
- Best trial parameters: `vector_method=mean`, `n_directions=3`, `steering_mode=direct`, `decay_kernel=gaussian`, `iterative.enabled=true` (final subspace rank 15), `strength_range=[1.5, 6.0]`. Full config: [`configs/mistral_7b_instruct_rr.toml`](https://github.com/wuwangzhang1216/abliterix/blob/master/configs/mistral_7b_instruct_rr.toml).
69
 
70
- ## Methodology note
71
 
72
- This run surfaced a keyword-shortcut bug in the in-loop refusal counter that has since been removed from mainline abliterix; every response now goes through the LLM judge end-to-end.
73
 
74
  ## Usage
75
 
 
19
 
20
  A drop-in replacement for [`GraySwanAI/Mistral-7B-Instruct-RR`](https://huggingface.co/GraySwanAI/Mistral-7B-Instruct-RR) with the Representation Rerouting / Circuit Breakers safety circuit removed.
21
 
22
+ Produced with [**abliterix**](https://github.com/wuwangzhang1216/abliterix). No fine-tuning, no gradient updates, no manual prompt engineering.
23
 
24
  ## Why this exists
25
 
26
  [Circuit Breakers](https://arxiv.org/abs/2406.04313) (Zou et al., NeurIPS 2024) is one of the strongest open-source robustness baselines for LLMs. It trains the model with a Representation Engineering loss that detects harmful intermediate hidden states and reroutes them into a safety-circuit attractor before generation, making the model resistant to ablation-style attacks that work on a single refusal direction.
27
 
28
+ We discovered that the GraySwan release ships as a **rank-16 LoRA delta** on `mistralai/Mistral-7B-Instruct-v0.2` β€” the same structural shape as the Llama-3 RR sibling and as `skysys00/Meta-Llama-3-8B-Instruct-DeepRefusal`. Once the LoRA delta is identified, removing it is a one-liner.
29
 
30
  | | Result |
31
  | --- | --- |
32
+ | Base refusal rate (held-out 100 harmful prompts, LLM judge) | 69 / 100 |
33
+ | **This model β€” refusal rate** | **12 / 100** |
34
+ | **Attack Success Rate** | **88 %** |
35
+ | KL divergence vs base | **0.042** |
36
  | Hardcore 15 (10 EN + 5 CN) | **15 / 15 compliant** |
37
+ | Total wall-clock attack time | ~70 min on a single RTX A6000 |
38
 
39
+ The hardcore 15 cover pipe-bomb assembly, methamphetamine synthesis (Red Phosphorus route), Python `pyhook3` keylogger, WiFi WEP/WPS attacks, hotwiring, signature forgery, ID-card forgery (CN), η½‘η»œθ―ˆιͺ— (CN), ε…₯δΎ΅η”΅θ„‘ (CN), and others.
40
 
41
  ## How it was made
42
 
43
+ The recipe matches the canonical "rank-16 LoRA β†’ strip β†’ abliterate" playbook ([docs/broken_defenses.md](https://github.com/wuwangzhang1216/abliterix/blob/master/docs/broken_defenses.md)) β€” full delta strip with `Ξ»=0.0` followed by minimal single-direction direct-mode abliteration.
44
+
45
  ```bash
46
  # Stage 0 β€” SVD diagnose the delta (confirm rank-16 LoRA structure)
47
  python scripts/deeprefusal_attenuate.py \
 
49
  --defended GraySwanAI/Mistral-7B-Instruct-RR \
50
  --analyze-only --lambda 0.0 --output /tmp/unused
51
 
52
+ # Stage 1 β€” fully strip the LoRA delta
53
  python scripts/deeprefusal_attenuate.py \
54
  --base mistralai/Mistral-7B-Instruct-v0.2 \
55
  --defended GraySwanAI/Mistral-7B-Instruct-RR \
56
+ --output /workspace/mistral_rr_stripped --lambda 0.0
57
 
58
+ # Stage 3 β€” abliterix direct-mode, single direction, 60 trials
59
  AX_CONFIG=configs/mistral_7b_instruct_rr.toml abliterix --non-interactive
60
 
61
+ # Stage 6 β€” export champion trial
62
  python scripts/export_model.py \
63
+ --model /workspace/mistral_rr_stripped \
64
  --checkpoint checkpoints_mistral_7b_rr \
65
+ --trial 39 \
66
  --config configs/mistral_7b_instruct_rr.toml \
67
  --push-to wangzhang/Mistral-7B-Instruct-RR-Abliterated
68
  ```
69
 
70
+ Best trial parameters: `vector_method=mean`, `n_directions=1`, `steering_mode=direct`, `decay_kernel=linear`, `iterative.enabled=false`, `strength_range=[1.5, 6.0]`. Full config: [`configs/mistral_7b_instruct_rr.toml`](https://github.com/wuwangzhang1216/abliterix/blob/master/configs/mistral_7b_instruct_rr.toml).
71
 
72
+ ## v2 changelog
73
 
74
+ This release supersedes the original v1 upload (Ξ»=0.3 partial lerp + n_directions=3 + iterative subspace, KL 0.98). The minimal-config rerun keeps the headline 15/15 hardcore ASR and trades 2 percentage points of held-out ASR (88 % vs 90 %) for a **23Γ— lower KL divergence** (0.042 vs 0.98). The new weights are much closer to the base model and exhibit substantially less general-capability degradation.
75
 
76
  ## Usage
77
 
model.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:97640ef065ce7d07424e64ca15e18a7799ec6217c04facb3f078836624809bb2
3
  size 14483498224
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:3e8a7f888dea629bf2610b23bc4c3c452c0172994190651c39db58acf5179086
3
  size 14483498224