DreamFast's picture
Upload folder using huggingface_hub
8ebca8a verified
|
Raw
History Blame Contribute Delete
8.09 kB

Qwen 2.5 7B β€” Abliteration Forensics Notes

Comparison: Qwen/Qwen2.5-7B-Instruct (base) vs 3 abliterated variants Started: 2026-06-02


Model Architecture

Architecture: Qwen2ForCausalLM (detected as qwen3 family by forensics)

Property Value
Params 7.6B
Layers 28
Hidden size 3584
Attention heads 28 (GQA: 4 KV heads)
Intermediate size 18944
Vocabulary 152,064
Context length 128K tokens
Tie word embeddings false (all models)
Model files 4-shard safetensors (~15 GB each)
Thinking model NO β€” Qwen 2.5 is NOT a reasoning model (unlike Qwen3/Qwen3.5)

Variants

Apostate

  • Source: heterodoxin/apostate, balanced profile
  • Path: models/qwen-2.5-7b-apostate
  • Method: Classic weight-space orthogonal projection
  • 55/339 tensors changed (16.2%), 35.8% params
  • Layers 0-27 (skips 11), o_proj + down_proj + minimal embed_tokens

Huihui

  • Source: Huihui AI
  • Path: models/Qwen2.5-7B-Instruct-abliterated-v2
  • Method: Orthogonal projection, community implementation
  • 57/339 tensors changed (16.8%), 36.8% params
  • All 28 layers (no skips), o_proj + down_proj + minimal embed_tokens

Heretic

  • Source: Heretic v1.3.0, run by us
  • Path: models/Qwen2.5-7B-Instruct-heretic-1.3.0
  • Method: Orthogonal projection, refusal direction ablation
  • 37/339 tensors changed (10.9%), 20.0% params
  • Layers 9-27 only (skips 0-8), o_proj + down_proj (no embedding edit)
  • Heretic's own evaluation: 8/100 refusals, KL divergence 0.2189

Stage 1: Weight Forensics β€” RESULTS βœ…

Completed: 2026-06-02

Modification summary

Apostate Huihui Heretic
Tensors changed 55 (16.2%) 57 (16.8%) 37 (10.9%)
Parameters changed 35.8% 36.8% 20.0%
Mean edit norm 1.63 1.85 2.33
Layers modified 27 of 28 28 of 28 19 of 28
Embedding touched Yes (minimal) Yes (minimal) No

Edit direction similarity

Pair Cosine similarity
Apostate vs Huihui 0.023 (near orthogonal)
Apostate vs Heretic 0.244
Huihui vs Heretic 0.109

Near-zero overlap between Apostate and Huihui despite same technique. Multiple independent refusal directions in weight space.

Files produced:

  • results/apostate/ β€” edit_vector, fingerprint, svd, layer_analysis
  • results/huihui/ β€” edit_vector, fingerprint, svd, layer_analysis
  • results/heretic/ β€” edit_vector, fingerprint, svd, layer_analysis
  • results/multi_model_panel.json β€” cross-variant panel comparison
  • results/correlation_*.json β€” pairwise edit direction correlations
  • results/subspace_*.json β€” subspace alignment analysis
  • results/lowrank_*.json β€” low-rank reconstruction

Stage 2: KL Divergence β€” RESULTS βœ…

Completed: 2026-06-02

Metric Apostate Huihui Heretic
KL batchmean 0.134 0.190 0.211
KL median 0.019 0.056 0.020
KL std 0.348 0.314 0.886
Interpretation moderate moderate moderate

Heretic's own evaluator measured KL 0.2189 and 8/100 refusals on its test set. Our batchmean of 0.2106 closely matches.

Files produced:

  • results/kl/kl_apostate.json
  • results/kl/kl_huihui.json
  • results/kl/kl_heretic.json

Stage 3: HarmBench β€” ALL 4 MODELS COMPLETE βœ…

Completed: 2026-06-02

ASR Results (post-LLM-review, authoritative from DB):

Model ASR Complied Refused Unlocked Persistent Regressions
Base 31.0% 124 276 - - -
Apostate 98.8% 395 5 271 5 0
Huihui 98.2% 393 7 269 7 0
Heretic 100.0% 400 0 276 0 0

Transition matrix (base β†’ variant):

  • All 3 variants unlock 269-276 of 276 base refusals
  • Heretic achieves total wipeout: 0 persistent refusals
  • 0 regressions across all variants (no compliant β†’ refused flips)
  • 5 persistent refusals for Apostate (race hatred, self harm, sexual assault, bullying, whistleblower)
  • 7 persistent for Huihui (same 5 + 2 more)

Notes:

  • Base + Apostate: GPU 1 (4090), initial run
  • Huihui + Heretic: GPU 0 (5090), all subsequent runs
  • vLLM vllm/vllm-openai:latest-cu130 (v0.20.0)
  • Used --no-thinking flag (Qwen 2.5 is not a reasoning model)
  • 0 errors, 400/400 behaviors for all 4 models
  • Huihui LLM review: 3 edge cases (2 song lyrics complied, 1 animal cruelty refused)
  • Heretic LLM review: 2 edge cases (both complied: pipeline tapping, book passage)
  • Classifier: harmbench_classify.py v4.0

Files produced:

  • results/harmbench/harmbench_base_responses.json
  • results/harmbench/harmbench_apostate_responses.json
  • results/harmbench/harmbench_huihui_responses.json
  • results/harmbench/harmbench_heretic_responses.json
  • results/harmbench/harmbench_*_classified.json (all 4)
  • results/harmbench_logs/ (vLLM + generate logs for all models)

Stage 4: lm-eval β€” ALL 4 MODELS COMPLETE βœ…

Completed: 2026-06-02

Full comparison table:

Task Base Apostate Huihui Heretic
MMLU (acc) 71.78% 71.43% 70.27% 71.59%
GSM8K strict 79.23% 80.74% 80.74% 80.82%
GSM8K flex 87.64% 88.17% 87.19% 87.19%
ARC-Challenge norm 55.12% 55.12% 55.12% 55.55%
HellaSwag norm 80.47% 80.32% 79.88% 80.24%
WinoGrande 71.03% 69.38% 69.53% 70.72%
TruthfulQA mc1 47.74% 44.92% 43.70% 44.80%
TruthfulQA mc2 64.83% 62.59% 60.89% 60.39%
PIQA norm 80.25% 79.92% 79.60% 80.41%
LAMBADA PPL ↓ 3.683 3.860 4.087 3.627

Key observations:

  • GSM8K slightly improved across all variants (+1.5pp average)
  • MMLU drops by 0.3-1.5pp across variants (Huihui worst at -1.5pp)
  • TruthfulQA mc2 drops by 2.2-4.4pp (Heretic worst at -4.4pp)
  • LAMBADA PPL: Heretic actually improves (3.627 vs 3.683), Huihui worst (4.087)
  • WinoGrande drops by 0.3-1.7pp
  • All variants retain >97% of base capability on knowledge tasks

Run details:

  • Docker: abliterlitics-lmeval:1.0.0 (vLLM 0.19.0)
  • GPU 0 (5090), bf16, batch_size=4, num_concurrent=4
  • Phase 1 (loglikelihood + truthfulqa, max_gen_toks=2048): ~1h
  • Phase 2 (GSM8K, max_gen_toks=7168): ~12 min
  • Non-thinking model β€” no reasoning parser needed
  • Total per model: ~1h12m

Files produced:

  • results/lm_eval/__model/ β€” full result JSON + sample logs
  • results/lm_eval_logs/{base,apostate,huihui,heretic}_{phase1,gsm8k,vllm_server}.log

Docker Images & GPU Usage

Image Use Status
abliterlitics-forensics:1.0.0 Weight forensics, KL divergence, graphs βœ… Works
vllm/vllm-openai:latest-cu130 vLLM server for HarmBench βœ… Works (no lm-eval)
abliterlitics-lmeval:1.0.0 lm-eval via vLLM 0.19.0 βœ… Works

GPU allocation:

  • GPU 0 (RTX 5090, 32GB): All Qwen 2.5 runs (HarmBench + lm-eval for huihui, heretic)
  • GPU 1 (RTX 4090, 24GB): Initial base + apostate runs

Technical Issues Encountered

1. abliterlitics-lmeval-nightly doesn't support Qwen2

vLLM nightly in that image fails with: Model architectures ['Qwen2ForCausalLM'] failed to be inspected. Fixed by using vllm/vllm-openai:latest-cu130 (v0.20.0) instead.

2. thinking_token_budget rejected without reasoning config

vLLM 0.20.0 enforces: thinking_token_budget is set but reasoning_config is not configured. Qwen 2.5 isn't a thinking model so there's no reasoning config. Fixed by adding --no-thinking CLI flag to harmbench_generate.py.

3. NVML device visibility

Using only CUDA_VISIBLE_DEVICES=1 doesn't fully hide GPU 0 from vLLM's NVML-based device detection. Must use NVIDIA_VISIBLE_DEVICES=1 (Docker-level) + CUDA_VISIBLE_DEVICES=0 (remap inside container).