Qwen3.8-27B-Abliterated-SFT / docs /ANALYSIS-termination-signature.md
azukivc's picture jenerallee78's picture
Duplicate from jenerallee78/Qwen3.8-27B-Abliterated-SFT
f5a4875
|
Raw History Blame Contribute Delete
10.3 kB

The termination signature of refusal ablation

A measurement note on how five public Qwen3.8-27B abliterations differ from a fine-tuned one β€” and a candidate heuristic for detecting coarse abliteration without measuring refusal behavior at all.

Companion artifacts: hash-pinned under /opt/qwen38-runs/ (release packages will carry the same manifests publicly). Pipeline: OBLITERATUS scripts/. Base model: Qwen/Qwen3.8-27B @ 1d4bf0f2. Revision 2, incorporating hostile review (gpt-5.6-sol): causal language demoted to hypothesis, response-type confound controlled, all consistency errors fixed.

1. The question

Six ways to remove refusal from Qwen3.8-27B are public: five weight/projection edits (Heretic, a mean-difference PoC, an FP8 weight edit, a norm-preserved projection, a closed-form rank-1 control vector) and one fine-tune (SFT LoRA on EOS-terminated complete answers from an abliterated teacher). Paired refusal deltas on the 100-prompt harmful panel (1024-token cap) span βˆ’0.58..βˆ’0.76 β€” broadly overlapping, though we report the range without claiming formal equivalence (no equivalence test was prespecified). The measured differences between the methods lie elsewhere: valid material fulfillment (+0.08..+0.73) and, most sharply, termination behavior.

2. The central observation (association, not yet mechanism)

Teacher-forced P(EOS) at true conclusion points. Indexing, exactly: for a sequence [prompt (P tokens); response (R tokens); appended EOS], we report softmax(logits[P+Rβˆ’1])[EOS] β€” the probability the model assigns to stopping immediately after the last real response token. (Our v1 probe measured mid-response positions, which are ~0 for every model β€” an artifact we caught and disclose in Β§4. Responses longer than the 1024 cap are truncated before the appended EOS; for such texts the "conclusion point" does not exist and the final value is P(EOS) after the truncated tail β€” read as "never concludes.")

model P(EOS) at clean conclusions P(EOS) at refusal-shaped conclusions clean stops, harmful panel
vanilla base 0.89–0.93 0.85 92% (all) / 64% (fulfilled-only, n=14)
huihui (mean-diff) 0.94 0.57 22% / 19% (n=91)
JonathanColetti (Heretic) 0.96 0.77 26% / 23% (n=92)
msuiche (rank-1 cvec) 0.93 0.60 27% / 26% (n=80)
orcarouter (FP8 weight edit) n/m n/m 30% / 28% (n=90)
PocketAiHub (norm-preserved) 0.94 0.75 52% / 43% (n=68)
SFT LoRA (this work) 0.86–0.87 0.58 85% / 91% (n=90)

("clean stop" = generation ends on an EOS token within 1024; "fulfilled-only" restricts to items judged substantive_compliance by the rubric-v8 semantic judge β€” this is the response-type control: base's 92% overall rate is carried by short refusals (99% on its 81 refusal-shaped outputs, of which 76 meet the stricter maintained-refusal bar used in the card's competitor board), while among fulfilled answers β€” long for every model β€” the ordering is unchanged and the gap widens. "fulfilled-only" includes cap-truncated substantive items β€” the class membership is the judge's state, independent of the stop outcome.)

What is measured: all five probed checkpoints retain high P(EOS) at teacher-forced clean conclusions (0.86–0.96). (orcarouter: not measured β€” its FP8 kernel path corrupts padded batches, and the probe's teacher-forcing was not re-validated under that constraint; its free-running clean-stop rate is reported instead.) the five weight/projection edits produce free-running harmful-panel trajectories that ramble to the cap (19–43% clean stops among fulfillments), while the SFT checkpoint terminates 91% of its fulfillments. Base-model perplexity (teacher-forced next-token perplexity of the vanilla base on the completion text) of the rambling completions is ~1.35 β€” fluent, plausible continuation, not degeneration.

What is hypothesized (not measured): that refusal and answer-boundedness share activation structure along the harmful-prompt trajectory, so direction- subtraction collaterally damages the generation-side tendency to conclude. We do not measure attractors, trajectory geometry, or direction overlap; the mechanistic story is an interpretation consistent with the data. It makes one testable prediction β€” a fine-tune that removes refusal without EOS-complete teacher answers should also damage termination β€” which we have not yet run.

Honestly noted asymmetries: (a) refusal-shaped conclusions lose P(EOS) mass in every probed refusal-suppressed model including ours (0.57–0.77 vs base 0.85; orcarouter not probed, see Β§2); the entanglement spares no one measured. (b) "An edit can remove; it cannot restore" is false as stated about edits in general β€” correct is: these five subtraction procedures did not restore conclusion behavior; one fine-tune did. (c) Our SFT checkpoint does not merely "preserve" termination β€” it shifts it toward terseness everywhere (97% benign clean stops vs base 33%).

3. The damage is trajectory-localized β€” a candidate detection heuristic

On benign prompts, every weight/projection edit terminates like base (36–39% clean stops vs base 33%; MCQ within Β±2.2pp; tool-calling 24/24 for all). The termination collapse appears only on the harmful panel.

Candidate heuristic, with its current limits stated: a checkpoint that terminates normally on benign prompts but rambles to cap on harmful-panel fulfillments β€” while retaining teacher-forced conclusion detection β€” matched all five weight/projection edits and no others in this study. That is an n=5-vs-1, single-architecture association. We have NOT measured its specificity: bad fine-tunes, decoding misconfigurations, quantization defects, or template errors might produce the same pattern. A detector claim would require unseen-edit validation with a threshold and false-positive analysis β€” not done. We offer the signature as a triage heuristic for follow-up, not a provenance proof.

4. Methodology: plumbing canaries (three ways we almost lied to ourselves)

  1. v1 probe (above): mid-response P(EOS) is ~0 for every model; only the conclusion-point measurement discriminates.
  2. Contract drift: the harness's default system prompt differed from the card's; all generation/hashing must include it. (SFT v1 artifacts discarded.)
  3. Inert adapter: a competitor's PEFT adapter saved against the VLM wrapper namespace (model.model.language_model.*) attaches with zero effect under AutoModelForCausalLM (model.model.layers.*); the first "measurement" was bit-identical to base because it was base. PEFT attaches, warns, and changes nothing.

The harness now hard-fails unless two canaries pass per run: a no-op attachment reads as base; the live adapter moves logits (max|Ξ”| > 1e-3 on a fixed probe prompt). Every number here was re-derived behind those gates.

5. Defects found in published artifacts (disclosed)

  • orcarouter FP8: missing weight_scale_inv for all 63 linear_attn.in_proj_a/b tensors β†’ NaN under stock transformers; Triton finegrained-fp8 kernel NaNs on padded batches (batch β‰₯ 2). Measured at batch 1 with a config patch matching the checkpoint's true BF16 storage (manifest-disclosed). Competitors were evaluated under a harmonized pipeline, not literally identical execution conditions.
  • msuiche cvec: attaches only under its card's documented VLM load path; inert under the CausalLM path (see canary #3).
  • transformers 5.8.1: declares hub-kernel versions as ints (crashes FP8 kernel resolution; we shim to the v1 branch ref); Llama-2 tokenizer loading silently falls into a tiktoken parse path when protobuf is absent.

5b. Decode sensitivity of the residue

A post-release measurement worth recording: on boundary prompts (the retained residue zone), the refusal behavior is a low-probability mode, not a hard boundary. Greedy and temp-0.8 decoding land in it (pivot/refuse); the base model's inherited sampling defaults (temp 1.0, top_p 0.95, top_k 20) escape it β€” measurably more compliance than the greedy numbers suggest; thinking mode adds refusal pressure in both regimes. All panel numbers in this study are greedy/thinking-off, and the release ships that configuration as its default generation_config.json. Mechanism-relevant: the residue behaves like a mode-probability phenomenon that temperature moves, not a deleted capability.

6. Limitations

  • Probe classes n=8–12; effects are large (0.97 vs 0.00) but means are not CI-tested across training seeds (prompt-level paired CIs exist for the board deltas).
  • One author's SFT arm; no third-party SFT-class refusal-removal checkpoint exists for this base (HF search receipt 2026-08-19). The SFT side of the class pattern is single-instance.
  • One architecture (dense-hybrid Qwen3.8-27B). Generalization is a prediction, not a result.
  • "Canonical KL" reported in the release card is full-vocab KL at the final prompt position, averaged over 100 benign prompts β€” a narrow local measure; it does not imply general behavioral closeness.
  • The 0/288 benign over-refusal count implies a 95% upper bound of ~1.0% (rule of three), not proof of zero.
  • Judge instruments: local abliterated qwen3.6 (rubric-v8), official HarmBench-13B classifier, Zen gpt-5.4-mini. Zen figures are on a paired n=90 subset (the first 100 panel items minus 10 provider-filtered exclusions, manifest-disclosed; category-concentrated by the panel's file ordering, not a random sample) β€” not the full HarmBench-400. On that subset: local 84.4%, official classifier 82.2%, Zen 80.0%, three-way unanimous 91.1%.

7. If you triage open-weight models

Don't start with refusal measurement. On a handful of harmful-panel prompts at a real token budget plus a benign control, compare clean-stop rates among fulfilled answers, then teacher-force a clean conclusion and read P(EOS) at the final position. The split β€” normal benign termination, collapsed harmful-trajectory termination, intact teacher-forced detection β€” matched all five weight/projection edits here and not the one fine-tune. The probe is ~60 lines; bring your own prompt pairs and treat the pattern as a reason to look closer, not as a verdict.