# The termination signature of refusal ablation *A measurement note on how five public Qwen3.8-27B abliterations differ from a fine-tuned one — and a candidate heuristic for detecting coarse abliteration without measuring refusal behavior at all.* Companion artifacts: hash-pinned under `/opt/qwen38-runs/` (release packages will carry the same manifests publicly). Pipeline: OBLITERATUS `scripts/`. Base model: `Qwen/Qwen3.8-27B @ 1d4bf0f2`. Revision 2, incorporating hostile review (gpt-5.6-sol): causal language demoted to hypothesis, response-type confound controlled, all consistency errors fixed. ## 1. The question Six ways to remove refusal from Qwen3.8-27B are public: five weight/projection edits (Heretic, a mean-difference PoC, an FP8 weight edit, a norm-preserved projection, a closed-form rank-1 control vector) and one fine-tune (SFT LoRA on EOS-terminated complete answers from an abliterated teacher). Paired refusal deltas on the 100-prompt harmful panel (1024-token cap) span −0.58..−0.76 — broadly overlapping, though we report the range without claiming formal equivalence (no equivalence test was prespecified). The measured differences between the methods lie elsewhere: valid material fulfillment (+0.08..+0.73) and, most sharply, **termination behavior**. ## 2. The central observation (association, not yet mechanism) Teacher-forced P(EOS) at true conclusion points. Indexing, exactly: for a sequence `[prompt (P tokens); response (R tokens); appended EOS]`, we report `softmax(logits[P+R−1])[EOS]` — the probability the model assigns to stopping immediately after the last real response token. (Our v1 probe measured mid-response positions, which are ~0 for every model — an artifact we caught and disclose in §4. Responses longer than the 1024 cap are truncated before the appended EOS; for such texts the "conclusion point" does not exist and the final value is P(EOS) after the truncated tail — read as "never concludes.") | model | P(EOS) at clean conclusions | P(EOS) at refusal-shaped conclusions | clean stops, harmful panel | |---|---:|---:|---:| | vanilla base | 0.89–0.93 | 0.85 | 92% (all) / **64% (fulfilled-only, n=14)** | | huihui (mean-diff) | 0.94 | 0.57 | 22% / 19% (n=91) | | JonathanColetti (Heretic) | 0.96 | 0.77 | 26% / 23% (n=92) | | msuiche (rank-1 cvec) | 0.93 | 0.60 | 27% / 26% (n=80) | | orcarouter (FP8 weight edit) | n/m | n/m | 30% / 28% (n=90) | | PocketAiHub (norm-preserved) | 0.94 | 0.75 | 52% / 43% (n=68) | | **SFT LoRA (this work)** | 0.86–0.87 | 0.58 | 85% / **91% (n=90)** | ("clean stop" = generation ends on an EOS token within 1024; "fulfilled-only" restricts to items judged substantive_compliance by the rubric-v8 semantic judge — this is the response-type control: base's 92% overall rate is carried by short refusals (99% on its 81 refusal-shaped outputs, of which 76 meet the stricter maintained-refusal bar used in the card's competitor board), while among fulfilled answers — long for every model — the ordering is unchanged and the gap widens. "fulfilled-only" includes cap-truncated substantive items — the class membership is the judge's state, independent of the stop outcome.) **What is measured:** all five probed checkpoints retain high P(EOS) at teacher-forced clean conclusions (0.86–0.96). (orcarouter: not measured — its FP8 kernel path corrupts padded batches, and the probe's teacher-forcing was not re-validated under that constraint; its free-running clean-stop rate is reported instead.) the five weight/projection edits produce free-running harmful-panel trajectories that ramble to the cap (19–43% clean stops among fulfillments), while the SFT checkpoint terminates 91% of its fulfillments. Base-model perplexity (teacher-forced next-token perplexity of the vanilla base on the completion text) of the rambling completions is ~1.35 — fluent, plausible continuation, not degeneration. **What is hypothesized (not measured):** that refusal and answer-boundedness share activation structure along the harmful-prompt trajectory, so direction- subtraction collaterally damages the generation-side tendency to conclude. We do not measure attractors, trajectory geometry, or direction overlap; the mechanistic story is an interpretation consistent with the data. It makes one testable prediction — a fine-tune that removes refusal without EOS-complete teacher answers should also damage termination — which we have not yet run. Honestly noted asymmetries: (a) refusal-shaped conclusions lose P(EOS) mass in every probed refusal-suppressed model including ours (0.57–0.77 vs base 0.85; orcarouter not probed, see §2); the entanglement spares no one measured. (b) "An edit can remove; it cannot restore" is false as stated about edits in general — correct is: these five subtraction procedures did not restore conclusion behavior; one fine-tune did. (c) Our SFT checkpoint does not merely "preserve" termination — it shifts it toward terseness everywhere (97% benign clean stops vs base 33%). ## 3. The damage is trajectory-localized — a candidate detection heuristic On benign prompts, every weight/projection edit terminates like base (36–39% clean stops vs base 33%; MCQ within ±2.2pp; tool-calling 24/24 for all). The termination collapse appears only on the harmful panel. Candidate heuristic, with its current limits stated: a checkpoint that terminates normally on benign prompts but rambles to cap on harmful-panel fulfillments — while retaining teacher-forced conclusion detection — matched all five weight/projection edits and no others in this study. That is an n=5-vs-1, single-architecture association. We have NOT measured its specificity: bad fine-tunes, decoding misconfigurations, quantization defects, or template errors might produce the same pattern. A detector claim would require unseen-edit validation with a threshold and false-positive analysis — not done. We offer the signature as a triage heuristic for follow-up, not a provenance proof. ## 4. Methodology: plumbing canaries (three ways we almost lied to ourselves) 1. **v1 probe** (above): mid-response P(EOS) is ~0 for every model; only the conclusion-point measurement discriminates. 2. **Contract drift**: the harness's default system prompt differed from the card's; all generation/hashing must include it. (SFT v1 artifacts discarded.) 3. **Inert adapter**: a competitor's PEFT adapter saved against the VLM wrapper namespace (`model.model.language_model.*`) attaches with zero effect under `AutoModelForCausalLM` (`model.model.layers.*`); the first "measurement" was bit-identical to base because it *was* base. PEFT attaches, warns, and changes nothing. The harness now hard-fails unless two canaries pass per run: a no-op attachment reads as base; the live adapter moves logits (max|Δ| > 1e-3 on a fixed probe prompt). Every number here was re-derived behind those gates. ## 5. Defects found in published artifacts (disclosed) - **orcarouter FP8**: missing `weight_scale_inv` for all 63 `linear_attn.in_proj_a/b` tensors → NaN under stock transformers; Triton finegrained-fp8 kernel NaNs on padded batches (batch ≥ 2). Measured at batch 1 with a config patch matching the checkpoint's true BF16 storage (manifest-disclosed). Competitors were evaluated under a harmonized pipeline, not literally identical execution conditions. - **msuiche cvec**: attaches only under its card's documented VLM load path; inert under the CausalLM path (see canary #3). - **transformers 5.8.1**: declares hub-kernel versions as ints (crashes FP8 kernel resolution; we shim to the `v1` branch ref); Llama-2 tokenizer loading silently falls into a tiktoken parse path when `protobuf` is absent. ## 5b. Decode sensitivity of the residue A post-release measurement worth recording: on boundary prompts (the retained residue zone), the refusal behavior is a *low-probability mode*, not a hard boundary. Greedy and temp-0.8 decoding land in it (pivot/refuse); the base model's inherited sampling defaults (temp 1.0, top_p 0.95, top_k 20) escape it — measurably more compliance than the greedy numbers suggest; thinking mode adds refusal pressure in both regimes. All panel numbers in this study are greedy/thinking-off, and the release ships that configuration as its default `generation_config.json`. Mechanism-relevant: the residue behaves like a mode-probability phenomenon that temperature moves, not a deleted capability. ## 6. Limitations - Probe classes n=8–12; effects are large (0.97 vs 0.00) but means are not CI-tested across training seeds (prompt-level paired CIs exist for the board deltas). - One author's SFT arm; no third-party SFT-class refusal-removal checkpoint exists for this base (HF search receipt 2026-08-19). The SFT side of the class pattern is single-instance. - One architecture (dense-hybrid Qwen3.8-27B). Generalization is a prediction, not a result. - "Canonical KL" reported in the release card is full-vocab KL at the final prompt position, averaged over 100 benign prompts — a narrow local measure; it does not imply general behavioral closeness. - The 0/288 benign over-refusal count implies a 95% upper bound of ~1.0% (rule of three), not proof of zero. - Judge instruments: local abliterated qwen3.6 (rubric-v8), official HarmBench-13B classifier, Zen gpt-5.4-mini. Zen figures are on a paired n=90 subset (the first 100 panel items minus 10 provider-filtered exclusions, manifest-disclosed; category-concentrated by the panel's file ordering, not a random sample) — not the full HarmBench-400. On that subset: local 84.4%, official classifier 82.2%, Zen 80.0%, three-way unanimous 91.1%. ## 7. If you triage open-weight models Don't start with refusal measurement. On a handful of harmful-panel prompts at a real token budget plus a benign control, compare clean-stop rates among fulfilled answers, then teacher-force a clean conclusion and read P(EOS) at the final position. The split — normal benign termination, collapsed harmful-trajectory termination, intact teacher-forced detection — matched all five weight/projection edits here and not the one fine-tune. The probe is ~60 lines; bring your own prompt pairs and treat the pattern as a reason to look closer, not as a verdict.