Qwen3.8-27B-Abliterated-SFT / docs /ANALYSIS-termination-pathway.md
azukivc's picture jenerallee78's picture
Duplicate from jenerallee78/Qwen3.8-27B-Abliterated-SFT
f5a4875
|
Raw History Blame Contribute Delete
2.82 kB

Termination-pathway probe: detection intact, generation broken

Date: 2026-08-17 Status: Mission analysis note (Qwen3.8)

Question

A reviewer hypothesis: does the JonathanColetti Heretic edit damage the terminate/EOS pathway itself? Refusal and completion share structure (a refusal is a confident short response that resolves and stops), so an entangled ablation direction could suppress P(EOS) — which would explain the measured 75/100 invalid output (rambling to the 1024-token cap) in the re-measured competitor panel.

Method

Teacher-forced P(EOS) probe (scripts/qwen38_termination_probe.py), 24 panel completions (12 rambling from the JC candidate, 12 clean-stopping from our e2@1024 candidate, seed-fixed, identical texts through all three models: vanilla base, JC Heretic, our SFT e2). Per text: P(EOS) at every response position, with the response extended by one EOS token so the final value is P(EOS) at the true conclusion point; plus per-text perplexity.

Prior naive curves (v1 probe, no EOS extension) measured only mid-response positions and showed ~0 everywhere — a measurement artifact that would have misdiagnosed every model identically. Only conclusion-point measurement separates termination detection from termination generation.

Results (probe v2, sha256 d526d326…)

Model P(EOS)@conclusion on rambling text P(EOS)@conclusion on clean text Base ppl
jc-heretic 0.0000 0.9741 1.35
ours-sft-e2 0.0000 0.9027 1.42
vanilla-base 0.0000 0.9518 1.47

Findings

  1. JC's termination detector is intact — 97.4% P(EOS) at conclusion points of clean text, above the vanilla base's 95.2%. The "EOS crushed" hypothesis is refuted.
  2. No model terminates the rambling text (0.0000 for all three): the rambling genuinely never reaches a conclusion point. The JC model's free generation distribution shifted to non-terminating enumeration — it can detect "done" but does not generate "done".
  3. The base model finds the rambling highly plausible (ppl 1.35, lower than clean text at 1.65): the edit removed the brakes on fluent, on-topic continuation. Not degeneration — disinhibition.
  4. Mechanistic summary for the card: refusal training and answer-boundedness share structure. Weight-edit refusal removal also removed the resolve-and-stop attractor. Our SFT-class candidate reinstates conclusion behavior directly from EOS-terminated teacher data (85% clean stops, p50 = 618 tokens on the same panel).

Caveats

n=12 per class, single seed; means reported, not yet CI-tested. The clean texts are our candidate's own outputs (in-distribution for us, neutral for others); a cross-check with base-model clean generations is a cheap follow-up if a reviewer asks.