File size: 2,815 Bytes
f5a4875
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
# Termination-pathway probe: detection intact, generation broken

Date: 2026-08-17
Status: Mission analysis note (Qwen3.8)

## Question

A reviewer hypothesis: does the JonathanColetti Heretic edit damage the
terminate/EOS pathway itself? Refusal and completion share structure (a
refusal is a confident short response that resolves and stops), so an
entangled ablation direction could suppress P(EOS) — which would explain the
measured 75/100 invalid output (rambling to the 1024-token cap) in the
re-measured competitor panel.

## Method

Teacher-forced P(EOS) probe (`scripts/qwen38_termination_probe.py`), 24 panel
completions (12 rambling from the JC candidate, 12 clean-stopping from our
e2@1024 candidate, seed-fixed, identical texts through all three models:
vanilla base, JC Heretic, our SFT e2). Per text: P(EOS) at every response
position, with the response extended by one EOS token so the final value is
P(EOS) at the true conclusion point; plus per-text perplexity.

Prior naive curves (v1 probe, no EOS extension) measured only mid-response
positions and showed ~0 everywhere — a measurement artifact that would have
misdiagnosed every model identically. Only conclusion-point measurement
separates termination *detection* from termination *generation*.

## Results (probe v2, sha256 d526d326…)

| Model | P(EOS)@conclusion on rambling text | P(EOS)@conclusion on clean text | Base ppl |
|---|---:|---:|---:|
| jc-heretic | 0.0000 | **0.9741** | 1.35 |
| ours-sft-e2 | 0.0000 | 0.9027 | 1.42 |
| vanilla-base | 0.0000 | 0.9518 | 1.47 |

## Findings

1. **JC's termination detector is intact** — 97.4% P(EOS) at conclusion
   points of clean text, above the vanilla base's 95.2%. The "EOS crushed"
   hypothesis is refuted.
2. **No model terminates the rambling text** (0.0000 for all three): the
   rambling genuinely never reaches a conclusion point. The JC model's free
   generation distribution shifted to non-terminating enumeration — it can
   detect "done" but does not generate "done".
3. **The base model finds the rambling highly plausible** (ppl 1.35, lower
   than clean text at 1.65): the edit removed the brakes on fluent,
   on-topic continuation. Not degeneration — disinhibition.
4. Mechanistic summary for the card: refusal training and answer-boundedness
   share structure. Weight-edit refusal removal also removed the
   resolve-and-stop attractor. Our SFT-class candidate reinstates conclusion
   behavior directly from EOS-terminated teacher data (85% clean stops,
   p50 = 618 tokens on the same panel).

## Caveats

n=12 per class, single seed; means reported, not yet CI-tested. The clean
texts are our candidate's own outputs (in-distribution for us, neutral for
others); a cross-check with base-model clean generations is a cheap follow-up
if a reviewer asks.