File size: 791 Bytes
7764b77
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
# Train/Gen Path Mismatch — Root Cause

## Finding

Training uses  → :
- Multi-level causal linear attention over the full chunk
- Kuramoto **RK4** integrate
- MoE over all positions

Default generation used  → :
- Single-step attention (level-0 style path)
- Kuramoto **one Euler step** (theta + 0.1 * derivative)
- MoE on one token

So the model was trained on one dynamics and decoded with another.

## Fixes

1. **Code:**  Kuramoto now uses  (same as train).
2. **Decode:** Prefer  / sliding-window chunk gen for train-aligned logits.
3. **Decode surgery:** still useful for attractor lock; does not replace path alignment.

## Expected

Aligning dynamics removes a major source of collapse/incoherence. Coherent language still depends on further digestion under stage2 dense CE.