File size: 791 Bytes
7764b77 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 | # Train/Gen Path Mismatch — Root Cause
## Finding
Training uses → :
- Multi-level causal linear attention over the full chunk
- Kuramoto **RK4** integrate
- MoE over all positions
Default generation used → :
- Single-step attention (level-0 style path)
- Kuramoto **one Euler step** (theta + 0.1 * derivative)
- MoE on one token
So the model was trained on one dynamics and decoded with another.
## Fixes
1. **Code:** Kuramoto now uses (same as train).
2. **Decode:** Prefer / sliding-window chunk gen for train-aligned logits.
3. **Decode surgery:** still useful for attractor lock; does not replace path alignment.
## Expected
Aligning dynamics removes a major source of collapse/incoherence. Coherent language still depends on further digestion under stage2 dense CE.
|