|
Download docs/TRAIN_GEN_MISMATCH.md from thefinalboss/fractus-cte: direct link, hf CLI and curl.
- Browser
- Download file 791 Bytes
-
https://huggingface.co/thefinalboss/fractus-cte/resolve/8e46f2cf53e63440225486f98e83043aec096172/docs/TRAIN_GEN_MISMATCH.md
- Command line
-
hf download hf://thefinalboss/fractus-cte@8e46f2cf53e63440225486f98e83043aec096172/docs/TRAIN_GEN_MISMATCH.md
-
curl -L -o TRAIN_GEN_MISMATCH.md https://huggingface.co/thefinalboss/fractus-cte/resolve/8e46f2cf53e63440225486f98e83043aec096172/docs/TRAIN_GEN_MISMATCH.md
791 Bytes
Train/Gen Path Mismatch — Root Cause
Finding
Training uses → :
- Multi-level causal linear attention over the full chunk
- Kuramoto RK4 integrate
- MoE over all positions
Default generation used → :
- Single-step attention (level-0 style path)
- Kuramoto one Euler step (theta + 0.1 * derivative)
- MoE on one token
So the model was trained on one dynamics and decoded with another.
Fixes
- Code: Kuramoto now uses (same as train).
- Decode: Prefer / sliding-window chunk gen for train-aligned logits.
- Decode surgery: still useful for attractor lock; does not replace path alignment.
Expected
Aligning dynamics removes a major source of collapse/incoherence. Coherent language still depends on further digestion under stage2 dense CE.