|
Download docs/KURAMOTO_BOTTLENECK_AND_FIX.md from thefinalboss/fractus-cte: direct link, hf CLI and curl.
- Browser
- Download file 3.11 kB
-
https://huggingface.co/thefinalboss/fractus-cte/resolve/main/docs/KURAMOTO_BOTTLENECK_AND_FIX.md
- Command line
-
hf download hf://thefinalboss/fractus-cte/docs/KURAMOTO_BOTTLENECK_AND_FIX.md
-
curl -L -o KURAMOTO_BOTTLENECK_AND_FIX.md https://huggingface.co/thefinalboss/fractus-cte/resolve/main/docs/KURAMOTO_BOTTLENECK_AND_FIX.md
3.11 kB
| # Kuramoto bottleneck and fix (next training) | |
| **Updated:** 2026-08-20 04:47 UTC | |
| **Status:** DIAGNOSED + PREP READY (apply on pod before resume) | |
| ## What was measured | |
| Agent probe (2026-08-19) on Kuramoto tensors (mmap, low RAM): | |
| 1. **Degenerate circular mean** of the phase fan β atan2(0,0)-style collapse on a fresh model path. | |
| 2. **Post-RK4 routing arc ~25Β° / 360Β°** β with 128 experts spaced ~2.8Β°, only a thin slice of the expert ring is reachable β structurally dead experts. | |
| 3. **Order parameter r β 0.0101** on fresh dynamics β same order as prod logs (0.01β0.03). | |
| 4. **Not only probe no_grad** β soft dynamics / narrow omega band are in the oscillator setup itself. | |
| Matches earlier mid-train observations: phases stuck, omega std ~0.02β0.03, high dead-expert rate on probe windows. | |
| ## Will more tokens alone fix it? | |
| **Unlikely as the sole plan.** | |
| Finishing phase2 without opening the routing mostly trains the already-open expert subset. Dead experts stay dark. | |
| Training helps **after** the route can move: wider omega, softer gates, real load-balance in the loss, Kuramoto not under no_grad in tick_chunk_train. | |
| ## Fix recipe (next pod) | |
| ### A. On checkpoint load (existing freeze weights) | |
| ```bash | |
| export GATE_TEMP=2.5 | |
| export OMEGA_SCALE=4.0 | |
| export OMEGA_NOISE=0.01 | |
| export OMEGA_CLAMP=0.5 | |
| export LB_COEF=0.05 | |
| python scripts/prep_kuramoto_fix_resume.py | |
| # backups: checkpoints/fractus_1b_gpu{i}_pre_kuramoto_fix.pt | |
| # writes: checkpoints/fractus_1b_gpu{i}.pt + KURAMOTO_FIX_APPLIED.json | |
| ``` | |
| `fractus/kuramoto_fix.py` β `apply_kuramoto_routing_fix(engine)`: | |
| - sets MoE temperature | |
| - omega β clamp(omega * scale + noise) | |
| ### B. During train (must stay true) | |
| | Requirement | Why | | |
| |-------------|-----| | |
| | loss = ce + LB_COEF * lb | LB not detached | | |
| | Kuramoto in autograd path | omega can keep adapting | | |
| | GATE_TEMP applied each load | runtime attr, not always in ckpt | | |
| | START_TOKEN from FROZEN_RESUME_MANIFEST.json | do not restart phase2 at 0 | | |
| ### C. New models only | |
| KuramotoLayer default omega init widened: uniform(-0.35, 0.35) was (-0.05, 0.05). | |
| ## Success checks (after fix + short smoke) | |
| 1. omega.std() per block clearly above pre-fix (logged in KURAMOTO_FIX_APPLIED.json). | |
| 2. Offline probe without resetting live phases: higher expert-usage entropy, fewer zero-load experts. | |
| 3. Routing arc / phase diversity up vs ~25Β° regime. | |
| 4. TF loss still trends down (fix must not destroy the freeze). | |
| ## Relation to freeze | |
| Freeze remains valid: ~15β17M tokens/GPU on phase2, weights on HF. | |
| Fix = open-heart on load, then continue digestion. Not a full reinit. | |
| ## Files | |
| | Path | Role | | |
| |------|------| | |
| | docs/KURAMOTO_BOTTLENECK_AND_FIX.md | this note | | |
| | fractus/kuramoto_fix.py | apply helper | | |
| | scripts/prep_kuramoto_fix_resume.py | offline prep all GPU ckpts | | |
| | fractus/nn/phase_ode.py | wider omega init for new builds | | |
| | scripts/fast4gpu_surgery.py | historical mid-train surgery (same family) | | |
| ## One line | |
| Dead experts are a routing-arc problem; scale omega + soften gates + real LB, then finish the pass β do not wait for tokens alone. | |