fractus-cte / docs /KURAMOTO_BOTTLENECK_AND_FIX.md
thefinalboss's picture
Upload docs/KURAMOTO_BOTTLENECK_AND_FIX.md with huggingface_hub
2bd6293 verified
|
Raw History Blame Contribute Delete
3.11 kB

Kuramoto bottleneck and fix (next training)

Updated: 2026-08-20 04:47 UTC
Status: DIAGNOSED + PREP READY (apply on pod before resume)

What was measured

Agent probe (2026-08-19) on Kuramoto tensors (mmap, low RAM):

  1. Degenerate circular mean of the phase fan — atan2(0,0)-style collapse on a fresh model path.
  2. Post-RK4 routing arc ~25° / 360° — with 128 experts spaced ~2.8°, only a thin slice of the expert ring is reachable → structurally dead experts.
  3. Order parameter r ≈ 0.0101 on fresh dynamics — same order as prod logs (0.01–0.03).
  4. Not only probe no_grad — soft dynamics / narrow omega band are in the oscillator setup itself.

Matches earlier mid-train observations: phases stuck, omega std ~0.02–0.03, high dead-expert rate on probe windows.

Will more tokens alone fix it?

Unlikely as the sole plan.

Finishing phase2 without opening the routing mostly trains the already-open expert subset. Dead experts stay dark.

Training helps after the route can move: wider omega, softer gates, real load-balance in the loss, Kuramoto not under no_grad in tick_chunk_train.

Fix recipe (next pod)

A. On checkpoint load (existing freeze weights)

export GATE_TEMP=2.5
export OMEGA_SCALE=4.0
export OMEGA_NOISE=0.01
export OMEGA_CLAMP=0.5
export LB_COEF=0.05

python scripts/prep_kuramoto_fix_resume.py
# backups: checkpoints/fractus_1b_gpu{i}_pre_kuramoto_fix.pt
# writes:  checkpoints/fractus_1b_gpu{i}.pt + KURAMOTO_FIX_APPLIED.json

fractus/kuramoto_fix.py → apply_kuramoto_routing_fix(engine):

  • sets MoE temperature
  • omega ← clamp(omega * scale + noise)

B. During train (must stay true)

Requirement Why
loss = ce + LB_COEF * lb LB not detached
Kuramoto in autograd path omega can keep adapting
GATE_TEMP applied each load runtime attr, not always in ckpt
START_TOKEN from FROZEN_RESUME_MANIFEST.json do not restart phase2 at 0

C. New models only

KuramotoLayer default omega init widened: uniform(-0.35, 0.35) was (-0.05, 0.05).

Success checks (after fix + short smoke)

  1. omega.std() per block clearly above pre-fix (logged in KURAMOTO_FIX_APPLIED.json).
  2. Offline probe without resetting live phases: higher expert-usage entropy, fewer zero-load experts.
  3. Routing arc / phase diversity up vs ~25° regime.
  4. TF loss still trends down (fix must not destroy the freeze).

Relation to freeze

Freeze remains valid: ~15–17M tokens/GPU on phase2, weights on HF.
Fix = open-heart on load, then continue digestion. Not a full reinit.

Files

Path Role
docs/KURAMOTO_BOTTLENECK_AND_FIX.md this note
fractus/kuramoto_fix.py apply helper
scripts/prep_kuramoto_fix_resume.py offline prep all GPU ckpts
fractus/nn/phase_ode.py wider omega init for new builds
scripts/fast4gpu_surgery.py historical mid-train surgery (same family)

One line

Dead experts are a routing-arc problem; scale omega + soften gates + real LB, then finish the pass — do not wait for tokens alone.