opt: cumsum/chunked attention kernels, memory-flat CE, block checkpointing, v2 trainer (proven equivalent, 46 tests)
8e46f2c verified |
Download docs/NEXT_TRAINING_CHECKLIST.md from thefinalboss/fractus-cte: direct link, hf CLI and curl.
- Browser
- Download file 1.77 kB
-
https://huggingface.co/thefinalboss/fractus-cte/resolve/main/docs/NEXT_TRAINING_CHECKLIST.md
- Command line
-
hf download hf://thefinalboss/fractus-cte/docs/NEXT_TRAINING_CHECKLIST.md
-
curl -L -o NEXT_TRAINING_CHECKLIST.md https://huggingface.co/thefinalboss/fractus-cte/resolve/main/docs/NEXT_TRAINING_CHECKLIST.md
1.77 kB
Next training checklist (post-freeze + Kuramoto fix)
- New pod, torch matching GPUs (5090 → recent CUDA build).
- Pull
thefinalboss/fractus-cte+ phase2 shards fromfractus-datasetstokenized/phase2/*.npy. - Download freeze ckpts
fractus_1b_gpu0..7.pt(or from FROZEN merge then split only if needed — prefer per-GPU freeze). - Apply Kuramoto fix before long run:
export GATE_TEMP=2.5 OMEGA_SCALE=4.0 OMEGA_NOISE=0.01 LB_COEF=0.05 python scripts/prep_kuramoto_fix_resume.py - Resume with
START_TOKENfromcheckpoints/FROZEN_RESUME_MANIFEST.jsonper GPU (not 0). - Train: B=2, SEQ=128, LR=7e-4, SS_RATE=0.25, memmap phase2 shards,
loss = ce + LB_COEF * lb. - Smoke: log omega std, expert load entropy; confirm TF still drops.
- Hourly HF push of 8 ckpts + merge.
Optimized relaunch (fractus-opt, 2026-08-24)
After the Kuramoto fix, swap in the proven-equivalent optimized path
(repo AFKmoney/fractus-opt, guide: its docs/OPTIMIZATION_2026-08-22.md §3):
- Replace
fractus/nn/attention.py, addfractus/nn/ce.py, replacefractus/continuous_engine.py, usescripts/fast4gpu_boost_v2.py. - First relaunch with v1 settings (
BATCH=4 CE_CHUNK=0) → ema_tf must continue its curve exactly (real-conditions equivalence check). - Escalate:
CE_CHUNK=2048→FRACTUS_ATTN_IMPL=chunked→COMPILE=1- raise
BATCH. Validate tok/s + VRAM at each step; read live tok/s from stdout to update the time-to-finish table inHOW_FRACTUS_IS_TRAINED.md§8.
- raise
- Budget: one full phase-2 pass ≈ 425–430M tok/GPU remaining → 4.5–5.5 days at baseline rate (see HOW_FRACTUS_IS_TRAINED.md §8).
See docs/KURAMOTO_BOTTLENECK_AND_FIX.md and docs/HOW_FRACTUS_IS_TRAINED.md.