opt: cumsum/chunked attention kernels, memory-flat CE, block checkpointing, v2 trainer (proven equivalent, 46 tests)
8e46f2c verified |
Download docs/NEXT_TRAINING_CHECKLIST.md from thefinalboss/fractus-cte: direct link, hf CLI and curl.
- Browser
- Download file 1.77 kB
-
https://huggingface.co/thefinalboss/fractus-cte/resolve/main/docs/NEXT_TRAINING_CHECKLIST.md
- Command line
-
hf download hf://thefinalboss/fractus-cte/docs/NEXT_TRAINING_CHECKLIST.md
-
curl -L -o NEXT_TRAINING_CHECKLIST.md https://huggingface.co/thefinalboss/fractus-cte/resolve/main/docs/NEXT_TRAINING_CHECKLIST.md
1.77 kB
| # Next training checklist (post-freeze + Kuramoto fix) | |
| 1. New pod, torch matching GPUs (5090 β recent CUDA build). | |
| 2. Pull `thefinalboss/fractus-cte` + phase2 shards from `fractus-datasets` `tokenized/phase2/*.npy`. | |
| 3. Download freeze ckpts `fractus_1b_gpu0..7.pt` (or from FROZEN merge then split only if needed β prefer per-GPU freeze). | |
| 4. **Apply Kuramoto fix before long run:** | |
| ```bash | |
| export GATE_TEMP=2.5 OMEGA_SCALE=4.0 OMEGA_NOISE=0.01 LB_COEF=0.05 | |
| python scripts/prep_kuramoto_fix_resume.py | |
| ``` | |
| 5. Resume with `START_TOKEN` from `checkpoints/FROZEN_RESUME_MANIFEST.json` per GPU (not 0). | |
| 6. Train: B=2, SEQ=128, LR=7e-4, SS_RATE=0.25, memmap phase2 shards, `loss = ce + LB_COEF * lb`. | |
| 7. Smoke: log omega std, expert load entropy; confirm TF still drops. | |
| 8. Hourly HF push of 8 ckpts + merge. | |
| ## Optimized relaunch (fractus-opt, 2026-08-24) | |
| After the Kuramoto fix, swap in the proven-equivalent optimized path | |
| (repo **AFKmoney/fractus-opt**, guide: its `docs/OPTIMIZATION_2026-08-22.md` Β§3): | |
| 1. Replace `fractus/nn/attention.py`, add `fractus/nn/ce.py`, | |
| replace `fractus/continuous_engine.py`, use `scripts/fast4gpu_boost_v2.py`. | |
| 2. First relaunch with v1 settings (`BATCH=4 CE_CHUNK=0`) β ema_tf must | |
| continue its curve exactly (real-conditions equivalence check). | |
| 3. Escalate: `CE_CHUNK=2048` β `FRACTUS_ATTN_IMPL=chunked` β `COMPILE=1` | |
| + raise `BATCH`. Validate tok/s + VRAM at each step; read live tok/s | |
| from stdout to update the time-to-finish table in | |
| `HOW_FRACTUS_IS_TRAINED.md` Β§8. | |
| 4. Budget: one full phase-2 pass β 425β430M tok/GPU remaining β | |
| **4.5β5.5 days at baseline rate** (see HOW_FRACTUS_IS_TRAINED.md Β§8). | |
| See `docs/KURAMOTO_BOTTLENECK_AND_FIX.md` and `docs/HOW_FRACTUS_IS_TRAINED.md`. | |