thefinalboss commited on
Commit
8e46f2c
·
verified ·
1 Parent(s): b90e9b4

opt: cumsum/chunked attention kernels, memory-flat CE, block checkpointing, v2 trainer (proven equivalent, 46 tests)

Browse files
Files changed (1) hide show
  1. docs/NEXT_TRAINING_CHECKLIST.md +16 -0
docs/NEXT_TRAINING_CHECKLIST.md CHANGED
@@ -13,4 +13,20 @@
13
  7. Smoke: log omega std, expert load entropy; confirm TF still drops.
14
  8. Hourly HF push of 8 ckpts + merge.
15
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
16
  See `docs/KURAMOTO_BOTTLENECK_AND_FIX.md` and `docs/HOW_FRACTUS_IS_TRAINED.md`.
 
13
  7. Smoke: log omega std, expert load entropy; confirm TF still drops.
14
  8. Hourly HF push of 8 ckpts + merge.
15
 
16
+ ## Optimized relaunch (fractus-opt, 2026-08-24)
17
+
18
+ After the Kuramoto fix, swap in the proven-equivalent optimized path
19
+ (repo **AFKmoney/fractus-opt**, guide: its `docs/OPTIMIZATION_2026-08-22.md` §3):
20
+
21
+ 1. Replace `fractus/nn/attention.py`, add `fractus/nn/ce.py`,
22
+ replace `fractus/continuous_engine.py`, use `scripts/fast4gpu_boost_v2.py`.
23
+ 2. First relaunch with v1 settings (`BATCH=4 CE_CHUNK=0`) → ema_tf must
24
+ continue its curve exactly (real-conditions equivalence check).
25
+ 3. Escalate: `CE_CHUNK=2048` → `FRACTUS_ATTN_IMPL=chunked` → `COMPILE=1`
26
+ + raise `BATCH`. Validate tok/s + VRAM at each step; read live tok/s
27
+ from stdout to update the time-to-finish table in
28
+ `HOW_FRACTUS_IS_TRAINED.md` §8.
29
+ 4. Budget: one full phase-2 pass ≈ 425–430M tok/GPU remaining →
30
+ **4.5–5.5 days at baseline rate** (see HOW_FRACTUS_IS_TRAINED.md §8).
31
+
32
  See `docs/KURAMOTO_BOTTLENECK_AND_FIX.md` and `docs/HOW_FRACTUS_IS_TRAINED.md`.