opt: cumsum/chunked attention kernels, memory-flat CE, block checkpointing, v2 trainer (proven equivalent, 46 tests)
Browse files- docs/TRAINING_LOG_1B.md +19 -0
docs/TRAINING_LOG_1B.md
CHANGED
|
@@ -182,5 +182,24 @@ OUTPUT: commands commands commands commands commands commands ...
|
|
| 182 |
|
| 183 |
---
|
| 184 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 185 |
*Fractus CTE — continuous thought, phase-routed experts, progressive growth.*
|
| 186 |
*Training log only. Not a claim of finished language competence.*
|
|
|
|
| 182 |
|
| 183 |
---
|
| 184 |
|
| 185 |
+
## 2026-08-22/23 — Optimization session (training PAUSED)
|
| 186 |
+
|
| 187 |
+
No token digestion in this window: code-surgery session instead, on
|
| 188 |
+
**AFKmoney/fractus-opt**.
|
| 189 |
+
|
| 190 |
+
- Phase-2 status: started 2026-08-17T23:26 UTC; paused since ~2026-08-20
|
| 191 |
+
with <1% of the pass consumed (~0.7M tok/GPU at last sync).
|
| 192 |
+
**Remaining ≈ 425–430M tok/GPU → one full pass ≈ 4.5–5.5 days at the
|
| 193 |
+
measured baseline 900–1100 tok/s/GPU** (see HOW_FRACTUS_IS_TRAINED.md §8
|
| 194 |
+
for the math and post-opt scenarios).
|
| 195 |
+
- Kernels: attention cumsum + chunked proven equivalent to the einsum
|
| 196 |
+
reference (fwd/grad/carry); chunked measured ×15–25 at 1B shapes,
|
| 197 |
+
memory-flat where the reference crashes.
|
| 198 |
+
- Head: memory-flat chunked CE (loss identical within fp32 rounding).
|
| 199 |
+
- Data: zero-copy int32 fetch (~27 GB RAM/pod saved).
|
| 200 |
+
- Tests: 44/44 (28 repo + 14 equivalence proofs + 2 v2-loop smoke).
|
| 201 |
+
|
| 202 |
+
---
|
| 203 |
+
|
| 204 |
*Fractus CTE — continuous thought, phase-routed experts, progressive growth.*
|
| 205 |
*Training log only. Not a claim of finished language competence.*
|