v3 trainer + English docs + production x8 results
Browse files
docs/TRAINING_OPTIMIZATION.md
CHANGED
|
@@ -3,6 +3,17 @@
|
|
| 3 |
**Date:** 2026-08-06
|
| 4 |
**Measured on:** Ryzen 5 5500U, 12 threads
|
| 5 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 6 |
## Profile: where does the time go?
|
| 7 |
|
| 8 |
### Per-iteration breakdown (d=512, 4 blocks, 16 experts, chunk_len=32)
|
|
@@ -107,7 +118,7 @@ The CTE supports batch_size > 1 in tick_chunk_train.
|
|
| 107 |
| 500M | ~4 days | Competent model |
|
| 108 |
| 1.76B (Chinchilla) | ~14 days | Full Chinchilla |
|
| 109 |
|
| 110 |
-
With progressive growth (warm start from
|
| 111 |
in ~1/4 of Chinchilla → **~3-4 days for a usable 1B model**.
|
| 112 |
|
| 113 |
## Remaining optimization levers
|
|
@@ -121,12 +132,12 @@ in ~1/4 of Chinchilla → **~3-4 days for a usable 1B model**.
|
|
| 121 |
|
| 122 |
## Training commands
|
| 123 |
|
| 124 |
-
### CPU progressive growth
|
| 125 |
```bash
|
| 126 |
python scripts/train_progressive.py --paliers 0,1,2,3 --accumulation-steps 8
|
| 127 |
```
|
| 128 |
|
| 129 |
-
### GPU 1B training (from
|
| 130 |
```bash
|
| 131 |
python scripts/train_1b_gpu.py \
|
| 132 |
--checkpoint checkpoints/fractus_palier3.pt \
|
|
|
|
| 3 |
**Date:** 2026-08-06
|
| 4 |
**Measured on:** Ryzen 5 5500U, 12 threads
|
| 5 |
|
| 6 |
+
> **Status update 2026-08-26:** this was the analysis phase. The implemented
|
| 7 |
+
> follow-up (attention kernels cumsum/chunked proven equivalent, memory-flat
|
| 8 |
+
> chunked CE, zero-copy data pipeline, v2/v3 trainers) lives in
|
| 9 |
+
> **AFKmoney/fractus-opt** — see `OPTIMIZATION_2026-08-22.md` there for
|
| 10 |
+
> proofs, deployment and measured results, including the live production
|
| 11 |
+
> run on 8×RTX 5090 (~1,570 tok/s/GPU sustained, phase-2 finish ≈3 days).
|
| 12 |
+
> Levers below resolved:
|
| 13 |
+
> *Fused CE kernel* → done (`fractus/nn/ce.py`); *torch.compile* → flag in v2
|
| 14 |
+
> trainer; *vocabulary reduction* → still open (breaks open-heart checkpoint
|
| 15 |
+
> compatibility → new phase).
|
| 16 |
+
|
| 17 |
## Profile: where does the time go?
|
| 18 |
|
| 19 |
### Per-iteration breakdown (d=512, 4 blocks, 16 experts, chunk_len=32)
|
|
|
|
| 118 |
| 500M | ~4 days | Competent model |
|
| 119 |
| 1.76B (Chinchilla) | ~14 days | Full Chinchilla |
|
| 120 |
|
| 121 |
+
With progressive growth (warm start from stage 3), the model converges
|
| 122 |
in ~1/4 of Chinchilla → **~3-4 days for a usable 1B model**.
|
| 123 |
|
| 124 |
## Remaining optimization levers
|
|
|
|
| 132 |
|
| 133 |
## Training commands
|
| 134 |
|
| 135 |
+
### CPU progressive growth (stage list; `--paliers` is the literal CLI flag)
|
| 136 |
```bash
|
| 137 |
python scripts/train_progressive.py --paliers 0,1,2,3 --accumulation-steps 8
|
| 138 |
```
|
| 139 |
|
| 140 |
+
### GPU 1B training (from stage-3 checkpoint)
|
| 141 |
```bash
|
| 142 |
python scripts/train_1b_gpu.py \
|
| 143 |
--checkpoint checkpoints/fractus_palier3.pt \
|