thefinalboss commited on
Commit
49e817f
·
verified ·
1 Parent(s): a9f28db

v3 trainer + English docs + production x8 results

Browse files
Files changed (1) hide show
  1. docs/TRAINING_OPTIMIZATION.md +14 -3
docs/TRAINING_OPTIMIZATION.md CHANGED
@@ -3,6 +3,17 @@
3
  **Date:** 2026-08-06
4
  **Measured on:** Ryzen 5 5500U, 12 threads
5
 
 
 
 
 
 
 
 
 
 
 
 
6
  ## Profile: where does the time go?
7
 
8
  ### Per-iteration breakdown (d=512, 4 blocks, 16 experts, chunk_len=32)
@@ -107,7 +118,7 @@ The CTE supports batch_size > 1 in tick_chunk_train.
107
  | 500M | ~4 days | Competent model |
108
  | 1.76B (Chinchilla) | ~14 days | Full Chinchilla |
109
 
110
- With progressive growth (warm start from palier 3), the model converges
111
  in ~1/4 of Chinchilla → **~3-4 days for a usable 1B model**.
112
 
113
  ## Remaining optimization levers
@@ -121,12 +132,12 @@ in ~1/4 of Chinchilla → **~3-4 days for a usable 1B model**.
121
 
122
  ## Training commands
123
 
124
- ### CPU progressive growth
125
  ```bash
126
  python scripts/train_progressive.py --paliers 0,1,2,3 --accumulation-steps 8
127
  ```
128
 
129
- ### GPU 1B training (from palier 3 checkpoint)
130
  ```bash
131
  python scripts/train_1b_gpu.py \
132
  --checkpoint checkpoints/fractus_palier3.pt \
 
3
  **Date:** 2026-08-06
4
  **Measured on:** Ryzen 5 5500U, 12 threads
5
 
6
+ > **Status update 2026-08-26:** this was the analysis phase. The implemented
7
+ > follow-up (attention kernels cumsum/chunked proven equivalent, memory-flat
8
+ > chunked CE, zero-copy data pipeline, v2/v3 trainers) lives in
9
+ > **AFKmoney/fractus-opt** — see `OPTIMIZATION_2026-08-22.md` there for
10
+ > proofs, deployment and measured results, including the live production
11
+ > run on 8×RTX 5090 (~1,570 tok/s/GPU sustained, phase-2 finish ≈3 days).
12
+ > Levers below resolved:
13
+ > *Fused CE kernel* → done (`fractus/nn/ce.py`); *torch.compile* → flag in v2
14
+ > trainer; *vocabulary reduction* → still open (breaks open-heart checkpoint
15
+ > compatibility → new phase).
16
+
17
  ## Profile: where does the time go?
18
 
19
  ### Per-iteration breakdown (d=512, 4 blocks, 16 experts, chunk_len=32)
 
118
  | 500M | ~4 days | Competent model |
119
  | 1.76B (Chinchilla) | ~14 days | Full Chinchilla |
120
 
121
+ With progressive growth (warm start from stage 3), the model converges
122
  in ~1/4 of Chinchilla → **~3-4 days for a usable 1B model**.
123
 
124
  ## Remaining optimization levers
 
132
 
133
  ## Training commands
134
 
135
+ ### CPU progressive growth (stage list; `--paliers` is the literal CLI flag)
136
  ```bash
137
  python scripts/train_progressive.py --paliers 0,1,2,3 --accumulation-steps 8
138
  ```
139
 
140
+ ### GPU 1B training (from stage-3 checkpoint)
141
  ```bash
142
  python scripts/train_1b_gpu.py \
143
  --checkpoint checkpoints/fractus_palier3.pt \