thefinalboss commited on
Commit
c957cb0
·
verified ·
1 Parent(s): 49e817f

v3 trainer + English docs + production x8 results

Browse files
Files changed (1) hide show
  1. docs/HOW_FRACTUS_IS_TRAINED.md +12 -7
docs/HOW_FRACTUS_IS_TRAINED.md CHANGED
@@ -1,6 +1,6 @@
1
  # How Fractus Is Trained
2
 
3
- **Updated:** 2026-08-24 (adds optimized v2 path + time-to-finish math)
4
  **Author:** Philippe-Antoine Robert
5
  **Model repo:** https://huggingface.co/thefinalboss/fractus-cte
6
  **Dataset:** https://huggingface.co/datasets/thefinalboss/fractus-datasets
@@ -164,12 +164,17 @@ Throughput ~900-1100 tok/s/GPU at B=2. One full phase-2 pass ~4-5 days wall-cloc
164
  | 900–1100 (baseline, measured) | **4.5 – 5.5 days** |
165
  | ~1300 (post-opt, ×1.3) | ~3.5 days |
166
  | ~2000 (×2) | ~2.3 days |
167
- | ~3000 (×3) | ~1.5 days |
168
-
169
- Post-optimization rates are NOT yet measured on GPU — the pod bench
170
- (benchmarks/bench_attention.py + fast4gpu_boost_v2.py stdout) gives the
171
- real number within minutes of relaunch. "Dataset complet" = one pass;
172
- multiply linearly for additional epochs.
 
 
 
 
 
173
 
174
  ---
175
 
 
1
  # How Fractus Is Trained
2
 
3
+ **Updated:** 2026-08-26 (production x8 run live — measured rates replace estimates; adds optimized v2/v3 path + time-to-finish math)
4
  **Author:** Philippe-Antoine Robert
5
  **Model repo:** https://huggingface.co/thefinalboss/fractus-cte
6
  **Dataset:** https://huggingface.co/datasets/thefinalboss/fractus-datasets
 
164
  | 900–1100 (baseline, measured) | **4.5 – 5.5 days** |
165
  | ~1300 (post-opt, ×1.3) | ~3.5 days |
166
  | ~2000 (×2) | ~2.3 days |
167
+ | **~1570 (MEASURED production x8, 2026-08-26)** | **~3.1 days** |
168
+
169
+ **Production status 2026-08-26:** phase-2 resumed on 8×RTX 5090 with the
170
+ optimized stack (v3 lineage: chunked kernel + BLOCK_CKPT + CE_CHUNK=2048,
171
+ B=8). Sustained **1,554–1,604 tok/s/GPU ≈ 12,600 aggregate**, VRAM
172
+ 18.9 GB/32 GB, lb = 14.028 stable. Resume honored gpu0–5 at the x6 positions
173
+ (1.25–1.40M), gpu6=655,360 / gpu7=768,000. Hourly safety sync to
174
+ `checkpoints/x8run/` + `checkpoints/X8_MANIFEST.json`. Projected finish:
175
+ **≈3.1 days from launch**. Deployment automated via
176
+ `scripts/pod_deploy_x8.sh`; full measurements in
177
+ `OPTIMIZATION_2026-08-22.md` §4.
178
 
179
  ---
180