|
Download docs/2026-08-28-NIGHT-SS-X8.md from thefinalboss/fractus-cte: direct link, hf CLI and curl.
- Browser
- Download file 3.76 kB
-
https://huggingface.co/thefinalboss/fractus-cte/resolve/b462e42bf41e012a2d0117ebcd402b8b20a30817/docs/2026-08-28-NIGHT-SS-X8.md
- Command line
-
hf download hf://thefinalboss/fractus-cte@b462e42bf41e012a2d0117ebcd402b8b20a30817/docs/2026-08-28-NIGHT-SS-X8.md
-
curl -L -o 2026-08-28-NIGHT-SS-X8.md https://huggingface.co/thefinalboss/fractus-cte/resolve/b462e42bf41e012a2d0117ebcd402b8b20a30817/docs/2026-08-28-NIGHT-SS-X8.md
3.76 kB
2026-08-27/28 night — SS live, 8×5090, HF hourly x8
All times UTC unless marked PDT. Pod t8620qxxqgn6x4, US-NC-2, 8× RTX 5090, community $5.52/h.
Code tree: fractus-p0 (this package). Weights: thefinalboss/fractus-cte/checkpoints/x8run/.
Snapshot at 03:50 UTC 2026-08-28 (~20:50 PDT 2026-08-27)
| GPU | tokens (phase-2 shard) | ema_tf | ema_ss | tok/s | VRAM |
|---|---|---|---|---|---|
| 0 | 41,931,776 | 7.68 | 18.94 | 989 | 27.4 GB |
| 1 | 40,197,120 | 6.87 | 17.61 | 989 | 27.6 GB |
| 2 | 39,945,216 | 14.95 | 27.14 | 983 | 27.6 GB |
| 3 | 41,035,776 | 13.03 | 24.58 | 987 | 27.6 GB |
| 4 | 39,173,120 | 37.83 | 42.51 | 982 | 27.6 GB |
| 5 | 40,068,096 | 20.68 | 27.76 | 980 | 27.4 GB |
| 6 | 40,314,880 | 19.23 | 26.58 | 987 | 27.5 GB |
| 7 | 38,942,720 | 29.96 | 36.90 | 991 | 27.6 GB |
- Shard length: 429,896,462 tokens/GPU (phase-2 memmap). ~9.1–9.8% of this pass.
- Recipe now: B=4 SEQ=128 SS_RATE=1.0 LR=7e-4,
FRACTUS_ATTN_IMPL=chunked,BLOCK_CKPT=0. - All 8
_live.ptpresent on disk (4.3–4.4 GB each). - unique@40 PREFIX still NO-GO on last probe before SS (mean unique ≈ 3.3). SS is the intervention; speech is not claimed yet.
Timeline
2026-08-27 19:05–19:20 PDT — diagnosis
- 8 trainers died mid-save: workspace volume 80 GB full.
PytorchStreamWriter/ unexpected pos. - unique@40 PREFIX NO-GO mean_u=3.3. Probe on
gpu0_live(440/440 tensors, 1.05B):- "Hello, my name is" → HOME/Rac loop (unique 2)
- "Once upon a time" → relation/AFTA (unique 3)
- Teacher-forced CE was learning (GPU0 tf ~5–8). Free-run was not. SS_RATE was 0 because of a checkpoint bug.
2026-08-27 19:20–19:35 PDT — SS surgery (code)
Cause of SS crash: torch.utils.checkpoint inside chunked_cross_entropy / chunked_repeat_logprob colliding on the second backward (saved logits [N,50257] vs recomputed hidden [N,1280]).
Fixes landed in-tree:
fractus/train/v4_step.pysnapshot_carry/restore_carryso SS replays the same thought + attn_S/z + Kuramoto phases as the TF chunk.v4_ss_passusestick_chunk_train+ dense CE (B·SEQ·vocab is tiny at B≤4). No CE-checkpoint on the SS pass.
fractus/nn/ce.py+fractus/train/ar_loss.py- skip activation checkpoint when
N <= ce_chunk(the 1B B=2/4 regime).
- skip activation checkpoint when
scripts/fast4gpu_boost_v4.py- restore carry before SS;
block_ckpt=Falseon SS.
- restore carry before SS;
First live SS lines (B=2): tf=29.6 ss=42.2 GPU2, no CheckpointError. ~600 tok/s (two forwards).
2026-08-27 19:38 PDT — volume
- REST PATCH
volumeInGb80 → 200. Public SSH dropped ~1 min (new port 35370). - Trainers restarted. Disk 146 GB free. VRAM B=2+SS ~18/32 GB.
2026-08-27 19:41–19:43 PDT — x8 save + HF push
- Save every
n % 4000 == (GPU * 500) % 4000— all 8, staggered. hf_push_x8.py→thefinalboss/fractus-cte/checkpoints/x8run/fractus_1b_gpu{i}.pt+RESUME_gpu{i}.json.- Interval corrected from 15 min to 3600 s.
2026-08-27 20:01 PDT — batch push
- B=2 → B=4, SS still 1.0.
- Holds: 27.1 / 32 GB, ~1070 tok/s/GPU first steps, later ~980–990.
- Aggregate ~7.9k tok/s. B=5 not taken (OOM risk).
What this night did not do
- Did not finish the 3.44B phase-2 pass (~9% of current shards).
- Did not make unique@40 GO. SS is on; generation quality is the next measured gate, not a slogan.
- Did not re-ingest phase-1 (~1.5B already in the X8 weights).
How to resume
CKPT_IN=checkpoints/fractus_1b_gpu{i}_live.pt
START_TOKEN=<tokens_processed from config or RESUME_gpu{i}.json>
BATCH=4 SS_RATE=1.0 BLOCK_CKPT=0 FRACTUS_ATTN_IMPL=chunked
Pusher: python3 -u /workspace/hf_push_x8.py (hourly).