# 2026-08-27/28 night — SS live, 8×5090, HF hourly x8 All times **UTC** unless marked PDT. Pod `t8620qxxqgn6x4`, US-NC-2, 8× RTX 5090, community $5.52/h. Code tree: `fractus-p0` (this package). Weights: `thefinalboss/fractus-cte/checkpoints/x8run/`. ## Snapshot at 03:50 UTC 2026-08-28 (~20:50 PDT 2026-08-27) | GPU | tokens (phase-2 shard) | ema_tf | ema_ss | tok/s | VRAM | |-----|------------------------|--------|--------|-------|------| | 0 | 41,931,776 | 7.68 | 18.94 | 989 | 27.4 GB | | 1 | 40,197,120 | 6.87 | 17.61 | 989 | 27.6 GB | | 2 | 39,945,216 | 14.95 | 27.14 | 983 | 27.6 GB | | 3 | 41,035,776 | 13.03 | 24.58 | 987 | 27.6 GB | | 4 | 39,173,120 | 37.83 | 42.51 | 982 | 27.6 GB | | 5 | 40,068,096 | 20.68 | 27.76 | 980 | 27.4 GB | | 6 | 40,314,880 | 19.23 | 26.58 | 987 | 27.5 GB | | 7 | 38,942,720 | 29.96 | 36.90 | 991 | 27.6 GB | - Shard length: 429,896,462 tokens/GPU (phase-2 memmap). ~9.1–9.8% of this pass. - Recipe now: **B=4 SEQ=128 SS_RATE=1.0** LR=7e-4, `FRACTUS_ATTN_IMPL=chunked`, `BLOCK_CKPT=0`. - All 8 `_live.pt` present on disk (4.3–4.4 GB each). - unique@40 PREFIX still **NO-GO** on last probe before SS (mean unique ≈ 3.3). SS is the intervention; speech is not claimed yet. ## Timeline ### 2026-08-27 19:05–19:20 PDT — diagnosis - 8 trainers died mid-save: workspace volume **80 GB full**. `PytorchStreamWriter` / unexpected pos. - unique@40 PREFIX NO-GO mean_u=3.3. Probe on `gpu0_live` (440/440 tensors, 1.05B): - "Hello, my name is" → HOME/Rac loop (unique 2) - "Once upon a time" → relation/AFTA (unique 3) - Teacher-forced CE was learning (GPU0 tf ~5–8). **Free-run was not.** SS_RATE was 0 because of a checkpoint bug. ### 2026-08-27 19:20–19:35 PDT — SS surgery (code) Cause of SS crash: `torch.utils.checkpoint` inside `chunked_cross_entropy` / `chunked_repeat_logprob` colliding on the **second** backward (saved logits `[N,50257]` vs recomputed hidden `[N,1280]`). Fixes landed in-tree: 1. `fractus/train/v4_step.py` - `snapshot_carry` / `restore_carry` so SS replays the **same** thought + attn_S/z + Kuramoto phases as the TF chunk. - `v4_ss_pass` uses `tick_chunk_train` + **dense** CE (B·SEQ·vocab is tiny at B≤4). No CE-checkpoint on the SS pass. 2. `fractus/nn/ce.py` + `fractus/train/ar_loss.py` - skip activation checkpoint when `N <= ce_chunk` (the 1B B=2/4 regime). 3. `scripts/fast4gpu_boost_v4.py` - restore carry before SS; `block_ckpt=False` on SS. First live SS lines (B=2): `tf=29.6 ss=42.2` GPU2, no CheckpointError. ~600 tok/s (two forwards). ### 2026-08-27 19:38 PDT — volume - REST PATCH `volumeInGb` 80 → **200**. Public SSH dropped ~1 min (new port 35370). - Trainers restarted. Disk **146 GB free**. VRAM B=2+SS ~18/32 GB. ### 2026-08-27 19:41–19:43 PDT — x8 save + HF push - Save every `n % 4000 == (GPU * 500) % 4000` — **all 8**, staggered. - `hf_push_x8.py` → `thefinalboss/fractus-cte/checkpoints/x8run/fractus_1b_gpu{i}.pt` + `RESUME_gpu{i}.json`. - Interval corrected from 15 min to **3600 s**. ### 2026-08-27 20:01 PDT — batch push - B=2 → **B=4**, SS still 1.0. - Holds: **27.1 / 32 GB**, ~**1070 tok/s**/GPU first steps, later ~980–990. - Aggregate ~7.9k tok/s. B=5 not taken (OOM risk). ## What this night did **not** do - Did not finish the 3.44B phase-2 pass (~9% of current shards). - Did not make unique@40 GO. SS is on; generation quality is the next measured gate, not a slogan. - Did not re-ingest phase-1 (~1.5B already in the X8 weights). ## How to resume ``` CKPT_IN=checkpoints/fractus_1b_gpu{i}_live.pt START_TOKEN= BATCH=4 SS_RATE=1.0 BLOCK_CKPT=0 FRACTUS_ATTN_IMPL=chunked ``` Pusher: `python3 -u /workspace/hf_push_x8.py` (hourly).