fractus-cte / docs /2026-08-28-NIGHT-SS-X8.md
thefinalboss's picture
Upload docs/2026-08-28-NIGHT-SS-X8.md with huggingface_hub
b462e42 verified
|
Raw History Blame
3.76 kB

2026-08-27/28 night — SS live, 8×5090, HF hourly x8

All times UTC unless marked PDT. Pod t8620qxxqgn6x4, US-NC-2, 8× RTX 5090, community $5.52/h. Code tree: fractus-p0 (this package). Weights: thefinalboss/fractus-cte/checkpoints/x8run/.

Snapshot at 03:50 UTC 2026-08-28 (~20:50 PDT 2026-08-27)

GPU tokens (phase-2 shard) ema_tf ema_ss tok/s VRAM
0 41,931,776 7.68 18.94 989 27.4 GB
1 40,197,120 6.87 17.61 989 27.6 GB
2 39,945,216 14.95 27.14 983 27.6 GB
3 41,035,776 13.03 24.58 987 27.6 GB
4 39,173,120 37.83 42.51 982 27.6 GB
5 40,068,096 20.68 27.76 980 27.4 GB
6 40,314,880 19.23 26.58 987 27.5 GB
7 38,942,720 29.96 36.90 991 27.6 GB
  • Shard length: 429,896,462 tokens/GPU (phase-2 memmap). ~9.1–9.8% of this pass.
  • Recipe now: B=4 SEQ=128 SS_RATE=1.0 LR=7e-4, FRACTUS_ATTN_IMPL=chunked, BLOCK_CKPT=0.
  • All 8 _live.pt present on disk (4.3–4.4 GB each).
  • unique@40 PREFIX still NO-GO on last probe before SS (mean unique ≈ 3.3). SS is the intervention; speech is not claimed yet.

Timeline

2026-08-27 19:05–19:20 PDT — diagnosis

  • 8 trainers died mid-save: workspace volume 80 GB full. PytorchStreamWriter / unexpected pos.
  • unique@40 PREFIX NO-GO mean_u=3.3. Probe on gpu0_live (440/440 tensors, 1.05B):
    • "Hello, my name is" → HOME/Rac loop (unique 2)
    • "Once upon a time" → relation/AFTA (unique 3)
  • Teacher-forced CE was learning (GPU0 tf ~5–8). Free-run was not. SS_RATE was 0 because of a checkpoint bug.

2026-08-27 19:20–19:35 PDT — SS surgery (code)

Cause of SS crash: torch.utils.checkpoint inside chunked_cross_entropy / chunked_repeat_logprob colliding on the second backward (saved logits [N,50257] vs recomputed hidden [N,1280]).

Fixes landed in-tree:

  1. fractus/train/v4_step.py
    • snapshot_carry / restore_carry so SS replays the same thought + attn_S/z + Kuramoto phases as the TF chunk.
    • v4_ss_pass uses tick_chunk_train + dense CE (B·SEQ·vocab is tiny at B≤4). No CE-checkpoint on the SS pass.
  2. fractus/nn/ce.py + fractus/train/ar_loss.py
    • skip activation checkpoint when N <= ce_chunk (the 1B B=2/4 regime).
  3. scripts/fast4gpu_boost_v4.py
    • restore carry before SS; block_ckpt=False on SS.

First live SS lines (B=2): tf=29.6 ss=42.2 GPU2, no CheckpointError. ~600 tok/s (two forwards).

2026-08-27 19:38 PDT — volume

  • REST PATCH volumeInGb 80 → 200. Public SSH dropped ~1 min (new port 35370).
  • Trainers restarted. Disk 146 GB free. VRAM B=2+SS ~18/32 GB.

2026-08-27 19:41–19:43 PDT — x8 save + HF push

  • Save every n % 4000 == (GPU * 500) % 4000 — all 8, staggered.
  • hf_push_x8.py → thefinalboss/fractus-cte/checkpoints/x8run/fractus_1b_gpu{i}.pt + RESUME_gpu{i}.json.
  • Interval corrected from 15 min to 3600 s.

2026-08-27 20:01 PDT — batch push

  • B=2 → B=4, SS still 1.0.
  • Holds: 27.1 / 32 GB, ~1070 tok/s/GPU first steps, later ~980–990.
  • Aggregate ~7.9k tok/s. B=5 not taken (OOM risk).

What this night did not do

  • Did not finish the 3.44B phase-2 pass (~9% of current shards).
  • Did not make unique@40 GO. SS is on; generation quality is the next measured gate, not a slogan.
  • Did not re-ingest phase-1 (~1.5B already in the X8 weights).

How to resume

CKPT_IN=checkpoints/fractus_1b_gpu{i}_live.pt
START_TOKEN=<tokens_processed from config or RESUME_gpu{i}.json>
BATCH=4 SS_RATE=1.0 BLOCK_CKPT=0 FRACTUS_ATTN_IMPL=chunked

Pusher: python3 -u /workspace/hf_push_x8.py (hourly).