fractus-cte / docs /2026-08-28-NIGHT-SS-X8.md
thefinalboss's picture
Upload docs/2026-08-28-NIGHT-SS-X8.md with huggingface_hub
b462e42 verified
|
Raw History Blame
3.76 kB
# 2026-08-27/28 night β€” SS live, 8Γ—5090, HF hourly x8
All times **UTC** unless marked PDT. Pod `t8620qxxqgn6x4`, US-NC-2, 8Γ— RTX 5090, community $5.52/h.
Code tree: `fractus-p0` (this package). Weights: `thefinalboss/fractus-cte/checkpoints/x8run/`.
## Snapshot at 03:50 UTC 2026-08-28 (~20:50 PDT 2026-08-27)
| GPU | tokens (phase-2 shard) | ema_tf | ema_ss | tok/s | VRAM |
|-----|------------------------|--------|--------|-------|------|
| 0 | 41,931,776 | 7.68 | 18.94 | 989 | 27.4 GB |
| 1 | 40,197,120 | 6.87 | 17.61 | 989 | 27.6 GB |
| 2 | 39,945,216 | 14.95 | 27.14 | 983 | 27.6 GB |
| 3 | 41,035,776 | 13.03 | 24.58 | 987 | 27.6 GB |
| 4 | 39,173,120 | 37.83 | 42.51 | 982 | 27.6 GB |
| 5 | 40,068,096 | 20.68 | 27.76 | 980 | 27.4 GB |
| 6 | 40,314,880 | 19.23 | 26.58 | 987 | 27.5 GB |
| 7 | 38,942,720 | 29.96 | 36.90 | 991 | 27.6 GB |
- Shard length: 429,896,462 tokens/GPU (phase-2 memmap). ~9.1–9.8% of this pass.
- Recipe now: **B=4 SEQ=128 SS_RATE=1.0** LR=7e-4, `FRACTUS_ATTN_IMPL=chunked`, `BLOCK_CKPT=0`.
- All 8 `_live.pt` present on disk (4.3–4.4 GB each).
- unique@40 PREFIX still **NO-GO** on last probe before SS (mean unique β‰ˆ 3.3). SS is the intervention; speech is not claimed yet.
## Timeline
### 2026-08-27 19:05–19:20 PDT β€” diagnosis
- 8 trainers died mid-save: workspace volume **80 GB full**. `PytorchStreamWriter` / unexpected pos.
- unique@40 PREFIX NO-GO mean_u=3.3. Probe on `gpu0_live` (440/440 tensors, 1.05B):
- "Hello, my name is" β†’ HOME/Rac loop (unique 2)
- "Once upon a time" β†’ relation/AFTA (unique 3)
- Teacher-forced CE was learning (GPU0 tf ~5–8). **Free-run was not.** SS_RATE was 0 because of a checkpoint bug.
### 2026-08-27 19:20–19:35 PDT β€” SS surgery (code)
Cause of SS crash: `torch.utils.checkpoint` inside `chunked_cross_entropy` / `chunked_repeat_logprob` colliding on the **second** backward (saved logits `[N,50257]` vs recomputed hidden `[N,1280]`).
Fixes landed in-tree:
1. `fractus/train/v4_step.py`
- `snapshot_carry` / `restore_carry` so SS replays the **same** thought + attn_S/z + Kuramoto phases as the TF chunk.
- `v4_ss_pass` uses `tick_chunk_train` + **dense** CE (BΒ·SEQΒ·vocab is tiny at B≀4). No CE-checkpoint on the SS pass.
2. `fractus/nn/ce.py` + `fractus/train/ar_loss.py`
- skip activation checkpoint when `N <= ce_chunk` (the 1B B=2/4 regime).
3. `scripts/fast4gpu_boost_v4.py`
- restore carry before SS; `block_ckpt=False` on SS.
First live SS lines (B=2): `tf=29.6 ss=42.2` GPU2, no CheckpointError. ~600 tok/s (two forwards).
### 2026-08-27 19:38 PDT β€” volume
- REST PATCH `volumeInGb` 80 β†’ **200**. Public SSH dropped ~1 min (new port 35370).
- Trainers restarted. Disk **146 GB free**. VRAM B=2+SS ~18/32 GB.
### 2026-08-27 19:41–19:43 PDT β€” x8 save + HF push
- Save every `n % 4000 == (GPU * 500) % 4000` β€” **all 8**, staggered.
- `hf_push_x8.py` β†’ `thefinalboss/fractus-cte/checkpoints/x8run/fractus_1b_gpu{i}.pt` + `RESUME_gpu{i}.json`.
- Interval corrected from 15 min to **3600 s**.
### 2026-08-27 20:01 PDT β€” batch push
- B=2 β†’ **B=4**, SS still 1.0.
- Holds: **27.1 / 32 GB**, ~**1070 tok/s**/GPU first steps, later ~980–990.
- Aggregate ~7.9k tok/s. B=5 not taken (OOM risk).
## What this night did **not** do
- Did not finish the 3.44B phase-2 pass (~9% of current shards).
- Did not make unique@40 GO. SS is on; generation quality is the next measured gate, not a slogan.
- Did not re-ingest phase-1 (~1.5B already in the X8 weights).
## How to resume
```
CKPT_IN=checkpoints/fractus_1b_gpu{i}_live.pt
START_TOKEN=<tokens_processed from config or RESUME_gpu{i}.json>
BATCH=4 SS_RATE=1.0 BLOCK_CKPT=0 FRACTUS_ATTN_IMPL=chunked
```
Pusher: `python3 -u /workspace/hf_push_x8.py` (hourly).