thefinalboss commited on
Commit
6701bd6
·
verified ·
1 Parent(s): 833d5b0

docs: QuickPod resume log

Browse files
Files changed (1) hide show
  1. docs/2026-08-28-QUICKPOD-RESUME.md +62 -0
docs/2026-08-28-QUICKPOD-RESUME.md ADDED
@@ -0,0 +1,62 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # QuickPod resume — 2026-08-28
2
+
3
+ Exact-token resume of Fractus-1B boost_v4 after the RunPod 8×5090 died.
4
+
5
+ ## What happened
6
+
7
+ - RunPod SSH (`66.222.154.19:35370`) refused mid-pass.
8
+ - Last good HF sync: `checkpoints/x8run/` at **2026-08-28T14:40:54Z**.
9
+ - GPU0 `tokens_processed = 79,922,176` / shard `429,896,462` (~18.6% of this Phase-2 pass).
10
+ - New machine: QuickPod 8× RTX 5090, empty disk.
11
+
12
+ ## What we did (do not skip these)
13
+
14
+ 1. Snapshot `thefinalboss/fractus-cte` code + `scripts/fast4gpu_boost_v4.py`.
15
+ 2. Download 8× `checkpoints/x8run/fractus_1b_gpu{i}.pt` + 8× `RESUME_gpu{i}.json`.
16
+ 3. Download Phase-2 memmaps from `thefinalboss/fractus-datasets`
17
+ `tokenized/phase2/shard_phase2_gpu{i}.npy` → symlink `data/shard_gpu{i}.npy`.
18
+ 4. Stock image torch **2.2.1+cu121 has no kernels for Blackwell 5090**
19
+ (`CUDA error: no kernel image is available`). Upgrade: **torch 2.11.0+cu128**.
20
+ 5. `set_p0_routing` missing on the published `continuous_engine.py` — added a flag setter so v4 can boot. Weights already carry the earlier routing surgery.
21
+ 6. Live `unique@40` probe crashed the first launch (`generate_greedy_prefix` not in that generate_aligned). **PROBE_EVERY=0**. Probes stay offline.
22
+ 7. Relaunch one process per GPU:
23
+
24
+ ```
25
+ CUDA_VISIBLE_DEVICES=$i GPU_ID=$i START_TOKEN=<RESUME.start_token_next> \
26
+ BATCH=4 SEQ=128 CE_CHUNK=2048 FRACTUS_ATTN_IMPL=chunked BLOCK_CKPT=1 \
27
+ SS_RATE=1.0 SS_PROB_START=0.2 SS_PROB_END=0.5 P0=1 COMPILE=0 PROBE_EVERY=0 \
28
+ CKPT_IN=checkpoints/x8run/fractus_1b_gpu$i.pt \
29
+ CKPT_OUT=checkpoints/x8run/fractus_1b_gpu$i.pt \
30
+ SHARD=data/shard_gpu$i.npy \
31
+ python -u scripts/fast4gpu_boost_v4.py
32
+ ```
33
+
34
+ Hourly HF push of `checkpoints/x8run/*`.
35
+
36
+ ## Resume offsets used 2026-08-28
37
+
38
+ | GPU | start_token_next |
39
+ |-----|------------------|
40
+ | 0 | 79,922,176 |
41
+ | 1 | 78,443,520 |
42
+ | 2 | 76,420,096 |
43
+ | 3 | 77,746,176 |
44
+ | 4 | 76,160,000 |
45
+ | 5 | 77,310,976 |
46
+ | 6 | 77,793,280 |
47
+ | 7 | 76,677,120 |
48
+
49
+ Shard length ≈ 429.9M int32 tokens / GPU.
50
+
51
+ ## Honest metrics after resume (first minutes)
52
+
53
+ - Throughput ~840–940 tok/s/GPU, climbing.
54
+ - `ema_tf` on GPU0/1 settled toward mid-teens after the first inflated steps (do **not** read the first 40-step EMA as the run).
55
+ - Speech gate remains **unique@40 greedy PREFIX** (no ban, no temperature). Last offline probe before the pod death: 3 / 12 / 1 / 6 on the four standard prompts. That is **NO-GO**.
56
+ - Finishing this Phase-2 pass from ~80M @ ~900 tok/s ≈ **4.1–4.5 days** wall if 8 GPUs stay up.
57
+
58
+ ## What is NOT true
59
+
60
+ - Low teacher-force CE ≠ the model speaks.
61
+ - Banned / temperature decode is not the metric.
62
+ - This pass is not “4.2B tokens from zero”. Phase-1 weights were already in the seed. This is the remainder of the 430M/GPU Phase-2 shards.