docs: QuickPod resume log
Browse files
docs/2026-08-28-QUICKPOD-RESUME.md
ADDED
|
@@ -0,0 +1,62 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# QuickPod resume — 2026-08-28
|
| 2 |
+
|
| 3 |
+
Exact-token resume of Fractus-1B boost_v4 after the RunPod 8×5090 died.
|
| 4 |
+
|
| 5 |
+
## What happened
|
| 6 |
+
|
| 7 |
+
- RunPod SSH (`66.222.154.19:35370`) refused mid-pass.
|
| 8 |
+
- Last good HF sync: `checkpoints/x8run/` at **2026-08-28T14:40:54Z**.
|
| 9 |
+
- GPU0 `tokens_processed = 79,922,176` / shard `429,896,462` (~18.6% of this Phase-2 pass).
|
| 10 |
+
- New machine: QuickPod 8× RTX 5090, empty disk.
|
| 11 |
+
|
| 12 |
+
## What we did (do not skip these)
|
| 13 |
+
|
| 14 |
+
1. Snapshot `thefinalboss/fractus-cte` code + `scripts/fast4gpu_boost_v4.py`.
|
| 15 |
+
2. Download 8× `checkpoints/x8run/fractus_1b_gpu{i}.pt` + 8× `RESUME_gpu{i}.json`.
|
| 16 |
+
3. Download Phase-2 memmaps from `thefinalboss/fractus-datasets`
|
| 17 |
+
`tokenized/phase2/shard_phase2_gpu{i}.npy` → symlink `data/shard_gpu{i}.npy`.
|
| 18 |
+
4. Stock image torch **2.2.1+cu121 has no kernels for Blackwell 5090**
|
| 19 |
+
(`CUDA error: no kernel image is available`). Upgrade: **torch 2.11.0+cu128**.
|
| 20 |
+
5. `set_p0_routing` missing on the published `continuous_engine.py` — added a flag setter so v4 can boot. Weights already carry the earlier routing surgery.
|
| 21 |
+
6. Live `unique@40` probe crashed the first launch (`generate_greedy_prefix` not in that generate_aligned). **PROBE_EVERY=0**. Probes stay offline.
|
| 22 |
+
7. Relaunch one process per GPU:
|
| 23 |
+
|
| 24 |
+
```
|
| 25 |
+
CUDA_VISIBLE_DEVICES=$i GPU_ID=$i START_TOKEN=<RESUME.start_token_next> \
|
| 26 |
+
BATCH=4 SEQ=128 CE_CHUNK=2048 FRACTUS_ATTN_IMPL=chunked BLOCK_CKPT=1 \
|
| 27 |
+
SS_RATE=1.0 SS_PROB_START=0.2 SS_PROB_END=0.5 P0=1 COMPILE=0 PROBE_EVERY=0 \
|
| 28 |
+
CKPT_IN=checkpoints/x8run/fractus_1b_gpu$i.pt \
|
| 29 |
+
CKPT_OUT=checkpoints/x8run/fractus_1b_gpu$i.pt \
|
| 30 |
+
SHARD=data/shard_gpu$i.npy \
|
| 31 |
+
python -u scripts/fast4gpu_boost_v4.py
|
| 32 |
+
```
|
| 33 |
+
|
| 34 |
+
Hourly HF push of `checkpoints/x8run/*`.
|
| 35 |
+
|
| 36 |
+
## Resume offsets used 2026-08-28
|
| 37 |
+
|
| 38 |
+
| GPU | start_token_next |
|
| 39 |
+
|-----|------------------|
|
| 40 |
+
| 0 | 79,922,176 |
|
| 41 |
+
| 1 | 78,443,520 |
|
| 42 |
+
| 2 | 76,420,096 |
|
| 43 |
+
| 3 | 77,746,176 |
|
| 44 |
+
| 4 | 76,160,000 |
|
| 45 |
+
| 5 | 77,310,976 |
|
| 46 |
+
| 6 | 77,793,280 |
|
| 47 |
+
| 7 | 76,677,120 |
|
| 48 |
+
|
| 49 |
+
Shard length ≈ 429.9M int32 tokens / GPU.
|
| 50 |
+
|
| 51 |
+
## Honest metrics after resume (first minutes)
|
| 52 |
+
|
| 53 |
+
- Throughput ~840–940 tok/s/GPU, climbing.
|
| 54 |
+
- `ema_tf` on GPU0/1 settled toward mid-teens after the first inflated steps (do **not** read the first 40-step EMA as the run).
|
| 55 |
+
- Speech gate remains **unique@40 greedy PREFIX** (no ban, no temperature). Last offline probe before the pod death: 3 / 12 / 1 / 6 on the four standard prompts. That is **NO-GO**.
|
| 56 |
+
- Finishing this Phase-2 pass from ~80M @ ~900 tok/s ≈ **4.1–4.5 days** wall if 8 GPUs stay up.
|
| 57 |
+
|
| 58 |
+
## What is NOT true
|
| 59 |
+
|
| 60 |
+
- Low teacher-force CE ≠ the model speaks.
|
| 61 |
+
- Banned / temperature decode is not the metric.
|
| 62 |
+
- This pass is not “4.2B tokens from zero”. Phase-1 weights were already in the seed. This is the remainder of the 430M/GPU Phase-2 shards.
|