File size: 6,036 Bytes
8ec3d47 5326f88 8ec3d47 963e28c 8ec3d47 f27efd5 8ec3d47 26916cf 5326f88 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 | # Fractus-1B Production Run Log (Complete)
**Updated:** 2026-09-02 23:55 UTC
English master log of the Fractus continuous-thought 1B training run: discoveries, bugs, surgeries, metrics, and generation probes.
---
## 1. Architecture
Fractus is not a standard decoder-only transformer training loop.
- ContinuousThoughtEngine (CTE): residual thought state across ticks
- Per-block linear attention with continuous carry (S, z)
- Kuramoto phase oscillators for routing dynamics
- PhaseRoutedMoE (von Mises gates, top-k experts, load-balance loss)
- Tied embedding / output head
- Target: d_model=1280, 16 layers, 128 experts, top-k=2
Checkpoints store weights + dynamic state; package fractus/ is the body.
---
## 2. Hardware and data
- 4x NVIDIA RTX 5090
- Independent shard per GPU (~1.057B tokens each, GPT-2 BPE)
- Batch B=2, seq_len=128 (256 tokens/step)
- Optimizer: SGD momentum 0.9
- Precision: bfloat16 autocast
---
## 3. Chronology of interventions
### Phase A β Initial distributed training
- Four independent GPU runs
- Early objective: last-position CE only
- Stable throughput, descending loss
- Generation collapsed to single-token loops
### Phase B β Routing surgery (weights kept)
Bugs found:
1. Kuramoto under torch.no_grad in tick_chunk_core (omega never trained)
2. lb_loss detached and never in training loss (dead experts)
3. Hard gate temperature
Fixes: Kuramoto gradients, CE+0.02*lb, gate temp 2.5, omega diversity, exact token resume
### Phase C β Stage2 dense CE
- tick_chunk_train returns logits for all positions
- Dense next-token CE over full chunk
- CE dropped hard (GPU1 into ~2 range)
### Phase D β Decode surgery (weights unchanged)
- Phase/thought noise, frequency penalty, cycle bans, forced escape tokens
- Broke single-token decode lock
- Did not produce coherent English by itself
### Phase E β Train/gen path mismatch
- Train: tick_chunk_core (causal attn + RK4 Kuramoto)
- Default gen: tick_single (simpler attn + Euler Kuramoto)
- Fix: align tick_single to RK4; prefer 100% tick_chunk generation
- Module: fractus/generate_aligned.py (generate_chunk, generate_window)
### Phase F β Fine phase loss recalibration
- Cumulative average CE misleading near 2.0
- Batch CE + EMA; real session tok/s
- LR 1e-3 -> 5e-4
### Phase G β Scheduled sampling (loss vs gen gap)
- Warm TF CE ~1.3-1.5 can coexist with broken free-run generation
- Free-run AR CE vs ground truth is catastrophic (exposure bias)
- Fix: sequential scheduled sampling
- Pass A: teacher-forced CE + LB
- Pass B (~30% steps): mix ~20% model samples into inputs
- Logs: tf/ema_tf and ss/ema_ss
---
## 4. Live SS metrics (at doc time)
- GPU 0: GPU 0: 190,809,600 tf=7.760 ema_tf=7.689 lb=14.027 670 tok/s [ss]
- GPU 1: GPU 1: 200,499,200 tf=1.292 ema_tf=1.334 lb=14.027 670 tok/s [ss]
- GPU 2: GPU 2: 201,446,400 tf=1.766 ema_tf=1.789 lb=14.026 656 tok/s [ss]
- GPU 3: GPU 3: 198,656,000 tf=5.662 ema_tf=5.998 lb=14.027 670 tok/s [ss]
---
## 5. Generation summary
| Stage | Behavior |
|-------|----------|
| Early mid-train | Single-token loops |
| After decode surgery | Multi-token diversity, non-sentences |
| After path alignment | WINDOW more diverse than CHUNK |
| Current | Lexical noise / short cycles; not coherent prose |
Expected while SS is still teaching free-run and only a fraction of each shard is consumed.
---
## 6. Operability principle
1. Keep .pt weights
2. Patch body (engine / trainer / decode)
3. Resume at recorded token offset
4. Document the surgery
Do not discard multi-day digestion for routing, objective, or decode bugs.
---
## 7. Key files
- fractus/continuous_engine.py β CTE body
- fractus/generate_aligned.py β train-aligned decode
- fractus/decode_surgery.py β optional decode stack
- scripts/fast4gpu_stage2_ss.py β current trainer
- scripts/fast4gpu_stage2_fine.py β fine phase trainer
- RESUME_MANIFEST_*.json β exact offsets
- checkpoints/fractus_1b_gpu0-3.pt β per-GPU brains
---
## 8. Related docs
- docs/COMPOSABILITY_AND_SURGERY.md β in-training surgery + multi-pt merge/grow
- docs/DISCOVERY_LOG.md
- docs/OPERABILITY_MIDTRAIN.md
- docs/DECODE_SURGERY.md
- docs/TRAIN_GEN_MISMATCH.md
- docs/GENERATE_ALIGNED.md
- docs/FINE_PHASE.md
- docs/LOSS_VS_GEN.md
- docs/TRUSTED_LOSS.md β which loss to trust for train vs gen
- docs/STAGE2_SS_GEN_PROBE.md (generation probe with outputs)
---
## 9. Conclusion
1. Learning is real under teacher-forced dense CE.
2. Generation coherence is not yet achieved.
3. Low TF loss does not imply clean free-run text.
4. Active remediation: scheduled sampling + aligned tick_chunk decode.
5. Keep rolling; re-probe generation after more SS tokens.
---
## 10. 2026-08-28 23:35 UTC β Decode surgery I (window + anti-copy)
- Space showed `RetailΓ32` / unique@32=1 on greedy length-1 `tick_chunk` + carry.
- Replaced that path in `fractus/generate_aligned.py`: sliding causal window 64 + mask previous token.
- Weights / trainers unchanged. QuickPod still on boost_v4 ~89M/430M GPU0.
- Post-op unique@32: 16 / 14 / 27 / 12 vs 1 / 1 / 1 / 2.
- Short cycles remain. Speech not claimed.
- Full note: `docs/2026-08-28-DECODE-WINDOW.md`
---
## 11. 2026-09-02 22:56β23:40 UTC β X8 unified checkpoint, sandbox live operation
- 8 QuickPod brains merged (`mean_float_params`) β `FRACTUS_1B_X8_MERGED.pt`
- Full README gate protocol reproduced in an external CPU sandbox
- Live operation on the same weights: greedy carry collapses to single-token
- Continuous state without forgetting shrinks: vocabulary overlap 91 % β 100 %
- Kuramoto order r β 0.003 everywhere; decode-surgery phase noise is a no-op
- Gate v2 proposed: greedy unique@40 (formal, unchanged) + sampled unique@48
(the "body" β GO today at 32.86) + continuous-state stability (overlap < 0.7,
unique β₯ 20 over 3 segments).
- Full journal + repro: `docs/X8_CIEL_OUVERT.md` Β· results:
`docs/X8_UNIFIED_PROBE.json`, `docs/OPERATION_CIEL_OUVERT_RESULTS.json` Β·
harness: `scripts/probe_ciel_ouvert.py`.
|