File size: 6,036 Bytes
8ec3d47
 
5326f88
8ec3d47
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
963e28c
 
8ec3d47
 
 
 
 
 
 
f27efd5
8ec3d47
 
 
 
 
 
 
 
 
 
 
26916cf
 
 
 
 
 
 
 
 
 
 
 
5326f88
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
# Fractus-1B Production Run Log (Complete)

**Updated:** 2026-09-02 23:55 UTC

English master log of the Fractus continuous-thought 1B training run: discoveries, bugs, surgeries, metrics, and generation probes.

---

## 1. Architecture

Fractus is not a standard decoder-only transformer training loop.

- ContinuousThoughtEngine (CTE): residual thought state across ticks
- Per-block linear attention with continuous carry (S, z)
- Kuramoto phase oscillators for routing dynamics
- PhaseRoutedMoE (von Mises gates, top-k experts, load-balance loss)
- Tied embedding / output head
- Target: d_model=1280, 16 layers, 128 experts, top-k=2

Checkpoints store weights + dynamic state; package fractus/ is the body.

---

## 2. Hardware and data

- 4x NVIDIA RTX 5090
- Independent shard per GPU (~1.057B tokens each, GPT-2 BPE)
- Batch B=2, seq_len=128 (256 tokens/step)
- Optimizer: SGD momentum 0.9
- Precision: bfloat16 autocast

---

## 3. Chronology of interventions

### Phase A β€” Initial distributed training
- Four independent GPU runs
- Early objective: last-position CE only
- Stable throughput, descending loss
- Generation collapsed to single-token loops

### Phase B β€” Routing surgery (weights kept)
Bugs found:
1. Kuramoto under torch.no_grad in tick_chunk_core (omega never trained)
2. lb_loss detached and never in training loss (dead experts)
3. Hard gate temperature

Fixes: Kuramoto gradients, CE+0.02*lb, gate temp 2.5, omega diversity, exact token resume

### Phase C β€” Stage2 dense CE
- tick_chunk_train returns logits for all positions
- Dense next-token CE over full chunk
- CE dropped hard (GPU1 into ~2 range)

### Phase D β€” Decode surgery (weights unchanged)
- Phase/thought noise, frequency penalty, cycle bans, forced escape tokens
- Broke single-token decode lock
- Did not produce coherent English by itself

### Phase E β€” Train/gen path mismatch
- Train: tick_chunk_core (causal attn + RK4 Kuramoto)
- Default gen: tick_single (simpler attn + Euler Kuramoto)
- Fix: align tick_single to RK4; prefer 100% tick_chunk generation
- Module: fractus/generate_aligned.py (generate_chunk, generate_window)

### Phase F β€” Fine phase loss recalibration
- Cumulative average CE misleading near 2.0
- Batch CE + EMA; real session tok/s
- LR 1e-3 -> 5e-4

### Phase G β€” Scheduled sampling (loss vs gen gap)
- Warm TF CE ~1.3-1.5 can coexist with broken free-run generation
- Free-run AR CE vs ground truth is catastrophic (exposure bias)
- Fix: sequential scheduled sampling
  - Pass A: teacher-forced CE + LB
  - Pass B (~30% steps): mix ~20% model samples into inputs
- Logs: tf/ema_tf and ss/ema_ss

---

## 4. Live SS metrics (at doc time)

- GPU 0: GPU 0:  190,809,600 tf=7.760 ema_tf=7.689 lb=14.027 670 tok/s [ss]
- GPU 1: GPU 1:  200,499,200 tf=1.292 ema_tf=1.334 lb=14.027 670 tok/s [ss]
- GPU 2: GPU 2:  201,446,400 tf=1.766 ema_tf=1.789 lb=14.026 656 tok/s [ss]
- GPU 3: GPU 3:  198,656,000 tf=5.662 ema_tf=5.998 lb=14.027 670 tok/s [ss]

---

## 5. Generation summary

| Stage | Behavior |
|-------|----------|
| Early mid-train | Single-token loops |
| After decode surgery | Multi-token diversity, non-sentences |
| After path alignment | WINDOW more diverse than CHUNK |
| Current | Lexical noise / short cycles; not coherent prose |

Expected while SS is still teaching free-run and only a fraction of each shard is consumed.

---

## 6. Operability principle

1. Keep .pt weights
2. Patch body (engine / trainer / decode)
3. Resume at recorded token offset
4. Document the surgery

Do not discard multi-day digestion for routing, objective, or decode bugs.

---

## 7. Key files

- fractus/continuous_engine.py β€” CTE body
- fractus/generate_aligned.py β€” train-aligned decode
- fractus/decode_surgery.py β€” optional decode stack
- scripts/fast4gpu_stage2_ss.py β€” current trainer
- scripts/fast4gpu_stage2_fine.py β€” fine phase trainer
- RESUME_MANIFEST_*.json β€” exact offsets
- checkpoints/fractus_1b_gpu0-3.pt β€” per-GPU brains

---

## 8. Related docs

- docs/COMPOSABILITY_AND_SURGERY.md β€” in-training surgery + multi-pt merge/grow

- docs/DISCOVERY_LOG.md
- docs/OPERABILITY_MIDTRAIN.md
- docs/DECODE_SURGERY.md
- docs/TRAIN_GEN_MISMATCH.md
- docs/GENERATE_ALIGNED.md
- docs/FINE_PHASE.md
- docs/LOSS_VS_GEN.md
- docs/TRUSTED_LOSS.md β€” which loss to trust for train vs gen
- docs/STAGE2_SS_GEN_PROBE.md (generation probe with outputs)

---

## 9. Conclusion

1. Learning is real under teacher-forced dense CE.
2. Generation coherence is not yet achieved.
3. Low TF loss does not imply clean free-run text.
4. Active remediation: scheduled sampling + aligned tick_chunk decode.
5. Keep rolling; re-probe generation after more SS tokens.

---

## 10. 2026-08-28 23:35 UTC β€” Decode surgery I (window + anti-copy)

- Space showed `RetailΓ—32` / unique@32=1 on greedy length-1 `tick_chunk` + carry.
- Replaced that path in `fractus/generate_aligned.py`: sliding causal window 64 + mask previous token.
- Weights / trainers unchanged. QuickPod still on boost_v4 ~89M/430M GPU0.
- Post-op unique@32: 16 / 14 / 27 / 12 vs 1 / 1 / 1 / 2.
- Short cycles remain. Speech not claimed.
- Full note: `docs/2026-08-28-DECODE-WINDOW.md`

---

## 11. 2026-09-02 22:56–23:40 UTC β€” X8 unified checkpoint, sandbox live operation

- 8 QuickPod brains merged (`mean_float_params`) β†’ `FRACTUS_1B_X8_MERGED.pt`
- Full README gate protocol reproduced in an external CPU sandbox
- Live operation on the same weights: greedy carry collapses to single-token
- Continuous state without forgetting shrinks: vocabulary overlap 91 % β†’ 100 %
- Kuramoto order r β‰ˆ 0.003 everywhere; decode-surgery phase noise is a no-op
- Gate v2 proposed: greedy unique@40 (formal, unchanged) + sampled unique@48
  (the "body" β€” GO today at 32.86) + continuous-state stability (overlap < 0.7,
  unique β‰₯ 20 over 3 segments).
- Full journal + repro: `docs/X8_CIEL_OUVERT.md` Β· results:
  `docs/X8_UNIFIED_PROBE.json`, `docs/OPERATION_CIEL_OUVERT_RESULTS.json` Β·
  harness: `scripts/probe_ciel_ouvert.py`.