File size: 9,932 Bytes
84a4afa 5326f88 84a4afa 5326f88 84a4afa 045ee27 5326f88 84a4afa | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 | # Fractus Discovery Log — Bugs, Optimizations, Emergent Features
**Updated:** 2026-09-02 (adds sandbox live-operation discoveries on the X8 unified checkpoint)
This document records what we found *while* running Fractus-1B — things that go beyond the original design notes. Empirical, not marketing.
---
## 1. Critical training bugs found (and fixed)
### 1.1 Kuramoto was not learning during training
**Symptom:** `omega` stayed near init (±0.05), phase order parameter r ≈ 0.01–0.03, soft dynamics.
**Root cause:** In `CTEBlock.tick_chunk_core`, Kuramoto integration ran under `torch.no_grad()`:
Comment in code even said "clock, not learned". So CE loss never reached `omega` / coupling.
**Fix:** Remove `no_grad` around Kuramoto; keep phase *state* detached for carry, keep differentiable `theta` for MoE routing so parameters receive gradients.
**Status:** Fixed in `fractus/continuous_engine.py` (surgery 2026-08-16).
### 1.2 Load-balance loss was computed then thrown away
**Symptom:** Offline probe showed ~70% experts dead (90+/128), only 2–3 experts active per block.
**Root causes:**
1. `tick_chunk_train` did `total_lb + lb.detach()` — no gradient through LB
2. `fast4gpu.py` optimized **only** cross-entropy — never added `lb_loss` to the loss
**Fix:**
- Keep LB in the graph
- `tick_chunk_train` returns `(logits, lb_loss)`
- Surgery trainer: `loss = CE + 0.02 * lb`
**Status:** Fixed; live `lb≈14` on all GPUs after resume.
### 1.3 Probe false alarm: "all phases are zero"
**Symptom:** A probe script reported all Kuramoto phases at 0.0.
**Root cause:** Script called `reset_thought()` before measuring — which zeros state buffers.
**Reality in checkpoints:** phases nonzero, std ≈ 1.8 across all 16 blocks × 4 GPUs.
**Lesson:** Never diagnose dynamical state after an intentional reset.
---
## 2. Optimizations discovered in production
| Optimization | What we learned |
|--------------|-----------------|
| 4-GPU independent shards + mean-merge | Works; unified `.pt` generates; different lexical attractors than single shard |
| Resume from exact token offset | Manifest-driven `start_token` preserves progress across pod reboot |
| Gate temperature ↑ (1.0 → 2.5) | Softens von Mises routing; more experts can enter the top-k mix |
| Omega scale ×4 on resume | Restores phase-rate diversity without wiping weights |
| `tick_vec` multimodal path | Vision patches can drive CTE without touching token embedding |
| CPU eyes prototype (CIFAR) | Small CTE+PatchEmbed learns real images offline while 1B trains on GPU |
---
## 3. Features that emerged beyond the original plan
### 3.1 Operable / open-heart model
Weights (`.pt`) + body (`fractus/` code) are separable. We can:
- merge brains
- change routing temperature
- inject LB pressure
- add vision front-end
without a full retrain from zero.
### 3.2 Infinite-ish checkpoint fusion (same architecture)
Compatible checkpoints can be mean-merged and trained again:
```
train → merge → train → merge → ...
```
Constraint: same shapes (d_model, layers, experts, etc.). Divergent merges can soup skills; shard-merge is the proven path.
### 3.3 Mid-training generation behavior
At loss ~33→25, generation pipeline works but outputs **word-level repetition collapse** (real tokens stuck in loops: "Colorado", "Population", "Fate", "thinks"…). Documents that Fractus emits lexical tokens before coherent sentences.
### 3.4 Parallel modality track
Text 1B digestion and vision prototype can run in parallel (GPU text + CPU eyes) without stopping the main run.
### 3.5 Routing pathology as first-class debug target
Expert-hit histograms + phase order parameter `r` are necessary metrics. Loss alone hides "model learns with 3 experts".
---
## 4. Live metrics after routing surgery (resume)
Resume offsets preserved from pre-crash run (~176–187M tokens/GPU).
| GPU | Resume start | Snapshot tokens | CE loss | lb |
|-----|--------------|-----------------|---------|-----|
| 0 | 176,281,600 | 176,614,400 | 36.5 | 14.026 |
| 1 | 186,137,600 | 186,444,800 | 12.4 | 14.024 |
| 2 | 186,854,400 | 187,212,800 | 34.7 | 14.026 |
| 3 | 183,833,600 | 184,166,400 | 41.0 | 14.027 |
Post-surgery CE can spike briefly then fall (routing distribution shift). GPU1 recovered fastest into the teens.
---
## 5. What this means for the Fractus thesis
Fractus is not only "another 1B trained on shards". The run forced discovery of:
1. **Silent non-learning** of the phase clock under `no_grad`
2. **Silent expert death** without LB in the loss
3. **Composable checkpoints** as a workflow
4. **Operability** (surgery without discarding digestion)
5. **A body that speaks before the brain converges** — decode-level surgery on the
same weights turns 2.14 unique/48 (greedy carry) into 32.86 unique/48 fluid,
repeat-free speech. The collapse is a sampling-regime attractor, not a weights
defect: the thesis holds above the training gate (see §7).
The architecture does more than the first README described — because production training exposed the dynamical bottlenecks.
---
## 6. Discoveries from the optimization session (2026-08-22/23)
### 6.1 The "attention matmul" was a disguised cumsum
The production kernel computed causal sums via
`einsum("tj,bjpq->btpq", tril_mask, outer)` — an O(C²·dH²) masked contraction
that IS mathematically a prefix sum. Replaced by `torch.cumsum(outer, dim=1)`
(O(C·dH²)), proven bit-close to the kept einsum reference (forward, gradients,
carried S/z). Lesson: profile semantics, not just FLOPs — the mask+matmul form
hid a linear-scan structure for months.
### 6.2 Discrete routing makes strict equivalence claims wrong at block level
topk over von Mises gates is DISCRETE: a float32 rounding difference near a
gate tie flips expert choice for isolated tokens (measure-zero boundary),
and depth amplifies it. Honest framing adopted: continuous stages are proven
bit-close; full-block equality is asserted statistically (median at rounding
scale, mismatch fraction bounded). Tests encode exactly this.
### 6.3 Two constructions under one seed have DIFFERENT weights
`manual_seed(s); A = Model(); B = Model()` gives two different networks — the
RNG stream continues across constructions. Any A/B equivalence test must clone
weights (`load_state_dict`) instead of re-seeding. This masqueraded as a
"routing divergence" for half a session before being identified.
### 6.4 CPU commit-limit segfaults are data, not noise
At 1B shapes the reference kernels materialize (G, C, dH, dH) and hit the
native commit limit on a RAM-constrained host (segfault, not MemoryError).
The benchmark harness now runs each cell in its own subprocess so one crash
documents itself without killing the table — and the crash boundary itself
measures the memory win of the chunked kernel.
---
## 7. Sandbox live operation on the X8 unified checkpoint (2026-09-02)
External CPU sandbox (fp32) loaded `FRACTUS_1B_X8_MERGED.pt` and ran the full
README gate protocol + a live speech operation on the continuous state.
Empirical findings:
### 7.1 The generation collapse is a greedy-regime attractor, not a distribution property
**Symptom:** greedy carry (native continuous mode) = 48 consecutive identical
tokens (2.14 unique/48). Same weights, top-150 sampling @ T 1.15 with ban 20:
32.86 unique/48, zero repeats, no 2-cycles.
**Root cause:** von Mises soft gates (architecture) + anti-copy/LB training
spread mass across the top-150 support; the argmax landscape still holds short
repetition attractors. The gate's greedy unique@40 measures the most
pathological regime of the model.
**Status:** Measured. Action: add a second birth metric — sampled
unique@48 (top-k 150, T 1.15, ban 20, continuous state). Gate v2 = greedy
(formal) + sampled (the "body") + state stability. See
docs/X8_CIEL_OUVERT.md §5.
### 7.2 The continuous (S,z) state is a memory of degradation, not a memory
**Symptom:** in one unreset thought stream, vocabulary overlap with earlier
segments: 91 % (segment 2), 100 % (segment 3). 192-token run = a ~12-token
motif macro-cycle (ban 20 breaks <20 loops, not 20+ motifs).
**Root cause:** `attn_S += outer.detach()` is unbounded, no decay — the readout
`q·S / q·z` becomes a function of the whole history and the effective
vocabulary shrinks with state age. A multiplicative leak on (S,z)
(λ 0.99, half-life ~69 ticks) makes continuity sustainable but does not invert
the shrinkage.
**Status:** Measured. Action: active forgetting (selective/importance-weighted
decay, or normalize S/z to code a mean). Continuity metric: overlap < 0.7 AND
unique ≥ 20 over 3+ segments.
### 7.3 Decode-surgery phase noise is a no-op on the tick path
**Symptom:** r ≈ 0.0028 in every configuration; perturbing
`blk.kuramoto_phases` never changes outputs.
**Root cause:** `CTEBlock.tick_single` computes a fresh θ from the hidden state
each tick and feeds it to the MoE; the stored `kuramoto_phases` buffer is read
only by the expert-hit counter. `decode_surgery.generate_with_surgery` mutates
the dead buffer.
**Status:** Verified in code. Action: apply phase noise to the θ the MoE
actually reads (one line) or delete the dead code. Kuramoto order must be
raised in training (LB under gradient), not in decode.
### 7.4 Publish gap: the load fix is not wired into the public path
**Symptom:** a fresh load of the public checkpoint runs at
`moe.temperature = 1.0` (code default) — the RAW config — while QuickPod runs
at 2.5 after `apply_kuramoto_routing_fix`. Nothing in the public code calls
the fix; `memory.py` docstring promises an `engine.inject_memory()` that does
not exist.
**Status:** Measured (probe RAW = 1.0 vs FIX = 2.5). Action: README line —
"apply `apply_kuramoto_routing_fix` at load"; implement or de-document
`inject_memory`.
---
---
*Keep this log updated when new probes, merges, or surgeries land.*
|