thefinalboss commited on
Commit
84a4afa
Β·
verified Β·
1 Parent(s): 1ff0c84

Upload docs/DISCOVERY_LOG.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. docs/DISCOVERY_LOG.md +123 -0
docs/DISCOVERY_LOG.md ADDED
@@ -0,0 +1,123 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Fractus Discovery Log β€” Bugs, Optimizations, Emergent Features
2
+
3
+ **Updated:** 2026-08-16 18:28 UTC
4
+
5
+ This document records what we found *while* running Fractus-1B β€” things that go beyond the original design notes. Empirical, not marketing.
6
+
7
+ ---
8
+
9
+ ## 1. Critical training bugs found (and fixed)
10
+
11
+ ### 1.1 Kuramoto was not learning during training
12
+
13
+ **Symptom:** `omega` stayed near init (Β±0.05), phase order parameter r β‰ˆ 0.01–0.03, soft dynamics.
14
+
15
+ **Root cause:** In `CTEBlock.tick_chunk_core`, Kuramoto integration ran under `torch.no_grad()`:
16
+
17
+
18
+
19
+ Comment in code even said "clock, not learned". So CE loss never reached `omega` / coupling.
20
+
21
+ **Fix:** Remove `no_grad` around Kuramoto; keep phase *state* detached for carry, keep differentiable `theta` for MoE routing so parameters receive gradients.
22
+
23
+ **Status:** Fixed in `fractus/continuous_engine.py` (surgery 2026-08-16).
24
+
25
+ ### 1.2 Load-balance loss was computed then thrown away
26
+
27
+ **Symptom:** Offline probe showed ~70% experts dead (90+/128), only 2–3 experts active per block.
28
+
29
+ **Root causes:**
30
+ 1. `tick_chunk_train` did `total_lb + lb.detach()` β€” no gradient through LB
31
+ 2. `fast4gpu.py` optimized **only** cross-entropy β€” never added `lb_loss` to the loss
32
+
33
+ **Fix:**
34
+ - Keep LB in the graph
35
+ - `tick_chunk_train` returns `(logits, lb_loss)`
36
+ - Surgery trainer: `loss = CE + 0.02 * lb`
37
+
38
+ **Status:** Fixed; live `lbβ‰ˆ14` on all GPUs after resume.
39
+
40
+ ### 1.3 Probe false alarm: "all phases are zero"
41
+
42
+ **Symptom:** A probe script reported all Kuramoto phases at 0.0.
43
+
44
+ **Root cause:** Script called `reset_thought()` before measuring β€” which zeros state buffers.
45
+
46
+ **Reality in checkpoints:** phases nonzero, std β‰ˆ 1.8 across all 16 blocks Γ— 4 GPUs.
47
+
48
+ **Lesson:** Never diagnose dynamical state after an intentional reset.
49
+
50
+ ---
51
+
52
+ ## 2. Optimizations discovered in production
53
+
54
+ | Optimization | What we learned |
55
+ |--------------|-----------------|
56
+ | 4-GPU independent shards + mean-merge | Works; unified `.pt` generates; different lexical attractors than single shard |
57
+ | Resume from exact token offset | Manifest-driven `start_token` preserves progress across pod reboot |
58
+ | Gate temperature ↑ (1.0 β†’ 2.5) | Softens von Mises routing; more experts can enter the top-k mix |
59
+ | Omega scale Γ—4 on resume | Restores phase-rate diversity without wiping weights |
60
+ | `tick_vec` multimodal path | Vision patches can drive CTE without touching token embedding |
61
+ | CPU eyes prototype (CIFAR) | Small CTE+PatchEmbed learns real images offline while 1B trains on GPU |
62
+
63
+ ---
64
+
65
+ ## 3. Features that emerged beyond the original plan
66
+
67
+ ### 3.1 Operable / open-heart model
68
+ Weights (`.pt`) + body (`fractus/` code) are separable. We can:
69
+ - merge brains
70
+ - change routing temperature
71
+ - inject LB pressure
72
+ - add vision front-end
73
+ without a full retrain from zero.
74
+
75
+ ### 3.2 Infinite-ish checkpoint fusion (same architecture)
76
+ Compatible checkpoints can be mean-merged and trained again:
77
+
78
+ ```
79
+ train β†’ merge β†’ train β†’ merge β†’ ...
80
+ ```
81
+
82
+ Constraint: same shapes (d_model, layers, experts, etc.). Divergent merges can soup skills; shard-merge is the proven path.
83
+
84
+ ### 3.3 Mid-training generation behavior
85
+ At loss ~33β†’25, generation pipeline works but outputs **word-level repetition collapse** (real tokens stuck in loops: "Colorado", "Population", "Fate", "thinks"…). Documents that Fractus emits lexical tokens before coherent sentences.
86
+
87
+ ### 3.4 Parallel modality track
88
+ Text 1B digestion and vision prototype can run in parallel (GPU text + CPU eyes) without stopping the main run.
89
+
90
+ ### 3.5 Routing pathology as first-class debug target
91
+ Expert-hit histograms + phase order parameter `r` are necessary metrics. Loss alone hides "model learns with 3 experts".
92
+
93
+ ---
94
+
95
+ ## 4. Live metrics after routing surgery (resume)
96
+
97
+ Resume offsets preserved from pre-crash run (~176–187M tokens/GPU).
98
+
99
+ | GPU | Resume start | Snapshot tokens | CE loss | lb |
100
+ |-----|--------------|-----------------|---------|-----|
101
+ | 0 | 176,281,600 | 176,614,400 | 36.5 | 14.026 |
102
+ | 1 | 186,137,600 | 186,444,800 | 12.4 | 14.024 |
103
+ | 2 | 186,854,400 | 187,212,800 | 34.7 | 14.026 |
104
+ | 3 | 183,833,600 | 184,166,400 | 41.0 | 14.027 |
105
+
106
+ Post-surgery CE can spike briefly then fall (routing distribution shift). GPU1 recovered fastest into the teens.
107
+
108
+ ---
109
+
110
+ ## 5. What this means for the Fractus thesis
111
+
112
+ Fractus is not only "another 1B trained on shards". The run forced discovery of:
113
+
114
+ 1. **Silent non-learning** of the phase clock under `no_grad`
115
+ 2. **Silent expert death** without LB in the loss
116
+ 3. **Composable checkpoints** as a workflow
117
+ 4. **Operability** (surgery without discarding digestion)
118
+
119
+ The architecture does more than the first README described β€” because production training exposed the dynamical bottlenecks.
120
+
121
+ ---
122
+
123
+ *Keep this log updated when new probes, merges, or surgeries land.*