Upload docs/DISCOVERY_LOG.md with huggingface_hub
Browse files- docs/DISCOVERY_LOG.md +123 -0
docs/DISCOVERY_LOG.md
ADDED
|
@@ -0,0 +1,123 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Fractus Discovery Log β Bugs, Optimizations, Emergent Features
|
| 2 |
+
|
| 3 |
+
**Updated:** 2026-08-16 18:28 UTC
|
| 4 |
+
|
| 5 |
+
This document records what we found *while* running Fractus-1B β things that go beyond the original design notes. Empirical, not marketing.
|
| 6 |
+
|
| 7 |
+
---
|
| 8 |
+
|
| 9 |
+
## 1. Critical training bugs found (and fixed)
|
| 10 |
+
|
| 11 |
+
### 1.1 Kuramoto was not learning during training
|
| 12 |
+
|
| 13 |
+
**Symptom:** `omega` stayed near init (Β±0.05), phase order parameter r β 0.01β0.03, soft dynamics.
|
| 14 |
+
|
| 15 |
+
**Root cause:** In `CTEBlock.tick_chunk_core`, Kuramoto integration ran under `torch.no_grad()`:
|
| 16 |
+
|
| 17 |
+
|
| 18 |
+
|
| 19 |
+
Comment in code even said "clock, not learned". So CE loss never reached `omega` / coupling.
|
| 20 |
+
|
| 21 |
+
**Fix:** Remove `no_grad` around Kuramoto; keep phase *state* detached for carry, keep differentiable `theta` for MoE routing so parameters receive gradients.
|
| 22 |
+
|
| 23 |
+
**Status:** Fixed in `fractus/continuous_engine.py` (surgery 2026-08-16).
|
| 24 |
+
|
| 25 |
+
### 1.2 Load-balance loss was computed then thrown away
|
| 26 |
+
|
| 27 |
+
**Symptom:** Offline probe showed ~70% experts dead (90+/128), only 2β3 experts active per block.
|
| 28 |
+
|
| 29 |
+
**Root causes:**
|
| 30 |
+
1. `tick_chunk_train` did `total_lb + lb.detach()` β no gradient through LB
|
| 31 |
+
2. `fast4gpu.py` optimized **only** cross-entropy β never added `lb_loss` to the loss
|
| 32 |
+
|
| 33 |
+
**Fix:**
|
| 34 |
+
- Keep LB in the graph
|
| 35 |
+
- `tick_chunk_train` returns `(logits, lb_loss)`
|
| 36 |
+
- Surgery trainer: `loss = CE + 0.02 * lb`
|
| 37 |
+
|
| 38 |
+
**Status:** Fixed; live `lbβ14` on all GPUs after resume.
|
| 39 |
+
|
| 40 |
+
### 1.3 Probe false alarm: "all phases are zero"
|
| 41 |
+
|
| 42 |
+
**Symptom:** A probe script reported all Kuramoto phases at 0.0.
|
| 43 |
+
|
| 44 |
+
**Root cause:** Script called `reset_thought()` before measuring β which zeros state buffers.
|
| 45 |
+
|
| 46 |
+
**Reality in checkpoints:** phases nonzero, std β 1.8 across all 16 blocks Γ 4 GPUs.
|
| 47 |
+
|
| 48 |
+
**Lesson:** Never diagnose dynamical state after an intentional reset.
|
| 49 |
+
|
| 50 |
+
---
|
| 51 |
+
|
| 52 |
+
## 2. Optimizations discovered in production
|
| 53 |
+
|
| 54 |
+
| Optimization | What we learned |
|
| 55 |
+
|--------------|-----------------|
|
| 56 |
+
| 4-GPU independent shards + mean-merge | Works; unified `.pt` generates; different lexical attractors than single shard |
|
| 57 |
+
| Resume from exact token offset | Manifest-driven `start_token` preserves progress across pod reboot |
|
| 58 |
+
| Gate temperature β (1.0 β 2.5) | Softens von Mises routing; more experts can enter the top-k mix |
|
| 59 |
+
| Omega scale Γ4 on resume | Restores phase-rate diversity without wiping weights |
|
| 60 |
+
| `tick_vec` multimodal path | Vision patches can drive CTE without touching token embedding |
|
| 61 |
+
| CPU eyes prototype (CIFAR) | Small CTE+PatchEmbed learns real images offline while 1B trains on GPU |
|
| 62 |
+
|
| 63 |
+
---
|
| 64 |
+
|
| 65 |
+
## 3. Features that emerged beyond the original plan
|
| 66 |
+
|
| 67 |
+
### 3.1 Operable / open-heart model
|
| 68 |
+
Weights (`.pt`) + body (`fractus/` code) are separable. We can:
|
| 69 |
+
- merge brains
|
| 70 |
+
- change routing temperature
|
| 71 |
+
- inject LB pressure
|
| 72 |
+
- add vision front-end
|
| 73 |
+
without a full retrain from zero.
|
| 74 |
+
|
| 75 |
+
### 3.2 Infinite-ish checkpoint fusion (same architecture)
|
| 76 |
+
Compatible checkpoints can be mean-merged and trained again:
|
| 77 |
+
|
| 78 |
+
```
|
| 79 |
+
train β merge β train β merge β ...
|
| 80 |
+
```
|
| 81 |
+
|
| 82 |
+
Constraint: same shapes (d_model, layers, experts, etc.). Divergent merges can soup skills; shard-merge is the proven path.
|
| 83 |
+
|
| 84 |
+
### 3.3 Mid-training generation behavior
|
| 85 |
+
At loss ~33β25, generation pipeline works but outputs **word-level repetition collapse** (real tokens stuck in loops: "Colorado", "Population", "Fate", "thinks"β¦). Documents that Fractus emits lexical tokens before coherent sentences.
|
| 86 |
+
|
| 87 |
+
### 3.4 Parallel modality track
|
| 88 |
+
Text 1B digestion and vision prototype can run in parallel (GPU text + CPU eyes) without stopping the main run.
|
| 89 |
+
|
| 90 |
+
### 3.5 Routing pathology as first-class debug target
|
| 91 |
+
Expert-hit histograms + phase order parameter `r` are necessary metrics. Loss alone hides "model learns with 3 experts".
|
| 92 |
+
|
| 93 |
+
---
|
| 94 |
+
|
| 95 |
+
## 4. Live metrics after routing surgery (resume)
|
| 96 |
+
|
| 97 |
+
Resume offsets preserved from pre-crash run (~176β187M tokens/GPU).
|
| 98 |
+
|
| 99 |
+
| GPU | Resume start | Snapshot tokens | CE loss | lb |
|
| 100 |
+
|-----|--------------|-----------------|---------|-----|
|
| 101 |
+
| 0 | 176,281,600 | 176,614,400 | 36.5 | 14.026 |
|
| 102 |
+
| 1 | 186,137,600 | 186,444,800 | 12.4 | 14.024 |
|
| 103 |
+
| 2 | 186,854,400 | 187,212,800 | 34.7 | 14.026 |
|
| 104 |
+
| 3 | 183,833,600 | 184,166,400 | 41.0 | 14.027 |
|
| 105 |
+
|
| 106 |
+
Post-surgery CE can spike briefly then fall (routing distribution shift). GPU1 recovered fastest into the teens.
|
| 107 |
+
|
| 108 |
+
---
|
| 109 |
+
|
| 110 |
+
## 5. What this means for the Fractus thesis
|
| 111 |
+
|
| 112 |
+
Fractus is not only "another 1B trained on shards". The run forced discovery of:
|
| 113 |
+
|
| 114 |
+
1. **Silent non-learning** of the phase clock under `no_grad`
|
| 115 |
+
2. **Silent expert death** without LB in the loss
|
| 116 |
+
3. **Composable checkpoints** as a workflow
|
| 117 |
+
4. **Operability** (surgery without discarding digestion)
|
| 118 |
+
|
| 119 |
+
The architecture does more than the first README described β because production training exposed the dynamical bottlenecks.
|
| 120 |
+
|
| 121 |
+
---
|
| 122 |
+
|
| 123 |
+
*Keep this log updated when new probes, merges, or surgeries land.*
|