thefinalboss commited on
Commit
6694761
Β·
verified Β·
1 Parent(s): 0077eee

Upload docs/COMPOSABILITY_AND_SURGERY.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. docs/COMPOSABILITY_AND_SURGERY.md +123 -0
docs/COMPOSABILITY_AND_SURGERY.md ADDED
@@ -0,0 +1,123 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Fractus Emergent Properties β€” In-Training Surgery and Checkpoint Composition
2
+
3
+ **Updated:** 2026-08-17 00:26 UTC
4
+
5
+ These properties were **discovered in production** while correcting Fractus mid-run.
6
+ They are not marketing claims; they were forced by real bugs, resumes, merges, and continued training.
7
+
8
+ ---
9
+
10
+ ## 1. In-training surgery (open-heart)
11
+
12
+ **Claim:** Fractus can be modified **during training** without throwing away learned weights.
13
+
14
+ ### What we did live
15
+ - Unfroze Kuramoto (removed no_grad) while runs continued from the same token offsets
16
+ - Wired load-balance into the loss mid-run
17
+ - Raised MoE gate temperature
18
+ - Switched from last-token CE to **dense full-chunk CE** (stage2)
19
+ - Added scheduled sampling
20
+ - Recalibrated LR and logging
21
+ - Fixed train/gen path mismatch (RK4 alignment, tick_chunk decode)
22
+
23
+ ### What was preserved
24
+ - All previously digested tokens (exact resume manifests)
25
+ - Model state_dict tensors (same shapes)
26
+ - Multi-day progress across pod restarts
27
+
28
+ ### Operational rule
29
+ 1. Save / keep current .pt
30
+ 2. Patch body (code / objective / routing)
31
+ 3. Reload weights
32
+ 4. Resume at recorded token offset
33
+ 5. Document the surgery
34
+
35
+ **This is mid-training operability**, not a full retrain.
36
+
37
+ See also: docs/OPERABILITY_MIDTRAIN.md
38
+
39
+ ---
40
+
41
+ ## 2. Parallel brains then merge
42
+
43
+ **Claim:** You can train **many different .pt files** (shards / GPUs / runs) and **merge** them into one Fractus.
44
+
45
+ ### Proven path (this run)
46
+ - 4 independent GPU runs on 4 data shards
47
+ - Each writes checkpoints/fractus_1b_gpu0.pt ... gpu3.pt
48
+ - **Mean-merge** of floating tensors -> unified checkpoint
49
+ - Example artifact: checkpoints/FRACTUS_1B_STAGE2_MERGED.pt
50
+ - Unified .pt still generates (different attractors than a single shard)
51
+
52
+ ### Workflow
53
+
54
+
55
+ Constraint: **same architecture shapes** (d_model, layers, experts, etc.).
56
+
57
+ ---
58
+
59
+ ## 3. Compose an already-trained Fractus with more checkpoints
60
+
61
+ **Claim:** Take a Fractus already trained, merge it with other compatible .pt files, then keep training β€” the system **accumulates** rather than restarting from zero.
62
+
63
+ ### Composition loop
64
+
65
+
66
+ ### What grow means here
67
+ Two related notions:
68
+
69
+ 1. **Compositional growth (proven in this run)**
70
+ - Knowledge / weights from multiple runs can be averaged into one brain
71
+ - Then training continues on the merged brain
72
+ - Digestion is not discarded
73
+
74
+ 2. **Architectural growth (design of Fractus / paliers)**
75
+ - Progressive growth: width, depth, experts can increase across paliers
76
+ - maybe_grow adds capacity when routing demands it
77
+ - Documented in the Chinchilla/growth notes
78
+
79
+ Merging many same-shape .pt files is **composition**.
80
+ Growing parameter count is **structural growth**.
81
+ Both are part of the Fractus operating model; this production run proved composition + mid-train surgery under fire.
82
+
83
+ ---
84
+
85
+ ## 4. Why this matters
86
+
87
+ Standard large training runs treat the job as fragile:
88
+ - change the objective -> often restart
89
+ - multi-GPU shard brains -> hard to recombine
90
+ - decode bugs -> blamed on needs more train only
91
+
92
+ Fractus production forced a different picture:
93
+
94
+ | Property | Observation |
95
+ |----------|-------------|
96
+ | Operable mid-train | Yes β€” weights kept, body patched |
97
+ | Multi-pt parallel train | Yes β€” 4 shards |
98
+ | Merge to one Fractus | Yes β€” mean-merge works |
99
+ | Resume after surgery | Yes β€” exact token offsets |
100
+ | Train then merge then train | Yes β€” intended workflow |
101
+
102
+ ---
103
+
104
+ ## 5. Limits (honest)
105
+
106
+ - Merge requires **matching tensor shapes**
107
+ - Naive mean-merge of very divergent specialists can **soup** skills
108
+ - Composition is not magic AGI; it is a **workflow property** of separable brain (.pt) + body (code)
109
+ - Generation coherence still lags teacher-forced loss (see docs/LOSS_VS_GEN.md)
110
+
111
+ ---
112
+
113
+ ## 6. Related documents
114
+
115
+ - docs/OPERABILITY_MIDTRAIN.md β€” open-heart principle and stage2 surgery
116
+ - docs/DISCOVERY_LOG.md β€” bugs and emergent features list
117
+ - docs/MASTER_RUN_LOG.md β€” full chronology
118
+ - docs/2026-08-12-fractus-chinchilla.md β€” progressive growth / paliers
119
+ - docs/LOSS_VS_GEN.md β€” why TF loss is not gen quality
120
+
121
+ ---
122
+
123
+ *Discovered by operating Fractus under real training pressure, not by design slide.*