Upload docs/COMPOSABILITY_AND_SURGERY.md with huggingface_hub
Browse files
docs/COMPOSABILITY_AND_SURGERY.md
ADDED
|
@@ -0,0 +1,123 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Fractus Emergent Properties β In-Training Surgery and Checkpoint Composition
|
| 2 |
+
|
| 3 |
+
**Updated:** 2026-08-17 00:26 UTC
|
| 4 |
+
|
| 5 |
+
These properties were **discovered in production** while correcting Fractus mid-run.
|
| 6 |
+
They are not marketing claims; they were forced by real bugs, resumes, merges, and continued training.
|
| 7 |
+
|
| 8 |
+
---
|
| 9 |
+
|
| 10 |
+
## 1. In-training surgery (open-heart)
|
| 11 |
+
|
| 12 |
+
**Claim:** Fractus can be modified **during training** without throwing away learned weights.
|
| 13 |
+
|
| 14 |
+
### What we did live
|
| 15 |
+
- Unfroze Kuramoto (removed no_grad) while runs continued from the same token offsets
|
| 16 |
+
- Wired load-balance into the loss mid-run
|
| 17 |
+
- Raised MoE gate temperature
|
| 18 |
+
- Switched from last-token CE to **dense full-chunk CE** (stage2)
|
| 19 |
+
- Added scheduled sampling
|
| 20 |
+
- Recalibrated LR and logging
|
| 21 |
+
- Fixed train/gen path mismatch (RK4 alignment, tick_chunk decode)
|
| 22 |
+
|
| 23 |
+
### What was preserved
|
| 24 |
+
- All previously digested tokens (exact resume manifests)
|
| 25 |
+
- Model state_dict tensors (same shapes)
|
| 26 |
+
- Multi-day progress across pod restarts
|
| 27 |
+
|
| 28 |
+
### Operational rule
|
| 29 |
+
1. Save / keep current .pt
|
| 30 |
+
2. Patch body (code / objective / routing)
|
| 31 |
+
3. Reload weights
|
| 32 |
+
4. Resume at recorded token offset
|
| 33 |
+
5. Document the surgery
|
| 34 |
+
|
| 35 |
+
**This is mid-training operability**, not a full retrain.
|
| 36 |
+
|
| 37 |
+
See also: docs/OPERABILITY_MIDTRAIN.md
|
| 38 |
+
|
| 39 |
+
---
|
| 40 |
+
|
| 41 |
+
## 2. Parallel brains then merge
|
| 42 |
+
|
| 43 |
+
**Claim:** You can train **many different .pt files** (shards / GPUs / runs) and **merge** them into one Fractus.
|
| 44 |
+
|
| 45 |
+
### Proven path (this run)
|
| 46 |
+
- 4 independent GPU runs on 4 data shards
|
| 47 |
+
- Each writes checkpoints/fractus_1b_gpu0.pt ... gpu3.pt
|
| 48 |
+
- **Mean-merge** of floating tensors -> unified checkpoint
|
| 49 |
+
- Example artifact: checkpoints/FRACTUS_1B_STAGE2_MERGED.pt
|
| 50 |
+
- Unified .pt still generates (different attractors than a single shard)
|
| 51 |
+
|
| 52 |
+
### Workflow
|
| 53 |
+
|
| 54 |
+
|
| 55 |
+
Constraint: **same architecture shapes** (d_model, layers, experts, etc.).
|
| 56 |
+
|
| 57 |
+
---
|
| 58 |
+
|
| 59 |
+
## 3. Compose an already-trained Fractus with more checkpoints
|
| 60 |
+
|
| 61 |
+
**Claim:** Take a Fractus already trained, merge it with other compatible .pt files, then keep training β the system **accumulates** rather than restarting from zero.
|
| 62 |
+
|
| 63 |
+
### Composition loop
|
| 64 |
+
|
| 65 |
+
|
| 66 |
+
### What grow means here
|
| 67 |
+
Two related notions:
|
| 68 |
+
|
| 69 |
+
1. **Compositional growth (proven in this run)**
|
| 70 |
+
- Knowledge / weights from multiple runs can be averaged into one brain
|
| 71 |
+
- Then training continues on the merged brain
|
| 72 |
+
- Digestion is not discarded
|
| 73 |
+
|
| 74 |
+
2. **Architectural growth (design of Fractus / paliers)**
|
| 75 |
+
- Progressive growth: width, depth, experts can increase across paliers
|
| 76 |
+
- maybe_grow adds capacity when routing demands it
|
| 77 |
+
- Documented in the Chinchilla/growth notes
|
| 78 |
+
|
| 79 |
+
Merging many same-shape .pt files is **composition**.
|
| 80 |
+
Growing parameter count is **structural growth**.
|
| 81 |
+
Both are part of the Fractus operating model; this production run proved composition + mid-train surgery under fire.
|
| 82 |
+
|
| 83 |
+
---
|
| 84 |
+
|
| 85 |
+
## 4. Why this matters
|
| 86 |
+
|
| 87 |
+
Standard large training runs treat the job as fragile:
|
| 88 |
+
- change the objective -> often restart
|
| 89 |
+
- multi-GPU shard brains -> hard to recombine
|
| 90 |
+
- decode bugs -> blamed on needs more train only
|
| 91 |
+
|
| 92 |
+
Fractus production forced a different picture:
|
| 93 |
+
|
| 94 |
+
| Property | Observation |
|
| 95 |
+
|----------|-------------|
|
| 96 |
+
| Operable mid-train | Yes β weights kept, body patched |
|
| 97 |
+
| Multi-pt parallel train | Yes β 4 shards |
|
| 98 |
+
| Merge to one Fractus | Yes β mean-merge works |
|
| 99 |
+
| Resume after surgery | Yes β exact token offsets |
|
| 100 |
+
| Train then merge then train | Yes β intended workflow |
|
| 101 |
+
|
| 102 |
+
---
|
| 103 |
+
|
| 104 |
+
## 5. Limits (honest)
|
| 105 |
+
|
| 106 |
+
- Merge requires **matching tensor shapes**
|
| 107 |
+
- Naive mean-merge of very divergent specialists can **soup** skills
|
| 108 |
+
- Composition is not magic AGI; it is a **workflow property** of separable brain (.pt) + body (code)
|
| 109 |
+
- Generation coherence still lags teacher-forced loss (see docs/LOSS_VS_GEN.md)
|
| 110 |
+
|
| 111 |
+
---
|
| 112 |
+
|
| 113 |
+
## 6. Related documents
|
| 114 |
+
|
| 115 |
+
- docs/OPERABILITY_MIDTRAIN.md β open-heart principle and stage2 surgery
|
| 116 |
+
- docs/DISCOVERY_LOG.md β bugs and emergent features list
|
| 117 |
+
- docs/MASTER_RUN_LOG.md β full chronology
|
| 118 |
+
- docs/2026-08-12-fractus-chinchilla.md β progressive growth / paliers
|
| 119 |
+
- docs/LOSS_VS_GEN.md β why TF loss is not gen quality
|
| 120 |
+
|
| 121 |
+
---
|
| 122 |
+
|
| 123 |
+
*Discovered by operating Fractus under real training pressure, not by design slide.*
|