thefinalboss commited on
Commit
badc05a
·
verified ·
1 Parent(s): f7467d9

Upload docs/OPERABILITY_MIDTRAIN.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. docs/OPERABILITY_MIDTRAIN.md +74 -0
docs/OPERABILITY_MIDTRAIN.md ADDED
@@ -0,0 +1,74 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Fractus Operability — Mid-Training Open-Heart Surgery
2
+
3
+ **Date:** 2026-08-16 19:03 UTC
4
+
5
+ ## Claim
6
+
7
+ Fractus can be **operated on during training** without discarding learned weights.
8
+ The brain (.pt) stays; the training objective and dynamical routing can be rewired; digestion continues from the same token offset.
9
+
10
+ This is not a full retrain. It is live surgery.
11
+
12
+ ## Surgeries performed in this run
13
+
14
+ ### Surgery A — Routing (stage 1 fix)
15
+
16
+ | Bug | Fix | Weights kept? |
17
+ |-----|-----|---------------|
18
+ | Kuramoto under torch.no_grad in tick_chunk_core | Gradients enabled on phase clock | Yes |
19
+ | lb_loss computed then detached / not in loss | CE + 0.02 * lb | Yes |
20
+ | Hard von Mises gates (temp=1) | gate temperature 2.5 | Yes |
21
+ | Near-zero omega diversity | controlled omega expand if std < 0.08 | Yes |
22
+
23
+ ### Surgery B — Sequence stage (stage 2)
24
+
25
+ | Before | After |
26
+ |--------|-------|
27
+ | CE on **last position only** (1 target / 128 tokens) | **Dense CE** on all positions (B*C targets) |
28
+ | Favours local lexical attractors | Forces token-to-token chaining |
29
+
30
+ Code change in ContinuousThoughtEngine.tick_chunk_train:
31
+
32
+ # before: last_logits = output_head(h[:, -1, :])
33
+ # after: logits = output_head(h) # (B, C, vocab)
34
+
35
+ Trainer: fast4gpu_stage2.py — resumes exact start_token from logs, keeps LB + soft gates.
36
+
37
+ ## Why this is possible
38
+
39
+ 1. **Separation of brain and body** — weights are tensors; objective and routing are code.
40
+ 2. **Compatible state_dict** — same shapes across surgeries; load strict=False only for batch-shaped buffers.
41
+ 3. **Manifest / log resume** — start_token preserved so shards are not rewound to zero.
42
+ 4. **Composable checkpoints** — 4 GPU shards can still mean-merge after surgery.
43
+
44
+ ## Live metrics after stage-2 surgery
45
+
46
+ | GPU | Tokens | CE (dense) | lb |
47
+ |-----|--------|------------|-----|
48
+ | 0 | 177,932,800 | 17.9 | 13.253 |
49
+ | 1 | 187,750,400 | 4.2 | 14.027 |
50
+ | 2 | 188,556,800 | 6.7 | 14.019 |
51
+ | 3 | 185,536,000 | 17.6 | 14.027 |
52
+
53
+ Loss at stage-2 start can differ in scale from last-position-only CE (different objective).
54
+ What matters: descent continues, lb stays active (~14), GPUs stay full, no weight wipe.
55
+
56
+ ## Relation to generation collapse
57
+
58
+ Mid-training generation showed word-level repetition (Colorado, Fate, Population loops).
59
+ Diagnosis: stage-1 lexical attractors; routing locks to same expert each tick.
60
+ Stage-2 dense CE is the mid-training intervention aimed at sequence chaining without resetting digestion.
61
+
62
+ ## Operational rule
63
+
64
+ If a dynamical bottleneck is found (dead experts, frozen clock, sparse target):
65
+
66
+ 1. Save / keep current .pt
67
+ 2. Patch body (engine / trainer)
68
+ 3. Reload weights
69
+ 4. Resume at recorded token offset
70
+ 5. Document the surgery
71
+
72
+ Do not throw away multi-day digestion for a routing or objective bug.
73
+
74
+ *Fractus is operable. This log is proof from production, not a design slide.*