thefinalboss commited on
Commit
42c2442
Β·
verified Β·
1 Parent(s): 5a517b0

Upload docs/CPU_MINI_MERGE_AND_DIMENSIONS.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. docs/CPU_MINI_MERGE_AND_DIMENSIONS.md +136 -0
docs/CPU_MINI_MERGE_AND_DIMENSIONS.md ADDED
@@ -0,0 +1,136 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Small Fractus on CPU, merges, and dimension rules
2
+
3
+ **Updated:** 2026-08-20 07:12 UTC
4
+ **Repo:** thefinalboss/fractus-cte
5
+
6
+ This note states **what is allowed** and **what is not magic** when training many small Fractus models on CPU and merging them β€” especially versus the frozen **Fractus-1B** phase-2 run.
7
+
8
+ ---
9
+
10
+ ## 1. Two tracks (do not confuse them)
11
+
12
+ | Track | Hardware | What it advances |
13
+ |-------|----------|------------------|
14
+ | **Fractus-1B phase2** | 8Γ— GPU (e.g. RTX 5090) | Token counter on the ~3.44B / ~4B corpus, 1B weights, production brain |
15
+ | **Mini-Fractus lab** | CPU (or 1 small GPU) | Ideas, routing experiments, decode, merge *practice*, diversity of seeds |
16
+
17
+ **CPU β‰  8Γ—5090.**
18
+ Training many small models on CPU **does not** move the 1B phase-2 progress (~4% of the pass at freeze). It does **not** replace the multi-day 1B digestion.
19
+
20
+ Both tracks are valid. Only one advances the **1B corpus counter**.
21
+
22
+ ---
23
+
24
+ ## 2. Mean-merge: when it is clean
25
+
26
+ Fractus checkpoints can be **mean-merged** when they share the **same tensor layout**.
27
+
28
+ ### Clean merge (supported pattern)
29
+
30
+ - Same architecture hyperparameters (see Β§3).
31
+ - Same parameter names and shapes in `state_dict` / `model_state`.
32
+ - Average floating-point tensors across N runs (optionally keep non-float buffers from one parent).
33
+ - Used in production for **8Γ— GPU** β†’ `FRACTUS_1B_*_MERGED.pt`.
34
+
35
+ ```text
36
+ merge(gpu0, gpu1, …, gpu7) β†’ one 1B brain βœ… same shapes
37
+ merge(mini_a, mini_b, mini_c) β†’ one mini brain βœ… if identical mini config
38
+ ```
39
+
40
+ ### Not automatic
41
+
42
+ ```text
43
+ merge(mini_64d, fractus_1b_1280d) β†’ ❌ shape mismatch
44
+ ```
45
+
46
+ A **badly sized** mini-Fractus **cannot** be mean-merged into the 1B. There is no free reshape that preserves a trained dynamical system.
47
+
48
+ ---
49
+
50
+ ## 3. Dimension contract (1B reference)
51
+
52
+ Production **Fractus-1B** (phase2 freeze / boost) uses approximately:
53
+
54
+ | Hyperparameter | 1B value |
55
+ |----------------|----------|
56
+ | `d_model` | 1280 |
57
+ | `n_layers` | 16 |
58
+ | `n_heads` | 20 |
59
+ | `d_head` | 64 |
60
+ | `n_oscillators` | 16 |
61
+ | `coupling_rank` | 8 |
62
+ | `n_experts` | 128 |
63
+ | `top_k` | 2 |
64
+ | `expert_d_ff` | 2048 |
65
+ | `siren_rank` | 64 |
66
+ | `vocab_size` | 50257 (GPT-2 BPE) |
67
+
68
+ **Rule:** any checkpoint you intend to mean-merge with **this** 1B must match these shapes **exactly**.
69
+
70
+ Mini configs for CPU lab **must document their own table**. Example lab tier (illustrative only β€” choose what fits RAM/CPU):
71
+
72
+ | Hyperparameter | Example mini | Mergeable into 1B? |
73
+ |----------------|--------------|--------------------|
74
+ | `d_model` | 256 or 512 | No |
75
+ | `n_layers` | 4–8 | No (unless equal 16) |
76
+ | `n_experts` | 32 | No (unless equal 128) |
77
+ | same as 1B table | full 1B | Yes, but not β€œsmall” on CPU |
78
+
79
+ **Same-arch mini swarm:** all minis share **one** fixed mini table β†’ merge **among themselves** βœ…
80
+ **Mini β†’ 1B:** not by mean-merge ❌
81
+
82
+ ---
83
+
84
+ ## 4. Mini β†’ 1B: what would be required (not implemented as magic)
85
+
86
+ If you ever want knowledge or structure from minis **into** the 1B, you need an explicit scheme, for example:
87
+
88
+ | Scheme | Idea | Status |
89
+ |--------|------|--------|
90
+ | **Mean-merge** | Average aligned tensors | Only identical shapes |
91
+ | **Interpolation / resize** | Pad or project layers | Research; can destroy dynamics |
92
+ | **Distillation** | 1B trains to match mini outputs on data | Possible; costs compute and a loss design |
93
+ | **Expert transplant** | Copy MoE expert slices if widths match | Only if expert ranks/dims match |
94
+ | **Vorax / pools** | Do **not** merge weights; write knowledge via organs | Preferred for *facts* without touching 1B shapes |
95
+
96
+ **Default recommendation while waiting for GPU pay cycle:**
97
+
98
+ 1. Keep **1B freeze** as the production brain (HF).
99
+ 2. On CPU, train **same-config minis**, merge **minis with minis**, probe routing / decode / Kuramoto fixes.
100
+ 3. Put durable knowledge into **Fractus-Vorax** (hash → pools), not into a fantasy mini→1B average.
101
+ 4. On next 8Γ—GPU run: resume 1B from `FROZEN_RESUME_MANIFEST` + Kuramoto prep (`docs/KURAMOTO_BOTTLENECK_AND_FIX.md`).
102
+
103
+ ---
104
+
105
+ ## 5. Practical CPU lab recipe
106
+
107
+ ```text
108
+ 1. Fix one MINI_CONFIG (write the table in the run README).
109
+ 2. Train N seeds or N data shards on CPU (short seq, small batch).
110
+ 3. Mean-merge the N minis β†’ mini_merged.pt
111
+ 4. Eval routing entropy, gen probes, Kuramoto r / dead experts on the mini.
112
+ 5. Document results; do not claim phase2 1B progress.
113
+ ```
114
+
115
+ Optional: apply the same **Kuramoto routing fix family** (wider Ο‰, gate temperature, LB in loss) on minis to validate the fix cheaply before the 1B resume β€” still not a substitute for applying it on the real 1B freeze at resume time.
116
+
117
+ ---
118
+
119
+ ## 6. One-page rules
120
+
121
+ 1. **CPU mini training β‰  1B phase2 progress.**
122
+ 2. **Mean-merge only for identical architecture shapes.**
123
+ 3. **Mini β†’ 1B mean-merge is invalid** without a separate, explicit transfer method.
124
+ 4. **Document every config table** next to every checkpoint family.
125
+ 5. **Production path** remains: freeze 1B β†’ Kuramoto prep β†’ resume phase2 on multi-GPU.
126
+ 6. **Knowledge path** without shape wars: Vorax ingestion / organs.
127
+
128
+ ---
129
+
130
+ ## Related docs
131
+
132
+ - `docs/HOW_FRACTUS_IS_TRAINED.md` β€” multi-GPU phase2 recipe
133
+ - `docs/KURAMOTO_BOTTLENECK_AND_FIX.md` β€” routing fix before next 1B resume
134
+ - `docs/NEXT_TRAINING_CHECKLIST.md` β€” pod checklist
135
+ - `docs/FROZEN_STATE.md` / `checkpoints/FROZEN_RESUME_MANIFEST.json` β€” resume offsets
136
+ - Fractus-Vorax β€” sealed 1B + write-side knowledge (no requirement to merge mini weights into 1B)