File size: 5,527 Bytes
42c2442
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
# Small Fractus on CPU, merges, and dimension rules

**Updated:** 2026-08-20 07:12 UTC  
**Repo:** thefinalboss/fractus-cte

This note states **what is allowed** and **what is not magic** when training many small Fractus models on CPU and merging them β€” especially versus the frozen **Fractus-1B** phase-2 run.

---

## 1. Two tracks (do not confuse them)

| Track | Hardware | What it advances |
|-------|----------|------------------|
| **Fractus-1B phase2** | 8Γ— GPU (e.g. RTX 5090) | Token counter on the ~3.44B / ~4B corpus, 1B weights, production brain |
| **Mini-Fractus lab** | CPU (or 1 small GPU) | Ideas, routing experiments, decode, merge *practice*, diversity of seeds |

**CPU β‰  8Γ—5090.**  
Training many small models on CPU **does not** move the 1B phase-2 progress (~4% of the pass at freeze). It does **not** replace the multi-day 1B digestion.

Both tracks are valid. Only one advances the **1B corpus counter**.

---

## 2. Mean-merge: when it is clean

Fractus checkpoints can be **mean-merged** when they share the **same tensor layout**.

### Clean merge (supported pattern)

- Same architecture hyperparameters (see Β§3).
- Same parameter names and shapes in `state_dict` / `model_state`.
- Average floating-point tensors across N runs (optionally keep non-float buffers from one parent).
- Used in production for **8Γ— GPU** β†’ `FRACTUS_1B_*_MERGED.pt`.

```text
merge(gpu0, gpu1, …, gpu7)  β†’  one 1B brain    βœ… same shapes
merge(mini_a, mini_b, mini_c) β†’ one mini brain βœ… if identical mini config
```

### Not automatic

```text
merge(mini_64d, fractus_1b_1280d)  β†’  ❌ shape mismatch
```

A **badly sized** mini-Fractus **cannot** be mean-merged into the 1B. There is no free reshape that preserves a trained dynamical system.

---

## 3. Dimension contract (1B reference)

Production **Fractus-1B** (phase2 freeze / boost) uses approximately:

| Hyperparameter | 1B value |
|----------------|----------|
| `d_model` | 1280 |
| `n_layers` | 16 |
| `n_heads` | 20 |
| `d_head` | 64 |
| `n_oscillators` | 16 |
| `coupling_rank` | 8 |
| `n_experts` | 128 |
| `top_k` | 2 |
| `expert_d_ff` | 2048 |
| `siren_rank` | 64 |
| `vocab_size` | 50257 (GPT-2 BPE) |

**Rule:** any checkpoint you intend to mean-merge with **this** 1B must match these shapes **exactly**.

Mini configs for CPU lab **must document their own table**. Example lab tier (illustrative only β€” choose what fits RAM/CPU):

| Hyperparameter | Example mini | Mergeable into 1B? |
|----------------|--------------|--------------------|
| `d_model` | 256 or 512 | No |
| `n_layers` | 4–8 | No (unless equal 16) |
| `n_experts` | 32 | No (unless equal 128) |
| same as 1B table | full 1B | Yes, but not β€œsmall” on CPU |

**Same-arch mini swarm:** all minis share **one** fixed mini table β†’ merge **among themselves** βœ…  
**Mini β†’ 1B:** not by mean-merge ❌

---

## 4. Mini β†’ 1B: what would be required (not implemented as magic)

If you ever want knowledge or structure from minis **into** the 1B, you need an explicit scheme, for example:

| Scheme | Idea | Status |
|--------|------|--------|
| **Mean-merge** | Average aligned tensors | Only identical shapes |
| **Interpolation / resize** | Pad or project layers | Research; can destroy dynamics |
| **Distillation** | 1B trains to match mini outputs on data | Possible; costs compute and a loss design |
| **Expert transplant** | Copy MoE expert slices if widths match | Only if expert ranks/dims match |
| **Vorax / pools** | Do **not** merge weights; write knowledge via organs | Preferred for *facts* without touching 1B shapes |

**Default recommendation while waiting for GPU pay cycle:**

1. Keep **1B freeze** as the production brain (HF).  
2. On CPU, train **same-config minis**, merge **minis with minis**, probe routing / decode / Kuramoto fixes.  
3. Put durable knowledge into **Fractus-Vorax** (hash → pools), not into a fantasy mini→1B average.  
4. On next 8Γ—GPU run: resume 1B from `FROZEN_RESUME_MANIFEST` + Kuramoto prep (`docs/KURAMOTO_BOTTLENECK_AND_FIX.md`).

---

## 5. Practical CPU lab recipe

```text
1. Fix one MINI_CONFIG (write the table in the run README).
2. Train N seeds or N data shards on CPU (short seq, small batch).
3. Mean-merge the N minis β†’ mini_merged.pt
4. Eval routing entropy, gen probes, Kuramoto r / dead experts on the mini.
5. Document results; do not claim phase2 1B progress.
```

Optional: apply the same **Kuramoto routing fix family** (wider Ο‰, gate temperature, LB in loss) on minis to validate the fix cheaply before the 1B resume β€” still not a substitute for applying it on the real 1B freeze at resume time.

---

## 6. One-page rules

1. **CPU mini training β‰  1B phase2 progress.**  
2. **Mean-merge only for identical architecture shapes.**  
3. **Mini β†’ 1B mean-merge is invalid** without a separate, explicit transfer method.  
4. **Document every config table** next to every checkpoint family.  
5. **Production path** remains: freeze 1B β†’ Kuramoto prep β†’ resume phase2 on multi-GPU.  
6. **Knowledge path** without shape wars: Vorax ingestion / organs.

---

## Related docs

- `docs/HOW_FRACTUS_IS_TRAINED.md` β€” multi-GPU phase2 recipe  
- `docs/KURAMOTO_BOTTLENECK_AND_FIX.md` β€” routing fix before next 1B resume  
- `docs/NEXT_TRAINING_CHECKLIST.md` β€” pod checklist  
- `docs/FROZEN_STATE.md` / `checkpoints/FROZEN_RESUME_MANIFEST.json` β€” resume offsets  
- Fractus-Vorax β€” sealed 1B + write-side knowledge (no requirement to merge mini weights into 1B)