Upload docs/CPU_MINI_MERGE_AND_DIMENSIONS.md with huggingface_hub
Browse files
docs/CPU_MINI_MERGE_AND_DIMENSIONS.md
ADDED
|
@@ -0,0 +1,136 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Small Fractus on CPU, merges, and dimension rules
|
| 2 |
+
|
| 3 |
+
**Updated:** 2026-08-20 07:12 UTC
|
| 4 |
+
**Repo:** thefinalboss/fractus-cte
|
| 5 |
+
|
| 6 |
+
This note states **what is allowed** and **what is not magic** when training many small Fractus models on CPU and merging them β especially versus the frozen **Fractus-1B** phase-2 run.
|
| 7 |
+
|
| 8 |
+
---
|
| 9 |
+
|
| 10 |
+
## 1. Two tracks (do not confuse them)
|
| 11 |
+
|
| 12 |
+
| Track | Hardware | What it advances |
|
| 13 |
+
|-------|----------|------------------|
|
| 14 |
+
| **Fractus-1B phase2** | 8Γ GPU (e.g. RTX 5090) | Token counter on the ~3.44B / ~4B corpus, 1B weights, production brain |
|
| 15 |
+
| **Mini-Fractus lab** | CPU (or 1 small GPU) | Ideas, routing experiments, decode, merge *practice*, diversity of seeds |
|
| 16 |
+
|
| 17 |
+
**CPU β 8Γ5090.**
|
| 18 |
+
Training many small models on CPU **does not** move the 1B phase-2 progress (~4% of the pass at freeze). It does **not** replace the multi-day 1B digestion.
|
| 19 |
+
|
| 20 |
+
Both tracks are valid. Only one advances the **1B corpus counter**.
|
| 21 |
+
|
| 22 |
+
---
|
| 23 |
+
|
| 24 |
+
## 2. Mean-merge: when it is clean
|
| 25 |
+
|
| 26 |
+
Fractus checkpoints can be **mean-merged** when they share the **same tensor layout**.
|
| 27 |
+
|
| 28 |
+
### Clean merge (supported pattern)
|
| 29 |
+
|
| 30 |
+
- Same architecture hyperparameters (see Β§3).
|
| 31 |
+
- Same parameter names and shapes in `state_dict` / `model_state`.
|
| 32 |
+
- Average floating-point tensors across N runs (optionally keep non-float buffers from one parent).
|
| 33 |
+
- Used in production for **8Γ GPU** β `FRACTUS_1B_*_MERGED.pt`.
|
| 34 |
+
|
| 35 |
+
```text
|
| 36 |
+
merge(gpu0, gpu1, β¦, gpu7) β one 1B brain β
same shapes
|
| 37 |
+
merge(mini_a, mini_b, mini_c) β one mini brain β
if identical mini config
|
| 38 |
+
```
|
| 39 |
+
|
| 40 |
+
### Not automatic
|
| 41 |
+
|
| 42 |
+
```text
|
| 43 |
+
merge(mini_64d, fractus_1b_1280d) β β shape mismatch
|
| 44 |
+
```
|
| 45 |
+
|
| 46 |
+
A **badly sized** mini-Fractus **cannot** be mean-merged into the 1B. There is no free reshape that preserves a trained dynamical system.
|
| 47 |
+
|
| 48 |
+
---
|
| 49 |
+
|
| 50 |
+
## 3. Dimension contract (1B reference)
|
| 51 |
+
|
| 52 |
+
Production **Fractus-1B** (phase2 freeze / boost) uses approximately:
|
| 53 |
+
|
| 54 |
+
| Hyperparameter | 1B value |
|
| 55 |
+
|----------------|----------|
|
| 56 |
+
| `d_model` | 1280 |
|
| 57 |
+
| `n_layers` | 16 |
|
| 58 |
+
| `n_heads` | 20 |
|
| 59 |
+
| `d_head` | 64 |
|
| 60 |
+
| `n_oscillators` | 16 |
|
| 61 |
+
| `coupling_rank` | 8 |
|
| 62 |
+
| `n_experts` | 128 |
|
| 63 |
+
| `top_k` | 2 |
|
| 64 |
+
| `expert_d_ff` | 2048 |
|
| 65 |
+
| `siren_rank` | 64 |
|
| 66 |
+
| `vocab_size` | 50257 (GPT-2 BPE) |
|
| 67 |
+
|
| 68 |
+
**Rule:** any checkpoint you intend to mean-merge with **this** 1B must match these shapes **exactly**.
|
| 69 |
+
|
| 70 |
+
Mini configs for CPU lab **must document their own table**. Example lab tier (illustrative only β choose what fits RAM/CPU):
|
| 71 |
+
|
| 72 |
+
| Hyperparameter | Example mini | Mergeable into 1B? |
|
| 73 |
+
|----------------|--------------|--------------------|
|
| 74 |
+
| `d_model` | 256 or 512 | No |
|
| 75 |
+
| `n_layers` | 4β8 | No (unless equal 16) |
|
| 76 |
+
| `n_experts` | 32 | No (unless equal 128) |
|
| 77 |
+
| same as 1B table | full 1B | Yes, but not βsmallβ on CPU |
|
| 78 |
+
|
| 79 |
+
**Same-arch mini swarm:** all minis share **one** fixed mini table β merge **among themselves** β
|
| 80 |
+
**Mini β 1B:** not by mean-merge β
|
| 81 |
+
|
| 82 |
+
---
|
| 83 |
+
|
| 84 |
+
## 4. Mini β 1B: what would be required (not implemented as magic)
|
| 85 |
+
|
| 86 |
+
If you ever want knowledge or structure from minis **into** the 1B, you need an explicit scheme, for example:
|
| 87 |
+
|
| 88 |
+
| Scheme | Idea | Status |
|
| 89 |
+
|--------|------|--------|
|
| 90 |
+
| **Mean-merge** | Average aligned tensors | Only identical shapes |
|
| 91 |
+
| **Interpolation / resize** | Pad or project layers | Research; can destroy dynamics |
|
| 92 |
+
| **Distillation** | 1B trains to match mini outputs on data | Possible; costs compute and a loss design |
|
| 93 |
+
| **Expert transplant** | Copy MoE expert slices if widths match | Only if expert ranks/dims match |
|
| 94 |
+
| **Vorax / pools** | Do **not** merge weights; write knowledge via organs | Preferred for *facts* without touching 1B shapes |
|
| 95 |
+
|
| 96 |
+
**Default recommendation while waiting for GPU pay cycle:**
|
| 97 |
+
|
| 98 |
+
1. Keep **1B freeze** as the production brain (HF).
|
| 99 |
+
2. On CPU, train **same-config minis**, merge **minis with minis**, probe routing / decode / Kuramoto fixes.
|
| 100 |
+
3. Put durable knowledge into **Fractus-Vorax** (hash β pools), not into a fantasy miniβ1B average.
|
| 101 |
+
4. On next 8ΓGPU run: resume 1B from `FROZEN_RESUME_MANIFEST` + Kuramoto prep (`docs/KURAMOTO_BOTTLENECK_AND_FIX.md`).
|
| 102 |
+
|
| 103 |
+
---
|
| 104 |
+
|
| 105 |
+
## 5. Practical CPU lab recipe
|
| 106 |
+
|
| 107 |
+
```text
|
| 108 |
+
1. Fix one MINI_CONFIG (write the table in the run README).
|
| 109 |
+
2. Train N seeds or N data shards on CPU (short seq, small batch).
|
| 110 |
+
3. Mean-merge the N minis β mini_merged.pt
|
| 111 |
+
4. Eval routing entropy, gen probes, Kuramoto r / dead experts on the mini.
|
| 112 |
+
5. Document results; do not claim phase2 1B progress.
|
| 113 |
+
```
|
| 114 |
+
|
| 115 |
+
Optional: apply the same **Kuramoto routing fix family** (wider Ο, gate temperature, LB in loss) on minis to validate the fix cheaply before the 1B resume β still not a substitute for applying it on the real 1B freeze at resume time.
|
| 116 |
+
|
| 117 |
+
---
|
| 118 |
+
|
| 119 |
+
## 6. One-page rules
|
| 120 |
+
|
| 121 |
+
1. **CPU mini training β 1B phase2 progress.**
|
| 122 |
+
2. **Mean-merge only for identical architecture shapes.**
|
| 123 |
+
3. **Mini β 1B mean-merge is invalid** without a separate, explicit transfer method.
|
| 124 |
+
4. **Document every config table** next to every checkpoint family.
|
| 125 |
+
5. **Production path** remains: freeze 1B β Kuramoto prep β resume phase2 on multi-GPU.
|
| 126 |
+
6. **Knowledge path** without shape wars: Vorax ingestion / organs.
|
| 127 |
+
|
| 128 |
+
---
|
| 129 |
+
|
| 130 |
+
## Related docs
|
| 131 |
+
|
| 132 |
+
- `docs/HOW_FRACTUS_IS_TRAINED.md` β multi-GPU phase2 recipe
|
| 133 |
+
- `docs/KURAMOTO_BOTTLENECK_AND_FIX.md` β routing fix before next 1B resume
|
| 134 |
+
- `docs/NEXT_TRAINING_CHECKLIST.md` β pod checklist
|
| 135 |
+
- `docs/FROZEN_STATE.md` / `checkpoints/FROZEN_RESUME_MANIFEST.json` β resume offsets
|
| 136 |
+
- Fractus-Vorax β sealed 1B + write-side knowledge (no requirement to merge mini weights into 1B)
|