Download docs/CPU_MINI_MERGE_AND_DIMENSIONS.md from thefinalboss/fractus-cte: direct link, hf CLI and curl.
- Browser
- Download file 5.53 kB
-
https://huggingface.co/thefinalboss/fractus-cte/resolve/8e46f2cf53e63440225486f98e83043aec096172/docs/CPU_MINI_MERGE_AND_DIMENSIONS.md
- Command line
-
hf download hf://thefinalboss/fractus-cte@8e46f2cf53e63440225486f98e83043aec096172/docs/CPU_MINI_MERGE_AND_DIMENSIONS.md
-
curl -L -o CPU_MINI_MERGE_AND_DIMENSIONS.md https://huggingface.co/thefinalboss/fractus-cte/resolve/8e46f2cf53e63440225486f98e83043aec096172/docs/CPU_MINI_MERGE_AND_DIMENSIONS.md
Small Fractus on CPU, merges, and dimension rules
Updated: 2026-08-20 07:12 UTC
Repo: thefinalboss/fractus-cte
This note states what is allowed and what is not magic when training many small Fractus models on CPU and merging them β especially versus the frozen Fractus-1B phase-2 run.
1. Two tracks (do not confuse them)
| Track | Hardware | What it advances |
|---|---|---|
| Fractus-1B phase2 | 8Γ GPU (e.g. RTX 5090) | Token counter on the ~3.44B / ~4B corpus, 1B weights, production brain |
| Mini-Fractus lab | CPU (or 1 small GPU) | Ideas, routing experiments, decode, merge practice, diversity of seeds |
CPU β 8Γ5090.
Training many small models on CPU does not move the 1B phase-2 progress (~4% of the pass at freeze). It does not replace the multi-day 1B digestion.
Both tracks are valid. Only one advances the 1B corpus counter.
2. Mean-merge: when it is clean
Fractus checkpoints can be mean-merged when they share the same tensor layout.
Clean merge (supported pattern)
- Same architecture hyperparameters (see Β§3).
- Same parameter names and shapes in
state_dict/model_state. - Average floating-point tensors across N runs (optionally keep non-float buffers from one parent).
- Used in production for 8Γ GPU β
FRACTUS_1B_*_MERGED.pt.
merge(gpu0, gpu1, β¦, gpu7) β one 1B brain β
same shapes
merge(mini_a, mini_b, mini_c) β one mini brain β
if identical mini config
Not automatic
merge(mini_64d, fractus_1b_1280d) β β shape mismatch
A badly sized mini-Fractus cannot be mean-merged into the 1B. There is no free reshape that preserves a trained dynamical system.
3. Dimension contract (1B reference)
Production Fractus-1B (phase2 freeze / boost) uses approximately:
| Hyperparameter | 1B value |
|---|---|
d_model |
1280 |
n_layers |
16 |
n_heads |
20 |
d_head |
64 |
n_oscillators |
16 |
coupling_rank |
8 |
n_experts |
128 |
top_k |
2 |
expert_d_ff |
2048 |
siren_rank |
64 |
vocab_size |
50257 (GPT-2 BPE) |
Rule: any checkpoint you intend to mean-merge with this 1B must match these shapes exactly.
Mini configs for CPU lab must document their own table. Example lab tier (illustrative only β choose what fits RAM/CPU):
| Hyperparameter | Example mini | Mergeable into 1B? |
|---|---|---|
d_model |
256 or 512 | No |
n_layers |
4β8 | No (unless equal 16) |
n_experts |
32 | No (unless equal 128) |
| same as 1B table | full 1B | Yes, but not βsmallβ on CPU |
Same-arch mini swarm: all minis share one fixed mini table β merge among themselves β
Mini β 1B: not by mean-merge β
4. Mini β 1B: what would be required (not implemented as magic)
If you ever want knowledge or structure from minis into the 1B, you need an explicit scheme, for example:
| Scheme | Idea | Status |
|---|---|---|
| Mean-merge | Average aligned tensors | Only identical shapes |
| Interpolation / resize | Pad or project layers | Research; can destroy dynamics |
| Distillation | 1B trains to match mini outputs on data | Possible; costs compute and a loss design |
| Expert transplant | Copy MoE expert slices if widths match | Only if expert ranks/dims match |
| Vorax / pools | Do not merge weights; write knowledge via organs | Preferred for facts without touching 1B shapes |
Default recommendation while waiting for GPU pay cycle:
- Keep 1B freeze as the production brain (HF).
- On CPU, train same-config minis, merge minis with minis, probe routing / decode / Kuramoto fixes.
- Put durable knowledge into Fractus-Vorax (hash β pools), not into a fantasy miniβ1B average.
- On next 8ΓGPU run: resume 1B from
FROZEN_RESUME_MANIFEST+ Kuramoto prep (docs/KURAMOTO_BOTTLENECK_AND_FIX.md).
5. Practical CPU lab recipe
1. Fix one MINI_CONFIG (write the table in the run README).
2. Train N seeds or N data shards on CPU (short seq, small batch).
3. Mean-merge the N minis β mini_merged.pt
4. Eval routing entropy, gen probes, Kuramoto r / dead experts on the mini.
5. Document results; do not claim phase2 1B progress.
Optional: apply the same Kuramoto routing fix family (wider Ο, gate temperature, LB in loss) on minis to validate the fix cheaply before the 1B resume β still not a substitute for applying it on the real 1B freeze at resume time.
6. One-page rules
- CPU mini training β 1B phase2 progress.
- Mean-merge only for identical architecture shapes.
- Mini β 1B mean-merge is invalid without a separate, explicit transfer method.
- Document every config table next to every checkpoint family.
- Production path remains: freeze 1B β Kuramoto prep β resume phase2 on multi-GPU.
- Knowledge path without shape wars: Vorax ingestion / organs.
Related docs
docs/HOW_FRACTUS_IS_TRAINED.mdβ multi-GPU phase2 recipedocs/KURAMOTO_BOTTLENECK_AND_FIX.mdβ routing fix before next 1B resumedocs/NEXT_TRAINING_CHECKLIST.mdβ pod checklistdocs/FROZEN_STATE.md/checkpoints/FROZEN_RESUME_MANIFEST.jsonβ resume offsets- Fractus-Vorax β sealed 1B + write-side knowledge (no requirement to merge mini weights into 1B)