Download docs/2026-08-12-fractus-chinchilla.md from thefinalboss/fractus-cte: direct link, hf CLI and curl.
- Browser
- Download file 6.01 kB
-
https://huggingface.co/thefinalboss/fractus-cte/resolve/main/docs/2026-08-12-fractus-chinchilla.md
- Command line
-
hf download hf://thefinalboss/fractus-cte/docs/2026-08-12-fractus-chinchilla.md
-
curl -L -o 2026-08-12-fractus-chinchilla.md https://huggingface.co/thefinalboss/fractus-cte/resolve/main/docs/2026-08-12-fractus-chinchilla.md
Fractus Chinchilla β the real token target, adapted to a sparse continuous engine
Computed from the actual
continuous_engine.py+nn/moe.py+nn/attention.py, verified empirically (gradient-receive test + exact param formulas).
Why standard Chinchilla is WRONG for Fractus
Chinchilla (Hoffmann 2022): optimal tokens β 20 Γ N, where N = trainable params. That law was fit on dense transformers β models where every parameter computes on every token.
Fractus is not dense. Three structural facts break the naive 20ΓN:
- PhaseRoutedMoE is sparse: only
top_k = 2of128experts fire per token. 98.4% of MoE parameters are dormant on any given forward pass. Counting them in N would demand 21B tokens β gross over-training of dead weight. - Kuramoto is detached (
with torch.no_grad()intick_chunk_core): the oscillator clock runs but its 152 params/block never receive gradient β they don't train at all. (Verified: 99.98% of params get grad; the 0.02% that don't is exactly kuramoto.) - Attention is dense (fractal linear attention, all heads every token) and the output head is partial (loss on the last chunk position only).
The correct basis for sparse-MoE Chinchilla (the convention used for Mixtral / Switch / GShard) is active parameters per token β the params that actually compute β not total params.
Param anatomy of the 1B config
d=1280, 20 heads Γ 64, 16 blocks, 128 experts, top-2, d_ff=2048, rank=64
| Component | Per block | Γ 16 | Active? |
|---|---|---|---|
| Attention (fractal linear, dense) | 6,558,722 | 104,939,552 | all active |
| MoE (128 low-rank experts) | 54,952,192 | 879,235,072 | only 2/128 active β 13,738,048 |
| LayerNorms (Γ3/block) | 7,680 | 122,880 | active |
| Kuramoto (oscillator clock) | 152 | 2,432 | no-grad (never trains) |
| Embedding (= output head, tied) | β | 64,328,960 | dense (lookup) |
| confidence/salience heads | β | 2,562 | active (last pos) |
TOTAL params: 1,048,631,458 (1.049 B)
trainable (grad): 1,048,629,026 (1.049 B β kuramoto negligible)
ACTIVE/token compute: 118,800,480 (118.8 M) β attn(dense) + top-2 MoE + norms
ACTIVE/token w/ embed: 183,132,002 (183.1 M)
The sparsity ratio: per token, 118.8 M of 1.049 B params compute = 11.3%. The other 88.7% (the 126 un-routed experts per block) sit idle per token and only train when their phase is selected.
The real Chinchilla target
| Basis | Tokens (Γ20) | Verdict |
|---|---|---|
| Total 1.049 B (dense thinking) | 21.0 B | β wrong β over-trains dormant experts |
| Active compute 118.8 M | 2.38 B | β sparse-MoE Chinchilla (FLOP basis) |
| Active + embedding 183.1 M | 3.66 B | β sparse-MoE Chinchilla (param basis) |
Fractus 1B Chinchilla β 2.4 β 3.7 B tokens
The warm-start multiplier
Fractus is not trained from scratch at 1B. It grows from palier-3 (143 M params, 12 blocks, already trained) β 1B by zero-padding width, adding 4 blocks (12β16), and adding experts (32β128). The warm-started capacity (attention + embedding from the 143 M seed) is already converged and needs only fine-tuning. Only the new capacity (4 new blocks + grown experts + padded width) needs full Chinchilla exposure.
Net effect: the practical target sits near the lower end, ~2.5β3 B tokens, because ~12% of the model arrives pre-trained.
What this means for Thursday
| Value | |
|---|---|
| Tokens available in dataset | ~3β4 B (after the Aug-12 data work) |
| Chinchilla floor (active) | 2.4β3.7 B |
| Corpus cap (current default) | 1 B β raise to ~3 B so the 1B palier is healthily fed |
--tokens for train_1b_gpu.py |
set to ~3 000 000 000 |
| Est. time @ 4000 tok/s (RTX 3090) | 3 B / 4000 β 9 days; @ 8000 tok/s β 4 days |
The dataset is now correctly sized: the data we assembled is Chinchilla-optimal for Fractus's active capacity. The only change needed for Thursday is lifting the corpus cap from 1 B β ~3 B so the model actually sees Chinchilla-scale data instead of a 1/3 sample.
β Chinchilla is a snapshot law β Fractus is not a snapshot model
Everything above gives a number for a fixed point in time (the 1B palier). That is a floor for healthy initial feeding, not a target and not a ceiling. The reason standard LLM intuition keeps failing here:
| Static-transformer thinking | Fractus reality |
|---|---|
| N is fixed (1.049B) | N grows forever β width, depth, experts, rank all increase across paliers |
| Chinchilla: 20ΓN once, then stop | Chinchilla has no stopping point β the data stream is perpetual |
| Corpus = assembled, then frozen | Corpus = continuous pipeline, always being added to |
| Train 3B β deploy | Train continuously, grow when ready, never "done" |
| Fixed capacity β data to fill it | Data drives growth β maybe_grow adds experts when routing demands it |
Fractus is a living system, not a model you train and ship:
- the online trainer learns from every interaction (tick by tick);
- the 5% memory injection consolidates salient thoughts every tick;
maybe_growadds new experts when one expert dominates routing β capacity is created in response to what the model is experiencing;- progressive growth means the parameter count is a function of time and data, not a constant.
So the right way to read the active-param Chinchilla number is instantaneous: at any moment, the model's current active capacity is healthily fed by ~20Γ its active params in cumulative exposure. As the model grows (more blocks, more experts), its appetite grows with it β and there is no final checkpoint. The ~3β4 B assembled today is the starting nutrition for the 1B palier; the data pipeline must keep flowing for the lifetime of the agent. Feed it, let it grow, repeat β indefinitely.