fractus-cte / docs /2026-08-12-fractus-chinchilla.md
thefinalboss's picture
Upload folder using huggingface_hub
6c223ff verified
|
Raw History Blame Contribute Delete
6.01 kB

Fractus Chinchilla β€” the real token target, adapted to a sparse continuous engine

Computed from the actual continuous_engine.py + nn/moe.py + nn/attention.py, verified empirically (gradient-receive test + exact param formulas).

Why standard Chinchilla is WRONG for Fractus

Chinchilla (Hoffmann 2022): optimal tokens β‰ˆ 20 Γ— N, where N = trainable params. That law was fit on dense transformers β€” models where every parameter computes on every token.

Fractus is not dense. Three structural facts break the naive 20Γ—N:

  1. PhaseRoutedMoE is sparse: only top_k = 2 of 128 experts fire per token. 98.4% of MoE parameters are dormant on any given forward pass. Counting them in N would demand 21B tokens β€” gross over-training of dead weight.
  2. Kuramoto is detached (with torch.no_grad() in tick_chunk_core): the oscillator clock runs but its 152 params/block never receive gradient β€” they don't train at all. (Verified: 99.98% of params get grad; the 0.02% that don't is exactly kuramoto.)
  3. Attention is dense (fractal linear attention, all heads every token) and the output head is partial (loss on the last chunk position only).

The correct basis for sparse-MoE Chinchilla (the convention used for Mixtral / Switch / GShard) is active parameters per token β€” the params that actually compute β€” not total params.

Param anatomy of the 1B config

d=1280, 20 heads Γ— 64, 16 blocks, 128 experts, top-2, d_ff=2048, rank=64

Component Per block Γ— 16 Active?
Attention (fractal linear, dense) 6,558,722 104,939,552 all active
MoE (128 low-rank experts) 54,952,192 879,235,072 only 2/128 active β†’ 13,738,048
LayerNorms (Γ—3/block) 7,680 122,880 active
Kuramoto (oscillator clock) 152 2,432 no-grad (never trains)
Embedding (= output head, tied) β€” 64,328,960 dense (lookup)
confidence/salience heads β€” 2,562 active (last pos)
TOTAL params:          1,048,631,458   (1.049 B)
trainable (grad):      1,048,629,026   (1.049 B β€” kuramoto negligible)
ACTIVE/token compute:    118,800,480   (118.8 M)   ← attn(dense) + top-2 MoE + norms
ACTIVE/token w/ embed:   183,132,002   (183.1 M)

The sparsity ratio: per token, 118.8 M of 1.049 B params compute = 11.3%. The other 88.7% (the 126 un-routed experts per block) sit idle per token and only train when their phase is selected.

The real Chinchilla target

Basis Tokens (Γ—20) Verdict
Total 1.049 B (dense thinking) 21.0 B ❌ wrong β€” over-trains dormant experts
Active compute 118.8 M 2.38 B βœ… sparse-MoE Chinchilla (FLOP basis)
Active + embedding 183.1 M 3.66 B βœ… sparse-MoE Chinchilla (param basis)

Fractus 1B Chinchilla β‰ˆ 2.4 – 3.7 B tokens

The warm-start multiplier

Fractus is not trained from scratch at 1B. It grows from palier-3 (143 M params, 12 blocks, already trained) β†’ 1B by zero-padding width, adding 4 blocks (12β†’16), and adding experts (32β†’128). The warm-started capacity (attention + embedding from the 143 M seed) is already converged and needs only fine-tuning. Only the new capacity (4 new blocks + grown experts + padded width) needs full Chinchilla exposure.

Net effect: the practical target sits near the lower end, ~2.5–3 B tokens, because ~12% of the model arrives pre-trained.

What this means for Thursday

Value
Tokens available in dataset ~3–4 B (after the Aug-12 data work)
Chinchilla floor (active) 2.4–3.7 B
Corpus cap (current default) 1 B ← raise to ~3 B so the 1B palier is healthily fed
--tokens for train_1b_gpu.py set to ~3 000 000 000
Est. time @ 4000 tok/s (RTX 3090) 3 B / 4000 β‰ˆ 9 days; @ 8000 tok/s β‰ˆ 4 days

The dataset is now correctly sized: the data we assembled is Chinchilla-optimal for Fractus's active capacity. The only change needed for Thursday is lifting the corpus cap from 1 B β†’ ~3 B so the model actually sees Chinchilla-scale data instead of a 1/3 sample.


⚠ Chinchilla is a snapshot law β€” Fractus is not a snapshot model

Everything above gives a number for a fixed point in time (the 1B palier). That is a floor for healthy initial feeding, not a target and not a ceiling. The reason standard LLM intuition keeps failing here:

Static-transformer thinking Fractus reality
N is fixed (1.049B) N grows forever β€” width, depth, experts, rank all increase across paliers
Chinchilla: 20Γ—N once, then stop Chinchilla has no stopping point β€” the data stream is perpetual
Corpus = assembled, then frozen Corpus = continuous pipeline, always being added to
Train 3B β†’ deploy Train continuously, grow when ready, never "done"
Fixed capacity β†’ data to fill it Data drives growth β€” maybe_grow adds experts when routing demands it

Fractus is a living system, not a model you train and ship:

  • the online trainer learns from every interaction (tick by tick);
  • the 5% memory injection consolidates salient thoughts every tick;
  • maybe_grow adds new experts when one expert dominates routing β€” capacity is created in response to what the model is experiencing;
  • progressive growth means the parameter count is a function of time and data, not a constant.

So the right way to read the active-param Chinchilla number is instantaneous: at any moment, the model's current active capacity is healthily fed by ~20Γ— its active params in cumulative exposure. As the model grows (more blocks, more experts), its appetite grows with it β€” and there is no final checkpoint. The ~3–4 B assembled today is the starting nutrition for the 1B palier; the data pipeline must keep flowing for the lifetime of the agent. Feed it, let it grow, repeat β€” indefinitely.