|
Download docs/EVOLUTION.md from thefinalboss/fractus-cte: direct link, hf CLI and curl.
- Browser
- Download file 5.3 kB
-
https://huggingface.co/thefinalboss/fractus-cte/resolve/b462e42bf41e012a2d0117ebcd402b8b20a30817/docs/EVOLUTION.md
- Command line
-
hf download hf://thefinalboss/fractus-cte@b462e42bf41e012a2d0117ebcd402b8b20a30817/docs/EVOLUTION.md
-
curl -L -o EVOLUTION.md https://huggingface.co/thefinalboss/fractus-cte/resolve/b462e42bf41e012a2d0117ebcd402b8b20a30817/docs/EVOLUTION.md
5.3 kB
Fractus — evolution log (timestamped)
Canonical chronology of the system, not the mythology.
Sister repos: thefinalboss/fractus-cte, fractus-p0, fractus-vorax, fractus-datasets.
Last updated: 2026-08-27.
Dates are UTC calendar days as recorded in docs, manifests, and HF lastModified. Token counts are manifest snapshots, not live GPU clocks.
2026-04 — seed (outside this repo)
- Preprint work on K3 elliptic CM / automorphy (separate mathematical line).
- Fractus as an idea: thought is continuous, accumulating, oscillatory — not one forward pass.
2026-07 — architectures in parallel
- Fractus CTE specified as multi-block continuous engine (linear attention + Kuramoto + PhaseRoutedMoE).
- Palimpseste-max: hypervector append-only memory, CPU, released on HF.
- Fractus-Vorax (late July / mid-August design): sealed CTE brain + organs that write knowledge (
.kn) without training the 1B again. - eQchip-Q / SOLARIS: other tracks, not this codebase.
2026-08-12 — training doctrine
docs/2026-08-12-fractus-chinchilla.md: Chinchilla-style epoch stacking is rejected for this MoE + HV setup. One real pass + Vorax ingest.
2026-08-13 → 2026-08-15 — first long 4× GPU burn
- 4× RTX 5090, ~1k tok/s/GPU class.
- Loss from chaotic high values down through the 30s.
- Discoveries (logged, not marketing):
- generation collapse to single-token attractors
- train graph ≠ length-1 generate graph
- mid-train surgery is possible without throwing the
.pt - mean-merge of same-shape shards produces a usable brain
- Hugging Face Space attempt: torch pin /
audioop/ Gradio stack failures on Python 3.13.
2026-08-16 — outage
- Pod / filesystem incident. Resume from last HF
.pt, not from token 0. - Migration toward 8× GPU recipe. Docs:
BOOST_AND_OUTAGE.md,BOOST_B4_RECOVERY.md.
2026-08-17 → 2026-08-18 — phase 2 + freeze + Vorax
- Phase-2 shards from
thefinalboss/fractus-datasets(tokenized npy memmap). - Anti re-ingest via manifest /
PHASE_SWITCH. - Freeze snapshot:
15–17M tokens/GPU on ~430M-token phase-2 shards (4% of that pass). - Vorax repo stands up: compiler → traces / Hebbian / spawn organs,
:corea sealed.pt. - Decision: finish the pass later on next GPU budget; do not pretend 4% = trained.
2026-08-19 — Kuramoto bottleneck named
- Agent probe: degenerate circular mean, post-RK4 routing arc ~25°/360°, (r \approx 0.01) even on a fresh path; ~70% dead experts on probe windows.
- Not “wait for more tokens.” Prep:
GATE_TEMP,OMEGA_SCALE, real LB in the loss, Kuramoto in autograd. - Docs:
KURAMOTO_BOTTLENECK_AND_FIX.md,CPU_MINI_MERGE_AND_DIMENSIONS.md(CPU minis ≠ 1B progress; mini→1B mean-merge is invalid).
2026-08-21 — wrong cluster
- QuickPod heterogeneous RTX (3090 / A4000 / 5060 Ti / A2000 12GB). 1B B=2 does not fit 12GB.
- Cluster deleted. Marketplace: no 8× offers at probe time. Credit stays on QuickPod; wait for identical 24GB+.
2026-08-22 → 2026-08-26 — open-heart kernels + x8
- Mission: faster phase-2 without breaking living checkpoints.
- Linear attention: masked einsum (O(C^2)) → cumsum / chunked.
chunked_cross_entropy,BLOCK_CKPT(1B in ~16GB class), atomic ckpt writes, int32 memmap.- Tests: attention/CE equivalence suite.
- White paper v2 +
arxiv/main.texlanded on the repo. - X8_MANIFEST (pushed):
status: x8_running, start tokens ~34.5–35.8M/GPU. - README of 26 Aug claimed ~1550–1600 tok/s/GPU aggregate ~12.6k on 8×5090 — treat as a run log, re-measure on the next pod.
2026-08-27 — P0 / v4 / DiffusionBlocks / lineage made explicit
- Package fractus-p0:
- P0 routing: atan2 encode from (h) (not LayerNorm-mean≈0), phase carry into RK4, per-token MoE phases, Switch LB. Not open-heart-equivalent (dynamics change, shapes do not).
- v4 trainer: SS ramp, anti-repeat λ=0.1, PREFIX unique@40 as language gate. CARRY length-1 is a separate hole.
- CPU mini: PREFIX counts
0 1 2 → 4 5 6…; CARRY locks4 4 4…. - Fractus-native DiffusionBlocks (Sakana/UTokyo ICLR 2026 mapped, not copied): one CTEBlock trained per step, noise band on residual, CE + MSE denoise, 0 new params. Mini 6/6 tests;
/σ²denoise detonated last block — removed.
- Explicit statement: Fractus is a linear-attention RNN, not a transformer. Open items listed in the root README (decay on (S), DDP vs merge, MI probe on (\bar\theta), matched PPL, eval split).
What “evolution” is not
- Not “AGI shipped.”
- Not “4.2B tokens digested” — a few percent of phase-2 plus earlier phase-1 shards.
- Not “mean-merge = 8× data efficiency.”
- Not “Kuramoto is solved” — P0 changes the drive; the MI test is still due.
Next (operational)
- Identical 24GB+ GPUs, x4 acceptable.
- Resume from X8 / freeze manifest — never token 0 by accident.
- Body = P0 + v4. Gate = PREFIX unique@40.
- Optional: (\gamma) decay on (S); MI((\bar\theta), token) vs tick; DDP if the fabric allows; matched transformer PPL on a held-out source split.
- Vorax for knowledge writes. Do not re-pretrain the brain for facts.