|
Download docs/EVOLUTION.md from thefinalboss/fractus-cte: direct link, hf CLI and curl.
- Browser
- Download file 6.45 kB
-
https://huggingface.co/thefinalboss/fractus-cte/resolve/main/docs/EVOLUTION.md
- Command line
-
hf download hf://thefinalboss/fractus-cte/docs/EVOLUTION.md
-
curl -L -o EVOLUTION.md https://huggingface.co/thefinalboss/fractus-cte/resolve/main/docs/EVOLUTION.md
6.45 kB
Fractus — evolution log (timestamped)
Canonical chronology of the system, not the mythology.
Sister repos: thefinalboss/fractus-cte, fractus-p0, fractus-vorax, fractus-datasets.
Last updated: 2026-08-28 03:50 UTC.
Dates are UTC calendar days as recorded in docs, manifests, and HF lastModified. Token counts are manifest snapshots, not live GPU clocks.
2026-04 — seed (outside this repo)
- Preprint work on K3 elliptic CM / automorphy (separate mathematical line).
- Fractus as an idea: thought is continuous, accumulating, oscillatory — not one forward pass.
2026-07 — architectures in parallel
- Fractus CTE specified as multi-block continuous engine (linear attention + Kuramoto + PhaseRoutedMoE).
- Palimpseste-max: hypervector append-only memory, CPU, released on HF.
- Fractus-Vorax (late July / mid-August design): sealed CTE brain + organs that write knowledge (
.kn) without training the 1B again. - eQchip-Q / SOLARIS: other tracks, not this codebase.
2026-08-12 — training doctrine
docs/2026-08-12-fractus-chinchilla.md: Chinchilla-style epoch stacking is rejected for this MoE + HV setup. One real pass + Vorax ingest.
2026-08-13 → 2026-08-15 — first long 4× GPU burn
- 4× RTX 5090, ~1k tok/s/GPU class.
- Loss from chaotic high values down through the 30s.
- Discoveries (logged, not marketing):
- generation collapse to single-token attractors
- train graph ≠ length-1 generate graph
- mid-train surgery is possible without throwing the
.pt - mean-merge of same-shape shards produces a usable brain
- Hugging Face Space attempt: torch pin /
audioop/ Gradio stack failures on Python 3.13.
2026-08-16 — outage
- Pod / filesystem incident. Resume from last HF
.pt, not from token 0. - Migration toward 8× GPU recipe. Docs:
BOOST_AND_OUTAGE.md,BOOST_B4_RECOVERY.md.
2026-08-17 → 2026-08-18 — phase 2 + freeze + Vorax
- Phase-2 shards from
thefinalboss/fractus-datasets(tokenized npy memmap). - Anti re-ingest via manifest /
PHASE_SWITCH. - Freeze snapshot:
15–17M tokens/GPU on ~430M-token phase-2 shards (4% of that pass). - Vorax repo stands up: compiler → traces / Hebbian / spawn organs,
:corea sealed.pt. - Decision: finish the pass later on next GPU budget; do not pretend 4% = trained.
2026-08-19 — Kuramoto bottleneck named
- Agent probe: degenerate circular mean, post-RK4 routing arc ~25°/360°, (r \approx 0.01) even on a fresh path; ~70% dead experts on probe windows.
- Not “wait for more tokens.” Prep:
GATE_TEMP,OMEGA_SCALE, real LB in the loss, Kuramoto in autograd. - Docs:
KURAMOTO_BOTTLENECK_AND_FIX.md,CPU_MINI_MERGE_AND_DIMENSIONS.md(CPU minis ≠ 1B progress; mini→1B mean-merge is invalid).
2026-08-21 — wrong cluster
- QuickPod heterogeneous RTX (3090 / A4000 / 5060 Ti / A2000 12GB). 1B B=2 does not fit 12GB.
- Cluster deleted. Marketplace: no 8× offers at probe time. Credit stays on QuickPod; wait for identical 24GB+.
2026-08-22 → 2026-08-26 — open-heart kernels + x8
- Mission: faster phase-2 without breaking living checkpoints.
- Linear attention: masked einsum (O(C^2)) → cumsum / chunked.
chunked_cross_entropy,BLOCK_CKPT(1B in ~16GB class), atomic ckpt writes, int32 memmap.- Tests: attention/CE equivalence suite.
- White paper v2 +
arxiv/main.texlanded on the repo. - X8_MANIFEST (pushed):
status: x8_running, start tokens ~34.5–35.8M/GPU. - README of 26 Aug claimed ~1550–1600 tok/s/GPU aggregate ~12.6k on 8×5090 — treat as a run log, re-measure on the next pod.
2026-08-27 — P0 / v4 / DiffusionBlocks / lineage made explicit
- Package fractus-p0:
- P0 routing: atan2 encode from (h) (not LayerNorm-mean≈0), phase carry into RK4, per-token MoE phases, Switch LB. Not open-heart-equivalent (dynamics change, shapes do not).
- v4 trainer: SS ramp, anti-repeat λ=0.1, PREFIX unique@40 as language gate. CARRY length-1 is a separate hole.
- CPU mini: PREFIX counts
0 1 2 → 4 5 6…; CARRY locks4 4 4…. - Fractus-native DiffusionBlocks (Sakana/UTokyo ICLR 2026 mapped, not copied): one CTEBlock trained per step, noise band on residual, CE + MSE denoise, 0 new params. Mini 6/6 tests;
/σ²denoise detonated last block — removed.
- Explicit statement: Fractus is a linear-attention RNN, not a transformer. Open items listed in the root README (decay on (S), DDP vs merge, MI probe on (\bar\theta), matched PPL, eval split).
What “evolution” is not
- Not “AGI shipped.”
- Not “4.2B tokens digested” — a few percent of phase-2 plus earlier phase-1 shards.
- Not “mean-merge = 8× data efficiency.”
- Not “Kuramoto is solved” — P0 changes the drive; the MI test is still due.
Next (operational)
- Identical 24GB+ GPUs, x4 acceptable.
- Resume from X8 / freeze manifest — never token 0 by accident.
- Body = P0 + v4. Gate = PREFIX unique@40.
- Optional: (\gamma) decay on (S); MI((\bar\theta), token) vs tick; DDP if the fabric allows; matched transformer PPL on a held-out source split.
- Vorax for knowledge writes. Do not re-pretrain the brain for facts.
2026-08-27 evening → 2026-08-28 03:50 UTC — SS on 8×5090 (this night)
Full log: docs/2026-08-28-NIGHT-SS-X8.md.
- Named the gap in production: CE teacher-force ≠ generation. unique@40 PREFIX NO-GO ~3.3 while tf on GPU0 was already ~5–8.
- Scheduled sampling enabled and surviving after checkpoint-metadata crash on the second backward.
- Carry snapshot/restore; SS dense CE; no CE-checkpoint when N≤ce_chunk.
- Volume 80→200 GB after disk-full trainer death.
- All 8 live checkpoints + hourly HF push to
thefinalboss/fractus-cte/checkpoints/x8run/. - B=4 + SS_RATE=1.0 live: ~27 GB, ~990 tok/s/GPU, ~40M tokens/GPU on 430M phase-2 shards.
- Speech not claimed. SS is the training-side fix; PREFIX unique@40 is the gate.
2026-08-28 16:35 PDT — Decode surgery I
Space output Retail×32 proved the length-1 carry decode is an attractor, not a sentence.
fractus/generate_aligned.py now decodes with a 64-token causal window and masks the previous token. Trainers not stopped.
unique@32 after: 16 / 14 / 27 / 12. Short cycles left. Gate still PREFIX unique@40.
docs/2026-08-28-DECODE-WINDOW.md