fractus-cte / docs /EVOLUTION.md
thefinalboss's picture
docs+code: decode surgery I window-64 anti-copy 2026-08-28
520beb2 verified
|
Raw History Blame Contribute Delete
6.45 kB
# Fractus — evolution log (timestamped)
Canonical chronology of the **system**, not the mythology.
Sister repos: `thefinalboss/fractus-cte`, `fractus-p0`, `fractus-vorax`, `fractus-datasets`.
Last updated: **2026-08-28 03:50 UTC**.
Dates are UTC calendar days as recorded in docs, manifests, and HF `lastModified`. Token counts are **manifest snapshots**, not live GPU clocks.
---
## 2026-04 — seed (outside this repo)
- Preprint work on K3 elliptic CM / automorphy (separate mathematical line).
- Fractus as an idea: thought is continuous, accumulating, oscillatory — not one forward pass.
## 2026-07 — architectures in parallel
- **Fractus CTE** specified as multi-block continuous engine (linear attention + Kuramoto + PhaseRoutedMoE).
- **Palimpseste-max**: hypervector append-only memory, CPU, released on HF.
- **Fractus-Vorax** (late July / mid-August design): sealed CTE brain + organs that *write* knowledge (`.kn`) without training the 1B again.
- eQchip-Q / SOLARIS: other tracks, not this codebase.
## 2026-08-12 — training doctrine
- `docs/2026-08-12-fractus-chinchilla.md`: Chinchilla-style epoch stacking is **rejected** for this MoE + HV setup. One real pass + Vorax ingest.
## 2026-08-13 → 2026-08-15 — first long 4× GPU burn
- 4× RTX 5090, ~1k tok/s/GPU class.
- Loss from chaotic high values down through the 30s.
- Discoveries (logged, not marketing):
- generation collapse to single-token attractors
- train graph ≠ length-1 generate graph
- mid-train surgery is possible without throwing the `.pt`
- mean-merge of same-shape shards produces a usable brain
- Hugging Face Space attempt: torch pin / `audioop` / Gradio stack failures on Python 3.13.
## 2026-08-16 — outage
- Pod / filesystem incident. Resume from last HF `.pt`, not from token 0.
- Migration toward 8× GPU recipe. Docs: `BOOST_AND_OUTAGE.md`, `BOOST_B4_RECOVERY.md`.
## 2026-08-17 → 2026-08-18 — phase 2 + freeze + Vorax
- Phase-2 shards from `thefinalboss/fractus-datasets` (tokenized npy memmap).
- Anti re-ingest via manifest / `PHASE_SWITCH`.
- **Freeze snapshot:** ~15–17M tokens/GPU on ~430M-token phase-2 shards (~4% of that pass).
- Vorax repo stands up: compiler → traces / Hebbian / spawn organs, `:core` a sealed `.pt`.
- Decision: finish the pass later on next GPU budget; do not pretend 4% = trained.
## 2026-08-19 — Kuramoto bottleneck named
- Agent probe: degenerate circular mean, post-RK4 routing arc ~**25°/360°**, \(r \approx 0.01\) even on a fresh path; ~70% dead experts on probe windows.
- Not “wait for more tokens.” Prep: `GATE_TEMP`, `OMEGA_SCALE`, real LB in the loss, Kuramoto **in** autograd.
- Docs: `KURAMOTO_BOTTLENECK_AND_FIX.md`, `CPU_MINI_MERGE_AND_DIMENSIONS.md` (CPU minis ≠ 1B progress; mini→1B mean-merge is invalid).
## 2026-08-21 — wrong cluster
- QuickPod heterogeneous RTX (3090 / A4000 / 5060 Ti / A2000 12GB). 1B B=2 does not fit 12GB.
- Cluster deleted. Marketplace: **no 8×** offers at probe time. Credit stays on QuickPod; wait for identical 24GB+.
## 2026-08-22 → 2026-08-26 — open-heart kernels + x8
- Mission: faster phase-2 **without breaking living checkpoints**.
- Linear attention: masked einsum \(O(C^2)\) → cumsum / **chunked**.
- `chunked_cross_entropy`, `BLOCK_CKPT` (1B in ~16GB class), atomic ckpt writes, int32 memmap.
- Tests: attention/CE equivalence suite.
- White paper v2 + `arxiv/main.tex` landed on the repo.
- **X8_MANIFEST** (pushed): `status: x8_running`, start tokens ~**34.5–35.8M**/GPU.
- README of 26 Aug claimed ~1550–1600 tok/s/GPU aggregate ~12.6k on 8×5090 — treat as a run log, re-measure on the next pod.
## 2026-08-27 — P0 / v4 / DiffusionBlocks / lineage made explicit
- Package [fractus-p0](https://huggingface.co/thefinalboss/fractus-p0):
- **P0 routing:** atan2 encode from \(h\) (not LayerNorm-mean≈0), phase carry into RK4, per-token MoE phases, Switch LB. **Not** open-heart-equivalent (dynamics change, shapes do not).
- **v4 trainer:** SS ramp, anti-repeat λ=0.1, PREFIX unique@40 as language gate. CARRY length-1 is a separate hole.
- CPU mini: PREFIX counts `0 1 2 → 4 5 6…`; CARRY locks `4 4 4…`.
- **Fractus-native DiffusionBlocks** (Sakana/UTokyo ICLR 2026 mapped, not copied): one CTEBlock trained per step, noise band on residual, CE + MSE denoise, 0 new params. Mini 6/6 tests; `/σ²` denoise detonated last block — removed.
- Explicit statement: Fractus is a **linear-attention RNN**, not a transformer. Open items listed in the root README (decay on \(S\), DDP vs merge, MI probe on \(\bar\theta\), matched PPL, eval split).
---
## What “evolution” is not
- Not “AGI shipped.”
- Not “4.2B tokens digested” — a few percent of phase-2 plus earlier phase-1 shards.
- Not “mean-merge = 8× data efficiency.”
- Not “Kuramoto is solved” — P0 changes the drive; the MI test is still due.
## Next (operational)
1. Identical 24GB+ GPUs, x4 acceptable.
2. Resume from X8 / freeze manifest — never token 0 by accident.
3. Body = P0 + v4. Gate = PREFIX unique@40.
4. Optional: \(\gamma\) decay on \(S\); MI(\(\bar\theta\), token) vs tick; DDP if the fabric allows; matched transformer PPL on a held-out **source split**.
5. Vorax for knowledge writes. Do not re-pretrain the brain for facts.
## 2026-08-27 evening → 2026-08-28 03:50 UTC — SS on 8×5090 (this night)
Full log: `docs/2026-08-28-NIGHT-SS-X8.md`.
- Named the gap in production: CE teacher-force ≠ generation. unique@40 PREFIX NO-GO ~3.3 while tf on GPU0 was already ~5–8.
- Scheduled sampling **enabled and surviving** after checkpoint-metadata crash on the second backward.
- Carry snapshot/restore; SS dense CE; no CE-checkpoint when N≤ce_chunk.
- Volume 80→200 GB after disk-full trainer death.
- All 8 live checkpoints + hourly HF push to `thefinalboss/fractus-cte/checkpoints/x8run/`.
- B=4 + SS_RATE=1.0 live: ~27 GB, ~990 tok/s/GPU, ~40M tokens/GPU on 430M phase-2 shards.
- Speech not claimed. SS is the training-side fix; PREFIX unique@40 is the gate.
## 2026-08-28 16:35 PDT — Decode surgery I
Space output `Retail×32` proved the length-1 carry decode is an attractor, not a sentence.
`fractus/generate_aligned.py` now decodes with a 64-token causal window and masks the previous token. Trainers not stopped.
unique@32 after: 16 / 14 / 27 / 12. Short cycles left. Gate still PREFIX unique@40.
`docs/2026-08-28-DECODE-WINDOW.md`