thefinalboss commited on
Commit
6a8bc11
·
verified ·
1 Parent(s): dbc0639

Upload docs/EVOLUTION.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. docs/EVOLUTION.md +96 -0
docs/EVOLUTION.md ADDED
@@ -0,0 +1,96 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Fractus — evolution log (timestamped)
2
+
3
+ Canonical chronology of the **system**, not the mythology.
4
+ Sister repos: `thefinalboss/fractus-cte`, `fractus-p0`, `fractus-vorax`, `fractus-datasets`.
5
+ Last updated: **2026-08-27**.
6
+
7
+ Dates are UTC calendar days as recorded in docs, manifests, and HF `lastModified`. Token counts are **manifest snapshots**, not live GPU clocks.
8
+
9
+ ---
10
+
11
+ ## 2026-04 — seed (outside this repo)
12
+
13
+ - Preprint work on K3 elliptic CM / automorphy (separate mathematical line).
14
+ - Fractus as an idea: thought is continuous, accumulating, oscillatory — not one forward pass.
15
+
16
+ ## 2026-07 — architectures in parallel
17
+
18
+ - **Fractus CTE** specified as multi-block continuous engine (linear attention + Kuramoto + PhaseRoutedMoE).
19
+ - **Palimpseste-max**: hypervector append-only memory, CPU, released on HF.
20
+ - **Fractus-Vorax** (late July / mid-August design): sealed CTE brain + organs that *write* knowledge (`.kn`) without training the 1B again.
21
+ - eQchip-Q / SOLARIS: other tracks, not this codebase.
22
+
23
+ ## 2026-08-12 — training doctrine
24
+
25
+ - `docs/2026-08-12-fractus-chinchilla.md`: Chinchilla-style epoch stacking is **rejected** for this MoE + HV setup. One real pass + Vorax ingest.
26
+
27
+ ## 2026-08-13 → 2026-08-15 — first long 4× GPU burn
28
+
29
+ - 4× RTX 5090, ~1k tok/s/GPU class.
30
+ - Loss from chaotic high values down through the 30s.
31
+ - Discoveries (logged, not marketing):
32
+ - generation collapse to single-token attractors
33
+ - train graph ≠ length-1 generate graph
34
+ - mid-train surgery is possible without throwing the `.pt`
35
+ - mean-merge of same-shape shards produces a usable brain
36
+ - Hugging Face Space attempt: torch pin / `audioop` / Gradio stack failures on Python 3.13.
37
+
38
+ ## 2026-08-16 — outage
39
+
40
+ - Pod / filesystem incident. Resume from last HF `.pt`, not from token 0.
41
+ - Migration toward 8× GPU recipe. Docs: `BOOST_AND_OUTAGE.md`, `BOOST_B4_RECOVERY.md`.
42
+
43
+ ## 2026-08-17 → 2026-08-18 — phase 2 + freeze + Vorax
44
+
45
+ - Phase-2 shards from `thefinalboss/fractus-datasets` (tokenized npy memmap).
46
+ - Anti re-ingest via manifest / `PHASE_SWITCH`.
47
+ - **Freeze snapshot:** ~15–17M tokens/GPU on ~430M-token phase-2 shards (~4% of that pass).
48
+ - Vorax repo stands up: compiler → traces / Hebbian / spawn organs, `:core` a sealed `.pt`.
49
+ - Decision: finish the pass later on next GPU budget; do not pretend 4% = trained.
50
+
51
+ ## 2026-08-19 — Kuramoto bottleneck named
52
+
53
+ - Agent probe: degenerate circular mean, post-RK4 routing arc ~**25°/360°**, \(r \approx 0.01\) even on a fresh path; ~70% dead experts on probe windows.
54
+ - Not “wait for more tokens.” Prep: `GATE_TEMP`, `OMEGA_SCALE`, real LB in the loss, Kuramoto **in** autograd.
55
+ - Docs: `KURAMOTO_BOTTLENECK_AND_FIX.md`, `CPU_MINI_MERGE_AND_DIMENSIONS.md` (CPU minis ≠ 1B progress; mini→1B mean-merge is invalid).
56
+
57
+ ## 2026-08-21 — wrong cluster
58
+
59
+ - QuickPod heterogeneous RTX (3090 / A4000 / 5060 Ti / A2000 12GB). 1B B=2 does not fit 12GB.
60
+ - Cluster deleted. Marketplace: **no 8×** offers at probe time. Credit stays on QuickPod; wait for identical 24GB+.
61
+
62
+ ## 2026-08-22 → 2026-08-26 — open-heart kernels + x8
63
+
64
+ - Mission: faster phase-2 **without breaking living checkpoints**.
65
+ - Linear attention: masked einsum \(O(C^2)\) → cumsum / **chunked**.
66
+ - `chunked_cross_entropy`, `BLOCK_CKPT` (1B in ~16GB class), atomic ckpt writes, int32 memmap.
67
+ - Tests: attention/CE equivalence suite.
68
+ - White paper v2 + `arxiv/main.tex` landed on the repo.
69
+ - **X8_MANIFEST** (pushed): `status: x8_running`, start tokens ~**34.5–35.8M**/GPU.
70
+ - README of 26 Aug claimed ~1550–1600 tok/s/GPU aggregate ~12.6k on 8×5090 — treat as a run log, re-measure on the next pod.
71
+
72
+ ## 2026-08-27 — P0 / v4 / DiffusionBlocks / lineage made explicit
73
+
74
+ - Package [fractus-p0](https://huggingface.co/thefinalboss/fractus-p0):
75
+ - **P0 routing:** atan2 encode from \(h\) (not LayerNorm-mean≈0), phase carry into RK4, per-token MoE phases, Switch LB. **Not** open-heart-equivalent (dynamics change, shapes do not).
76
+ - **v4 trainer:** SS ramp, anti-repeat λ=0.1, PREFIX unique@40 as language gate. CARRY length-1 is a separate hole.
77
+ - CPU mini: PREFIX counts `0 1 2 → 4 5 6…`; CARRY locks `4 4 4…`.
78
+ - **Fractus-native DiffusionBlocks** (Sakana/UTokyo ICLR 2026 mapped, not copied): one CTEBlock trained per step, noise band on residual, CE + MSE denoise, 0 new params. Mini 6/6 tests; `/σ²` denoise detonated last block — removed.
79
+ - Explicit statement: Fractus is a **linear-attention RNN**, not a transformer. Open items listed in the root README (decay on \(S\), DDP vs merge, MI probe on \(\bar\theta\), matched PPL, eval split).
80
+
81
+ ---
82
+
83
+ ## What “evolution” is not
84
+
85
+ - Not “AGI shipped.”
86
+ - Not “4.2B tokens digested” — a few percent of phase-2 plus earlier phase-1 shards.
87
+ - Not “mean-merge = 8× data efficiency.”
88
+ - Not “Kuramoto is solved” — P0 changes the drive; the MI test is still due.
89
+
90
+ ## Next (operational)
91
+
92
+ 1. Identical 24GB+ GPUs, x4 acceptable.
93
+ 2. Resume from X8 / freeze manifest — never token 0 by accident.
94
+ 3. Body = P0 + v4. Gate = PREFIX unique@40.
95
+ 4. Optional: \(\gamma\) decay on \(S\); MI(\(\bar\theta\), token) vs tick; DDP if the fabric allows; matched transformer PPL on a held-out **source split**.
96
+ 5. Vorax for knowledge writes. Do not re-pretrain the brain for facts.