thefinalboss commited on
Commit
833d5b0
Β·
verified Β·
1 Parent(s): 3e6c564

docs: QuickPod 8x5090 exact-token resume 2026-08-28

Browse files
Files changed (1) hide show
  1. README.md +424 -80
README.md CHANGED
@@ -1,127 +1,471 @@
1
  ---
2
  license: mit
3
  language:
4
- - en
5
- - fr
6
  tags:
7
- - continuous-thought
8
- - linear-attention
9
- - kuramoto
10
- - moe
11
- - rnn
 
 
 
 
12
  library_name: pytorch
 
 
 
 
 
13
  ---
14
 
15
- # Fractus-CTE
16
 
17
- **Continuous Thought Engine.** A dynamical system with a fixed-size recurrent state, not a transformer.
18
 
19
- Fractus belongs to the linear-attention RNN family (Katharopoulos 2020, RetNet, RWKV, GLA, DeltaNet, Mamba/S6) **plus** phase-routed MoE and a persistent thought state across chunks. Depth is time. The checkpoint is a living state (weights + phases + attention memory), not a frozen function.
20
 
21
- | | Transformer | Fractus |
22
- |---|---|---|
23
- | Computation | one forward per prompt | continuous ticks / chunks |
24
- | Memory | KV cache grows with length | fixed-size \(S, z\), thought state |
25
- | Routing | (optional) learned softmax MoE | Kuramoto phases on a circle |
26
- | Knowledge after train | fine-tune / RAG | **Vorax** organs, append-only `.kn` |
27
- | Operability | replace the blob | open-heart: code changes, `.pt` shapes stay |
 
 
 
 
 
 
 
 
 
 
 
 
28
 
29
- **Author:** Philippe-Antoine Robert ([thefinalboss](https://huggingface.co/thefinalboss))
30
- **License:** MIT
31
- **Sibling repos:** [fractus-p0](https://huggingface.co/thefinalboss/fractus-p0) (routing + v4 + DiffusionBlocks body) Β· [fractus-vorax](https://huggingface.co/thefinalboss/fractus-vorax) (sealed brain + ingest) Β· [fractus-datasets](https://huggingface.co/datasets/thefinalboss/fractus-datasets)
 
 
 
 
32
 
33
- Full chronology: [`docs/EVOLUTION.md`](docs/EVOLUTION.md) Β· franΓ§ais [`docs/EVOLUTION.fr.md`](docs/EVOLUTION.fr.md)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
34
 
35
  ---
36
 
37
- ## Architecture (1B production config)
38
 
39
- | | |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
40
  |---|---|
41
- | Params | ~1.165B |
42
- | `d_model` | 1280 |
43
- | `n_layers` | 16 CTEBlocks |
44
- | Attention | causal **linear** (cumsum / chunked). Env `FRACTUS_ATTN_IMPL` |
45
- | Oscillators / block | 16, coupling rank 8 |
46
- | Experts / block | 128, top-2, phase-gated |
47
- | Vocab | 50257 GPT-2 BPE |
48
- | Train objective | dense next-token CE on the chunk + Switch load-balance (+ optional SS / anti-repeat in v4) |
 
 
 
 
 
 
 
 
 
 
49
 
50
- Each block: attention β†’ Kuramoto β†’ PhaseRoutedMoE. \(S,z\) and phases are **per-block** and carried across chunk boundaries.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
51
 
52
  ---
53
 
54
- ## What is true right now (2026-08-27)
55
 
56
- **Done**
57
- - Multi-GPU data-parallel shards, exact `START_TOKEN` resume, hourly mean-merge of same-shape `.pt`
58
- - Open-heart speed work (22–26 Aug): chunked linear attention, `chunked_cross_entropy`, `BLOCK_CKPT`, atomic ckpts β€” **shapes unchanged**
59
- - Phase-2 corpus on HF; freeze then x8 resume (~35M tokens/GPU on ~430M-token shards at last pushed manifest)
60
- - P0 routing body (atan2 encode, phase carry, per-token MoE, Switch LB) in [fractus-p0](https://huggingface.co/thefinalboss/fractus-p0)
61
- - PREFIX vs CARRY decode hole **named and measured** on a CPU mini
62
- - Fractus-native DiffusionBlocks prototype (one CTEBlock / step, 0 new params) β€” experiment, not default
 
 
 
 
 
63
 
64
- **Open (do not skip)**
65
- - Attention state \(S\) still looks like \(S_t = S_{t-1} + k\otimes v\) β€” **no learned decay** yet (RWKV/RetNet/Mamba all added one)
66
- - Mean-merge of independent MoE shards β‰  DDP; expert #k is not aligned across GPUs
67
- - Kuramoto decision is still a **circle** (need \(\mathrm{MI}(\bar\theta, \text{token})\) vs tick)
68
- - No published **matched-compute PPL** vs a vanilla transformer on a clean held-out
69
- - CARRY length-1 generation is not the trained graph (PREFIX is the language gate)
70
- - Eval must be split by **source**, not by line (identity text in the corpus)
71
 
72
- Default next run: **v4 + P0 body + resume offsets + 24GB+ identical GPUs (x4 is enough)**.
73
- DiffusionBlocks and DDP are options. Finishing the pass without a baseline PPL is not a paper.
 
 
 
 
 
 
 
 
74
 
75
  ---
76
 
77
- ## Load
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
78
 
79
- ```python
80
- from fractus.continuous_engine import ContinuousThoughtEngine
81
- eng = ContinuousThoughtEngine.from_pretrained("checkpoints/FRACTUS_1B_PHASE2_FROZEN_MERGED.pt")
82
- # .pt = weights + live state. fractus/ = the body that ticks.
83
  ```
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
84
 
85
- Same-shape checkpoints can be mean-merged. Different `d_model` / layer / expert counts **cannot**. See `docs/CPU_MINI_MERGE_AND_DIMENSIONS.md`.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
86
 
87
  ---
88
 
89
- ## Train (resume, never from token 0 unless you mean it)
 
 
90
 
91
  ```bash
92
- # offsets: checkpoints/X8_MANIFEST.json or FROZEN_RESUME_MANIFEST.json
93
- CUDA_VISIBLE_DEVICES=0 GPU_ID=0 START_TOKEN=<manifest> \
94
- BATCH=2 CE_CHUNK=2048 FRACTUS_ATTN_IMPL=chunked BLOCK_CKPT=1 \
95
- python -u scripts/fast4gpu_boost_v3.py
96
  ```
97
 
98
- P0 + v4 launcher lives in [fractus-p0](https://huggingface.co/thefinalboss/fractus-p0) (`scripts/fast4gpu_boost_v4.py`).
99
- PREFIX `unique@40` is the language gate. Do not raise `REPEAT_COEF` above 0.1.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
100
 
101
  ---
102
 
103
- ## Docs worth reading first
104
 
105
- | Doc | Why |
106
- |---|---|
107
- | `docs/EVOLUTION.md` | timestamped history |
108
- | `docs/HOW_FRACTUS_IS_TRAINED.md` | phase-2 recipe |
109
- | `docs/KURAMOTO_BOTTLENECK_AND_FIX.md` | dead experts / 25Β° arc |
110
- | `docs/OPTIMIZATION_2026-08-22.md` | open-heart kernels |
111
- | `docs/TRAIN_GEN_MISMATCH.md` / p0 `V4_AR.md` | PREFIX vs CARRY |
112
- | `docs/TRUSTED_LOSS.md` | loss β‰  generation |
113
- | `Fractus_White_Paper_v2.md` | architecture write-up |
114
 
115
  ---
116
 
117
- ## Citation
118
 
119
  ```
120
- @misc{robert2026fractus,
121
- title={Fractus: a Continuous Thought Engine},
122
- author={Robert, Philippe-Antoine},
123
- year={2026},
124
- howpublished={https://huggingface.co/thefinalboss/fractus-cte},
125
- note={MIT. Dynamical system; linear-attention RNN lineage + phase-routed MoE.}
126
- }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
127
  ```
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
  license: mit
3
  language:
4
+ - en
5
+ - fr
6
  tags:
7
+ - continuous-thought-engine
8
+ - cognitive-agent
9
+ - mixture-of-experts
10
+ - kuramoto
11
+ - self-modifying
12
+ - progressive-growth
13
+ - decentralized-ai
14
+ - personal-ai
15
+ - neuroscience
16
  library_name: pytorch
17
+ pipeline_tag: text-generation
18
+ models:
19
+ - thefinalboss/fractus-cte
20
+ datasets:
21
+ - thefinalboss/fractus-datasets
22
  ---
23
 
24
+ # Fractus CTE
25
 
26
+ **A living AI that thinks continuously, remembers forever, and grows on its own.**
27
 
28
+ **Fractus is NOT a transformer.** It's a Continuous Cognitive Agent β€” a dynamical system that maintains a persistent thought state, advances it tick by tick through 16 blocks, and routes via Kuramoto oscillator phases. The checkpoint is a living seed: it never freezes, grows at runtime, and trains forever.
29
 
30
+ > **Status (2026-08-28):** Training live again on **QuickPod 8Γ— RTX 5090** after a RunPod death. Exact-token resume from `checkpoints/x8run` (GPU0 at 79.9M / 430M of this Phase-2 pass). Trainer: `scripts/fast4gpu_boost_v4.py`, SS_RATE=1.0, BATCH=4. Speech gate is **unique@40 greedy PREFIX** β€” last offline probe was NO-GO (3/12/1/6). Not a finished assistant. See [Live Training](#live-training-status) and [`docs/2026-08-28-QUICKPOD-RESUME.md`](docs/2026-08-28-QUICKPOD-RESUME.md).
31
+
32
+ ## Quick Start
33
+
34
+ ```bash
35
+ git clone https://github.com/AFKmoney/fractus-cte.git
36
+ cd fractus-cte && pip install torch numpy tokenizers
37
+
38
+ # Load the trained 1B and generate
39
+ python -c "
40
+ from fractus.continuous_engine import ContinuousThoughtEngine
41
+ engine = ContinuousThoughtEngine.from_pretrained('checkpoints/fractus_1b_gpu3.pt')
42
+ import torch
43
+ logits, confidence = engine.tick(torch.tensor([42]))
44
+ print(f'Fractus is thinking. Confidence: {confidence.item():.2f}')
45
+ "
46
+ ```
47
+
48
+ The `.pt` checkpoint contains the full model (weights + dynamic state). You need this repo's code to run it β€” Fractus is a custom architecture, not a transformer. Checkpoints are on [HF](https://huggingface.co/thefinalboss/fractus-cte).
49
 
50
+ ![License](https://img.shields.io/badge/license-MIT-blue)
51
+ ![Python](https://img.shields.io/badge/python-3.13-green)
52
+ ![PyTorch](https://img.shields.io/badge/PyTorch-2.11-orange)
53
+ ![Params](https://img.shields.io/badge/params-1.05B-red)
54
+ ![Active](https://img.shields.io/badge/active%20params-119M-yellow)
55
+ ![Status](https://img.shields.io/badge/status-training%20live-brightgreen)
56
+ ![Datasets](https://img.shields.io/badge/datasets-4.2B%20tokens-purple)
57
 
58
+ ---
59
+
60
+ ## What is Fractus?
61
+
62
+ Fractus is not a chatbot. It's not GPT. It's not a transformer.
63
+
64
+ Fractus is a **Continuous Cognitive Agent** β€” an AI that works like a brain, not a calculator. Instead of processing input β†’ output in one pass, Fractus **ticks** like a biological system: it maintains a persistent thought state, advances it through multiple blocks of processing, remembers everything across sessions, and can grow new capacity by itself.
65
+
66
+ ### What makes it different from GPT/Claude?
67
+
68
+ | | GPT-4 / Claude | Fractus |
69
+ |---|---|---|
70
+ | **Thinking** | One pass, done | Continuous ticks (like a heartbeat) |
71
+ | **Memory** | Forgets when context window fills | Remembers forever (survives restarts) |
72
+ | **Learning** | Retrain from scratch ($$$) | Learns from every interaction |
73
+ | **Growth** | Fixed size forever | Grows new experts at runtime |
74
+ | **Mental states** | One mode always | Shifts between cognitive modes |
75
+ | **Where it runs** | Corporate cloud | Your machine |
76
+ | **Training** | Fixed, done once | Perpetual, never stops |
77
 
78
  ---
79
 
80
+ ## Architecture (1.05B total / ~119M active per token)
81
 
82
+ ```
83
+ ContinuousThoughtEngine
84
+ β”œβ”€β”€ d_model=1280, 16 layers, 16 heads
85
+ β”œβ”€β”€ FractalLinearAttention multi-level causal linear attention (O(L)),
86
+ β”‚ carry state (S, z) persists across chunks
87
+ β”œβ”€β”€ Kuramoto phase clock RK4-integrated oscillators β†’ phase vectors
88
+ β”‚ (learned end-to-end; feeds routing)
89
+ β”œβ”€β”€ PhaseRoutedMoE 128 experts, top-2 active per token,
90
+ β”‚ von Mises gate over phases, Farey-sequence
91
+ β”‚ expert phases, load-balance loss in objective
92
+ β”œβ”€β”€ Low-rank experts W = scale Β· U@Vα΅€ (rank 64) β†’ 64Γ— less compute
93
+ β”œβ”€β”€ Tied embedding/head GPT-2 BPE vocab (50257)
94
+ └── Persistent thought state residual stream carried across ticks
95
+ ```
96
+
97
+ **Parameter accounting:** ~1.05B total, but only ~118.8M are active per token (dense attention + top-2 of 128 sparse experts) β€” ~183.1M including the tied embedding. An 11.3% sparsity ratio. This is the basis of the adapted scaling target below.
98
+
99
+ ### The 12 Building Blocks
100
+
101
+ | Block | What it does |
102
  |---|---|
103
+ | **Continuous Thought Engine** | The brain β€” thinks tick by tick through 16 blocks |
104
+ | **Persistent Memory** | Vector bank surviving sessions, cosine recall, 5% blend injection, salience head gates impact |
105
+ | **Cognitive Modes** | Mental states discovered unsupervised (k-means on Kuramoto phase features): focused, creative, exploratory |
106
+ | **RAG Knowledge Base** | Learns facts instantly β€” no retraining needed |
107
+ | **Cognitive Plugins** | Hot-swappable modes: analyst, coder, creative, teacher |
108
+ | **MetaCognition** | Decides its own actions: retrieve, learn, generate |
109
+ | **Progressive Growth** | Grows from 6M to 1B+ params, palier by palier (`maybe_grow`: width + depth + experts) |
110
+ | **Self-Modification** | Adds new experts at runtime when routing is imbalanced (zero-init, placed near the dominant expert) |
111
+ | **PhaseRoutedMoE** | Sparse experts routed by oscillator phases |
112
+ | **Kuramoto Clock** | A dynamical system that drives routing decisions |
113
+ | **Online Trainer** | Learns continuously, one chunk at a time |
114
+ | **Vision (`tick_vec`)** | Multimodal input path β€” image patches drive the engine directly, bypassing the token embedding (CIFAR "eyes" prototype trained while the text run continued) |
115
+
116
+ ---
117
+
118
+ ## Live Training Status (28 August 2026)
119
+
120
+ **Hardware:** QuickPod 8Γ— RTX 5090. One independent Python process per GPU. No DDP / no gradient sync. Consolidation = mean-merge of the 8 `.pt` when we need a single brain.
121
 
122
+ **This pass (Phase 2 shards):**
123
+ - 8 GPT-2 BPE int32 memmaps, `~429,896,462` tokens/GPU β€” dataset: [thefinalboss/fractus-datasets](https://huggingface.co/datasets/thefinalboss/fractus-datasets) `tokenized/phase2/shard_phase2_gpu{i}.npy`
124
+ - Resume point **2026-08-28 14:40 UTC** after RunPod SSH death: GPU0 **79,922,176** tokens (~18.6% of this pass). Other GPUs 76.2M–78.4M. Exact JSON: `checkpoints/x8run/RESUME_gpu{i}.json`
125
+ - Trainer: `scripts/fast4gpu_boost_v4.py`
126
+ - Config: **BATCH=4**, SEQ=128, LR 7e-4, SGD+momentum 0.9, bf16, TF32, compile off, `FRACTUS_ATTN_IMPL=chunked`, `BLOCK_CKPT=1`, **SS_RATE=1.0**, SS_PROB 0.2β†’0.5 over 50M tokens, REPEAT_COEF=0.1, P0 on, **PROBE_EVERY=0** (live unique@40 probe crashed the first QuickPod launch β€” probes are offline)
127
+ - Signals: `tf`/`ema_tf`, `ss`/`ema_ss`, `rep`/`ema_rep`, `lb` (~14)
128
+ - Throughput after resume: ~840–940 tok/s/GPU. Remainder of this pass β‰ˆ **4.1–4.5 days** if 8 GPUs stay up
129
+ - **Torch on RTX 5090:** image default `2.2.1+cu121` has **no Blackwell kernels**. Required: **torch β‰₯ 2.11 + cu128**
130
+
131
+ **Crash recovery:** hourly upload of `checkpoints/x8run/fractus_1b_gpu{i}.pt` + `RESUME_gpu{i}.json`. That folder is the source of truth. See [`docs/2026-08-28-QUICKPOD-RESUME.md`](docs/2026-08-28-QUICKPOD-RESUME.md).
132
+
133
+ **Speech gate (not the loss):** `unique@40` greedy PREFIX β€” 40 generated tokens, count distinct ids, no ban, no temperature. Last offline numbers: Hello=3, Once upon a time=12, meaning of life=1, Fractus is=6. **NO-GO.** Teacher-force CE going down does **not** mean Fractus speaks. Do not cite banned-decode diversity as speech.
134
+
135
+ ## Vision β€” the CIFAR "eyes" prototype
136
+
137
+ Fractus is not text-only. The engine exposes **`tick_vec(obs_vec)`**: a multimodal entry point that accepts a precomputed `(B, d_model)` vector and injects it directly into the residual thought state β€” **bypassing the token embedding entirely**. Any modality that can be embedded into `d_model` dimensions can drive continuous thought: images, audio, sensor streams.
138
+
139
+ ```
140
+ image β†’ PatchEmbed β†’ (B, d_model) patch vectors
141
+ β”‚
142
+ β–Ό
143
+ engine.tick_vec(patch) ← no tokenizer involved
144
+ β”‚
145
+ thought state advances through all 16 blocks
146
+ (same Kuramoto routing + MoE as text)
147
+ ```
148
+
149
+ **Proof of concept β€” CIFAR-10 eyes:** a small CTE + PatchEmbed stack (`fractus/nn/vision.py`) was trained on CPU on real CIFAR-10 images (`fractus_eyes_cifar_final.pt`), demonstrating that image patches can drive the continuous-thought loop. Crucially, this ran as a **parallel track**: the 8-GPU text run was never interrupted for eyes work β€” GPU text digestion and CPU vision learning proceed simultaneously on the same living system.
150
+
151
+ This mirrors the biological premise: eyes evolved as a peripheral sense feeding a central dynamical brain, not as a core feature of it. The text 1B is the brain; vision is a front-end that plugs into `tick_vec`.
152
+
153
+ ---
154
+
155
+ ### Adapted Chinchilla target
156
+
157
+ Standard Chinchilla (20 tokens/param) assumes dense transformers where all parameters train on every token. Fractus violates that: only ~119M of 1.05B params are active per token, and the Kuramoto clock trains end-to-end now. The adapted target is **20 Γ— active params β‰ˆ 2.4–3.7B tokens**; with a warm-started checkpoint, the practical target is **~2.5–3B tokens** β€” which is why the phase-2 corpus is 3.44B. Since Fractus grows continuously (`maybe_grow`), this is an instantaneous cumulative-exposure requirement that grows with the model, not a final-stop condition. Details: [`docs/2026-08-12-fractus-chinchilla.md`](docs/2026-08-12-fractus-chinchilla.md).
158
 
159
  ---
160
 
161
+ ## The Surgeries (mid-training interventions, no weight wipes)
162
 
163
+ A defining discovery of this run: **the brain (.pt) and the code are separable.** Bottlenecks were fixed by live surgery β€” save the checkpoint, patch the code, reload weights, resume at the exact recorded token offset. Multi-day digestion is never thrown away.
164
+
165
+ | Phase | Intervention | Result |
166
+ |---|---|---|
167
+ | **A β€” Initial** | Last-position-only CE, 4 independent GPUs | Loss fell; generation collapsed into single-token loops |
168
+ | **B β€” Routing surgery** | Kuramoto was frozen under `no_grad` (order parameter r β‰ˆ 0.01–0.03); LB loss was detached β†’ ~70% of experts dead. Fixed: gradients enabled (state kept detached for carry), CE + 0.02Β·lb, gate temp 1.0 β†’ 2.5, omega scale Γ—4 | lb β‰ˆ 14 live on all GPUs; experts alive |
169
+ | **C β€” Dense CE** | Replaced last-position CE (1 target / 128 tokens) with CE over all positions | Sharp loss drop; token-to-token chaining enforced |
170
+ | **D β€” Decode surgery** | Phase/thought noise, frequency penalties, cycle bans, forced escape tokens | Loop lock broken; still no coherent English |
171
+ | **E β€” Train/gen mismatch** | Training used causal attention + RK4 Kuramoto; generation used a simpler single-tick Euler path. Aligned decode via `generate_aligned.py` | Decode now matches the training path |
172
+ | **F β€” Loss recalibration** | Cumulative-average CE was misleading near 2.0 β†’ batch CE + EMA; LR 1e-3 β†’ 5e-4 | Trustworthy metrics |
173
+ | **G β€” Scheduled sampling** | Two-pass training: TF CE + LB, plus SS steps mixing model samples into inputs | Live; now SS_RATE=1.0 on boost_v4 |
174
+ | **H β€” Pod death / exact resume** | RunPod SSH died 2026-08-28. Reloaded 8Γ— `.pt` + manifests on a new QuickPod 8Γ—5090 at the recorded `start_token_next`. Torch upgraded 2.2β†’2.11+cu128 for Blackwell | Weights kept. Pass continues from ~80M/430M |
175
 
176
+ Full logs: [`docs/MASTER_RUN_LOG.md`](docs/MASTER_RUN_LOG.md), [`docs/DISCOVERY_LOG.md`](docs/DISCOVERY_LOG.md), [`docs/OPERABILITY_MIDTRAIN.md`](docs/OPERABILITY_MIDTRAIN.md).
 
 
 
 
 
 
177
 
178
+ ### Composable checkpoints
179
+
180
+ Because shapes stay compatible and manifests record token offsets, the following operations are proven on this run:
181
+ - **Parallel independent training** β€” N GPUs on separate shards
182
+ - **Mean-merge** β€” per-GPU checkpoints averaged into one unified model that still generates (424/440 tensors on the 4-GPU merge; stateful buffers reset)
183
+ - **Iterative fusion** β€” train β†’ merge β†’ train cycles; the model absorbs compatible checkpoints and keeps going
184
+ - **Exact-token resume** β€” across pod reboots and driver crashes
185
+ - **Compositional growth β‰  structural growth** β€” weight averaging vs `maybe_grow` paliers
186
+
187
+ Details: [`docs/COMPOSABILITY_AND_SURGERY.md`](docs/COMPOSABILITY_AND_SURGERY.md).
188
 
189
  ---
190
 
191
+ ## Trusted Loss (reading the numbers)
192
+
193
+ Three metrics, three different questions:
194
+
195
+ | Metric | What it measures | Where |
196
+ |---|---|---|
197
+ | `ema_tf` | CE with ground-truth history + continuous internal state β€” is the model digesting data? | live |
198
+ | `ema_ss` | CE after scheduled sampling mixes model-generated tokens in β€” partial free-run robustness | live |
199
+ | **AR** | Warm on 32 true tokens, greedily free-run 32 steps, CE vs truth β€” actual generation quality | offline |
200
+
201
+ **A single CE number cannot represent both teacher-forced learning and free-run generation.** Live ema_tf after the QuickPod resume is mid-teens on the lead GPUs (first steps inflate EMA β€” wait). AR stays the offline generation metric and last unique@40 was NO-GO. Random-vocab CE β‰ˆ 10.82; AR in the thousands (random guessing over the GPT-2 vocab is ~10.82): free-running compounds every error, and teacher forcing always supplies the correct past. Trust `ema_tf`/`ema_ss` for learning progress, AR + text probes for generation progress. The convergence signal is AR falling toward the ss/tf order of magnitude, plus readable output. Details: [`docs/TRUSTED_LOSS.md`](docs/TRUSTED_LOSS.md).
202
+
203
+ ---
204
+
205
+ ## Research Results (Honest)
206
+
207
+ **Refuted:**
208
+ - **EDT** (Expert Decoupled Training) β€” all 5 variants ~19–20% worse than from-scratch; MSE objective misaligned with CE, router ignores half the experts.
209
+ - **Forward-Forward** (Hinton 2022) β€” NLL rose from 124 to 221; local objectives can't replace global backprop here.
210
+
211
+ **Validated:**
212
+ - **Progressive growth** β€” warm start converges faster; trained through palier 3 (350M, loss 23.0, ~5 days CPU).
213
+ - **Sparse low-rank MoE** β€” 2/128 experts = 64Γ— less compute.
214
+ - **Open-heart operability** β€” live surgery on a training model works (see above).
215
+ - **Routing pathology as a first-class debug target** β€” expert-hit histograms and the Kuramoto order parameter catch failures that loss curves hide (they caught the frozen clock and the dead experts).
216
+
217
+ **Training optimizations (measured):** tied head (~1.1Γ—), head-partial training (~2Γ—), sparse gathered low-rank MoE (8Γ— at 16 experts, 64Γ— at 128), detached-state Kuramoto, gradient accumulation (~1.4Γ—), SGD+momentum over AdamW (~1.37Γ—), batching (335 β†’ 1345 tok/s at B=8), bf16 (~2Γ—). Combined CPU: ~336Γ— over naive. Measured on CPU: 707 tok/s single-stream, **1345 tok/s batched**.
218
+
219
+ ---
220
+
221
+ ## Datasets (4.2 Billion Tokens)
222
+
223
+ Source of truth: [huggingface.co/datasets/thefinalboss/fractus-datasets](https://huggingface.co/datasets/thefinalboss/fractus-datasets).
224
+
225
+ | Dataset | Tokens | Content |
226
+ |---|---|---|
227
+ | **neuro-paradigms-1b** | **~1B** | 100 neuroscience β†’ software architecture paradigms (300 chunked files) |
228
+ | **neuro-code-math** | **~900M** | Neuro-inspired coding, mathematics, algorithms (incl. 40 applied-neuroscience topics) |
229
+ | **cognitive-skills** | **~780M** | Coding, reasoning, speaking, thinking, understanding |
230
+ | **fractus-generated-corpus** | **340M** | Bilingual FR/EN generated by Fractus ontology engine |
231
+ | **paradigms-full** | **191M** | 140 paradigms (neuroscience, CS, architecture) |
232
+ | **gutenberg-esoteric** | **~58M** | 487 public-domain esoteric / masonic / hermetic books |
233
+ | **neuro-arch-full** | **86M** | 60 neuroscience paradigms (neuro-software-architecture) |
234
+ | **all-github-repos** | **54M+** | 80+ repos (public + private, secret-filtered) |
235
+ | **mega-corpus-v3** | **20M** | Literature, philosophy, occult, masonry, science, medicine |
236
+ | **wordnet** | **3M** | 117K dictionary synset entries |
237
+ | **Total** | **~4.2B** | |
238
+
239
+ Tokenized streams: Phase 1 = ~1.52B tokens (`.pt` files); Phase 2 = ~3.44B tokens (8 GPT-2 BPE int32 memmap shards). Phases are kept separate to avoid re-ingesting the same ordered stream.
240
+
241
+ ---
242
+
243
+ ## Applied Neuroscience β€” the theoretical core
244
+
245
+ Fractus is a neuroscience-grounded architecture: real brain mechanisms are mapped to software/AI patterns, and that mapping is itself training data. Every entry below is present in the dataset β€” verified by file listing, not just claimed.
246
+
247
+ ### 100 neuroscience β†’ software-architecture paradigms (`neuro_paradigms_1b`, 300 chunked files)
248
+
249
+ Each paradigm maps a biological mechanism to an engineering pattern (e.g. *adenosine sleep pressure* β†’ cache-stampede recovery; *myelin sheath* β†’ caching; *hippocampal replay* β†’ trajectory consolidation).
250
+
251
+ <details><summary><b>show all 100 paradigms</b></summary>
252
 
 
 
 
 
253
  ```
254
+ adenosine_sleep_pressure amygdala_prefrontal_topdown anterior_cingulate_conflict_monitor
255
+ apoptosis_self_destructing_service arc_gene_plasticity_marker astrocyte_tripartite_synapse
256
+ axon_initial_segment_trigger basal_ganglia_loop_arbitration bdnf_growth_factor_scaling
257
+ bergmann_glia_purkinje binaural_cross_correlation_localization brainstem_vital_functions
258
+ broca_area_api_generator calcium_transmitter_coupling camp_second_messenger_amplifier
259
+ cerebellar_forward_model cholinergic_attentional_filter circadian_gene_expression
260
+ climbing_fiber_error_broadcast cochlear_compressive_nonlinearity cortical_area_specialization
261
+ cortical_minicolumn_pipeline cortico_cortical_pathways corticotropin_releasing_hormone
262
+ cortisol_slow_stress_recovery critical_period_learning_rate dendritic_compartmentalization
263
+ endocannabinoid_retrograde enteric_glia_gut_brain ependymal_cell_barrier
264
+ fusiform_face_service_registry gaba_inhibitory_bus gap_junction_electrical_sync
265
+ ghrelin_hunger_signal glomerular_convergence_gateway glutamate_excitatory_bus
266
+ glycine_coagonist_modulator granule_cell_inhibitory_relay hair_cell_banks_event_clusters
267
+ hippocampal_4ec_loop_replay histamine_wakefulness_keeper hox_gene_service_specialization
268
+ hypercolumn_module_federation hypercolumn_sharding hypothalamus_homeostasis
269
+ insula_interoception_monitor ip3_inositol_cascade k_complex_event_trigger
270
+ kcc2_chloride_shift_inhibitor leptin_satiety_signal locus_coeruleus_ne_global_signal
271
+ melatonin_circadian_scheduler microglia_active_surveillance mitral_tufted_cell_dual_path
272
+ morphogen_gradient_config muller_glia_retina_repair myelin_sheath_caching
273
+ neural_crest_migration_deploy neuropeptide_y_stress_buffer ng2_glia_pool_renewal
274
+ nitric_oxide_gas_signal node_of_ranvier_bypass nrem_slow_wave_cleanup
275
+ nucleus_accumbens_reward_routing oligodendrocyte_myelination_dynamic orexin_stability_keeper
276
+ orientation_column_indexing oscillatory_phase_locking_io oxytocin_trust_protocol
277
+ parahippocampal_place_topology parallel_fiber_fanout_aggregation pineal_circadian_release
278
+ pinwheel_central_layout pituitary_master_gland posterior_parietal_integration
279
+ prolactin_parental_care quantal_release_batching radial_glia_neural_stem
280
+ radial_glial_scaffold raphe_serotonin_rate_limit rem_paradoxical_processing
281
+ replay_consolidation_trajectory reticular_activating_system retinotopic_data_layout
282
+ satellite_glial_ganglion schwann_cell_peripheral_repair sleep_pressure_forced_maintenance
283
+ sleep_spindle_memory_transfer slow_oscillation_sync subplate_wait_state
284
+ suprachiasmatic_clock synaptic_vesicle_pool synaptogenesis_service_wiring
285
+ tanycyte_metabolic_sensor temporal_pole_semantic_cache thalamocortical_loop_api
286
+ tonotopic_stream_partitioning vasopressin_loyalty_aware_routing vta_dopamine_rpe_scheduler
287
+ wernicke_area_api_parser
288
+ ```
289
+ </details>
290
+
291
+ ### 40 applied-neuroscience topics (`neuro_code_math/applied_neuroscience/`)
292
 
293
+ Deep dives on computational neuroscience theories β€” the science Fractus's design draws from.
294
+
295
+ <details><summary><b>show all 40 topics</b></summary>
296
+
297
+ ```
298
+ active_inference axonal_computation basal_ganglia_circuits bayesian_brain
299
+ cerebellar_computation consolidation cortical_minicolumns cross_frequency_coupling
300
+ dendritic_computation dopamine_reward entorhinal_grid_cells free_energy_principle
301
+ gamma_oscillations global_workspace_theory head_direction_cells hierarchical_processing
302
+ higher_order_theories hippocampal_formation homeostatic_plasticity integrated_information_theory
303
+ long_term_depression long_term_potentiation metaplasticity neural_coding
304
+ neural_decoding neural_manifolds neuromodulation place_cells
305
+ population_coding predictive_coding predictive_processing rate_coding
306
+ serotonin_modulation sharp_wave_ripples sleep_replay sparse_coding
307
+ spike_timing_dependent_plasticity temporal_coding thalamic_reticular_nucleus theta_oscillations
308
+ ```
309
+ </details>
310
+
311
+ ### Foundational researchers & concepts honored in the corpus
312
+
313
+ **Hebb** (Hebbian learning), **Bi & Poo** (STDP timing curves), **Friston** (free energy / active inference), **BuzsΓ‘ki** (hippocampal sharp-wave ripples, replay), **Moser & Moser** (grid cells), **Hodgkin & Huxley** (axon dynamics), **Izhikevich** (spike models), **Tononi** (integrated information), **Baars/Dehaene** (global workspace), **O'Keefe** (place cells), **Kandel** (memory consolidation), plus neuromodulators (dopamine RPE, serotonin, oxytocin, vasopressin) and glial biology (astrocytes, microglia, oligodendrocytes, Schwann cells).
314
 
315
  ---
316
 
317
+ ## How to Use
318
+
319
+ ### Install
320
 
321
  ```bash
322
+ git clone https://github.com/AFKmoney/fractus-cte.git
323
+ cd fractus-cte
324
+ pip install torch numpy tokenizers matplotlib fastapi uvicorn pydantic
 
325
  ```
326
 
327
+ ### Run tests
328
+
329
+ ```bash
330
+ pytest tests/ -q
331
+ ```
332
+
333
+ ### Train on CPU (progressive growth)
334
+
335
+ ```bash
336
+ python scripts/train_progressive.py --paliers 0,1,2,3 --accumulation-steps 8
337
+ ```
338
+
339
+ ### Train the 1B (GPU, sharded)
340
+
341
+ ```bash
342
+ # Phase-2 1B on 8 GPUs β€” resume from HF x8run (do not restart at token 0)
343
+ # 1) torch >= 2.11+cu128 on RTX 5090
344
+ # 2) download checkpoints/x8run/*.pt + RESUME_gpu*.json
345
+ # 3) download tokenized/phase2/shard_phase2_gpu{i}.npy, symlink to data/shard_gpu{i}.npy
346
+ CUDA_VISIBLE_DEVICES=$i GPU_ID=$i START_TOKEN=<from RESUME_gpu$i.json> \
347
+ BATCH=4 SEQ=128 CE_CHUNK=2048 FRACTUS_ATTN_IMPL=chunked BLOCK_CKPT=1 \
348
+ SS_RATE=1.0 P0=1 COMPILE=0 PROBE_EVERY=0 \
349
+ CKPT_IN=checkpoints/x8run/fractus_1b_gpu$i.pt \
350
+ CKPT_OUT=checkpoints/x8run/fractus_1b_gpu$i.pt \
351
+ SHARD=data/shard_gpu$i.npy \
352
+ python -u scripts/fast4gpu_boost_v4.py
353
+ ```
354
+ Speech metric (offline, not in the trainer loop): unique@40 greedy PREFIX. Gate β‰ˆ mean unique β‰₯ 20 and not a single-token lock.
355
+
356
+ ### Use the agent
357
+
358
+ ```python
359
+ from fractus.continuous_engine import ContinuousThoughtEngine
360
+ from fractus.memory import PersistentMemory
361
+ from fractus.tokenizer import FractusTokenizer
362
+
363
+ # Build the brain
364
+ engine = ContinuousThoughtEngine(
365
+ vocab_size=50257, d_model=128, n_heads=2, d_head=64,
366
+ n_layers=2, n_levels=2, n_oscillators=8, coupling_rank=4,
367
+ n_experts=4, top_k=2, expert_d_ff=128, siren_rank=32)
368
+
369
+ # Give it memory
370
+ memory = PersistentMemory(d_model=128, path="~/.fractus/memory.pt")
371
+ engine.attach_memory(memory)
372
+
373
+ # Think
374
+ engine.reset_thought(batch_size=1)
375
+ logits, confidence = engine.tick(torch.tensor([42]))
376
+ print(f"Confidence: {confidence.item():.2f}")
377
+
378
+ # Vision: drive thought with an image patch vector (no tokenizer)
379
+ patch_vec = torch.randn(1, 128) # any (B, d_model) embedding
380
+ logits, confidence = engine.tick_vec(patch_vec)
381
+ ```
382
 
383
  ---
384
 
385
+ ## The Growth Path
386
 
387
+ | Stage | Size | Blocks | Experts | What it can do |
388
+ |---|---|---|---|---|
389
+ | Palier 0 | 6.6M | 1 | 4 | Learn basic patterns |
390
+ | Palier 1 | 25M | 2 | 8 | Simple text generation |
391
+ | Palier 2 | 120M | 4 | 16 | Coherent fragments |
392
+ | Palier 3 | 350M | 8 | 32 | Decent text quality |
393
+ | **Palier 4** | **1B** | **16** | **128** | **Full language model (training now)** |
394
+
395
+ Each stage inherits the previous one's knowledge via zero-padded warm starts. The model never starts from zero.
396
 
397
  ---
398
 
399
+ ## Repository Layout
400
 
401
  ```
402
+ fractus-cte/
403
+ β”œβ”€β”€ fractus/
404
+ β”‚ β”œβ”€β”€ continuous_engine.py ← The brain (CTE + CTEBlock)
405
+ β”‚ β”œβ”€β”€ memory.py ← Cross-session persistent memory
406
+ β”‚ β”œβ”€β”€ cognitive_modes.py ← Unsupervised mental state detection
407
+ β”‚ β”œβ”€β”€ grow.py ← Progressive growth (width + depth + experts)
408
+ β”‚ β”œβ”€β”€ rag.py ← Knowledge base + plugins + metacognition
409
+ β”‚ β”œβ”€β”€ tokenizer.py ← GPT-2 BPE tokenizer
410
+ β”‚ β”œβ”€β”€ nn/
411
+ β”‚ β”‚ β”œβ”€β”€ moe.py ← PhaseRoutedMoE (sparse, low-rank)
412
+ β”‚ β”‚ β”œβ”€β”€ attention.py ← Multi-level causal linear attention + carry
413
+ β”‚ β”‚ β”œβ”€β”€ phase_ode.py ← Kuramoto RK4 oscillators
414
+ β”‚ β”‚ └── lazy_siren.py ← Low-rank weight storage
415
+ β”‚ └── train/online.py ← Online trainer
416
+ β”œβ”€β”€ tests/
417
+ β”œβ”€β”€ scripts/ fast_tokenize, shard_corpus, launch_4gpu,
418
+ β”‚ fast4gpu_boost_v4, generate_aligned, hourly HF x8run sync
419
+ β”œβ”€β”€ checkpoints/ per-GPU + merged + recovery alias
420
+ β”œβ”€β”€ docs/ run logs, surgeries, trusted loss, scaling
421
+ β”œβ”€β”€ space/ HF Space demo
422
+ β”œβ”€β”€ Fractus_White_Paper_v2.md White paper v2.0
423
+ └── arxiv/ LaTeX source for arXiv submission
424
  ```
425
+
426
+ ---
427
+
428
+ ## Key Concepts
429
+
430
+ **Tick**: one step of thinking. The engine processes an observation, updates its thought state through all blocks, and optionally emits output.
431
+
432
+ **Thought state**: a vector that persists across ticks β€” the engine's "consciousness."
433
+
434
+ **Chunk**: 32 tokens processed in one forward pass. The thought state and per-block attention state carry between chunks.
435
+
436
+ **Expert**: a small low-rank network (`W = scaleΒ·U@Vα΅€`) that specializes in certain thoughts. Only 2 of 128 active per token.
437
+
438
+ **Kuramoto clock**: coupled oscillators producing phase vectors that route tokens to experts β€” learned end-to-end since the routing surgery.
439
+
440
+ **`tick_vec`**: the multimodal tick β€” feed any precomputed `(B, d_model)` vector (image patches, embeddings from any encoder) straight into the thought state without tokens. This is how vision plugs in.
441
+
442
+ **Surgery**: patching code around a preserved checkpoint mid-training, resuming at the exact token offset. Weights are never wiped for a routing, objective, or decode bug.
443
+
444
+ ---
445
+
446
+ ## Limitations (stated plainly)
447
+
448
+ - Generation is not yet coherent English β€” word-level repetition loops / lexical noise. Exposure bias is being addressed by scheduled sampling; AR is the metric to watch.
449
+ - GPT-2 vocab dominates parameters (81% at d=768) β€” vocab reduction is a known lever.
450
+ - "Remembers forever" and "grows on its own" describe the architecture's design; no independent benchmarks are provided.
451
+ - This is a research artifact and a live training run, not a production assistant.
452
+
453
+ ---
454
+
455
+ ## License
456
+
457
+ MIT. Fractus belongs to you, not to a corporation.
458
+
459
+ ## Author
460
+
461
+ **Philippe-Antoine Robert** β€” 2026 β€” rpa.tu@proton.me
462
+
463
+ ## Links
464
+
465
+ - **GitHub:** [github.com/AFKmoney/fractus-cte](https://github.com/AFKmoney/fractus-cte)
466
+ - **HuggingFace Model:** [huggingface.co/thefinalboss/fractus-cte](https://huggingface.co/thefinalboss/fractus-cte)
467
+ - **HuggingFace Datasets:** [huggingface.co/datasets/thefinalboss/fractus-datasets](https://huggingface.co/datasets/thefinalboss/fractus-datasets)
468
+ - **White Paper:** [Fractus_White_Paper_v2.md](Fractus_White_Paper_v2.md) / [PDF](Fractus_White_Paper.pdf)
469
+ - **Run logs:** [MASTER_RUN_LOG](docs/MASTER_RUN_LOG.md) Β· [TRAINING_LOG_1B](docs/TRAINING_LOG_1B.md) Β· [DISCOVERY_LOG](docs/DISCOVERY_LOG.md)
470
+ - **arXiv source:** [arxiv/main.tex](arxiv/main.tex)
471
+