docs: QuickPod 8x5090 exact-token resume 2026-08-28
Browse files
README.md
CHANGED
|
@@ -1,127 +1,471 @@
|
|
| 1 |
---
|
| 2 |
license: mit
|
| 3 |
language:
|
| 4 |
-
- en
|
| 5 |
-
- fr
|
| 6 |
tags:
|
| 7 |
-
- continuous-thought
|
| 8 |
-
-
|
| 9 |
-
-
|
| 10 |
-
-
|
| 11 |
-
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 12 |
library_name: pytorch
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 13 |
---
|
| 14 |
|
| 15 |
-
# Fractus
|
| 16 |
|
| 17 |
-
**
|
| 18 |
|
| 19 |
-
Fractus
|
| 20 |
|
| 21 |
-
|
| 22 |
-
|
| 23 |
-
|
| 24 |
-
|
| 25 |
-
|
| 26 |
-
|
| 27 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 28 |
|
| 29 |
-
|
| 30 |
-
|
| 31 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 32 |
|
| 33 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 34 |
|
| 35 |
---
|
| 36 |
|
| 37 |
-
## Architecture (
|
| 38 |
|
| 39 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 40 |
|---|---|
|
| 41 |
-
|
|
| 42 |
-
|
|
| 43 |
-
|
|
| 44 |
-
|
|
| 45 |
-
|
|
| 46 |
-
|
|
| 47 |
-
|
|
| 48 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 49 |
|
| 50 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 51 |
|
| 52 |
---
|
| 53 |
|
| 54 |
-
##
|
| 55 |
|
| 56 |
-
**
|
| 57 |
-
|
| 58 |
-
|
| 59 |
-
-
|
| 60 |
-
|
| 61 |
-
|
| 62 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 63 |
|
| 64 |
-
|
| 65 |
-
- Attention state \(S\) still looks like \(S_t = S_{t-1} + k\otimes v\) β **no learned decay** yet (RWKV/RetNet/Mamba all added one)
|
| 66 |
-
- Mean-merge of independent MoE shards β DDP; expert #k is not aligned across GPUs
|
| 67 |
-
- Kuramoto decision is still a **circle** (need \(\mathrm{MI}(\bar\theta, \text{token})\) vs tick)
|
| 68 |
-
- No published **matched-compute PPL** vs a vanilla transformer on a clean held-out
|
| 69 |
-
- CARRY length-1 generation is not the trained graph (PREFIX is the language gate)
|
| 70 |
-
- Eval must be split by **source**, not by line (identity text in the corpus)
|
| 71 |
|
| 72 |
-
|
| 73 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 74 |
|
| 75 |
---
|
| 76 |
|
| 77 |
-
##
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 78 |
|
| 79 |
-
```python
|
| 80 |
-
from fractus.continuous_engine import ContinuousThoughtEngine
|
| 81 |
-
eng = ContinuousThoughtEngine.from_pretrained("checkpoints/FRACTUS_1B_PHASE2_FROZEN_MERGED.pt")
|
| 82 |
-
# .pt = weights + live state. fractus/ = the body that ticks.
|
| 83 |
```
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 84 |
|
| 85 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 86 |
|
| 87 |
---
|
| 88 |
|
| 89 |
-
##
|
|
|
|
|
|
|
| 90 |
|
| 91 |
```bash
|
| 92 |
-
|
| 93 |
-
|
| 94 |
-
|
| 95 |
-
python -u scripts/fast4gpu_boost_v3.py
|
| 96 |
```
|
| 97 |
|
| 98 |
-
|
| 99 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 100 |
|
| 101 |
---
|
| 102 |
|
| 103 |
-
##
|
| 104 |
|
| 105 |
-
|
|
| 106 |
-
|---|---|
|
| 107 |
-
|
|
| 108 |
-
|
|
| 109 |
-
|
|
| 110 |
-
|
|
| 111 |
-
|
|
| 112 |
-
|
| 113 |
-
|
| 114 |
|
| 115 |
---
|
| 116 |
|
| 117 |
-
##
|
| 118 |
|
| 119 |
```
|
| 120 |
-
|
| 121 |
-
|
| 122 |
-
|
| 123 |
-
|
| 124 |
-
|
| 125 |
-
|
| 126 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 127 |
```
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
---
|
| 2 |
license: mit
|
| 3 |
language:
|
| 4 |
+
- en
|
| 5 |
+
- fr
|
| 6 |
tags:
|
| 7 |
+
- continuous-thought-engine
|
| 8 |
+
- cognitive-agent
|
| 9 |
+
- mixture-of-experts
|
| 10 |
+
- kuramoto
|
| 11 |
+
- self-modifying
|
| 12 |
+
- progressive-growth
|
| 13 |
+
- decentralized-ai
|
| 14 |
+
- personal-ai
|
| 15 |
+
- neuroscience
|
| 16 |
library_name: pytorch
|
| 17 |
+
pipeline_tag: text-generation
|
| 18 |
+
models:
|
| 19 |
+
- thefinalboss/fractus-cte
|
| 20 |
+
datasets:
|
| 21 |
+
- thefinalboss/fractus-datasets
|
| 22 |
---
|
| 23 |
|
| 24 |
+
# Fractus CTE
|
| 25 |
|
| 26 |
+
**A living AI that thinks continuously, remembers forever, and grows on its own.**
|
| 27 |
|
| 28 |
+
**Fractus is NOT a transformer.** It's a Continuous Cognitive Agent β a dynamical system that maintains a persistent thought state, advances it tick by tick through 16 blocks, and routes via Kuramoto oscillator phases. The checkpoint is a living seed: it never freezes, grows at runtime, and trains forever.
|
| 29 |
|
| 30 |
+
> **Status (2026-08-28):** Training live again on **QuickPod 8Γ RTX 5090** after a RunPod death. Exact-token resume from `checkpoints/x8run` (GPU0 at 79.9M / 430M of this Phase-2 pass). Trainer: `scripts/fast4gpu_boost_v4.py`, SS_RATE=1.0, BATCH=4. Speech gate is **unique@40 greedy PREFIX** β last offline probe was NO-GO (3/12/1/6). Not a finished assistant. See [Live Training](#live-training-status) and [`docs/2026-08-28-QUICKPOD-RESUME.md`](docs/2026-08-28-QUICKPOD-RESUME.md).
|
| 31 |
+
|
| 32 |
+
## Quick Start
|
| 33 |
+
|
| 34 |
+
```bash
|
| 35 |
+
git clone https://github.com/AFKmoney/fractus-cte.git
|
| 36 |
+
cd fractus-cte && pip install torch numpy tokenizers
|
| 37 |
+
|
| 38 |
+
# Load the trained 1B and generate
|
| 39 |
+
python -c "
|
| 40 |
+
from fractus.continuous_engine import ContinuousThoughtEngine
|
| 41 |
+
engine = ContinuousThoughtEngine.from_pretrained('checkpoints/fractus_1b_gpu3.pt')
|
| 42 |
+
import torch
|
| 43 |
+
logits, confidence = engine.tick(torch.tensor([42]))
|
| 44 |
+
print(f'Fractus is thinking. Confidence: {confidence.item():.2f}')
|
| 45 |
+
"
|
| 46 |
+
```
|
| 47 |
+
|
| 48 |
+
The `.pt` checkpoint contains the full model (weights + dynamic state). You need this repo's code to run it β Fractus is a custom architecture, not a transformer. Checkpoints are on [HF](https://huggingface.co/thefinalboss/fractus-cte).
|
| 49 |
|
| 50 |
+

|
| 51 |
+

|
| 52 |
+

|
| 53 |
+

|
| 54 |
+

|
| 55 |
+

|
| 56 |
+

|
| 57 |
|
| 58 |
+
---
|
| 59 |
+
|
| 60 |
+
## What is Fractus?
|
| 61 |
+
|
| 62 |
+
Fractus is not a chatbot. It's not GPT. It's not a transformer.
|
| 63 |
+
|
| 64 |
+
Fractus is a **Continuous Cognitive Agent** β an AI that works like a brain, not a calculator. Instead of processing input β output in one pass, Fractus **ticks** like a biological system: it maintains a persistent thought state, advances it through multiple blocks of processing, remembers everything across sessions, and can grow new capacity by itself.
|
| 65 |
+
|
| 66 |
+
### What makes it different from GPT/Claude?
|
| 67 |
+
|
| 68 |
+
| | GPT-4 / Claude | Fractus |
|
| 69 |
+
|---|---|---|
|
| 70 |
+
| **Thinking** | One pass, done | Continuous ticks (like a heartbeat) |
|
| 71 |
+
| **Memory** | Forgets when context window fills | Remembers forever (survives restarts) |
|
| 72 |
+
| **Learning** | Retrain from scratch ($$$) | Learns from every interaction |
|
| 73 |
+
| **Growth** | Fixed size forever | Grows new experts at runtime |
|
| 74 |
+
| **Mental states** | One mode always | Shifts between cognitive modes |
|
| 75 |
+
| **Where it runs** | Corporate cloud | Your machine |
|
| 76 |
+
| **Training** | Fixed, done once | Perpetual, never stops |
|
| 77 |
|
| 78 |
---
|
| 79 |
|
| 80 |
+
## Architecture (1.05B total / ~119M active per token)
|
| 81 |
|
| 82 |
+
```
|
| 83 |
+
ContinuousThoughtEngine
|
| 84 |
+
βββ d_model=1280, 16 layers, 16 heads
|
| 85 |
+
βββ FractalLinearAttention multi-level causal linear attention (O(L)),
|
| 86 |
+
β carry state (S, z) persists across chunks
|
| 87 |
+
βββ Kuramoto phase clock RK4-integrated oscillators β phase vectors
|
| 88 |
+
β (learned end-to-end; feeds routing)
|
| 89 |
+
βββ PhaseRoutedMoE 128 experts, top-2 active per token,
|
| 90 |
+
β von Mises gate over phases, Farey-sequence
|
| 91 |
+
β expert phases, load-balance loss in objective
|
| 92 |
+
βββ Low-rank experts W = scale Β· U@Vα΅ (rank 64) β 64Γ less compute
|
| 93 |
+
βββ Tied embedding/head GPT-2 BPE vocab (50257)
|
| 94 |
+
βββ Persistent thought state residual stream carried across ticks
|
| 95 |
+
```
|
| 96 |
+
|
| 97 |
+
**Parameter accounting:** ~1.05B total, but only ~118.8M are active per token (dense attention + top-2 of 128 sparse experts) β ~183.1M including the tied embedding. An 11.3% sparsity ratio. This is the basis of the adapted scaling target below.
|
| 98 |
+
|
| 99 |
+
### The 12 Building Blocks
|
| 100 |
+
|
| 101 |
+
| Block | What it does |
|
| 102 |
|---|---|
|
| 103 |
+
| **Continuous Thought Engine** | The brain β thinks tick by tick through 16 blocks |
|
| 104 |
+
| **Persistent Memory** | Vector bank surviving sessions, cosine recall, 5% blend injection, salience head gates impact |
|
| 105 |
+
| **Cognitive Modes** | Mental states discovered unsupervised (k-means on Kuramoto phase features): focused, creative, exploratory |
|
| 106 |
+
| **RAG Knowledge Base** | Learns facts instantly β no retraining needed |
|
| 107 |
+
| **Cognitive Plugins** | Hot-swappable modes: analyst, coder, creative, teacher |
|
| 108 |
+
| **MetaCognition** | Decides its own actions: retrieve, learn, generate |
|
| 109 |
+
| **Progressive Growth** | Grows from 6M to 1B+ params, palier by palier (`maybe_grow`: width + depth + experts) |
|
| 110 |
+
| **Self-Modification** | Adds new experts at runtime when routing is imbalanced (zero-init, placed near the dominant expert) |
|
| 111 |
+
| **PhaseRoutedMoE** | Sparse experts routed by oscillator phases |
|
| 112 |
+
| **Kuramoto Clock** | A dynamical system that drives routing decisions |
|
| 113 |
+
| **Online Trainer** | Learns continuously, one chunk at a time |
|
| 114 |
+
| **Vision (`tick_vec`)** | Multimodal input path β image patches drive the engine directly, bypassing the token embedding (CIFAR "eyes" prototype trained while the text run continued) |
|
| 115 |
+
|
| 116 |
+
---
|
| 117 |
+
|
| 118 |
+
## Live Training Status (28 August 2026)
|
| 119 |
+
|
| 120 |
+
**Hardware:** QuickPod 8Γ RTX 5090. One independent Python process per GPU. No DDP / no gradient sync. Consolidation = mean-merge of the 8 `.pt` when we need a single brain.
|
| 121 |
|
| 122 |
+
**This pass (Phase 2 shards):**
|
| 123 |
+
- 8 GPT-2 BPE int32 memmaps, `~429,896,462` tokens/GPU β dataset: [thefinalboss/fractus-datasets](https://huggingface.co/datasets/thefinalboss/fractus-datasets) `tokenized/phase2/shard_phase2_gpu{i}.npy`
|
| 124 |
+
- Resume point **2026-08-28 14:40 UTC** after RunPod SSH death: GPU0 **79,922,176** tokens (~18.6% of this pass). Other GPUs 76.2Mβ78.4M. Exact JSON: `checkpoints/x8run/RESUME_gpu{i}.json`
|
| 125 |
+
- Trainer: `scripts/fast4gpu_boost_v4.py`
|
| 126 |
+
- Config: **BATCH=4**, SEQ=128, LR 7e-4, SGD+momentum 0.9, bf16, TF32, compile off, `FRACTUS_ATTN_IMPL=chunked`, `BLOCK_CKPT=1`, **SS_RATE=1.0**, SS_PROB 0.2β0.5 over 50M tokens, REPEAT_COEF=0.1, P0 on, **PROBE_EVERY=0** (live unique@40 probe crashed the first QuickPod launch β probes are offline)
|
| 127 |
+
- Signals: `tf`/`ema_tf`, `ss`/`ema_ss`, `rep`/`ema_rep`, `lb` (~14)
|
| 128 |
+
- Throughput after resume: ~840β940 tok/s/GPU. Remainder of this pass β **4.1β4.5 days** if 8 GPUs stay up
|
| 129 |
+
- **Torch on RTX 5090:** image default `2.2.1+cu121` has **no Blackwell kernels**. Required: **torch β₯ 2.11 + cu128**
|
| 130 |
+
|
| 131 |
+
**Crash recovery:** hourly upload of `checkpoints/x8run/fractus_1b_gpu{i}.pt` + `RESUME_gpu{i}.json`. That folder is the source of truth. See [`docs/2026-08-28-QUICKPOD-RESUME.md`](docs/2026-08-28-QUICKPOD-RESUME.md).
|
| 132 |
+
|
| 133 |
+
**Speech gate (not the loss):** `unique@40` greedy PREFIX β 40 generated tokens, count distinct ids, no ban, no temperature. Last offline numbers: Hello=3, Once upon a time=12, meaning of life=1, Fractus is=6. **NO-GO.** Teacher-force CE going down does **not** mean Fractus speaks. Do not cite banned-decode diversity as speech.
|
| 134 |
+
|
| 135 |
+
## Vision β the CIFAR "eyes" prototype
|
| 136 |
+
|
| 137 |
+
Fractus is not text-only. The engine exposes **`tick_vec(obs_vec)`**: a multimodal entry point that accepts a precomputed `(B, d_model)` vector and injects it directly into the residual thought state β **bypassing the token embedding entirely**. Any modality that can be embedded into `d_model` dimensions can drive continuous thought: images, audio, sensor streams.
|
| 138 |
+
|
| 139 |
+
```
|
| 140 |
+
image β PatchEmbed β (B, d_model) patch vectors
|
| 141 |
+
β
|
| 142 |
+
βΌ
|
| 143 |
+
engine.tick_vec(patch) β no tokenizer involved
|
| 144 |
+
β
|
| 145 |
+
thought state advances through all 16 blocks
|
| 146 |
+
(same Kuramoto routing + MoE as text)
|
| 147 |
+
```
|
| 148 |
+
|
| 149 |
+
**Proof of concept β CIFAR-10 eyes:** a small CTE + PatchEmbed stack (`fractus/nn/vision.py`) was trained on CPU on real CIFAR-10 images (`fractus_eyes_cifar_final.pt`), demonstrating that image patches can drive the continuous-thought loop. Crucially, this ran as a **parallel track**: the 8-GPU text run was never interrupted for eyes work β GPU text digestion and CPU vision learning proceed simultaneously on the same living system.
|
| 150 |
+
|
| 151 |
+
This mirrors the biological premise: eyes evolved as a peripheral sense feeding a central dynamical brain, not as a core feature of it. The text 1B is the brain; vision is a front-end that plugs into `tick_vec`.
|
| 152 |
+
|
| 153 |
+
---
|
| 154 |
+
|
| 155 |
+
### Adapted Chinchilla target
|
| 156 |
+
|
| 157 |
+
Standard Chinchilla (20 tokens/param) assumes dense transformers where all parameters train on every token. Fractus violates that: only ~119M of 1.05B params are active per token, and the Kuramoto clock trains end-to-end now. The adapted target is **20 Γ active params β 2.4β3.7B tokens**; with a warm-started checkpoint, the practical target is **~2.5β3B tokens** β which is why the phase-2 corpus is 3.44B. Since Fractus grows continuously (`maybe_grow`), this is an instantaneous cumulative-exposure requirement that grows with the model, not a final-stop condition. Details: [`docs/2026-08-12-fractus-chinchilla.md`](docs/2026-08-12-fractus-chinchilla.md).
|
| 158 |
|
| 159 |
---
|
| 160 |
|
| 161 |
+
## The Surgeries (mid-training interventions, no weight wipes)
|
| 162 |
|
| 163 |
+
A defining discovery of this run: **the brain (.pt) and the code are separable.** Bottlenecks were fixed by live surgery β save the checkpoint, patch the code, reload weights, resume at the exact recorded token offset. Multi-day digestion is never thrown away.
|
| 164 |
+
|
| 165 |
+
| Phase | Intervention | Result |
|
| 166 |
+
|---|---|---|
|
| 167 |
+
| **A β Initial** | Last-position-only CE, 4 independent GPUs | Loss fell; generation collapsed into single-token loops |
|
| 168 |
+
| **B β Routing surgery** | Kuramoto was frozen under `no_grad` (order parameter r β 0.01β0.03); LB loss was detached β ~70% of experts dead. Fixed: gradients enabled (state kept detached for carry), CE + 0.02Β·lb, gate temp 1.0 β 2.5, omega scale Γ4 | lb β 14 live on all GPUs; experts alive |
|
| 169 |
+
| **C β Dense CE** | Replaced last-position CE (1 target / 128 tokens) with CE over all positions | Sharp loss drop; token-to-token chaining enforced |
|
| 170 |
+
| **D β Decode surgery** | Phase/thought noise, frequency penalties, cycle bans, forced escape tokens | Loop lock broken; still no coherent English |
|
| 171 |
+
| **E β Train/gen mismatch** | Training used causal attention + RK4 Kuramoto; generation used a simpler single-tick Euler path. Aligned decode via `generate_aligned.py` | Decode now matches the training path |
|
| 172 |
+
| **F β Loss recalibration** | Cumulative-average CE was misleading near 2.0 β batch CE + EMA; LR 1e-3 β 5e-4 | Trustworthy metrics |
|
| 173 |
+
| **G β Scheduled sampling** | Two-pass training: TF CE + LB, plus SS steps mixing model samples into inputs | Live; now SS_RATE=1.0 on boost_v4 |
|
| 174 |
+
| **H β Pod death / exact resume** | RunPod SSH died 2026-08-28. Reloaded 8Γ `.pt` + manifests on a new QuickPod 8Γ5090 at the recorded `start_token_next`. Torch upgraded 2.2β2.11+cu128 for Blackwell | Weights kept. Pass continues from ~80M/430M |
|
| 175 |
|
| 176 |
+
Full logs: [`docs/MASTER_RUN_LOG.md`](docs/MASTER_RUN_LOG.md), [`docs/DISCOVERY_LOG.md`](docs/DISCOVERY_LOG.md), [`docs/OPERABILITY_MIDTRAIN.md`](docs/OPERABILITY_MIDTRAIN.md).
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 177 |
|
| 178 |
+
### Composable checkpoints
|
| 179 |
+
|
| 180 |
+
Because shapes stay compatible and manifests record token offsets, the following operations are proven on this run:
|
| 181 |
+
- **Parallel independent training** β N GPUs on separate shards
|
| 182 |
+
- **Mean-merge** β per-GPU checkpoints averaged into one unified model that still generates (424/440 tensors on the 4-GPU merge; stateful buffers reset)
|
| 183 |
+
- **Iterative fusion** β train β merge β train cycles; the model absorbs compatible checkpoints and keeps going
|
| 184 |
+
- **Exact-token resume** β across pod reboots and driver crashes
|
| 185 |
+
- **Compositional growth β structural growth** β weight averaging vs `maybe_grow` paliers
|
| 186 |
+
|
| 187 |
+
Details: [`docs/COMPOSABILITY_AND_SURGERY.md`](docs/COMPOSABILITY_AND_SURGERY.md).
|
| 188 |
|
| 189 |
---
|
| 190 |
|
| 191 |
+
## Trusted Loss (reading the numbers)
|
| 192 |
+
|
| 193 |
+
Three metrics, three different questions:
|
| 194 |
+
|
| 195 |
+
| Metric | What it measures | Where |
|
| 196 |
+
|---|---|---|
|
| 197 |
+
| `ema_tf` | CE with ground-truth history + continuous internal state β is the model digesting data? | live |
|
| 198 |
+
| `ema_ss` | CE after scheduled sampling mixes model-generated tokens in β partial free-run robustness | live |
|
| 199 |
+
| **AR** | Warm on 32 true tokens, greedily free-run 32 steps, CE vs truth β actual generation quality | offline |
|
| 200 |
+
|
| 201 |
+
**A single CE number cannot represent both teacher-forced learning and free-run generation.** Live ema_tf after the QuickPod resume is mid-teens on the lead GPUs (first steps inflate EMA β wait). AR stays the offline generation metric and last unique@40 was NO-GO. Random-vocab CE β 10.82; AR in the thousands (random guessing over the GPT-2 vocab is ~10.82): free-running compounds every error, and teacher forcing always supplies the correct past. Trust `ema_tf`/`ema_ss` for learning progress, AR + text probes for generation progress. The convergence signal is AR falling toward the ss/tf order of magnitude, plus readable output. Details: [`docs/TRUSTED_LOSS.md`](docs/TRUSTED_LOSS.md).
|
| 202 |
+
|
| 203 |
+
---
|
| 204 |
+
|
| 205 |
+
## Research Results (Honest)
|
| 206 |
+
|
| 207 |
+
**Refuted:**
|
| 208 |
+
- **EDT** (Expert Decoupled Training) β all 5 variants ~19β20% worse than from-scratch; MSE objective misaligned with CE, router ignores half the experts.
|
| 209 |
+
- **Forward-Forward** (Hinton 2022) β NLL rose from 124 to 221; local objectives can't replace global backprop here.
|
| 210 |
+
|
| 211 |
+
**Validated:**
|
| 212 |
+
- **Progressive growth** β warm start converges faster; trained through palier 3 (350M, loss 23.0, ~5 days CPU).
|
| 213 |
+
- **Sparse low-rank MoE** β 2/128 experts = 64Γ less compute.
|
| 214 |
+
- **Open-heart operability** β live surgery on a training model works (see above).
|
| 215 |
+
- **Routing pathology as a first-class debug target** β expert-hit histograms and the Kuramoto order parameter catch failures that loss curves hide (they caught the frozen clock and the dead experts).
|
| 216 |
+
|
| 217 |
+
**Training optimizations (measured):** tied head (~1.1Γ), head-partial training (~2Γ), sparse gathered low-rank MoE (8Γ at 16 experts, 64Γ at 128), detached-state Kuramoto, gradient accumulation (~1.4Γ), SGD+momentum over AdamW (~1.37Γ), batching (335 β 1345 tok/s at B=8), bf16 (~2Γ). Combined CPU: ~336Γ over naive. Measured on CPU: 707 tok/s single-stream, **1345 tok/s batched**.
|
| 218 |
+
|
| 219 |
+
---
|
| 220 |
+
|
| 221 |
+
## Datasets (4.2 Billion Tokens)
|
| 222 |
+
|
| 223 |
+
Source of truth: [huggingface.co/datasets/thefinalboss/fractus-datasets](https://huggingface.co/datasets/thefinalboss/fractus-datasets).
|
| 224 |
+
|
| 225 |
+
| Dataset | Tokens | Content |
|
| 226 |
+
|---|---|---|
|
| 227 |
+
| **neuro-paradigms-1b** | **~1B** | 100 neuroscience β software architecture paradigms (300 chunked files) |
|
| 228 |
+
| **neuro-code-math** | **~900M** | Neuro-inspired coding, mathematics, algorithms (incl. 40 applied-neuroscience topics) |
|
| 229 |
+
| **cognitive-skills** | **~780M** | Coding, reasoning, speaking, thinking, understanding |
|
| 230 |
+
| **fractus-generated-corpus** | **340M** | Bilingual FR/EN generated by Fractus ontology engine |
|
| 231 |
+
| **paradigms-full** | **191M** | 140 paradigms (neuroscience, CS, architecture) |
|
| 232 |
+
| **gutenberg-esoteric** | **~58M** | 487 public-domain esoteric / masonic / hermetic books |
|
| 233 |
+
| **neuro-arch-full** | **86M** | 60 neuroscience paradigms (neuro-software-architecture) |
|
| 234 |
+
| **all-github-repos** | **54M+** | 80+ repos (public + private, secret-filtered) |
|
| 235 |
+
| **mega-corpus-v3** | **20M** | Literature, philosophy, occult, masonry, science, medicine |
|
| 236 |
+
| **wordnet** | **3M** | 117K dictionary synset entries |
|
| 237 |
+
| **Total** | **~4.2B** | |
|
| 238 |
+
|
| 239 |
+
Tokenized streams: Phase 1 = ~1.52B tokens (`.pt` files); Phase 2 = ~3.44B tokens (8 GPT-2 BPE int32 memmap shards). Phases are kept separate to avoid re-ingesting the same ordered stream.
|
| 240 |
+
|
| 241 |
+
---
|
| 242 |
+
|
| 243 |
+
## Applied Neuroscience β the theoretical core
|
| 244 |
+
|
| 245 |
+
Fractus is a neuroscience-grounded architecture: real brain mechanisms are mapped to software/AI patterns, and that mapping is itself training data. Every entry below is present in the dataset β verified by file listing, not just claimed.
|
| 246 |
+
|
| 247 |
+
### 100 neuroscience β software-architecture paradigms (`neuro_paradigms_1b`, 300 chunked files)
|
| 248 |
+
|
| 249 |
+
Each paradigm maps a biological mechanism to an engineering pattern (e.g. *adenosine sleep pressure* β cache-stampede recovery; *myelin sheath* β caching; *hippocampal replay* β trajectory consolidation).
|
| 250 |
+
|
| 251 |
+
<details><summary><b>show all 100 paradigms</b></summary>
|
| 252 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 253 |
```
|
| 254 |
+
adenosine_sleep_pressure amygdala_prefrontal_topdown anterior_cingulate_conflict_monitor
|
| 255 |
+
apoptosis_self_destructing_service arc_gene_plasticity_marker astrocyte_tripartite_synapse
|
| 256 |
+
axon_initial_segment_trigger basal_ganglia_loop_arbitration bdnf_growth_factor_scaling
|
| 257 |
+
bergmann_glia_purkinje binaural_cross_correlation_localization brainstem_vital_functions
|
| 258 |
+
broca_area_api_generator calcium_transmitter_coupling camp_second_messenger_amplifier
|
| 259 |
+
cerebellar_forward_model cholinergic_attentional_filter circadian_gene_expression
|
| 260 |
+
climbing_fiber_error_broadcast cochlear_compressive_nonlinearity cortical_area_specialization
|
| 261 |
+
cortical_minicolumn_pipeline cortico_cortical_pathways corticotropin_releasing_hormone
|
| 262 |
+
cortisol_slow_stress_recovery critical_period_learning_rate dendritic_compartmentalization
|
| 263 |
+
endocannabinoid_retrograde enteric_glia_gut_brain ependymal_cell_barrier
|
| 264 |
+
fusiform_face_service_registry gaba_inhibitory_bus gap_junction_electrical_sync
|
| 265 |
+
ghrelin_hunger_signal glomerular_convergence_gateway glutamate_excitatory_bus
|
| 266 |
+
glycine_coagonist_modulator granule_cell_inhibitory_relay hair_cell_banks_event_clusters
|
| 267 |
+
hippocampal_4ec_loop_replay histamine_wakefulness_keeper hox_gene_service_specialization
|
| 268 |
+
hypercolumn_module_federation hypercolumn_sharding hypothalamus_homeostasis
|
| 269 |
+
insula_interoception_monitor ip3_inositol_cascade k_complex_event_trigger
|
| 270 |
+
kcc2_chloride_shift_inhibitor leptin_satiety_signal locus_coeruleus_ne_global_signal
|
| 271 |
+
melatonin_circadian_scheduler microglia_active_surveillance mitral_tufted_cell_dual_path
|
| 272 |
+
morphogen_gradient_config muller_glia_retina_repair myelin_sheath_caching
|
| 273 |
+
neural_crest_migration_deploy neuropeptide_y_stress_buffer ng2_glia_pool_renewal
|
| 274 |
+
nitric_oxide_gas_signal node_of_ranvier_bypass nrem_slow_wave_cleanup
|
| 275 |
+
nucleus_accumbens_reward_routing oligodendrocyte_myelination_dynamic orexin_stability_keeper
|
| 276 |
+
orientation_column_indexing oscillatory_phase_locking_io oxytocin_trust_protocol
|
| 277 |
+
parahippocampal_place_topology parallel_fiber_fanout_aggregation pineal_circadian_release
|
| 278 |
+
pinwheel_central_layout pituitary_master_gland posterior_parietal_integration
|
| 279 |
+
prolactin_parental_care quantal_release_batching radial_glia_neural_stem
|
| 280 |
+
radial_glial_scaffold raphe_serotonin_rate_limit rem_paradoxical_processing
|
| 281 |
+
replay_consolidation_trajectory reticular_activating_system retinotopic_data_layout
|
| 282 |
+
satellite_glial_ganglion schwann_cell_peripheral_repair sleep_pressure_forced_maintenance
|
| 283 |
+
sleep_spindle_memory_transfer slow_oscillation_sync subplate_wait_state
|
| 284 |
+
suprachiasmatic_clock synaptic_vesicle_pool synaptogenesis_service_wiring
|
| 285 |
+
tanycyte_metabolic_sensor temporal_pole_semantic_cache thalamocortical_loop_api
|
| 286 |
+
tonotopic_stream_partitioning vasopressin_loyalty_aware_routing vta_dopamine_rpe_scheduler
|
| 287 |
+
wernicke_area_api_parser
|
| 288 |
+
```
|
| 289 |
+
</details>
|
| 290 |
+
|
| 291 |
+
### 40 applied-neuroscience topics (`neuro_code_math/applied_neuroscience/`)
|
| 292 |
|
| 293 |
+
Deep dives on computational neuroscience theories β the science Fractus's design draws from.
|
| 294 |
+
|
| 295 |
+
<details><summary><b>show all 40 topics</b></summary>
|
| 296 |
+
|
| 297 |
+
```
|
| 298 |
+
active_inference axonal_computation basal_ganglia_circuits bayesian_brain
|
| 299 |
+
cerebellar_computation consolidation cortical_minicolumns cross_frequency_coupling
|
| 300 |
+
dendritic_computation dopamine_reward entorhinal_grid_cells free_energy_principle
|
| 301 |
+
gamma_oscillations global_workspace_theory head_direction_cells hierarchical_processing
|
| 302 |
+
higher_order_theories hippocampal_formation homeostatic_plasticity integrated_information_theory
|
| 303 |
+
long_term_depression long_term_potentiation metaplasticity neural_coding
|
| 304 |
+
neural_decoding neural_manifolds neuromodulation place_cells
|
| 305 |
+
population_coding predictive_coding predictive_processing rate_coding
|
| 306 |
+
serotonin_modulation sharp_wave_ripples sleep_replay sparse_coding
|
| 307 |
+
spike_timing_dependent_plasticity temporal_coding thalamic_reticular_nucleus theta_oscillations
|
| 308 |
+
```
|
| 309 |
+
</details>
|
| 310 |
+
|
| 311 |
+
### Foundational researchers & concepts honored in the corpus
|
| 312 |
+
|
| 313 |
+
**Hebb** (Hebbian learning), **Bi & Poo** (STDP timing curves), **Friston** (free energy / active inference), **BuzsΓ‘ki** (hippocampal sharp-wave ripples, replay), **Moser & Moser** (grid cells), **Hodgkin & Huxley** (axon dynamics), **Izhikevich** (spike models), **Tononi** (integrated information), **Baars/Dehaene** (global workspace), **O'Keefe** (place cells), **Kandel** (memory consolidation), plus neuromodulators (dopamine RPE, serotonin, oxytocin, vasopressin) and glial biology (astrocytes, microglia, oligodendrocytes, Schwann cells).
|
| 314 |
|
| 315 |
---
|
| 316 |
|
| 317 |
+
## How to Use
|
| 318 |
+
|
| 319 |
+
### Install
|
| 320 |
|
| 321 |
```bash
|
| 322 |
+
git clone https://github.com/AFKmoney/fractus-cte.git
|
| 323 |
+
cd fractus-cte
|
| 324 |
+
pip install torch numpy tokenizers matplotlib fastapi uvicorn pydantic
|
|
|
|
| 325 |
```
|
| 326 |
|
| 327 |
+
### Run tests
|
| 328 |
+
|
| 329 |
+
```bash
|
| 330 |
+
pytest tests/ -q
|
| 331 |
+
```
|
| 332 |
+
|
| 333 |
+
### Train on CPU (progressive growth)
|
| 334 |
+
|
| 335 |
+
```bash
|
| 336 |
+
python scripts/train_progressive.py --paliers 0,1,2,3 --accumulation-steps 8
|
| 337 |
+
```
|
| 338 |
+
|
| 339 |
+
### Train the 1B (GPU, sharded)
|
| 340 |
+
|
| 341 |
+
```bash
|
| 342 |
+
# Phase-2 1B on 8 GPUs β resume from HF x8run (do not restart at token 0)
|
| 343 |
+
# 1) torch >= 2.11+cu128 on RTX 5090
|
| 344 |
+
# 2) download checkpoints/x8run/*.pt + RESUME_gpu*.json
|
| 345 |
+
# 3) download tokenized/phase2/shard_phase2_gpu{i}.npy, symlink to data/shard_gpu{i}.npy
|
| 346 |
+
CUDA_VISIBLE_DEVICES=$i GPU_ID=$i START_TOKEN=<from RESUME_gpu$i.json> \
|
| 347 |
+
BATCH=4 SEQ=128 CE_CHUNK=2048 FRACTUS_ATTN_IMPL=chunked BLOCK_CKPT=1 \
|
| 348 |
+
SS_RATE=1.0 P0=1 COMPILE=0 PROBE_EVERY=0 \
|
| 349 |
+
CKPT_IN=checkpoints/x8run/fractus_1b_gpu$i.pt \
|
| 350 |
+
CKPT_OUT=checkpoints/x8run/fractus_1b_gpu$i.pt \
|
| 351 |
+
SHARD=data/shard_gpu$i.npy \
|
| 352 |
+
python -u scripts/fast4gpu_boost_v4.py
|
| 353 |
+
```
|
| 354 |
+
Speech metric (offline, not in the trainer loop): unique@40 greedy PREFIX. Gate β mean unique β₯ 20 and not a single-token lock.
|
| 355 |
+
|
| 356 |
+
### Use the agent
|
| 357 |
+
|
| 358 |
+
```python
|
| 359 |
+
from fractus.continuous_engine import ContinuousThoughtEngine
|
| 360 |
+
from fractus.memory import PersistentMemory
|
| 361 |
+
from fractus.tokenizer import FractusTokenizer
|
| 362 |
+
|
| 363 |
+
# Build the brain
|
| 364 |
+
engine = ContinuousThoughtEngine(
|
| 365 |
+
vocab_size=50257, d_model=128, n_heads=2, d_head=64,
|
| 366 |
+
n_layers=2, n_levels=2, n_oscillators=8, coupling_rank=4,
|
| 367 |
+
n_experts=4, top_k=2, expert_d_ff=128, siren_rank=32)
|
| 368 |
+
|
| 369 |
+
# Give it memory
|
| 370 |
+
memory = PersistentMemory(d_model=128, path="~/.fractus/memory.pt")
|
| 371 |
+
engine.attach_memory(memory)
|
| 372 |
+
|
| 373 |
+
# Think
|
| 374 |
+
engine.reset_thought(batch_size=1)
|
| 375 |
+
logits, confidence = engine.tick(torch.tensor([42]))
|
| 376 |
+
print(f"Confidence: {confidence.item():.2f}")
|
| 377 |
+
|
| 378 |
+
# Vision: drive thought with an image patch vector (no tokenizer)
|
| 379 |
+
patch_vec = torch.randn(1, 128) # any (B, d_model) embedding
|
| 380 |
+
logits, confidence = engine.tick_vec(patch_vec)
|
| 381 |
+
```
|
| 382 |
|
| 383 |
---
|
| 384 |
|
| 385 |
+
## The Growth Path
|
| 386 |
|
| 387 |
+
| Stage | Size | Blocks | Experts | What it can do |
|
| 388 |
+
|---|---|---|---|---|
|
| 389 |
+
| Palier 0 | 6.6M | 1 | 4 | Learn basic patterns |
|
| 390 |
+
| Palier 1 | 25M | 2 | 8 | Simple text generation |
|
| 391 |
+
| Palier 2 | 120M | 4 | 16 | Coherent fragments |
|
| 392 |
+
| Palier 3 | 350M | 8 | 32 | Decent text quality |
|
| 393 |
+
| **Palier 4** | **1B** | **16** | **128** | **Full language model (training now)** |
|
| 394 |
+
|
| 395 |
+
Each stage inherits the previous one's knowledge via zero-padded warm starts. The model never starts from zero.
|
| 396 |
|
| 397 |
---
|
| 398 |
|
| 399 |
+
## Repository Layout
|
| 400 |
|
| 401 |
```
|
| 402 |
+
fractus-cte/
|
| 403 |
+
βββ fractus/
|
| 404 |
+
β βββ continuous_engine.py β The brain (CTE + CTEBlock)
|
| 405 |
+
β βββ memory.py β Cross-session persistent memory
|
| 406 |
+
β βββ cognitive_modes.py β Unsupervised mental state detection
|
| 407 |
+
β βββ grow.py β Progressive growth (width + depth + experts)
|
| 408 |
+
β βββ rag.py β Knowledge base + plugins + metacognition
|
| 409 |
+
β βββ tokenizer.py β GPT-2 BPE tokenizer
|
| 410 |
+
β βββ nn/
|
| 411 |
+
β β βββ moe.py β PhaseRoutedMoE (sparse, low-rank)
|
| 412 |
+
β β βββ attention.py β Multi-level causal linear attention + carry
|
| 413 |
+
β β βββ phase_ode.py β Kuramoto RK4 oscillators
|
| 414 |
+
β β βββ lazy_siren.py β Low-rank weight storage
|
| 415 |
+
β βββ train/online.py β Online trainer
|
| 416 |
+
βββ tests/
|
| 417 |
+
βββ scripts/ fast_tokenize, shard_corpus, launch_4gpu,
|
| 418 |
+
β fast4gpu_boost_v4, generate_aligned, hourly HF x8run sync
|
| 419 |
+
βββ checkpoints/ per-GPU + merged + recovery alias
|
| 420 |
+
βββ docs/ run logs, surgeries, trusted loss, scaling
|
| 421 |
+
βββ space/ HF Space demo
|
| 422 |
+
βββ Fractus_White_Paper_v2.md White paper v2.0
|
| 423 |
+
βββ arxiv/ LaTeX source for arXiv submission
|
| 424 |
```
|
| 425 |
+
|
| 426 |
+
---
|
| 427 |
+
|
| 428 |
+
## Key Concepts
|
| 429 |
+
|
| 430 |
+
**Tick**: one step of thinking. The engine processes an observation, updates its thought state through all blocks, and optionally emits output.
|
| 431 |
+
|
| 432 |
+
**Thought state**: a vector that persists across ticks β the engine's "consciousness."
|
| 433 |
+
|
| 434 |
+
**Chunk**: 32 tokens processed in one forward pass. The thought state and per-block attention state carry between chunks.
|
| 435 |
+
|
| 436 |
+
**Expert**: a small low-rank network (`W = scaleΒ·U@Vα΅`) that specializes in certain thoughts. Only 2 of 128 active per token.
|
| 437 |
+
|
| 438 |
+
**Kuramoto clock**: coupled oscillators producing phase vectors that route tokens to experts β learned end-to-end since the routing surgery.
|
| 439 |
+
|
| 440 |
+
**`tick_vec`**: the multimodal tick β feed any precomputed `(B, d_model)` vector (image patches, embeddings from any encoder) straight into the thought state without tokens. This is how vision plugs in.
|
| 441 |
+
|
| 442 |
+
**Surgery**: patching code around a preserved checkpoint mid-training, resuming at the exact token offset. Weights are never wiped for a routing, objective, or decode bug.
|
| 443 |
+
|
| 444 |
+
---
|
| 445 |
+
|
| 446 |
+
## Limitations (stated plainly)
|
| 447 |
+
|
| 448 |
+
- Generation is not yet coherent English β word-level repetition loops / lexical noise. Exposure bias is being addressed by scheduled sampling; AR is the metric to watch.
|
| 449 |
+
- GPT-2 vocab dominates parameters (81% at d=768) β vocab reduction is a known lever.
|
| 450 |
+
- "Remembers forever" and "grows on its own" describe the architecture's design; no independent benchmarks are provided.
|
| 451 |
+
- This is a research artifact and a live training run, not a production assistant.
|
| 452 |
+
|
| 453 |
+
---
|
| 454 |
+
|
| 455 |
+
## License
|
| 456 |
+
|
| 457 |
+
MIT. Fractus belongs to you, not to a corporation.
|
| 458 |
+
|
| 459 |
+
## Author
|
| 460 |
+
|
| 461 |
+
**Philippe-Antoine Robert** β 2026 β rpa.tu@proton.me
|
| 462 |
+
|
| 463 |
+
## Links
|
| 464 |
+
|
| 465 |
+
- **GitHub:** [github.com/AFKmoney/fractus-cte](https://github.com/AFKmoney/fractus-cte)
|
| 466 |
+
- **HuggingFace Model:** [huggingface.co/thefinalboss/fractus-cte](https://huggingface.co/thefinalboss/fractus-cte)
|
| 467 |
+
- **HuggingFace Datasets:** [huggingface.co/datasets/thefinalboss/fractus-datasets](https://huggingface.co/datasets/thefinalboss/fractus-datasets)
|
| 468 |
+
- **White Paper:** [Fractus_White_Paper_v2.md](Fractus_White_Paper_v2.md) / [PDF](Fractus_White_Paper.pdf)
|
| 469 |
+
- **Run logs:** [MASTER_RUN_LOG](docs/MASTER_RUN_LOG.md) Β· [TRAINING_LOG_1B](docs/TRAINING_LOG_1B.md) Β· [DISCOVERY_LOG](docs/DISCOVERY_LOG.md)
|
| 470 |
+
- **arXiv source:** [arxiv/main.tex](arxiv/main.tex)
|
| 471 |
+
|