File size: 17,776 Bytes
2ba757f
 
 
 
 
 
 
 
 
 
 
 
 
 
1f1d55d
2ba757f
 
 
 
1f1d55d
 
2ba757f
 
3f2e79e
 
 
 
2ba757f
 
 
 
 
454e3ba
2ba757f
3f2e79e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1f1d55d
3f2e79e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
454e3ba
1f1d55d
 
 
 
 
6c223ff
 
1f1d55d
 
 
6c223ff
1f1d55d
6c223ff
1f1d55d
6c223ff
 
1f1d55d
 
 
 
 
6c223ff
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3f2e79e
 
 
 
 
 
 
 
 
 
 
 
 
 
1f1d55d
3f2e79e
 
 
 
 
 
 
 
 
 
baa4bdb
 
 
 
 
 
 
 
 
 
5a8812e
baa4bdb
 
 
 
 
5a8812e
 
baa4bdb
 
 
 
 
 
 
 
 
 
 
3f2e79e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1f1d55d
3f2e79e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1f1d55d
 
 
3f2e79e
 
 
1f1d55d
3f2e79e
1f1d55d
3f2e79e
1f1d55d
3f2e79e
1f1d55d
3f2e79e
 
 
 
 
1f1d55d
 
 
 
 
3f2e79e
 
 
 
 
 
 
 
 
 
 
 
 
 
1f1d55d
 
3f2e79e
 
baa4bdb
 
 
bc10efd
 
 
 
0cd746e
 
 
 
0077eee
 
 
 
f52d2cd
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
---
license: mit
language:
  - en
  - fr
tags:
  - continuous-thought-engine
  - cognitive-agent
  - mixture-of-experts
  - kuramoto
  - self-modifying
  - progressive-growth
  - decentralized-ai
  - personal-ai
  - neuroscience
library_name: pytorch
pipeline_tag: text-generation
models:
  - thefinalboss/fractus-cte
datasets:
  - thefinalboss/fractus-datasets
---

# Fractus CTE

**A living AI that thinks continuously, remembers forever, and grows on its own.**

![License](https://img.shields.io/badge/license-MIT-blue)
![Python](https://img.shields.io/badge/python-3.13-green)
![PyTorch](https://img.shields.io/badge/PyTorch-2.9-orange)
![Params](https://img.shields.io/badge/params-1.05B-red)
![Status](https://img.shields.io/badge/status-active-brightgreen)
![Datasets](https://img.shields.io/badge/datasets-4.15B%20tokens-purple)

---

## What is Fractus?

Fractus is not a chatbot. It's not GPT. It's not a transformer.

Fractus is a **Continuous Cognitive Agent** β€” an AI that works like a brain, not a calculator. Instead of processing input β†’ output in one pass, Fractus **ticks** like a biological system: it maintains a persistent thought state, advances it through multiple blocks of processing, remembers everything across sessions, and can grow new capacity by itself.

### What makes it different from GPT/Claude?

| | GPT-4 / Claude | Fractus |
|---|---|---|
| **Thinking** | One pass, done | Continuous ticks (like a heartbeat) |
| **Memory** | Forgets when context window fills | Remembers forever (survives restarts) |
| **Learning** | Retrain from scratch ($$$) | Learns from every interaction |
| **Growth** | Fixed size forever | Grows new experts at runtime |
| **Mental states** | One mode always | Shifts between cognitive modes |
| **Where it runs** | Corporate cloud | Your machine |
| **Training** | Fixed, done once | Perpetual, never stops |

---

## The 12 Building Blocks

| Block | What it does |
|---|---|
| **Continuous Thought Engine** | The brain β€” thinks tick by tick through 16 blocks |
| **Persistent Memory** | Remembers you across sessions, never forgets |
| **Cognitive Modes** | Shifts mental states (focused, creative, exploratory...) |
| **RAG Knowledge Base** | Learns facts instantly β€” no retraining needed |
| **Cognitive Plugins** | Hot-swappable modes: analyst, coder, creative, teacher |
| **MetaCognition** | Decides its own actions: retrieve, learn, generate |
| **Progressive Growth** | Grows from 6M to 1B+ params, palier by palier |
| **Self-Modification** | Adds new experts at runtime when it needs them |
| **PhaseRoutedMoE** | Sparse experts routed by oscillator phases |
| **Kuramoto Clock** | A dynamical system that drives routing decisions |
| **Online Trainer** | Learns continuously, one chunk at a time |
| **HF Space** | Live chat demo with shared memory |

---

## Datasets (4.15 Billion Tokens)

Fractus is trained on a massive, diverse corpus available at [huggingface.co/datasets/thefinalboss/fractus-datasets](https://huggingface.co/datasets/thefinalboss/fractus-datasets):

| Dataset | Tokens | Content |
|---|---|---|
| **neuro-paradigms-1b** | **~1B** | 100 neuroscience β†’ software architecture paradigms (300 chunked files) |
| **neuro-code-math** | **~900M** | Neuro-inspired coding, mathematics, algorithms (incl. 40 applied-neuroscience topics) |
| **cognitive-skills** | **~780M** | Coding, reasoning, speaking, thinking, understanding |
| **fractus-generated-corpus** | **340M** | Bilingual FR/EN generated by Fractus ontology engine |
| **paradigms-full** | **191M** | 140 paradigms (neuroscience, CS, architecture) |
| **gutenberg-esoteric** | **~58M** | 487 public-domain esoteric / masonic / hermetic books |
| **neuro-arch-full** | **86M** | 60 neuroscience paradigms (neuro-software-architecture) |
| **all-github-repos** | **54M+** | 80+ of your GitHub repos (public + private, secret-filtered) |
| **mega-corpus-v3** | **20M** | Literature, philosophy, occult, masonry, science, medicine |
| **wordnet** | **3M** | 117K dictionary synset entries |
| **Total** | **~4.2B** | |

The corpus covers neuroscience, software architecture, philosophy, psychology, literature, esoteric traditions, programming, medicine, and Fractus's own source code.

---

## Applied Neuroscience β€” the theoretical core

Fractus is a neuroscience-grounded architecture: real brain mechanisms are mapped to software/AI patterns, and that mapping is itself training data. Every entry below is present in the dataset (`neuro_paradigms_1b/`, `neuro_code_math/applied_neuroscience/`, and the `*.pt` files in `datasets/`) β€” verified by file listing, not just claimed.

### 100 neuroscience β†’ software-architecture paradigms (`neuro_paradigms_1b`, 300 chunked files)

Each paradigm maps a biological mechanism to an engineering pattern (e.g. *adenosine sleep pressure* β†’ cache-stampede recovery; *myelin sheath* β†’ caching; *hippocampal replay* β†’ trajectory consolidation).

<details><summary><b>show all 100 paradigms</b></summary>

```
adenosine_sleep_pressure            amygdala_prefrontal_topdown       anterior_cingulate_conflict_monitor
apoptosis_self_destructing_service  arc_gene_plasticity_marker        astrocyte_tripartite_synapse
axon_initial_segment_trigger        basal_ganglia_loop_arbitration    bdnf_growth_factor_scaling
bergmann_glia_purkinje              binaural_cross_correlation_localization   brainstem_vital_functions
broca_area_api_generator            calcium_transmitter_coupling      camp_second_messenger_amplifier
cerebellar_forward_model            cholinergic_attentional_filter    circadian_gene_expression
climbing_fiber_error_broadcast      cochlear_compressive_nonlinearity cortical_area_specialization
cortical_minicolumn_pipeline        cortico_cortical_pathways         corticotropin_releasing_hormone
cortisol_slow_stress_recovery       critical_period_learning_rate     dendritic_compartmentalization
endocannabinoid_retrograde          enteric_glia_gut_brain            ependymal_cell_barrier
fusiform_face_service_registry      gaba_inhibitory_bus               gap_junction_electrical_sync
ghrelin_hunger_signal               glomerular_convergence_gateway    glutamate_excitatory_bus
glycine_coagonist_modulator         granule_cell_inhibitory_relay     hair_cell_banks_event_clusters
hippocampal_4ec_loop_replay         histamine_wakefulness_keeper      hox_gene_service_specialization
hypercolumn_module_federation       hypercolumn_sharding              hypothalamus_homeostasis
insula_interoception_monitor        ip3_inositol_cascade              k_complex_event_trigger
kcc2_chloride_shift_inhibitor       leptin_satiety_signal             locus_coeruleus_ne_global_signal
melatonin_circadian_scheduler       microglia_active_surveillance     mitral_tufted_cell_dual_path
morphogen_gradient_config           muller_glia_retina_repair         myelin_sheath_caching
neural_crest_migration_deploy       neuropeptide_y_stress_buffer      ng2_glia_pool_renewal
nitric_oxide_gas_signal             node_of_ranvier_bypass            nrem_slow_wave_cleanup
nucleus_accumbens_reward_routing    oligodendrocyte_myelination_dynamic      orexin_stability_keeper
orientation_column_indexing         oscillatory_phase_locking_io      oxytocin_trust_protocol
parahippocampal_place_topology      parallel_fiber_fanout_aggregation pineal_circadian_release
pinwheel_central_layout             pituitary_master_gland            posterior_parietal_integration
prolactin_parental_care             quantal_release_batching          radial_glia_neural_stem
radial_glial_scaffold               raphe_serotonin_rate_limit        rem_paradoxical_processing
replay_consolidation_trajectory     reticular_activating_system       retinotopic_data_layout
satellite_glial_ganglion            schwann_cell_peripheral_repair    sleep_pressure_forced_maintenance
sleep_spindle_memory_transfer       slow_oscillation_sync             subplate_wait_state
suprachiasmatic_clock               synaptic_vesicle_pool             synaptogenesis_service_wiring
tanycyte_metabolic_sensor           temporal_pole_semantic_cache      thalamocortical_loop_api
tonotopic_stream_partitioning       vasopressin_loyalty_aware_routing vta_dopamine_rpe_scheduler
wernicke_area_api_parser
```
</details>

### 40 applied-neuroscience topics (`neuro_code_math/applied_neuroscience/`)

Deep dives on computational neuroscience theories β€” the science Fractus's design draws from.

<details><summary><b>show all 40 topics</b></summary>

```
active_inference          axonal_computation         basal_ganglia_circuits     bayesian_brain
cerebellar_computation    consolidation              cortical_minicolumns       cross_frequency_coupling
dendritic_computation     dopamine_reward            entorhinal_grid_cells      free_energy_principle
gamma_oscillations        global_workspace_theory    head_direction_cells       hierarchical_processing
higher_order_theories     hippocampal_formation      homeostatic_plasticity     integrated_information_theory
long_term_depression      long_term_potentiation     metaplasticity             neural_coding
neural_decoding           neural_manifolds           neuromodulation            place_cells
population_coding         predictive_coding          predictive_processing      rate_coding
serotonin_modulation      sharp_wave_ripples         sleep_replay               sparse_coding
spike_timing_dependent_plasticity   temporal_coding  thalamic_reticular_nucleus theta_oscillations
```
</details>

### Foundational researchers & concepts honored in the corpus

**Hebb** (Hebbian learning), **Bi & Poo** (STDP timing curves), **Friston** (free energy / active inference), **BuzsΓ‘ki** (hippocampal sharp-wave ripples, replay), **Moser & Moser** (grid cells), **Hodgkin & Huxley** (axon dynamics), **Izhikevich** (spike models), **Tononi** (integrated information), **Baars/Dehaene** (global workspace), **O'Keefe** (place cells), **Kandel** (memory consolidation), plus neuromodulators (dopamine RPE, serotonin, oxytocin, vasopressin) and glial biology (astrocytes, microglia, oligodendrocytes, Schwann cells).

### Source files (all verified present)

| File | Content |
|---|---|
| `neuro_paradigms_1b/*.jsonl.gz` (300) | 100 paradigms Γ— 3 chunks, instruction+response+citations |
| `neuro_code_math/applied_neuroscience__*.jsonl` (40) | Computational neuroscience deep-dives |
| `datasets/neuro_arch_full.pt` | 60 neuroscience-grounded architecture paradigms |
| `datasets/neuro_software_architecture.pt` | Same family, alternate cut |
| `datasets/paradigms_full.pt` / `paradigms_dataset.pt` | 140 foundational + neuroscience paradigms |
| `datasets/fractus_generated_corpus.pt` | Fractus ontology engine (neuroscience β†’ AI) |

---

## How to Use

### Install

```bash
git clone https://github.com/AFKmoney/fractus-cte.git
cd fractus-cte
pip install torch numpy tokenizers matplotlib fastapi uvicorn pydantic
```

### Run tests

```bash
pytest tests/ -q
# β†’ 28 passed
```

### Train on CPU (progressive growth)

```bash
python scripts/train_progressive.py --paliers 0,1,2,3 --accumulation-steps 8
```

### Train on GPU (1B scale)

Fractus 1B training is designed to be **interruptible and resumable** β€” not a one-shot pretrain. The full corpus (~4.23B tokens) is the starting nutrition for the 1B palier. You can stop when budget ends, use the model, and continue later. That is how Fractus is meant to work.

**Current production setup (multi-GPU, full corpus):**

1. Shard the full corpus across GPUs:
```bash
python scripts/shard_corpus.py --src data/full_corpus.pt --out-dir data --n-shards 4
```

2. Launch one process per GPU (B=2, seq=128, torch.compile):
```bash
bash scripts/launch_4gpu.sh
# or manually:
# GPU_ID=0 CUDA_VISIBLE_DEVICES=0 python -u scripts/train_1b_multi_gpu.py
# GPU_ID=1 CUDA_VISIBLE_DEVICES=1 python -u scripts/train_1b_multi_gpu.py
# ...
```

3. Each GPU trains independently on its shard from the same `palier0` β†’ grow to 1B. Checkpoints are saved per GPU (`checkpoints/fractus_1b_gpu{0-3}.pt`).

4. Merge the 4 shard-trained weights into one model that has seen the full corpus:
```bash
# produces checkpoints/fractus_1b_merged.pt (weight average)
```

**Status (Aug 2026 run):** 4Γ— RTX 5090, full 4.23B corpus sharded, B=2 + compile, loss on best GPU ~86 after ~11M tokens and still falling. Merged checkpoint on the Hub: `checkpoints/fractus_1b_merged.pt`. Training can be stopped and resumed at any time.

> **Note:** A single-GPU full-corpus 1B run is not the supported path. The 1B model + full 4.23B corpus is trained by sharding across GPUs, then merging. One card can resume from a checkpoint later, but the production recipe is multi-GPU shard β†’ merge.

### Use the agent

```python
from fractus.continuous_engine import ContinuousThoughtEngine
from fractus.memory import PersistentMemory
from fractus.tokenizer import FractusTokenizer

# Build the brain
engine = ContinuousThoughtEngine(
    vocab_size=50257, d_model=128, n_heads=2, d_head=64,
    n_layers=2, n_levels=2, n_oscillators=8, coupling_rank=4,
    n_experts=4, top_k=2, expert_d_ff=128, siren_rank=32)

# Give it memory
memory = PersistentMemory(d_model=128, path="~/.fractus/memory.pt")
engine.attach_memory(memory)

# Think
engine.reset_thought(batch_size=1)
logits, confidence = engine.tick(torch.tensor([42]))
print(f"Confidence: {confidence.item():.2f}")
```

---

## The Growth Path

| Stage | Size | Blocks | Experts | What it can do |
|---|---|---|---|---|
| Palier 0 | 6.6M | 1 | 4 | Learn basic patterns |
| Palier 1 | 25M | 2 | 8 | Simple text generation |
| Palier 2 | 120M | 4 | 16 | Coherent fragments |
| Palier 3 | 350M | 8 | 32 | Decent text quality |
| **Palier 4** | **1B** | **16** | **128** | **Full language model** |

Each stage inherits the previous one's knowledge. The model never starts from zero.

---

## Architecture (for developers)

```
fractus-cte/
β”œβ”€β”€ fractus/
β”‚   β”œβ”€β”€ continuous_engine.py      ← The brain (CTE + CTEBlock)
β”‚   β”‚   β”œβ”€β”€ CTEBlock              One block: attention + Kuramoto + MoE
β”‚   β”‚   └── ContinuousThoughtEngine  Stacks N blocks, carries thought state
β”‚   β”œβ”€β”€ memory.py                 ← Cross-session persistent memory
β”‚   β”œβ”€β”€ cognitive_modes.py        ← Unsupervised mental state detection
β”‚   β”œβ”€β”€ grow.py                   ← Progressive growth operator (width + depth + experts)
β”‚   β”œβ”€β”€ rag.py                    ← Knowledge base + plugins + metacognition
β”‚   β”œβ”€β”€ tokenizer.py              ← GPT-2 BPE tokenizer
β”‚   β”œβ”€β”€ nn/
β”‚   β”‚   β”œβ”€β”€ moe.py                ← PhaseRoutedMoE (sparse, low-rank, differentiable)
β”‚   β”‚   β”œβ”€β”€ attention.py          ← Multi-level causal linear attention
β”‚   β”‚   β”œβ”€β”€ phase_ode.py          ← Kuramoto RK4 oscillators
β”‚   β”‚   └── lazy_siren.py         ← Low-rank weight storage
β”‚   └── train/
β”‚       └── online.py             ← Online trainer (SGD/AdamW, accumulation)
β”œβ”€β”€ tests/                        28 tests
β”œβ”€β”€ scripts/                      Training + corpus + GPU scripts
β”œβ”€β”€ space/                        HF Space demo
β”œβ”€β”€ docs/                         Optimization analysis
β”œβ”€β”€ Fractus_White_Paper.pdf       Technical white paper v2.0
└── arxiv/                        LaTeX source for arXiv submission
```

---

## Key Concepts

**Tick**: one step of thinking. The engine processes an observation, updates its thought state through all blocks, and optionally emits output.

**Thought state**: a vector that persists across ticks β€” the engine's "consciousness."

**Chunk**: 32 tokens processed in one forward pass. The thought state and per-block attention state carry between chunks.

**Expert**: a small neural network (low-rank `W = scaleΒ·U@V^T`) that specializes in certain thoughts. Only 2 of 128 active per token (sparse routing).

**Kuramoto**: coupled oscillators producing phase vectors that route tokens to experts. The engine's "internal clock."

---

## Research Results (Honest)

- **EDT** (Expert Decoupled Training): refuted. 5 variants, all ~19% worse.
- **Forward-Forward** (Hinton 2022): refuted. Local learning can't replace global backprop.
- **Progressive growth**: works. Warm start converges faster.
- **Sparse MoE low-rank**: works. 2/128 experts = 64x less compute.
- **1345 tok/s on CPU**: measured (batch=8 + SGD + all optimizations).

---

## License

MIT. Fractus belongs to you, not to a corporation.

## Author

**Philippe-Antoine Robert** β€” 2026 β€” rpa.tu@proton.me

## Links

- **GitHub:** [github.com/AFKmoney/fractus-cte](https://github.com/AFKmoney/fractus-cte)
- **HuggingFace Model:** [huggingface.co/thefinalboss/fractus-cte](https://huggingface.co/thefinalboss/fractus-cte)
- **HuggingFace Datasets:** [huggingface.co/datasets/thefinalboss/fractus-datasets](https://huggingface.co/datasets/thefinalboss/fractus-datasets)
- **White Paper:** [Fractus_White_Paper.pdf](Fractus_White_Paper.pdf)
- **arXiv source:** [arxiv/main.tex](arxiv/main.tex)

## Training log (1B live run)
See [docs/TRAINING_LOG_1B.md](docs/TRAINING_LOG_1B.md) for loss curves, rates, and mid-training generation probes (including raw failure-mode outputs).


## Discovery log
See [docs/DISCOVERY_LOG.md](docs/DISCOVERY_LOG.md) for bugs found in production, routing surgery, and emergent features.


## Mid-training operability
See [docs/OPERABILITY_MIDTRAIN.md](docs/OPERABILITY_MIDTRAIN.md).


## Master run log
Full production chronology: [docs/MASTER_RUN_LOG.md](docs/MASTER_RUN_LOG.md)


## Composability
In-training surgery and checkpoint merge/grow: [docs/COMPOSABILITY_AND_SURGERY.md](docs/COMPOSABILITY_AND_SURGERY.md)