---
license: mit
language:
- en
- fr
tags:
- continuous-thought-engine
- cognitive-agent
- mixture-of-experts
- kuramoto
- self-modifying
- progressive-growth
- decentralized-ai
- personal-ai
- neuroscience
library_name: pytorch
pipeline_tag: text-generation
models:
- thefinalboss/fractus-cte
datasets:
- thefinalboss/fractus-datasets
---
# Fractus CTE
**A living AI that thinks continuously, remembers forever, and grows on its own.**






---
## What is Fractus?
Fractus is not a chatbot. It's not GPT. It's not a transformer.
Fractus is a **Continuous Cognitive Agent** — an AI that works like a brain, not a calculator. Instead of processing input → output in one pass, Fractus **ticks** like a biological system: it maintains a persistent thought state, advances it through multiple blocks of processing, remembers everything across sessions, and can grow new capacity by itself.
### What makes it different from GPT/Claude?
| | GPT-4 / Claude | Fractus |
|---|---|---|
| **Thinking** | One pass, done | Continuous ticks (like a heartbeat) |
| **Memory** | Forgets when context window fills | Remembers forever (survives restarts) |
| **Learning** | Retrain from scratch ($$$) | Learns from every interaction |
| **Growth** | Fixed size forever | Grows new experts at runtime |
| **Mental states** | One mode always | Shifts between cognitive modes |
| **Where it runs** | Corporate cloud | Your machine |
| **Training** | Fixed, done once | Perpetual, never stops |
---
## The 12 Building Blocks
| Block | What it does |
|---|---|
| **Continuous Thought Engine** | The brain — thinks tick by tick through 16 blocks |
| **Persistent Memory** | Remembers you across sessions, never forgets |
| **Cognitive Modes** | Shifts mental states (focused, creative, exploratory...) |
| **RAG Knowledge Base** | Learns facts instantly — no retraining needed |
| **Cognitive Plugins** | Hot-swappable modes: analyst, coder, creative, teacher |
| **MetaCognition** | Decides its own actions: retrieve, learn, generate |
| **Progressive Growth** | Grows from 6M to 1B+ params, palier by palier |
| **Self-Modification** | Adds new experts at runtime when it needs them |
| **PhaseRoutedMoE** | Sparse experts routed by oscillator phases |
| **Kuramoto Clock** | A dynamical system that drives routing decisions |
| **Online Trainer** | Learns continuously, one chunk at a time |
| **HF Space** | Live chat demo with shared memory |
---
## Datasets (4.15 Billion Tokens)
Fractus is trained on a massive, diverse corpus available at [huggingface.co/datasets/thefinalboss/fractus-datasets](https://huggingface.co/datasets/thefinalboss/fractus-datasets):
| Dataset | Tokens | Content |
|---|---|---|
| **neuro-paradigms-1b** | **~1B** | 100 neuroscience → software architecture paradigms (300 chunked files) |
| **neuro-code-math** | **~900M** | Neuro-inspired coding, mathematics, algorithms (incl. 40 applied-neuroscience topics) |
| **cognitive-skills** | **~780M** | Coding, reasoning, speaking, thinking, understanding |
| **fractus-generated-corpus** | **340M** | Bilingual FR/EN generated by Fractus ontology engine |
| **paradigms-full** | **191M** | 140 paradigms (neuroscience, CS, architecture) |
| **gutenberg-esoteric** | **~58M** | 487 public-domain esoteric / masonic / hermetic books |
| **neuro-arch-full** | **86M** | 60 neuroscience paradigms (neuro-software-architecture) |
| **all-github-repos** | **54M+** | 80+ of your GitHub repos (public + private, secret-filtered) |
| **mega-corpus-v3** | **20M** | Literature, philosophy, occult, masonry, science, medicine |
| **wordnet** | **3M** | 117K dictionary synset entries |
| **Total** | **~4.2B** | |
The corpus covers neuroscience, software architecture, philosophy, psychology, literature, esoteric traditions, programming, medicine, and Fractus's own source code.
---
## Applied Neuroscience — the theoretical core
Fractus is a neuroscience-grounded architecture: real brain mechanisms are mapped to software/AI patterns, and that mapping is itself training data. Every entry below is present in the dataset (`neuro_paradigms_1b/`, `neuro_code_math/applied_neuroscience/`, and the `*.pt` files in `datasets/`) — verified by file listing, not just claimed.
### 100 neuroscience → software-architecture paradigms (`neuro_paradigms_1b`, 300 chunked files)
Each paradigm maps a biological mechanism to an engineering pattern (e.g. *adenosine sleep pressure* → cache-stampede recovery; *myelin sheath* → caching; *hippocampal replay* → trajectory consolidation).
show all 100 paradigms
```
adenosine_sleep_pressure amygdala_prefrontal_topdown anterior_cingulate_conflict_monitor
apoptosis_self_destructing_service arc_gene_plasticity_marker astrocyte_tripartite_synapse
axon_initial_segment_trigger basal_ganglia_loop_arbitration bdnf_growth_factor_scaling
bergmann_glia_purkinje binaural_cross_correlation_localization brainstem_vital_functions
broca_area_api_generator calcium_transmitter_coupling camp_second_messenger_amplifier
cerebellar_forward_model cholinergic_attentional_filter circadian_gene_expression
climbing_fiber_error_broadcast cochlear_compressive_nonlinearity cortical_area_specialization
cortical_minicolumn_pipeline cortico_cortical_pathways corticotropin_releasing_hormone
cortisol_slow_stress_recovery critical_period_learning_rate dendritic_compartmentalization
endocannabinoid_retrograde enteric_glia_gut_brain ependymal_cell_barrier
fusiform_face_service_registry gaba_inhibitory_bus gap_junction_electrical_sync
ghrelin_hunger_signal glomerular_convergence_gateway glutamate_excitatory_bus
glycine_coagonist_modulator granule_cell_inhibitory_relay hair_cell_banks_event_clusters
hippocampal_4ec_loop_replay histamine_wakefulness_keeper hox_gene_service_specialization
hypercolumn_module_federation hypercolumn_sharding hypothalamus_homeostasis
insula_interoception_monitor ip3_inositol_cascade k_complex_event_trigger
kcc2_chloride_shift_inhibitor leptin_satiety_signal locus_coeruleus_ne_global_signal
melatonin_circadian_scheduler microglia_active_surveillance mitral_tufted_cell_dual_path
morphogen_gradient_config muller_glia_retina_repair myelin_sheath_caching
neural_crest_migration_deploy neuropeptide_y_stress_buffer ng2_glia_pool_renewal
nitric_oxide_gas_signal node_of_ranvier_bypass nrem_slow_wave_cleanup
nucleus_accumbens_reward_routing oligodendrocyte_myelination_dynamic orexin_stability_keeper
orientation_column_indexing oscillatory_phase_locking_io oxytocin_trust_protocol
parahippocampal_place_topology parallel_fiber_fanout_aggregation pineal_circadian_release
pinwheel_central_layout pituitary_master_gland posterior_parietal_integration
prolactin_parental_care quantal_release_batching radial_glia_neural_stem
radial_glial_scaffold raphe_serotonin_rate_limit rem_paradoxical_processing
replay_consolidation_trajectory reticular_activating_system retinotopic_data_layout
satellite_glial_ganglion schwann_cell_peripheral_repair sleep_pressure_forced_maintenance
sleep_spindle_memory_transfer slow_oscillation_sync subplate_wait_state
suprachiasmatic_clock synaptic_vesicle_pool synaptogenesis_service_wiring
tanycyte_metabolic_sensor temporal_pole_semantic_cache thalamocortical_loop_api
tonotopic_stream_partitioning vasopressin_loyalty_aware_routing vta_dopamine_rpe_scheduler
wernicke_area_api_parser
```
### 40 applied-neuroscience topics (`neuro_code_math/applied_neuroscience/`)
Deep dives on computational neuroscience theories — the science Fractus's design draws from.
show all 40 topics
```
active_inference axonal_computation basal_ganglia_circuits bayesian_brain
cerebellar_computation consolidation cortical_minicolumns cross_frequency_coupling
dendritic_computation dopamine_reward entorhinal_grid_cells free_energy_principle
gamma_oscillations global_workspace_theory head_direction_cells hierarchical_processing
higher_order_theories hippocampal_formation homeostatic_plasticity integrated_information_theory
long_term_depression long_term_potentiation metaplasticity neural_coding
neural_decoding neural_manifolds neuromodulation place_cells
population_coding predictive_coding predictive_processing rate_coding
serotonin_modulation sharp_wave_ripples sleep_replay sparse_coding
spike_timing_dependent_plasticity temporal_coding thalamic_reticular_nucleus theta_oscillations
```
### Foundational researchers & concepts honored in the corpus
**Hebb** (Hebbian learning), **Bi & Poo** (STDP timing curves), **Friston** (free energy / active inference), **Buzsáki** (hippocampal sharp-wave ripples, replay), **Moser & Moser** (grid cells), **Hodgkin & Huxley** (axon dynamics), **Izhikevich** (spike models), **Tononi** (integrated information), **Baars/Dehaene** (global workspace), **O'Keefe** (place cells), **Kandel** (memory consolidation), plus neuromodulators (dopamine RPE, serotonin, oxytocin, vasopressin) and glial biology (astrocytes, microglia, oligodendrocytes, Schwann cells).
### Source files (all verified present)
| File | Content |
|---|---|
| `neuro_paradigms_1b/*.jsonl.gz` (300) | 100 paradigms × 3 chunks, instruction+response+citations |
| `neuro_code_math/applied_neuroscience__*.jsonl` (40) | Computational neuroscience deep-dives |
| `datasets/neuro_arch_full.pt` | 60 neuroscience-grounded architecture paradigms |
| `datasets/neuro_software_architecture.pt` | Same family, alternate cut |
| `datasets/paradigms_full.pt` / `paradigms_dataset.pt` | 140 foundational + neuroscience paradigms |
| `datasets/fractus_generated_corpus.pt` | Fractus ontology engine (neuroscience → AI) |
---
## How to Use
### Install
```bash
git clone https://github.com/AFKmoney/fractus-cte.git
cd fractus-cte
pip install torch numpy tokenizers matplotlib fastapi uvicorn pydantic
```
### Run tests
```bash
pytest tests/ -q
# → 28 passed
```
### Train on CPU (progressive growth)
```bash
python scripts/train_progressive.py --paliers 0,1,2,3 --accumulation-steps 8
```
### Train on GPU (1B scale)
Fractus 1B training is designed to be **interruptible and resumable** — not a one-shot pretrain. The full corpus (~4.23B tokens) is the starting nutrition for the 1B palier. You can stop when budget ends, use the model, and continue later. That is how Fractus is meant to work.
**Current production setup (multi-GPU, full corpus):**
1. Shard the full corpus across GPUs:
```bash
python scripts/shard_corpus.py --src data/full_corpus.pt --out-dir data --n-shards 4
```
2. Launch one process per GPU (B=2, seq=128, torch.compile):
```bash
bash scripts/launch_4gpu.sh
# or manually:
# GPU_ID=0 CUDA_VISIBLE_DEVICES=0 python -u scripts/train_1b_multi_gpu.py
# GPU_ID=1 CUDA_VISIBLE_DEVICES=1 python -u scripts/train_1b_multi_gpu.py
# ...
```
3. Each GPU trains independently on its shard from the same `palier0` → grow to 1B. Checkpoints are saved per GPU (`checkpoints/fractus_1b_gpu{0-3}.pt`).
4. Merge the 4 shard-trained weights into one model that has seen the full corpus:
```bash
# produces checkpoints/fractus_1b_merged.pt (weight average)
```
**Status (Aug 2026 run):** 4× RTX 5090, full 4.23B corpus sharded, B=2 + compile, loss on best GPU ~86 after ~11M tokens and still falling. Merged checkpoint on the Hub: `checkpoints/fractus_1b_merged.pt`. Training can be stopped and resumed at any time.
> **Note:** A single-GPU full-corpus 1B run is not the supported path. The 1B model + full 4.23B corpus is trained by sharding across GPUs, then merging. One card can resume from a checkpoint later, but the production recipe is multi-GPU shard → merge.
### Use the agent
```python
from fractus.continuous_engine import ContinuousThoughtEngine
from fractus.memory import PersistentMemory
from fractus.tokenizer import FractusTokenizer
# Build the brain
engine = ContinuousThoughtEngine(
vocab_size=50257, d_model=128, n_heads=2, d_head=64,
n_layers=2, n_levels=2, n_oscillators=8, coupling_rank=4,
n_experts=4, top_k=2, expert_d_ff=128, siren_rank=32)
# Give it memory
memory = PersistentMemory(d_model=128, path="~/.fractus/memory.pt")
engine.attach_memory(memory)
# Think
engine.reset_thought(batch_size=1)
logits, confidence = engine.tick(torch.tensor([42]))
print(f"Confidence: {confidence.item():.2f}")
```
---
## The Growth Path
| Stage | Size | Blocks | Experts | What it can do |
|---|---|---|---|---|
| Palier 0 | 6.6M | 1 | 4 | Learn basic patterns |
| Palier 1 | 25M | 2 | 8 | Simple text generation |
| Palier 2 | 120M | 4 | 16 | Coherent fragments |
| Palier 3 | 350M | 8 | 32 | Decent text quality |
| **Palier 4** | **1B** | **16** | **128** | **Full language model** |
Each stage inherits the previous one's knowledge. The model never starts from zero.
---
## Architecture (for developers)
```
fractus-cte/
├── fractus/
│ ├── continuous_engine.py ← The brain (CTE + CTEBlock)
│ │ ├── CTEBlock One block: attention + Kuramoto + MoE
│ │ └── ContinuousThoughtEngine Stacks N blocks, carries thought state
│ ├── memory.py ← Cross-session persistent memory
│ ├── cognitive_modes.py ← Unsupervised mental state detection
│ ├── grow.py ← Progressive growth operator (width + depth + experts)
│ ├── rag.py ← Knowledge base + plugins + metacognition
│ ├── tokenizer.py ← GPT-2 BPE tokenizer
│ ├── nn/
│ │ ├── moe.py ← PhaseRoutedMoE (sparse, low-rank, differentiable)
│ │ ├── attention.py ← Multi-level causal linear attention
│ │ ├── phase_ode.py ← Kuramoto RK4 oscillators
│ │ └── lazy_siren.py ← Low-rank weight storage
│ └── train/
│ └── online.py ← Online trainer (SGD/AdamW, accumulation)
├── tests/ 28 tests
├── scripts/ Training + corpus + GPU scripts
├── space/ HF Space demo
├── docs/ Optimization analysis
├── Fractus_White_Paper.pdf Technical white paper v2.0
└── arxiv/ LaTeX source for arXiv submission
```
---
## Key Concepts
**Tick**: one step of thinking. The engine processes an observation, updates its thought state through all blocks, and optionally emits output.
**Thought state**: a vector that persists across ticks — the engine's "consciousness."
**Chunk**: 32 tokens processed in one forward pass. The thought state and per-block attention state carry between chunks.
**Expert**: a small neural network (low-rank `W = scale·U@V^T`) that specializes in certain thoughts. Only 2 of 128 active per token (sparse routing).
**Kuramoto**: coupled oscillators producing phase vectors that route tokens to experts. The engine's "internal clock."
---
## Research Results (Honest)
- **EDT** (Expert Decoupled Training): refuted. 5 variants, all ~19% worse.
- **Forward-Forward** (Hinton 2022): refuted. Local learning can't replace global backprop.
- **Progressive growth**: works. Warm start converges faster.
- **Sparse MoE low-rank**: works. 2/128 experts = 64x less compute.
- **1345 tok/s on CPU**: measured (batch=8 + SGD + all optimizations).
---
## License
MIT. Fractus belongs to you, not to a corporation.
## Author
**Philippe-Antoine Robert** — 2026 — rpa.tu@proton.me
## Links
- **GitHub:** [github.com/AFKmoney/fractus-cte](https://github.com/AFKmoney/fractus-cte)
- **HuggingFace Model:** [huggingface.co/thefinalboss/fractus-cte](https://huggingface.co/thefinalboss/fractus-cte)
- **HuggingFace Datasets:** [huggingface.co/datasets/thefinalboss/fractus-datasets](https://huggingface.co/datasets/thefinalboss/fractus-datasets)
- **White Paper:** [Fractus_White_Paper.pdf](Fractus_White_Paper.pdf)
- **arXiv source:** [arxiv/main.tex](arxiv/main.tex)
## Training log (1B live run)
See [docs/TRAINING_LOG_1B.md](docs/TRAINING_LOG_1B.md) for loss curves, rates, and mid-training generation probes (including raw failure-mode outputs).
## Discovery log
See [docs/DISCOVERY_LOG.md](docs/DISCOVERY_LOG.md) for bugs found in production, routing surgery, and emergent features.