--- license: mit language: - en - fr tags: - continuous-thought-engine - cognitive-agent - mixture-of-experts - kuramoto - self-modifying - progressive-growth - decentralized-ai - personal-ai - neuroscience library_name: pytorch pipeline_tag: text-generation models: - thefinalboss/fractus-cte datasets: - thefinalboss/fractus-datasets --- # Fractus CTE **A living AI that thinks continuously, remembers forever, and grows on its own.** ![License](https://img.shields.io/badge/license-MIT-blue) ![Python](https://img.shields.io/badge/python-3.13-green) ![PyTorch](https://img.shields.io/badge/PyTorch-2.9-orange) ![Params](https://img.shields.io/badge/params-1.05B-red) ![Status](https://img.shields.io/badge/status-active-brightgreen) ![Datasets](https://img.shields.io/badge/datasets-4.15B%20tokens-purple) --- ## What is Fractus? Fractus is not a chatbot. It's not GPT. It's not a transformer. Fractus is a **Continuous Cognitive Agent** — an AI that works like a brain, not a calculator. Instead of processing input → output in one pass, Fractus **ticks** like a biological system: it maintains a persistent thought state, advances it through multiple blocks of processing, remembers everything across sessions, and can grow new capacity by itself. ### What makes it different from GPT/Claude? | | GPT-4 / Claude | Fractus | |---|---|---| | **Thinking** | One pass, done | Continuous ticks (like a heartbeat) | | **Memory** | Forgets when context window fills | Remembers forever (survives restarts) | | **Learning** | Retrain from scratch ($$$) | Learns from every interaction | | **Growth** | Fixed size forever | Grows new experts at runtime | | **Mental states** | One mode always | Shifts between cognitive modes | | **Where it runs** | Corporate cloud | Your machine | | **Training** | Fixed, done once | Perpetual, never stops | --- ## The 12 Building Blocks | Block | What it does | |---|---| | **Continuous Thought Engine** | The brain — thinks tick by tick through 16 blocks | | **Persistent Memory** | Remembers you across sessions, never forgets | | **Cognitive Modes** | Shifts mental states (focused, creative, exploratory...) | | **RAG Knowledge Base** | Learns facts instantly — no retraining needed | | **Cognitive Plugins** | Hot-swappable modes: analyst, coder, creative, teacher | | **MetaCognition** | Decides its own actions: retrieve, learn, generate | | **Progressive Growth** | Grows from 6M to 1B+ params, palier by palier | | **Self-Modification** | Adds new experts at runtime when it needs them | | **PhaseRoutedMoE** | Sparse experts routed by oscillator phases | | **Kuramoto Clock** | A dynamical system that drives routing decisions | | **Online Trainer** | Learns continuously, one chunk at a time | | **HF Space** | Live chat demo with shared memory | --- ## Datasets (4.15 Billion Tokens) Fractus is trained on a massive, diverse corpus available at [huggingface.co/datasets/thefinalboss/fractus-datasets](https://huggingface.co/datasets/thefinalboss/fractus-datasets): | Dataset | Tokens | Content | |---|---|---| | **neuro-paradigms-1b** | **~1B** | 100 neuroscience → software architecture paradigms (300 chunked files) | | **neuro-code-math** | **~900M** | Neuro-inspired coding, mathematics, algorithms (incl. 40 applied-neuroscience topics) | | **cognitive-skills** | **~780M** | Coding, reasoning, speaking, thinking, understanding | | **fractus-generated-corpus** | **340M** | Bilingual FR/EN generated by Fractus ontology engine | | **paradigms-full** | **191M** | 140 paradigms (neuroscience, CS, architecture) | | **gutenberg-esoteric** | **~58M** | 487 public-domain esoteric / masonic / hermetic books | | **neuro-arch-full** | **86M** | 60 neuroscience paradigms (neuro-software-architecture) | | **all-github-repos** | **54M+** | 80+ of your GitHub repos (public + private, secret-filtered) | | **mega-corpus-v3** | **20M** | Literature, philosophy, occult, masonry, science, medicine | | **wordnet** | **3M** | 117K dictionary synset entries | | **Total** | **~4.2B** | | The corpus covers neuroscience, software architecture, philosophy, psychology, literature, esoteric traditions, programming, medicine, and Fractus's own source code. --- ## Applied Neuroscience — the theoretical core Fractus is a neuroscience-grounded architecture: real brain mechanisms are mapped to software/AI patterns, and that mapping is itself training data. Every entry below is present in the dataset (`neuro_paradigms_1b/`, `neuro_code_math/applied_neuroscience/`, and the `*.pt` files in `datasets/`) — verified by file listing, not just claimed. ### 100 neuroscience → software-architecture paradigms (`neuro_paradigms_1b`, 300 chunked files) Each paradigm maps a biological mechanism to an engineering pattern (e.g. *adenosine sleep pressure* → cache-stampede recovery; *myelin sheath* → caching; *hippocampal replay* → trajectory consolidation).
show all 100 paradigms ``` adenosine_sleep_pressure amygdala_prefrontal_topdown anterior_cingulate_conflict_monitor apoptosis_self_destructing_service arc_gene_plasticity_marker astrocyte_tripartite_synapse axon_initial_segment_trigger basal_ganglia_loop_arbitration bdnf_growth_factor_scaling bergmann_glia_purkinje binaural_cross_correlation_localization brainstem_vital_functions broca_area_api_generator calcium_transmitter_coupling camp_second_messenger_amplifier cerebellar_forward_model cholinergic_attentional_filter circadian_gene_expression climbing_fiber_error_broadcast cochlear_compressive_nonlinearity cortical_area_specialization cortical_minicolumn_pipeline cortico_cortical_pathways corticotropin_releasing_hormone cortisol_slow_stress_recovery critical_period_learning_rate dendritic_compartmentalization endocannabinoid_retrograde enteric_glia_gut_brain ependymal_cell_barrier fusiform_face_service_registry gaba_inhibitory_bus gap_junction_electrical_sync ghrelin_hunger_signal glomerular_convergence_gateway glutamate_excitatory_bus glycine_coagonist_modulator granule_cell_inhibitory_relay hair_cell_banks_event_clusters hippocampal_4ec_loop_replay histamine_wakefulness_keeper hox_gene_service_specialization hypercolumn_module_federation hypercolumn_sharding hypothalamus_homeostasis insula_interoception_monitor ip3_inositol_cascade k_complex_event_trigger kcc2_chloride_shift_inhibitor leptin_satiety_signal locus_coeruleus_ne_global_signal melatonin_circadian_scheduler microglia_active_surveillance mitral_tufted_cell_dual_path morphogen_gradient_config muller_glia_retina_repair myelin_sheath_caching neural_crest_migration_deploy neuropeptide_y_stress_buffer ng2_glia_pool_renewal nitric_oxide_gas_signal node_of_ranvier_bypass nrem_slow_wave_cleanup nucleus_accumbens_reward_routing oligodendrocyte_myelination_dynamic orexin_stability_keeper orientation_column_indexing oscillatory_phase_locking_io oxytocin_trust_protocol parahippocampal_place_topology parallel_fiber_fanout_aggregation pineal_circadian_release pinwheel_central_layout pituitary_master_gland posterior_parietal_integration prolactin_parental_care quantal_release_batching radial_glia_neural_stem radial_glial_scaffold raphe_serotonin_rate_limit rem_paradoxical_processing replay_consolidation_trajectory reticular_activating_system retinotopic_data_layout satellite_glial_ganglion schwann_cell_peripheral_repair sleep_pressure_forced_maintenance sleep_spindle_memory_transfer slow_oscillation_sync subplate_wait_state suprachiasmatic_clock synaptic_vesicle_pool synaptogenesis_service_wiring tanycyte_metabolic_sensor temporal_pole_semantic_cache thalamocortical_loop_api tonotopic_stream_partitioning vasopressin_loyalty_aware_routing vta_dopamine_rpe_scheduler wernicke_area_api_parser ```
### 40 applied-neuroscience topics (`neuro_code_math/applied_neuroscience/`) Deep dives on computational neuroscience theories — the science Fractus's design draws from.
show all 40 topics ``` active_inference axonal_computation basal_ganglia_circuits bayesian_brain cerebellar_computation consolidation cortical_minicolumns cross_frequency_coupling dendritic_computation dopamine_reward entorhinal_grid_cells free_energy_principle gamma_oscillations global_workspace_theory head_direction_cells hierarchical_processing higher_order_theories hippocampal_formation homeostatic_plasticity integrated_information_theory long_term_depression long_term_potentiation metaplasticity neural_coding neural_decoding neural_manifolds neuromodulation place_cells population_coding predictive_coding predictive_processing rate_coding serotonin_modulation sharp_wave_ripples sleep_replay sparse_coding spike_timing_dependent_plasticity temporal_coding thalamic_reticular_nucleus theta_oscillations ```
### Foundational researchers & concepts honored in the corpus **Hebb** (Hebbian learning), **Bi & Poo** (STDP timing curves), **Friston** (free energy / active inference), **Buzsáki** (hippocampal sharp-wave ripples, replay), **Moser & Moser** (grid cells), **Hodgkin & Huxley** (axon dynamics), **Izhikevich** (spike models), **Tononi** (integrated information), **Baars/Dehaene** (global workspace), **O'Keefe** (place cells), **Kandel** (memory consolidation), plus neuromodulators (dopamine RPE, serotonin, oxytocin, vasopressin) and glial biology (astrocytes, microglia, oligodendrocytes, Schwann cells). ### Source files (all verified present) | File | Content | |---|---| | `neuro_paradigms_1b/*.jsonl.gz` (300) | 100 paradigms × 3 chunks, instruction+response+citations | | `neuro_code_math/applied_neuroscience__*.jsonl` (40) | Computational neuroscience deep-dives | | `datasets/neuro_arch_full.pt` | 60 neuroscience-grounded architecture paradigms | | `datasets/neuro_software_architecture.pt` | Same family, alternate cut | | `datasets/paradigms_full.pt` / `paradigms_dataset.pt` | 140 foundational + neuroscience paradigms | | `datasets/fractus_generated_corpus.pt` | Fractus ontology engine (neuroscience → AI) | --- ## How to Use ### Install ```bash git clone https://github.com/AFKmoney/fractus-cte.git cd fractus-cte pip install torch numpy tokenizers matplotlib fastapi uvicorn pydantic ``` ### Run tests ```bash pytest tests/ -q # → 28 passed ``` ### Train on CPU (progressive growth) ```bash python scripts/train_progressive.py --paliers 0,1,2,3 --accumulation-steps 8 ``` ### Train on GPU (1B scale) Fractus 1B training is designed to be **interruptible and resumable** — not a one-shot pretrain. The full corpus (~4.23B tokens) is the starting nutrition for the 1B palier. You can stop when budget ends, use the model, and continue later. That is how Fractus is meant to work. **Current production setup (multi-GPU, full corpus):** 1. Shard the full corpus across GPUs: ```bash python scripts/shard_corpus.py --src data/full_corpus.pt --out-dir data --n-shards 4 ``` 2. Launch one process per GPU (B=2, seq=128, torch.compile): ```bash bash scripts/launch_4gpu.sh # or manually: # GPU_ID=0 CUDA_VISIBLE_DEVICES=0 python -u scripts/train_1b_multi_gpu.py # GPU_ID=1 CUDA_VISIBLE_DEVICES=1 python -u scripts/train_1b_multi_gpu.py # ... ``` 3. Each GPU trains independently on its shard from the same `palier0` → grow to 1B. Checkpoints are saved per GPU (`checkpoints/fractus_1b_gpu{0-3}.pt`). 4. Merge the 4 shard-trained weights into one model that has seen the full corpus: ```bash # produces checkpoints/fractus_1b_merged.pt (weight average) ``` **Status (Aug 2026 run):** 4× RTX 5090, full 4.23B corpus sharded, B=2 + compile, loss on best GPU ~86 after ~11M tokens and still falling. Merged checkpoint on the Hub: `checkpoints/fractus_1b_merged.pt`. Training can be stopped and resumed at any time. > **Note:** A single-GPU full-corpus 1B run is not the supported path. The 1B model + full 4.23B corpus is trained by sharding across GPUs, then merging. One card can resume from a checkpoint later, but the production recipe is multi-GPU shard → merge. ### Use the agent ```python from fractus.continuous_engine import ContinuousThoughtEngine from fractus.memory import PersistentMemory from fractus.tokenizer import FractusTokenizer # Build the brain engine = ContinuousThoughtEngine( vocab_size=50257, d_model=128, n_heads=2, d_head=64, n_layers=2, n_levels=2, n_oscillators=8, coupling_rank=4, n_experts=4, top_k=2, expert_d_ff=128, siren_rank=32) # Give it memory memory = PersistentMemory(d_model=128, path="~/.fractus/memory.pt") engine.attach_memory(memory) # Think engine.reset_thought(batch_size=1) logits, confidence = engine.tick(torch.tensor([42])) print(f"Confidence: {confidence.item():.2f}") ``` --- ## The Growth Path | Stage | Size | Blocks | Experts | What it can do | |---|---|---|---|---| | Palier 0 | 6.6M | 1 | 4 | Learn basic patterns | | Palier 1 | 25M | 2 | 8 | Simple text generation | | Palier 2 | 120M | 4 | 16 | Coherent fragments | | Palier 3 | 350M | 8 | 32 | Decent text quality | | **Palier 4** | **1B** | **16** | **128** | **Full language model** | Each stage inherits the previous one's knowledge. The model never starts from zero. --- ## Architecture (for developers) ``` fractus-cte/ ├── fractus/ │ ├── continuous_engine.py ← The brain (CTE + CTEBlock) │ │ ├── CTEBlock One block: attention + Kuramoto + MoE │ │ └── ContinuousThoughtEngine Stacks N blocks, carries thought state │ ├── memory.py ← Cross-session persistent memory │ ├── cognitive_modes.py ← Unsupervised mental state detection │ ├── grow.py ← Progressive growth operator (width + depth + experts) │ ├── rag.py ← Knowledge base + plugins + metacognition │ ├── tokenizer.py ← GPT-2 BPE tokenizer │ ├── nn/ │ │ ├── moe.py ← PhaseRoutedMoE (sparse, low-rank, differentiable) │ │ ├── attention.py ← Multi-level causal linear attention │ │ ├── phase_ode.py ← Kuramoto RK4 oscillators │ │ └── lazy_siren.py ← Low-rank weight storage │ └── train/ │ └── online.py ← Online trainer (SGD/AdamW, accumulation) ├── tests/ 28 tests ├── scripts/ Training + corpus + GPU scripts ├── space/ HF Space demo ├── docs/ Optimization analysis ├── Fractus_White_Paper.pdf Technical white paper v2.0 └── arxiv/ LaTeX source for arXiv submission ``` --- ## Key Concepts **Tick**: one step of thinking. The engine processes an observation, updates its thought state through all blocks, and optionally emits output. **Thought state**: a vector that persists across ticks — the engine's "consciousness." **Chunk**: 32 tokens processed in one forward pass. The thought state and per-block attention state carry between chunks. **Expert**: a small neural network (low-rank `W = scale·U@V^T`) that specializes in certain thoughts. Only 2 of 128 active per token (sparse routing). **Kuramoto**: coupled oscillators producing phase vectors that route tokens to experts. The engine's "internal clock." --- ## Research Results (Honest) - **EDT** (Expert Decoupled Training): refuted. 5 variants, all ~19% worse. - **Forward-Forward** (Hinton 2022): refuted. Local learning can't replace global backprop. - **Progressive growth**: works. Warm start converges faster. - **Sparse MoE low-rank**: works. 2/128 experts = 64x less compute. - **1345 tok/s on CPU**: measured (batch=8 + SGD + all optimizations). --- ## License MIT. Fractus belongs to you, not to a corporation. ## Author **Philippe-Antoine Robert** — 2026 — rpa.tu@proton.me ## Links - **GitHub:** [github.com/AFKmoney/fractus-cte](https://github.com/AFKmoney/fractus-cte) - **HuggingFace Model:** [huggingface.co/thefinalboss/fractus-cte](https://huggingface.co/thefinalboss/fractus-cte) - **HuggingFace Datasets:** [huggingface.co/datasets/thefinalboss/fractus-datasets](https://huggingface.co/datasets/thefinalboss/fractus-datasets) - **White Paper:** [Fractus_White_Paper.pdf](Fractus_White_Paper.pdf) - **arXiv source:** [arxiv/main.tex](arxiv/main.tex) ## Training log (1B live run) See [docs/TRAINING_LOG_1B.md](docs/TRAINING_LOG_1B.md) for loss curves, rates, and mid-training generation probes (including raw failure-mode outputs). ## Discovery log See [docs/DISCOVERY_LOG.md](docs/DISCOVERY_LOG.md) for bugs found in production, routing surgery, and emergent features.