|
Download README.md from thefinalboss/fractus-cte: direct link, hf CLI and curl.
- Browser
- Download file 17.8 kB
-
https://huggingface.co/thefinalboss/fractus-cte/resolve/12e938bbf3793fc5c60d0988b2bd1034081c233b/README.md
- Command line
-
hf download hf://thefinalboss/fractus-cte@12e938bbf3793fc5c60d0988b2bd1034081c233b/README.md
-
curl -L -o README.md https://huggingface.co/thefinalboss/fractus-cte/resolve/12e938bbf3793fc5c60d0988b2bd1034081c233b/README.md
17.8 kB
| license: mit | |
| language: | |
| - en | |
| - fr | |
| tags: | |
| - continuous-thought-engine | |
| - cognitive-agent | |
| - mixture-of-experts | |
| - kuramoto | |
| - self-modifying | |
| - progressive-growth | |
| - decentralized-ai | |
| - personal-ai | |
| - neuroscience | |
| library_name: pytorch | |
| pipeline_tag: text-generation | |
| models: | |
| - thefinalboss/fractus-cte | |
| datasets: | |
| - thefinalboss/fractus-datasets | |
| # Fractus CTE | |
| **A living AI that thinks continuously, remembers forever, and grows on its own.** | |
|  | |
|  | |
|  | |
|  | |
|  | |
|  | |
| --- | |
| ## What is Fractus? | |
| Fractus is not a chatbot. It's not GPT. It's not a transformer. | |
| Fractus is a **Continuous Cognitive Agent** β an AI that works like a brain, not a calculator. Instead of processing input β output in one pass, Fractus **ticks** like a biological system: it maintains a persistent thought state, advances it through multiple blocks of processing, remembers everything across sessions, and can grow new capacity by itself. | |
| ### What makes it different from GPT/Claude? | |
| | | GPT-4 / Claude | Fractus | | |
| |---|---|---| | |
| | **Thinking** | One pass, done | Continuous ticks (like a heartbeat) | | |
| | **Memory** | Forgets when context window fills | Remembers forever (survives restarts) | | |
| | **Learning** | Retrain from scratch ($$$) | Learns from every interaction | | |
| | **Growth** | Fixed size forever | Grows new experts at runtime | | |
| | **Mental states** | One mode always | Shifts between cognitive modes | | |
| | **Where it runs** | Corporate cloud | Your machine | | |
| | **Training** | Fixed, done once | Perpetual, never stops | | |
| --- | |
| ## The 12 Building Blocks | |
| | Block | What it does | | |
| |---|---| | |
| | **Continuous Thought Engine** | The brain β thinks tick by tick through 16 blocks | | |
| | **Persistent Memory** | Remembers you across sessions, never forgets | | |
| | **Cognitive Modes** | Shifts mental states (focused, creative, exploratory...) | | |
| | **RAG Knowledge Base** | Learns facts instantly β no retraining needed | | |
| | **Cognitive Plugins** | Hot-swappable modes: analyst, coder, creative, teacher | | |
| | **MetaCognition** | Decides its own actions: retrieve, learn, generate | | |
| | **Progressive Growth** | Grows from 6M to 1B+ params, palier by palier | | |
| | **Self-Modification** | Adds new experts at runtime when it needs them | | |
| | **PhaseRoutedMoE** | Sparse experts routed by oscillator phases | | |
| | **Kuramoto Clock** | A dynamical system that drives routing decisions | | |
| | **Online Trainer** | Learns continuously, one chunk at a time | | |
| | **HF Space** | Live chat demo with shared memory | | |
| --- | |
| ## Datasets (4.15 Billion Tokens) | |
| Fractus is trained on a massive, diverse corpus available at [huggingface.co/datasets/thefinalboss/fractus-datasets](https://huggingface.co/datasets/thefinalboss/fractus-datasets): | |
| | Dataset | Tokens | Content | | |
| |---|---|---| | |
| | **neuro-paradigms-1b** | **~1B** | 100 neuroscience β software architecture paradigms (300 chunked files) | | |
| | **neuro-code-math** | **~900M** | Neuro-inspired coding, mathematics, algorithms (incl. 40 applied-neuroscience topics) | | |
| | **cognitive-skills** | **~780M** | Coding, reasoning, speaking, thinking, understanding | | |
| | **fractus-generated-corpus** | **340M** | Bilingual FR/EN generated by Fractus ontology engine | | |
| | **paradigms-full** | **191M** | 140 paradigms (neuroscience, CS, architecture) | | |
| | **gutenberg-esoteric** | **~58M** | 487 public-domain esoteric / masonic / hermetic books | | |
| | **neuro-arch-full** | **86M** | 60 neuroscience paradigms (neuro-software-architecture) | | |
| | **all-github-repos** | **54M+** | 80+ of your GitHub repos (public + private, secret-filtered) | | |
| | **mega-corpus-v3** | **20M** | Literature, philosophy, occult, masonry, science, medicine | | |
| | **wordnet** | **3M** | 117K dictionary synset entries | | |
| | **Total** | **~4.2B** | | | |
| The corpus covers neuroscience, software architecture, philosophy, psychology, literature, esoteric traditions, programming, medicine, and Fractus's own source code. | |
| --- | |
| ## Applied Neuroscience β the theoretical core | |
| Fractus is a neuroscience-grounded architecture: real brain mechanisms are mapped to software/AI patterns, and that mapping is itself training data. Every entry below is present in the dataset (`neuro_paradigms_1b/`, `neuro_code_math/applied_neuroscience/`, and the `*.pt` files in `datasets/`) β verified by file listing, not just claimed. | |
| ### 100 neuroscience β software-architecture paradigms (`neuro_paradigms_1b`, 300 chunked files) | |
| Each paradigm maps a biological mechanism to an engineering pattern (e.g. *adenosine sleep pressure* β cache-stampede recovery; *myelin sheath* β caching; *hippocampal replay* β trajectory consolidation). | |
| <details><summary><b>show all 100 paradigms</b></summary> | |
| ``` | |
| adenosine_sleep_pressure amygdala_prefrontal_topdown anterior_cingulate_conflict_monitor | |
| apoptosis_self_destructing_service arc_gene_plasticity_marker astrocyte_tripartite_synapse | |
| axon_initial_segment_trigger basal_ganglia_loop_arbitration bdnf_growth_factor_scaling | |
| bergmann_glia_purkinje binaural_cross_correlation_localization brainstem_vital_functions | |
| broca_area_api_generator calcium_transmitter_coupling camp_second_messenger_amplifier | |
| cerebellar_forward_model cholinergic_attentional_filter circadian_gene_expression | |
| climbing_fiber_error_broadcast cochlear_compressive_nonlinearity cortical_area_specialization | |
| cortical_minicolumn_pipeline cortico_cortical_pathways corticotropin_releasing_hormone | |
| cortisol_slow_stress_recovery critical_period_learning_rate dendritic_compartmentalization | |
| endocannabinoid_retrograde enteric_glia_gut_brain ependymal_cell_barrier | |
| fusiform_face_service_registry gaba_inhibitory_bus gap_junction_electrical_sync | |
| ghrelin_hunger_signal glomerular_convergence_gateway glutamate_excitatory_bus | |
| glycine_coagonist_modulator granule_cell_inhibitory_relay hair_cell_banks_event_clusters | |
| hippocampal_4ec_loop_replay histamine_wakefulness_keeper hox_gene_service_specialization | |
| hypercolumn_module_federation hypercolumn_sharding hypothalamus_homeostasis | |
| insula_interoception_monitor ip3_inositol_cascade k_complex_event_trigger | |
| kcc2_chloride_shift_inhibitor leptin_satiety_signal locus_coeruleus_ne_global_signal | |
| melatonin_circadian_scheduler microglia_active_surveillance mitral_tufted_cell_dual_path | |
| morphogen_gradient_config muller_glia_retina_repair myelin_sheath_caching | |
| neural_crest_migration_deploy neuropeptide_y_stress_buffer ng2_glia_pool_renewal | |
| nitric_oxide_gas_signal node_of_ranvier_bypass nrem_slow_wave_cleanup | |
| nucleus_accumbens_reward_routing oligodendrocyte_myelination_dynamic orexin_stability_keeper | |
| orientation_column_indexing oscillatory_phase_locking_io oxytocin_trust_protocol | |
| parahippocampal_place_topology parallel_fiber_fanout_aggregation pineal_circadian_release | |
| pinwheel_central_layout pituitary_master_gland posterior_parietal_integration | |
| prolactin_parental_care quantal_release_batching radial_glia_neural_stem | |
| radial_glial_scaffold raphe_serotonin_rate_limit rem_paradoxical_processing | |
| replay_consolidation_trajectory reticular_activating_system retinotopic_data_layout | |
| satellite_glial_ganglion schwann_cell_peripheral_repair sleep_pressure_forced_maintenance | |
| sleep_spindle_memory_transfer slow_oscillation_sync subplate_wait_state | |
| suprachiasmatic_clock synaptic_vesicle_pool synaptogenesis_service_wiring | |
| tanycyte_metabolic_sensor temporal_pole_semantic_cache thalamocortical_loop_api | |
| tonotopic_stream_partitioning vasopressin_loyalty_aware_routing vta_dopamine_rpe_scheduler | |
| wernicke_area_api_parser | |
| ``` | |
| </details> | |
| ### 40 applied-neuroscience topics (`neuro_code_math/applied_neuroscience/`) | |
| Deep dives on computational neuroscience theories β the science Fractus's design draws from. | |
| <details><summary><b>show all 40 topics</b></summary> | |
| ``` | |
| active_inference axonal_computation basal_ganglia_circuits bayesian_brain | |
| cerebellar_computation consolidation cortical_minicolumns cross_frequency_coupling | |
| dendritic_computation dopamine_reward entorhinal_grid_cells free_energy_principle | |
| gamma_oscillations global_workspace_theory head_direction_cells hierarchical_processing | |
| higher_order_theories hippocampal_formation homeostatic_plasticity integrated_information_theory | |
| long_term_depression long_term_potentiation metaplasticity neural_coding | |
| neural_decoding neural_manifolds neuromodulation place_cells | |
| population_coding predictive_coding predictive_processing rate_coding | |
| serotonin_modulation sharp_wave_ripples sleep_replay sparse_coding | |
| spike_timing_dependent_plasticity temporal_coding thalamic_reticular_nucleus theta_oscillations | |
| ``` | |
| </details> | |
| ### Foundational researchers & concepts honored in the corpus | |
| **Hebb** (Hebbian learning), **Bi & Poo** (STDP timing curves), **Friston** (free energy / active inference), **BuzsΓ‘ki** (hippocampal sharp-wave ripples, replay), **Moser & Moser** (grid cells), **Hodgkin & Huxley** (axon dynamics), **Izhikevich** (spike models), **Tononi** (integrated information), **Baars/Dehaene** (global workspace), **O'Keefe** (place cells), **Kandel** (memory consolidation), plus neuromodulators (dopamine RPE, serotonin, oxytocin, vasopressin) and glial biology (astrocytes, microglia, oligodendrocytes, Schwann cells). | |
| ### Source files (all verified present) | |
| | File | Content | | |
| |---|---| | |
| | `neuro_paradigms_1b/*.jsonl.gz` (300) | 100 paradigms Γ 3 chunks, instruction+response+citations | | |
| | `neuro_code_math/applied_neuroscience__*.jsonl` (40) | Computational neuroscience deep-dives | | |
| | `datasets/neuro_arch_full.pt` | 60 neuroscience-grounded architecture paradigms | | |
| | `datasets/neuro_software_architecture.pt` | Same family, alternate cut | | |
| | `datasets/paradigms_full.pt` / `paradigms_dataset.pt` | 140 foundational + neuroscience paradigms | | |
| | `datasets/fractus_generated_corpus.pt` | Fractus ontology engine (neuroscience β AI) | | |
| --- | |
| ## How to Use | |
| ### Install | |
| ```bash | |
| git clone https://github.com/AFKmoney/fractus-cte.git | |
| cd fractus-cte | |
| pip install torch numpy tokenizers matplotlib fastapi uvicorn pydantic | |
| ``` | |
| ### Run tests | |
| ```bash | |
| pytest tests/ -q | |
| # β 28 passed | |
| ``` | |
| ### Train on CPU (progressive growth) | |
| ```bash | |
| python scripts/train_progressive.py --paliers 0,1,2,3 --accumulation-steps 8 | |
| ``` | |
| ### Train on GPU (1B scale) | |
| Fractus 1B training is designed to be **interruptible and resumable** β not a one-shot pretrain. The full corpus (~4.23B tokens) is the starting nutrition for the 1B palier. You can stop when budget ends, use the model, and continue later. That is how Fractus is meant to work. | |
| **Current production setup (multi-GPU, full corpus):** | |
| 1. Shard the full corpus across GPUs: | |
| ```bash | |
| python scripts/shard_corpus.py --src data/full_corpus.pt --out-dir data --n-shards 4 | |
| ``` | |
| 2. Launch one process per GPU (B=2, seq=128, torch.compile): | |
| ```bash | |
| bash scripts/launch_4gpu.sh | |
| # or manually: | |
| # GPU_ID=0 CUDA_VISIBLE_DEVICES=0 python -u scripts/train_1b_multi_gpu.py | |
| # GPU_ID=1 CUDA_VISIBLE_DEVICES=1 python -u scripts/train_1b_multi_gpu.py | |
| # ... | |
| ``` | |
| 3. Each GPU trains independently on its shard from the same `palier0` β grow to 1B. Checkpoints are saved per GPU (`checkpoints/fractus_1b_gpu{0-3}.pt`). | |
| 4. Merge the 4 shard-trained weights into one model that has seen the full corpus: | |
| ```bash | |
| # produces checkpoints/fractus_1b_merged.pt (weight average) | |
| ``` | |
| **Status (Aug 2026 run):** 4Γ RTX 5090, full 4.23B corpus sharded, B=2 + compile, loss on best GPU ~86 after ~11M tokens and still falling. Merged checkpoint on the Hub: `checkpoints/fractus_1b_merged.pt`. Training can be stopped and resumed at any time. | |
| > **Note:** A single-GPU full-corpus 1B run is not the supported path. The 1B model + full 4.23B corpus is trained by sharding across GPUs, then merging. One card can resume from a checkpoint later, but the production recipe is multi-GPU shard β merge. | |
| ### Use the agent | |
| ```python | |
| from fractus.continuous_engine import ContinuousThoughtEngine | |
| from fractus.memory import PersistentMemory | |
| from fractus.tokenizer import FractusTokenizer | |
| # Build the brain | |
| engine = ContinuousThoughtEngine( | |
| vocab_size=50257, d_model=128, n_heads=2, d_head=64, | |
| n_layers=2, n_levels=2, n_oscillators=8, coupling_rank=4, | |
| n_experts=4, top_k=2, expert_d_ff=128, siren_rank=32) | |
| # Give it memory | |
| memory = PersistentMemory(d_model=128, path="~/.fractus/memory.pt") | |
| engine.attach_memory(memory) | |
| # Think | |
| engine.reset_thought(batch_size=1) | |
| logits, confidence = engine.tick(torch.tensor([42])) | |
| print(f"Confidence: {confidence.item():.2f}") | |
| ``` | |
| --- | |
| ## The Growth Path | |
| | Stage | Size | Blocks | Experts | What it can do | | |
| |---|---|---|---|---| | |
| | Palier 0 | 6.6M | 1 | 4 | Learn basic patterns | | |
| | Palier 1 | 25M | 2 | 8 | Simple text generation | | |
| | Palier 2 | 120M | 4 | 16 | Coherent fragments | | |
| | Palier 3 | 350M | 8 | 32 | Decent text quality | | |
| | **Palier 4** | **1B** | **16** | **128** | **Full language model** | | |
| Each stage inherits the previous one's knowledge. The model never starts from zero. | |
| --- | |
| ## Architecture (for developers) | |
| ``` | |
| fractus-cte/ | |
| βββ fractus/ | |
| β βββ continuous_engine.py β The brain (CTE + CTEBlock) | |
| β β βββ CTEBlock One block: attention + Kuramoto + MoE | |
| β β βββ ContinuousThoughtEngine Stacks N blocks, carries thought state | |
| β βββ memory.py β Cross-session persistent memory | |
| β βββ cognitive_modes.py β Unsupervised mental state detection | |
| β βββ grow.py β Progressive growth operator (width + depth + experts) | |
| β βββ rag.py β Knowledge base + plugins + metacognition | |
| β βββ tokenizer.py β GPT-2 BPE tokenizer | |
| β βββ nn/ | |
| β β βββ moe.py β PhaseRoutedMoE (sparse, low-rank, differentiable) | |
| β β βββ attention.py β Multi-level causal linear attention | |
| β β βββ phase_ode.py β Kuramoto RK4 oscillators | |
| β β βββ lazy_siren.py β Low-rank weight storage | |
| β βββ train/ | |
| β βββ online.py β Online trainer (SGD/AdamW, accumulation) | |
| βββ tests/ 28 tests | |
| βββ scripts/ Training + corpus + GPU scripts | |
| βββ space/ HF Space demo | |
| βββ docs/ Optimization analysis | |
| βββ Fractus_White_Paper.pdf Technical white paper v2.0 | |
| βββ arxiv/ LaTeX source for arXiv submission | |
| ``` | |
| --- | |
| ## Key Concepts | |
| **Tick**: one step of thinking. The engine processes an observation, updates its thought state through all blocks, and optionally emits output. | |
| **Thought state**: a vector that persists across ticks β the engine's "consciousness." | |
| **Chunk**: 32 tokens processed in one forward pass. The thought state and per-block attention state carry between chunks. | |
| **Expert**: a small neural network (low-rank `W = scaleΒ·U@V^T`) that specializes in certain thoughts. Only 2 of 128 active per token (sparse routing). | |
| **Kuramoto**: coupled oscillators producing phase vectors that route tokens to experts. The engine's "internal clock." | |
| --- | |
| ## Research Results (Honest) | |
| - **EDT** (Expert Decoupled Training): refuted. 5 variants, all ~19% worse. | |
| - **Forward-Forward** (Hinton 2022): refuted. Local learning can't replace global backprop. | |
| - **Progressive growth**: works. Warm start converges faster. | |
| - **Sparse MoE low-rank**: works. 2/128 experts = 64x less compute. | |
| - **1345 tok/s on CPU**: measured (batch=8 + SGD + all optimizations). | |
| --- | |
| ## License | |
| MIT. Fractus belongs to you, not to a corporation. | |
| ## Author | |
| **Philippe-Antoine Robert** β 2026 β rpa.tu@proton.me | |
| ## Links | |
| - **GitHub:** [github.com/AFKmoney/fractus-cte](https://github.com/AFKmoney/fractus-cte) | |
| - **HuggingFace Model:** [huggingface.co/thefinalboss/fractus-cte](https://huggingface.co/thefinalboss/fractus-cte) | |
| - **HuggingFace Datasets:** [huggingface.co/datasets/thefinalboss/fractus-datasets](https://huggingface.co/datasets/thefinalboss/fractus-datasets) | |
| - **White Paper:** [Fractus_White_Paper.pdf](Fractus_White_Paper.pdf) | |
| - **arXiv source:** [arxiv/main.tex](arxiv/main.tex) | |
| ## Training log (1B live run) | |
| See [docs/TRAINING_LOG_1B.md](docs/TRAINING_LOG_1B.md) for loss curves, rates, and mid-training generation probes (including raw failure-mode outputs). | |
| ## Discovery log | |
| See [docs/DISCOVERY_LOG.md](docs/DISCOVERY_LOG.md) for bugs found in production, routing surgery, and emergent features. | |
| ## Mid-training operability | |
| See [docs/OPERABILITY_MIDTRAIN.md](docs/OPERABILITY_MIDTRAIN.md). | |
| ## Master run log | |
| Full production chronology: [docs/MASTER_RUN_LOG.md](docs/MASTER_RUN_LOG.md) | |
| ## Composability | |
| In-training surgery and checkpoint merge/grow: [docs/COMPOSABILITY_AND_SURGERY.md](docs/COMPOSABILITY_AND_SURGERY.md) | |