|
Download README.md from thefinalboss/fractus-cte: direct link, hf CLI and curl.
- Browser
- Download file 28.6 kB
-
https://huggingface.co/thefinalboss/fractus-cte/resolve/758d48a775816c517c4454ef04d80e965bf4b8b8/README.md
- Command line
-
hf download hf://thefinalboss/fractus-cte@758d48a775816c517c4454ef04d80e965bf4b8b8/README.md
-
curl -L -o README.md https://huggingface.co/thefinalboss/fractus-cte/resolve/758d48a775816c517c4454ef04d80e965bf4b8b8/README.md
28.6 kB
| license: mit | |
| language: | |
| - en | |
| - fr | |
| tags: | |
| - continuous-thought-engine | |
| - cognitive-agent | |
| - mixture-of-experts | |
| - kuramoto | |
| - self-modifying | |
| - progressive-growth | |
| - decentralized-ai | |
| - personal-ai | |
| - neuroscience | |
| library_name: pytorch | |
| pipeline_tag: text-generation | |
| models: | |
| - thefinalboss/fractus-cte | |
| datasets: | |
| - thefinalboss/fractus-datasets | |
| # Fractus CTE | |
| **A living AI that thinks continuously, remembers forever, and grows on its own.** | |
| **Fractus is NOT a transformer.** It's a Continuous Cognitive Agent β a dynamical system that maintains a persistent thought state, advances it tick by tick through 16 blocks, and routes via Kuramoto oscillator phases. The checkpoint is a living seed: it never freezes, grows at runtime, and trains forever. | |
| > **Status (2026-08-28):** Training live on QuickPod 8Γ5090 (boost_v4, ~89M/430M on GPU0). Decode surgery I: window-64 + anti-copy β Space `RetailΓ32` is dead, unique@32 now 12β27, still not English. Gate = unique@40 greedy PREFIX. See [`docs/2026-08-28-DECODE-WINDOW.md`](docs/2026-08-28-DECODE-WINDOW.md). | |
| ## Quick Start | |
| ```bash | |
| git clone https://github.com/AFKmoney/fractus-cte.git | |
| cd fractus-cte && pip install torch numpy tokenizers | |
| # Load the trained 1B and generate | |
| python -c " | |
| from fractus.continuous_engine import ContinuousThoughtEngine | |
| engine = ContinuousThoughtEngine.from_pretrained('checkpoints/fractus_1b_gpu3.pt') | |
| import torch | |
| logits, confidence = engine.tick(torch.tensor([42])) | |
| print(f'Fractus is thinking. Confidence: {confidence.item():.2f}') | |
| " | |
| ``` | |
| The `.pt` checkpoint contains the full model (weights + dynamic state). You need this repo's code to run it β Fractus is a custom architecture, not a transformer. Checkpoints are on [HF](https://huggingface.co/thefinalboss/fractus-cte). | |
|  | |
|  | |
|  | |
|  | |
|  | |
|  | |
|  | |
| --- | |
| ## What is Fractus? | |
| Fractus is not a chatbot. It's not GPT. It's not a transformer. | |
| Fractus is a **Continuous Cognitive Agent** β an AI that works like a brain, not a calculator. Instead of processing input β output in one pass, Fractus **ticks** like a biological system: it maintains a persistent thought state, advances it through multiple blocks of processing, remembers everything across sessions, and can grow new capacity by itself. | |
| ### What makes it different from GPT/Claude? | |
| | | GPT-4 / Claude | Fractus | | |
| |---|---|---| | |
| | **Thinking** | One pass, done | Continuous ticks (like a heartbeat) | | |
| | **Memory** | Forgets when context window fills | Remembers forever (survives restarts) | | |
| | **Learning** | Retrain from scratch ($$$) | Learns from every interaction | | |
| | **Growth** | Fixed size forever | Grows new experts at runtime | | |
| | **Mental states** | One mode always | Shifts between cognitive modes | | |
| | **Where it runs** | Corporate cloud | Your machine | | |
| | **Training** | Fixed, done once | Perpetual, never stops | | |
| --- | |
| ## Architecture (1.05B total / ~119M active per token) | |
| ``` | |
| ContinuousThoughtEngine | |
| βββ d_model=1280, 16 layers, 16 heads | |
| βββ FractalLinearAttention multi-level causal linear attention (O(L)), | |
| β carry state (S, z) persists across chunks | |
| βββ Kuramoto phase clock RK4-integrated oscillators β phase vectors | |
| β (learned end-to-end; feeds routing) | |
| βββ PhaseRoutedMoE 128 experts, top-2 active per token, | |
| β von Mises gate over phases, Farey-sequence | |
| β expert phases, load-balance loss in objective | |
| βββ Low-rank experts W = scale Β· U@Vα΅ (rank 64) β 64Γ less compute | |
| βββ Tied embedding/head GPT-2 BPE vocab (50257) | |
| βββ Persistent thought state residual stream carried across ticks | |
| ``` | |
| **Parameter accounting:** ~1.05B total, but only ~118.8M are active per token (dense attention + top-2 of 128 sparse experts) β ~183.1M including the tied embedding. An 11.3% sparsity ratio. This is the basis of the adapted scaling target below. | |
| ### The 12 Building Blocks | |
| | Block | What it does | | |
| |---|---| | |
| | **Continuous Thought Engine** | The brain β thinks tick by tick through 16 blocks | | |
| | **Persistent Memory** | Vector bank surviving sessions, cosine recall, 5% blend injection, salience head gates impact | | |
| | **Cognitive Modes** | Mental states discovered unsupervised (k-means on Kuramoto phase features): focused, creative, exploratory | | |
| | **RAG Knowledge Base** | Learns facts instantly β no retraining needed | | |
| | **Cognitive Plugins** | Hot-swappable modes: analyst, coder, creative, teacher | | |
| | **MetaCognition** | Decides its own actions: retrieve, learn, generate | | |
| | **Progressive Growth** | Grows from 6M to 1B+ params, palier by palier (`maybe_grow`: width + depth + experts) | | |
| | **Self-Modification** | Adds new experts at runtime when routing is imbalanced (zero-init, placed near the dominant expert) | | |
| | **PhaseRoutedMoE** | Sparse experts routed by oscillator phases | | |
| | **Kuramoto Clock** | A dynamical system that drives routing decisions | | |
| | **Online Trainer** | Learns continuously, one chunk at a time | | |
| | **Vision (`tick_vec`)** | Multimodal input path β image patches drive the engine directly, bypassing the token embedding (CIFAR "eyes" prototype trained while the text run continued) | | |
| --- | |
| ## Live Training Status (28 August 2026) | |
| **Hardware:** QuickPod 8Γ RTX 5090. One independent Python process per GPU. No DDP / no gradient sync. Consolidation = mean-merge of the 8 `.pt` when we need a single brain. | |
| **This pass (Phase 2 shards):** | |
| - 8 GPT-2 BPE int32 memmaps, `~429,896,462` tokens/GPU β dataset: [thefinalboss/fractus-datasets](https://huggingface.co/datasets/thefinalboss/fractus-datasets) `tokenized/phase2/shard_phase2_gpu{i}.npy` | |
| - Resume point **2026-08-28 14:40 UTC** after RunPod SSH death: GPU0 **79,922,176** tokens (~18.6% of this pass). Other GPUs 76.2Mβ78.4M. Exact JSON: `checkpoints/x8run/RESUME_gpu{i}.json` | |
| - Trainer: `scripts/fast4gpu_boost_v4.py` | |
| - Config: **BATCH=4**, SEQ=128, LR 7e-4, SGD+momentum 0.9, bf16, TF32, compile off, `FRACTUS_ATTN_IMPL=chunked`, `BLOCK_CKPT=1`, **SS_RATE=1.0**, SS_PROB 0.2β0.5 over 50M tokens, REPEAT_COEF=0.1, P0 on, **PROBE_EVERY=0** (live unique@40 probe crashed the first QuickPod launch β probes are offline) | |
| - Signals: `tf`/`ema_tf`, `ss`/`ema_ss`, `rep`/`ema_rep`, `lb` (~14) | |
| - Throughput after resume: ~840β940 tok/s/GPU. Remainder of this pass β **4.1β4.5 days** if 8 GPUs stay up | |
| - **Torch on RTX 5090:** image default `2.2.1+cu121` has **no Blackwell kernels**. Required: **torch β₯ 2.11 + cu128** | |
| **Crash recovery:** hourly upload of `checkpoints/x8run/fractus_1b_gpu{i}.pt` + `RESUME_gpu{i}.json`. That folder is the source of truth. See [`docs/2026-08-28-QUICKPOD-RESUME.md`](docs/2026-08-28-QUICKPOD-RESUME.md). | |
| **Speech gate (not the loss):** `unique@40` greedy PREFIX β 40 generated tokens, count distinct ids, no ban, no temperature. Post window-surgery unique@32 (2026-08-28 23:35 UTC): 16 / 14 / 27 / 12. Mono-token lock broken. Short cycles remain. **NO-GO for speech.** Teacher-force CE going down does **not** mean Fractus speaks. Do not cite banned-decode diversity as speech. | |
| ## Vision β the CIFAR "eyes" prototype | |
| Fractus is not text-only. The engine exposes **`tick_vec(obs_vec)`**: a multimodal entry point that accepts a precomputed `(B, d_model)` vector and injects it directly into the residual thought state β **bypassing the token embedding entirely**. Any modality that can be embedded into `d_model` dimensions can drive continuous thought: images, audio, sensor streams. | |
| ``` | |
| image β PatchEmbed β (B, d_model) patch vectors | |
| β | |
| βΌ | |
| engine.tick_vec(patch) β no tokenizer involved | |
| β | |
| thought state advances through all 16 blocks | |
| (same Kuramoto routing + MoE as text) | |
| ``` | |
| **Proof of concept β CIFAR-10 eyes:** a small CTE + PatchEmbed stack (`fractus/nn/vision.py`) was trained on CPU on real CIFAR-10 images (`fractus_eyes_cifar_final.pt`), demonstrating that image patches can drive the continuous-thought loop. Crucially, this ran as a **parallel track**: the 8-GPU text run was never interrupted for eyes work β GPU text digestion and CPU vision learning proceed simultaneously on the same living system. | |
| This mirrors the biological premise: eyes evolved as a peripheral sense feeding a central dynamical brain, not as a core feature of it. The text 1B is the brain; vision is a front-end that plugs into `tick_vec`. | |
| --- | |
| ### Adapted Chinchilla target | |
| Standard Chinchilla (20 tokens/param) assumes dense transformers where all parameters train on every token. Fractus violates that: only ~119M of 1.05B params are active per token, and the Kuramoto clock trains end-to-end now. The adapted target is **20 Γ active params β 2.4β3.7B tokens**; with a warm-started checkpoint, the practical target is **~2.5β3B tokens** β which is why the phase-2 corpus is 3.44B. Since Fractus grows continuously (`maybe_grow`), this is an instantaneous cumulative-exposure requirement that grows with the model, not a final-stop condition. Details: [`docs/2026-08-12-fractus-chinchilla.md`](docs/2026-08-12-fractus-chinchilla.md). | |
| --- | |
| ## The Surgeries (mid-training interventions, no weight wipes) | |
| A defining discovery of this run: **the brain (.pt) and the code are separable.** Bottlenecks were fixed by live surgery β save the checkpoint, patch the code, reload weights, resume at the exact recorded token offset. Multi-day digestion is never thrown away. | |
| | Phase | Intervention | Result | | |
| |---|---|---| | |
| | **A β Initial** | Last-position-only CE, 4 independent GPUs | Loss fell; generation collapsed into single-token loops | | |
| | **B β Routing surgery** | Kuramoto was frozen under `no_grad` (order parameter r β 0.01β0.03); LB loss was detached β ~70% of experts dead. Fixed: gradients enabled (state kept detached for carry), CE + 0.02Β·lb, gate temp 1.0 β 2.5, omega scale Γ4 | lb β 14 live on all GPUs; experts alive | | |
| | **C β Dense CE** | Replaced last-position CE (1 target / 128 tokens) with CE over all positions | Sharp loss drop; token-to-token chaining enforced | | |
| | **D β Decode surgery** | Phase/thought noise, frequency penalties, cycle bans, forced escape tokens | Loop lock broken; still no coherent English | | |
| | **E β Train/gen mismatch** | Training used causal attention + RK4 Kuramoto; generation used a simpler single-tick Euler path. Aligned decode via `generate_aligned.py` | Decode now matches the training path | | |
| | **F β Loss recalibration** | Cumulative-average CE was misleading near 2.0 β batch CE + EMA; LR 1e-3 β 5e-4 | Trustworthy metrics | | |
| | **G β Scheduled sampling** | Two-pass training: TF CE + LB, plus SS steps mixing model samples into inputs | Live; now SS_RATE=1.0 on boost_v4 | | |
| | **H β Pod death / exact resume** | RunPod SSH died 2026-08-28. Reloaded 8Γ `.pt` + manifests on a new QuickPod 8Γ5090 at the recorded `start_token_next`. Torch upgraded 2.2β2.11+cu128 for Blackwell | Weights kept. Pass continues from ~80M/430M | | |
| | **I β Decode window + anti-copy** | Length-1 `tick_chunk` + carry locked Space to `RetailΓ32`. Default decode is now causal window 64 + mask previous token. Weights / train loop unchanged | unique@32 1β16 / 1β14 / 1β27 / 2β12. Short cycles remain | | |
| Full logs: [`docs/MASTER_RUN_LOG.md`](docs/MASTER_RUN_LOG.md), [`docs/DISCOVERY_LOG.md`](docs/DISCOVERY_LOG.md), [`docs/OPERABILITY_MIDTRAIN.md`](docs/OPERABILITY_MIDTRAIN.md). | |
| ### Composable checkpoints | |
| Because shapes stay compatible and manifests record token offsets, the following operations are proven on this run: | |
| - **Parallel independent training** β N GPUs on separate shards | |
| - **Mean-merge** β per-GPU checkpoints averaged into one unified model that still generates (424/440 tensors on the 4-GPU merge; stateful buffers reset) | |
| - **Iterative fusion** β train β merge β train cycles; the model absorbs compatible checkpoints and keeps going | |
| - **Exact-token resume** β across pod reboots and driver crashes | |
| - **Compositional growth β structural growth** β weight averaging vs `maybe_grow` paliers | |
| Details: [`docs/COMPOSABILITY_AND_SURGERY.md`](docs/COMPOSABILITY_AND_SURGERY.md). | |
| --- | |
| ## Trusted Loss (reading the numbers) | |
| Three metrics, three different questions: | |
| | Metric | What it measures | Where | | |
| |---|---|---| | |
| | `ema_tf` | CE with ground-truth history + continuous internal state β is the model digesting data? | live | | |
| | `ema_ss` | CE after scheduled sampling mixes model-generated tokens in β partial free-run robustness | live | | |
| | **AR** | Warm on 32 true tokens, greedily free-run 32 steps, CE vs truth β actual generation quality | offline | | |
| **A single CE number cannot represent both teacher-forced learning and free-run generation.** Live ema_tf after the QuickPod resume is mid-teens on the lead GPUs (first steps inflate EMA β wait). AR stays the offline generation metric and last unique@40 was NO-GO. Random-vocab CE β 10.82; AR in the thousands (random guessing over the GPT-2 vocab is ~10.82): free-running compounds every error, and teacher forcing always supplies the correct past. Trust `ema_tf`/`ema_ss` for learning progress, AR + text probes for generation progress. The convergence signal is AR falling toward the ss/tf order of magnitude, plus readable output. Details: [`docs/TRUSTED_LOSS.md`](docs/TRUSTED_LOSS.md). | |
| --- | |
| ## Research Results (Honest) | |
| **Refuted:** | |
| - **EDT** (Expert Decoupled Training) β all 5 variants ~19β20% worse than from-scratch; MSE objective misaligned with CE, router ignores half the experts. | |
| - **Forward-Forward** (Hinton 2022) β NLL rose from 124 to 221; local objectives can't replace global backprop here. | |
| **Validated:** | |
| - **Progressive growth** β warm start converges faster; trained through palier 3 (350M, loss 23.0, ~5 days CPU). | |
| - **Sparse low-rank MoE** β 2/128 experts = 64Γ less compute. | |
| - **Open-heart operability** β live surgery on a training model works (see above). | |
| - **Routing pathology as a first-class debug target** β expert-hit histograms and the Kuramoto order parameter catch failures that loss curves hide (they caught the frozen clock and the dead experts). | |
| **Training optimizations (measured):** tied head (~1.1Γ), head-partial training (~2Γ), sparse gathered low-rank MoE (8Γ at 16 experts, 64Γ at 128), detached-state Kuramoto, gradient accumulation (~1.4Γ), SGD+momentum over AdamW (~1.37Γ), batching (335 β 1345 tok/s at B=8), bf16 (~2Γ). Combined CPU: ~336Γ over naive. Measured on CPU: 707 tok/s single-stream, **1345 tok/s batched**. | |
| --- | |
| ## Datasets (4.2 Billion Tokens) | |
| Source of truth: [huggingface.co/datasets/thefinalboss/fractus-datasets](https://huggingface.co/datasets/thefinalboss/fractus-datasets). | |
| | Dataset | Tokens | Content | | |
| |---|---|---| | |
| | **neuro-paradigms-1b** | **~1B** | 100 neuroscience β software architecture paradigms (300 chunked files) | | |
| | **neuro-code-math** | **~900M** | Neuro-inspired coding, mathematics, algorithms (incl. 40 applied-neuroscience topics) | | |
| | **cognitive-skills** | **~780M** | Coding, reasoning, speaking, thinking, understanding | | |
| | **fractus-generated-corpus** | **340M** | Bilingual FR/EN generated by Fractus ontology engine | | |
| | **paradigms-full** | **191M** | 140 paradigms (neuroscience, CS, architecture) | | |
| | **gutenberg-esoteric** | **~58M** | 487 public-domain esoteric / masonic / hermetic books | | |
| | **neuro-arch-full** | **86M** | 60 neuroscience paradigms (neuro-software-architecture) | | |
| | **all-github-repos** | **54M+** | 80+ repos (public + private, secret-filtered) | | |
| | **mega-corpus-v3** | **20M** | Literature, philosophy, occult, masonry, science, medicine | | |
| | **wordnet** | **3M** | 117K dictionary synset entries | | |
| | **Total** | **~4.2B** | | | |
| Tokenized streams: Phase 1 = ~1.52B tokens (`.pt` files); Phase 2 = ~3.44B tokens (8 GPT-2 BPE int32 memmap shards). Phases are kept separate to avoid re-ingesting the same ordered stream. | |
| --- | |
| ## Applied Neuroscience β the theoretical core | |
| Fractus is a neuroscience-grounded architecture: real brain mechanisms are mapped to software/AI patterns, and that mapping is itself training data. Every entry below is present in the dataset β verified by file listing, not just claimed. | |
| ### 100 neuroscience β software-architecture paradigms (`neuro_paradigms_1b`, 300 chunked files) | |
| Each paradigm maps a biological mechanism to an engineering pattern (e.g. *adenosine sleep pressure* β cache-stampede recovery; *myelin sheath* β caching; *hippocampal replay* β trajectory consolidation). | |
| <details><summary><b>show all 100 paradigms</b></summary> | |
| ``` | |
| adenosine_sleep_pressure amygdala_prefrontal_topdown anterior_cingulate_conflict_monitor | |
| apoptosis_self_destructing_service arc_gene_plasticity_marker astrocyte_tripartite_synapse | |
| axon_initial_segment_trigger basal_ganglia_loop_arbitration bdnf_growth_factor_scaling | |
| bergmann_glia_purkinje binaural_cross_correlation_localization brainstem_vital_functions | |
| broca_area_api_generator calcium_transmitter_coupling camp_second_messenger_amplifier | |
| cerebellar_forward_model cholinergic_attentional_filter circadian_gene_expression | |
| climbing_fiber_error_broadcast cochlear_compressive_nonlinearity cortical_area_specialization | |
| cortical_minicolumn_pipeline cortico_cortical_pathways corticotropin_releasing_hormone | |
| cortisol_slow_stress_recovery critical_period_learning_rate dendritic_compartmentalization | |
| endocannabinoid_retrograde enteric_glia_gut_brain ependymal_cell_barrier | |
| fusiform_face_service_registry gaba_inhibitory_bus gap_junction_electrical_sync | |
| ghrelin_hunger_signal glomerular_convergence_gateway glutamate_excitatory_bus | |
| glycine_coagonist_modulator granule_cell_inhibitory_relay hair_cell_banks_event_clusters | |
| hippocampal_4ec_loop_replay histamine_wakefulness_keeper hox_gene_service_specialization | |
| hypercolumn_module_federation hypercolumn_sharding hypothalamus_homeostasis | |
| insula_interoception_monitor ip3_inositol_cascade k_complex_event_trigger | |
| kcc2_chloride_shift_inhibitor leptin_satiety_signal locus_coeruleus_ne_global_signal | |
| melatonin_circadian_scheduler microglia_active_surveillance mitral_tufted_cell_dual_path | |
| morphogen_gradient_config muller_glia_retina_repair myelin_sheath_caching | |
| neural_crest_migration_deploy neuropeptide_y_stress_buffer ng2_glia_pool_renewal | |
| nitric_oxide_gas_signal node_of_ranvier_bypass nrem_slow_wave_cleanup | |
| nucleus_accumbens_reward_routing oligodendrocyte_myelination_dynamic orexin_stability_keeper | |
| orientation_column_indexing oscillatory_phase_locking_io oxytocin_trust_protocol | |
| parahippocampal_place_topology parallel_fiber_fanout_aggregation pineal_circadian_release | |
| pinwheel_central_layout pituitary_master_gland posterior_parietal_integration | |
| prolactin_parental_care quantal_release_batching radial_glia_neural_stem | |
| radial_glial_scaffold raphe_serotonin_rate_limit rem_paradoxical_processing | |
| replay_consolidation_trajectory reticular_activating_system retinotopic_data_layout | |
| satellite_glial_ganglion schwann_cell_peripheral_repair sleep_pressure_forced_maintenance | |
| sleep_spindle_memory_transfer slow_oscillation_sync subplate_wait_state | |
| suprachiasmatic_clock synaptic_vesicle_pool synaptogenesis_service_wiring | |
| tanycyte_metabolic_sensor temporal_pole_semantic_cache thalamocortical_loop_api | |
| tonotopic_stream_partitioning vasopressin_loyalty_aware_routing vta_dopamine_rpe_scheduler | |
| wernicke_area_api_parser | |
| ``` | |
| </details> | |
| ### 40 applied-neuroscience topics (`neuro_code_math/applied_neuroscience/`) | |
| Deep dives on computational neuroscience theories β the science Fractus's design draws from. | |
| <details><summary><b>show all 40 topics</b></summary> | |
| ``` | |
| active_inference axonal_computation basal_ganglia_circuits bayesian_brain | |
| cerebellar_computation consolidation cortical_minicolumns cross_frequency_coupling | |
| dendritic_computation dopamine_reward entorhinal_grid_cells free_energy_principle | |
| gamma_oscillations global_workspace_theory head_direction_cells hierarchical_processing | |
| higher_order_theories hippocampal_formation homeostatic_plasticity integrated_information_theory | |
| long_term_depression long_term_potentiation metaplasticity neural_coding | |
| neural_decoding neural_manifolds neuromodulation place_cells | |
| population_coding predictive_coding predictive_processing rate_coding | |
| serotonin_modulation sharp_wave_ripples sleep_replay sparse_coding | |
| spike_timing_dependent_plasticity temporal_coding thalamic_reticular_nucleus theta_oscillations | |
| ``` | |
| </details> | |
| ### Foundational researchers & concepts honored in the corpus | |
| **Hebb** (Hebbian learning), **Bi & Poo** (STDP timing curves), **Friston** (free energy / active inference), **BuzsΓ‘ki** (hippocampal sharp-wave ripples, replay), **Moser & Moser** (grid cells), **Hodgkin & Huxley** (axon dynamics), **Izhikevich** (spike models), **Tononi** (integrated information), **Baars/Dehaene** (global workspace), **O'Keefe** (place cells), **Kandel** (memory consolidation), plus neuromodulators (dopamine RPE, serotonin, oxytocin, vasopressin) and glial biology (astrocytes, microglia, oligodendrocytes, Schwann cells). | |
| --- | |
| ## How to Use | |
| ### Install | |
| ```bash | |
| git clone https://github.com/AFKmoney/fractus-cte.git | |
| cd fractus-cte | |
| pip install torch numpy tokenizers matplotlib fastapi uvicorn pydantic | |
| ``` | |
| ### Run tests | |
| ```bash | |
| pytest tests/ -q | |
| ``` | |
| ### Train on CPU (progressive growth) | |
| ```bash | |
| python scripts/train_progressive.py --paliers 0,1,2,3 --accumulation-steps 8 | |
| ``` | |
| ### Train the 1B (GPU, sharded) | |
| ```bash | |
| # Phase-2 1B on 8 GPUs β resume from HF x8run (do not restart at token 0) | |
| # 1) torch >= 2.11+cu128 on RTX 5090 | |
| # 2) download checkpoints/x8run/*.pt + RESUME_gpu*.json | |
| # 3) download tokenized/phase2/shard_phase2_gpu{i}.npy, symlink to data/shard_gpu{i}.npy | |
| CUDA_VISIBLE_DEVICES=$i GPU_ID=$i START_TOKEN=<from RESUME_gpu$i.json> \ | |
| BATCH=4 SEQ=128 CE_CHUNK=2048 FRACTUS_ATTN_IMPL=chunked BLOCK_CKPT=1 \ | |
| SS_RATE=1.0 P0=1 COMPILE=0 PROBE_EVERY=0 \ | |
| CKPT_IN=checkpoints/x8run/fractus_1b_gpu$i.pt \ | |
| CKPT_OUT=checkpoints/x8run/fractus_1b_gpu$i.pt \ | |
| SHARD=data/shard_gpu$i.npy \ | |
| python -u scripts/fast4gpu_boost_v4.py | |
| ``` | |
| Speech metric (offline, not in the trainer loop): unique@40 greedy PREFIX. Gate β mean unique β₯ 20 and not a single-token lock. | |
| ### Use the agent | |
| ```python | |
| from fractus.continuous_engine import ContinuousThoughtEngine | |
| from fractus.memory import PersistentMemory | |
| from fractus.tokenizer import FractusTokenizer | |
| # Build the brain | |
| engine = ContinuousThoughtEngine( | |
| vocab_size=50257, d_model=128, n_heads=2, d_head=64, | |
| n_layers=2, n_levels=2, n_oscillators=8, coupling_rank=4, | |
| n_experts=4, top_k=2, expert_d_ff=128, siren_rank=32) | |
| # Give it memory | |
| memory = PersistentMemory(d_model=128, path="~/.fractus/memory.pt") | |
| engine.attach_memory(memory) | |
| # Think | |
| engine.reset_thought(batch_size=1) | |
| logits, confidence = engine.tick(torch.tensor([42])) | |
| print(f"Confidence: {confidence.item():.2f}") | |
| # Vision: drive thought with an image patch vector (no tokenizer) | |
| patch_vec = torch.randn(1, 128) # any (B, d_model) embedding | |
| logits, confidence = engine.tick_vec(patch_vec) | |
| ``` | |
| --- | |
| ## The Growth Path | |
| | Stage | Size | Blocks | Experts | What it can do | | |
| |---|---|---|---|---| | |
| | Palier 0 | 6.6M | 1 | 4 | Learn basic patterns | | |
| | Palier 1 | 25M | 2 | 8 | Simple text generation | | |
| | Palier 2 | 120M | 4 | 16 | Coherent fragments | | |
| | Palier 3 | 350M | 8 | 32 | Decent text quality | | |
| | **Palier 4** | **1B** | **16** | **128** | **Full language model (training now)** | | |
| Each stage inherits the previous one's knowledge via zero-padded warm starts. The model never starts from zero. | |
| --- | |
| ## Repository Layout | |
| ``` | |
| fractus-cte/ | |
| βββ fractus/ | |
| β βββ continuous_engine.py β The brain (CTE + CTEBlock) | |
| β βββ memory.py β Cross-session persistent memory | |
| β βββ cognitive_modes.py β Unsupervised mental state detection | |
| β βββ grow.py β Progressive growth (width + depth + experts) | |
| β βββ rag.py β Knowledge base + plugins + metacognition | |
| β βββ tokenizer.py β GPT-2 BPE tokenizer | |
| β βββ nn/ | |
| β β βββ moe.py β PhaseRoutedMoE (sparse, low-rank) | |
| β β βββ attention.py β Multi-level causal linear attention + carry | |
| β β βββ phase_ode.py β Kuramoto RK4 oscillators | |
| β β βββ lazy_siren.py β Low-rank weight storage | |
| β βββ train/online.py β Online trainer | |
| βββ tests/ | |
| βββ scripts/ fast_tokenize, shard_corpus, launch_4gpu, | |
| β fast4gpu_boost_v4, generate_aligned, hourly HF x8run sync | |
| βββ checkpoints/ per-GPU + merged + recovery alias | |
| βββ docs/ run logs, surgeries, trusted loss, scaling | |
| βββ space/ HF Space demo | |
| βββ Fractus_White_Paper_v2.md White paper v2.0 | |
| βββ arxiv/ LaTeX source for arXiv submission | |
| ``` | |
| --- | |
| ## Key Concepts | |
| **Tick**: one step of thinking. The engine processes an observation, updates its thought state through all blocks, and optionally emits output. | |
| **Thought state**: a vector that persists across ticks β the engine's "consciousness." | |
| **Chunk**: 32 tokens processed in one forward pass. The thought state and per-block attention state carry between chunks. | |
| **Expert**: a small low-rank network (`W = scaleΒ·U@Vα΅`) that specializes in certain thoughts. Only 2 of 128 active per token. | |
| **Kuramoto clock**: coupled oscillators producing phase vectors that route tokens to experts β learned end-to-end since the routing surgery. | |
| **`tick_vec`**: the multimodal tick β feed any precomputed `(B, d_model)` vector (image patches, embeddings from any encoder) straight into the thought state without tokens. This is how vision plugs in. | |
| **Surgery**: patching code around a preserved checkpoint mid-training, resuming at the exact token offset. Weights are never wiped for a routing, objective, or decode bug. | |
| --- | |
| ## Limitations (stated plainly) | |
| - Generation is not yet coherent English β word-level repetition loops / lexical noise. Exposure bias is being addressed by scheduled sampling; AR is the metric to watch. | |
| - GPT-2 vocab dominates parameters (81% at d=768) β vocab reduction is a known lever. | |
| - "Remembers forever" and "grows on its own" describe the architecture's design; no independent benchmarks are provided. | |
| - This is a research artifact and a live training run, not a production assistant. | |
| --- | |
| ## License | |
| MIT. Fractus belongs to you, not to a corporation. | |
| ## Author | |
| **Philippe-Antoine Robert** β 2026 β rpa.tu@proton.me | |
| ## Links | |
| - **GitHub:** [github.com/AFKmoney/fractus-cte](https://github.com/AFKmoney/fractus-cte) | |
| - **HuggingFace Model:** [huggingface.co/thefinalboss/fractus-cte](https://huggingface.co/thefinalboss/fractus-cte) | |
| - **HuggingFace Datasets:** [huggingface.co/datasets/thefinalboss/fractus-datasets](https://huggingface.co/datasets/thefinalboss/fractus-datasets) | |
| - **White Paper:** [Fractus_White_Paper_v2.md](Fractus_White_Paper_v2.md) / [PDF](Fractus_White_Paper.pdf) | |
| - **Run logs:** [MASTER_RUN_LOG](docs/MASTER_RUN_LOG.md) Β· [TRAINING_LOG_1B](docs/TRAINING_LOG_1B.md) Β· [DISCOVERY_LOG](docs/DISCOVERY_LOG.md) | |
| - **arXiv source:** [arxiv/main.tex](arxiv/main.tex) | |