File size: 7,417 Bytes
3f2e79e | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 | # Fractus CTE
**A living AI that thinks continuously, remembers forever, and grows on its own.**
---
## What is Fractus?
Fractus is not a chatbot. It's not GPT. It's not a transformer.
Fractus is a **Continuous Cognitive Agent** β an AI that works like a brain, not a calculator. Instead of processing input β output in one pass, Fractus **ticks** like a biological system: it maintains a persistent thought state, advances it through multiple blocks of processing, remembers everything across sessions, and can grow new capacity by itself.
### What makes it different from GPT/Claude?
| | GPT-4 / Claude | Fractus |
|---|---|---|
| **Thinking** | One pass, done | Continuous ticks (like a heartbeat) |
| **Memory** | Forgets when context window fills | Remembers forever (survives restarts) |
| **Learning** | Retrain from scratch ($$$) | Learns from every interaction |
| **Growth** | Fixed size forever | Grows new experts at runtime |
| **Mental states** | One mode always | Shifts between cognitive modes |
| **Where it runs** | Corporate cloud | Your machine |
---
## The 12 Building Blocks
| Block | What it does |
|---|---|
| **Continuous Thought Engine** | The brain β thinks tick by tick through 16 blocks |
| **Persistent Memory** | Remembers you across sessions, never forgets |
| **Cognitive Modes** | Shifts mental states (focused, creative, exploratory...) |
| **RAG Knowledge Base** | Learns facts instantly β no retraining needed |
| **Cognitive Plugins** | Hot-swappable modes: analyst, coder, creative, teacher |
| **MetaCognition** | Decides its own actions: retrieve, learn, generate |
| **Progressive Growth** | Grows from 6M to 1B+ params, palier by palier |
| **Self-Modification** | Adds new experts at runtime when it needs them |
| **PhaseRoutedMoE** | Sparse experts routed by oscillator phases |
| **Kuramoto Clock** | A dynamical system that drives routing decisions |
| **Online Trainer** | Learns continuously, one chunk at a time |
| **HF Space** | Live chat demo with shared memory |
---
## How to Use
### Install
```bash
git clone https://github.com/AFKmoney/fractus-cte.git
cd fractus-cte
pip install torch numpy tokenizers matplotlib fastapi uvicorn pydantic
```
### Run tests
```bash
pytest tests/ -q
# β 28 passed (CTE + memory + MoE + multi-block + continuous thought)
```
### Build a corpus
```bash
python scripts/build_quality_corpus.py
```
### Train on CPU (progressive growth)
```bash
# Paliers 0-3: grows from 6M to 350M params
python scripts/train_progressive.py --paliers 0,1,2,3 --accumulation-steps 8
```
### Train on GPU (1B scale)
```bash
# Grows from palier 3 checkpoint to 1B, then trains
python scripts/train_1b_gpu.py \
--checkpoint checkpoints/fractus_palier3.pt \
--tokens 500000000 \
--batch-size 8 \
--bf16 \
--accumulation-steps 4
```
### Use the agent
```python
from fractus.continuous_engine import ContinuousThoughtEngine
from fractus.memory import PersistentMemory
from fractus.tokenizer import FractusTokenizer
# Build the brain
engine = ContinuousThoughtEngine(
vocab_size=50257, d_model=128, n_heads=2, d_head=64,
n_layers=2, n_levels=2, n_oscillators=8, coupling_rank=4,
n_experts=4, top_k=2, expert_d_ff=128, siren_rank=32)
# Give it memory
memory = PersistentMemory(d_model=128, path="~/.fractus/memory.pt")
engine.attach_memory(memory)
# Think
engine.reset_thought(batch_size=1)
logits, confidence = engine.tick(torch.tensor([42]))
print(f"Confidence: {confidence.item():.2f}")
```
---
## The Growth Path
Fractus grows like a brain β small at first, bigger over time:
| Stage | Size | Blocks | Experts | What it can do |
|---|---|---|---|---|
| Palier 0 | 6.6M | 1 | 4 | Learn basic patterns |
| Palier 1 | 25M | 2 | 8 | Simple text generation |
| Palier 2 | 120M | 4 | 16 | Coherent fragments |
| Palier 3 | 350M | 8 | 32 | Decent text quality |
| **Palier 4** | **1B** | **16** | **128** | **Full language model** |
Each stage inherits the previous one's knowledge. The model never starts from zero.
---
## Architecture (for developers)
```
fractus-cte/
βββ fractus/
β βββ continuous_engine.py β The brain (CTE + CTEBlock)
β β βββ CTEBlock One block: attention + Kuramoto + MoE
β β βββ ContinuousThoughtEngine Stacks N blocks, carries thought state
β βββ memory.py β Cross-session persistent memory
β βββ cognitive_modes.py β Unsupervised mental state detection
β βββ grow.py β Progressive growth operator
β βββ rag.py β Knowledge base + plugins + metacognition
β βββ tokenizer.py β GPT-2 BPE tokenizer
β βββ nn/
β β βββ moe.py β PhaseRoutedMoE (sparse, low-rank, differentiable)
β β βββ attention.py β Multi-level causal linear attention
β β βββ phase_ode.py β Kuramoto RK4 oscillators
β β βββ lazy_siren.py β Low-rank weight storage
β βββ train/
β βββ online.py β Online trainer (SGD/AdamW, accumulation)
βββ tests/ 28 tests
βββ scripts/ Training + corpus + GPU scripts
βββ space/ HF Space demo
βββ docs/ Optimization analysis
βββ Fractus_White_Paper.pdf Technical white paper v2.0
βββ arxiv/ LaTeX source for arXiv submission
```
### Key concepts
**Tick**: one step of thinking. The engine processes an observation, updates its thought state through all blocks, and optionally emits output.
**Thought state**: a vector `h β R^d_model` that persists across ticks. It's the engine's "consciousness" β it carries context forward.
**Chunk**: 32 tokens processed in one forward pass (for speed). The thought state and per-block attention state carry between chunks.
**Expert**: a small neural network (low-rank `W = scaleΒ·U@V^T`) that specializes in certain types of thoughts. Only 2 out of 128 are active per token (sparse routing).
**Kuramoto**: coupled oscillators that produce phase vectors. These phases route tokens to the right experts. Think of it as the engine's "internal clock" β different phase patterns = different cognitive modes.
---
## Research Results (Honest)
We tested alternative training methods. Both failed:
- **Expert Decoupled Training (EDT)**: claimed 189x speedup. Reality: 19% worse than standard training. The pre-training objective doesn't align with the final task.
- **Forward-Forward (Hinton 2022)**: local goodness signal. Reality: the model got worse. Local learning can't replace global backpropagation.
**What works**: standard gradient descent + our architectural optimizations = **1345 tokens/second on CPU** (was 4 tok/s before).
---
## License
MIT. Fractus belongs to you, not to a corporation.
## Author
**Philippe-Antoine Robert** β 2026 β rpa.tu@proton.me
## Links
- **GitHub:** [github.com/AFKmoney/fractus-cte](https://github.com/AFKmoney/fractus-cte)
- **HuggingFace:** [huggingface.co/thefinalboss/fractus-cte](https://huggingface.co/thefinalboss/fractus-cte)
- **White Paper:** [Fractus_White_Paper.pdf](Fractus_White_Paper.pdf)
- **arXiv source:** [arxiv/main.tex](arxiv/main.tex)
|