File size: 15,384 Bytes
a80f736
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
# Fractus White Paper v2.0

**A Continuous Thought Engine with Multi-Block Depth, Self-Modification, and Progressive Growth**

---

**Author:** Philippe-Antoine Robert
**Contact:** rpa.tu@proton.me
**Date:** August 6, 2026
**Version:** 2.0
**Repository:** github.com/AFKmoney/fractus-test
**Model Hub:** huggingface.co/thefinalboss/Fractus-1B
**License:** MIT

---

## Abstract

I present Fractus v2.0 β€” a continuous cognitive agent architecture that departs fundamentally from the transformer paradigm. Unlike static models that map input to output in a single forward pass, Fractus is a **dynamical system** that maintains a persistent thought state, advances it tick by tick through a multi-block residual stack, and emits output only when it has something confident to say.

This version introduces three structural advances over the original:

1. **Multi-block depth** β€” the Continuous Thought Engine (CTE) now stacks N blocks, each with its own attention state (S,z), Kuramoto oscillator phases, and PhaseRoutedMoE. The thought flows through the stack as a residual stream, with per-block state carried continuously across chunk boundaries.

2. **Progressive growth** β€” instead of training a large model from scratch, Fractus grows palier by palier (width + depth + experts), inheriting previous weights via zero-padding. Each palier starts warm and converges faster.

3. **Runtime self-modification** β€” the model detects routing imbalance and grows new experts while it runs, with zero-init stability validated.

I also report **negative results** honestly: Expert Decoupled Training (EDT) and the Forward-Forward algorithm were both tested and refuted β€” their objectives are misaligned with the final cross-entropy loss. The only training method that works is standard gradient descent, but the architectural optimizations I describe (sparse low-rank MoE, head-partial training, gradient accumulation, Kuramoto detachment) achieve **707 tokens/second on a consumer CPU** β€” a 177x improvement over the baseline.

---

## 1. Introduction

Contemporary large language models (GPT-4, Claude, Llama) share four fundamental limitations: they are static functions (one forward pass per output), stateless (no memory between conversations), generic (one monolithic network for all tasks), and centralized (training requires datacenter GPUs).

Fractus challenges each of these assumptions. The Continuous Thought Engine replaces the static function with a dynamical system. Persistent Memory gives the engine a cross-session memory bank. Expert Specialization forces each MoE expert to own a distinct skill domain. And the LazyStructuredSiren compression combined with progressive growth enables training on consumer hardware.

The question is not whether Fractus matches GPT-4 on benchmarks. It does not. The question is whether the paradigm of continuous, personal, decentralized AI is viable. This work demonstrates that it is β€” with measured, reproducible results.

---

## 2. Architecture

### 2.1 The Continuous Thought Engine (CTE)

The CTE is a dynamical system that maintains a persistent thought state `h ∈ R^{d_model}` and advances it tick by tick. The engine stacks `n_layers` blocks, each refining the thought:

```
h β†’ [Block 0: norm β†’ attn β†’ +residual β†’ norm β†’ kuramoto β†’ phases β†’ norm β†’ moe β†’ +residual]
  β†’ [Block 1: ... ]
  β†’ ...
  β†’ [Block N: ... ] β†’ thought_state = h_final
```

Each `CTEBlock` owns:
- **FractalLinearAttention** (Katharopoulos 2020) β€” multi-level causal linear attention with a persistent state (S, z) that accumulates across ticks and chunk boundaries.
- **Kuramoto oscillators** β€” a coupled dynamical system (low-rank RK4) that acts as a "consciousness clock," producing phase vectors that route MoE experts.
- **PhaseRoutedMoE** β€” a sparse mixture-of-experts with von Mises gate on Farey-distributed expert phases.

The thought state `h` is a **residual stream** β€” each block adds its transformation. The attention state (S, z) is **per-block** and **continuous across chunk boundaries** (verified: S grows monotonically across chunks, never reset).

### 2.2 PhaseRoutedMoE (Sparse, Low-Rank, Differentiable)

The MoE routes tokens via Kuramoto oscillator phases through a von Mises gate:

```
g_e = exp(ΞΊ Β· cos(ΞΈ_token βˆ’ ΞΈ_expert)) / Ξ£_e' g_e'
```

Expert phases are drawn from the Farey sequence F_{2E}, providing E angles in [0, 2Ο€) that are dense, non-collapsing, and deterministic. Only `top_k=2` experts are computed per token (gather-first sparse dispatch).

**Low-rank experts**: each expert weight matrix W is decomposed as `W = scale Β· U@Vα΅€` (rank r=64). The forward pass is two cheap matmuls that never materialize the full matrix. The sparse path gathers only the top-k experts' U/V factors, computing K experts instead of E. At 128 experts with top-k=2, this is 64x less compute.

**Self-modification**: `add_expert()` grows a new expert at runtime, placed near the dominant expert's phase (to capture overflow traffic), with zero-init (no forward perturbation). `maybe_grow()` triggers automatically when routing imbalance exceeds a threshold.

### 2.3 Multi-Block Depth

The original CTE (v1.0) had a single attention + Kuramoto + MoE block. This was the deepest limitation: depth 1 means the thought traverses one refinement then exits.

v2.0 introduces the `CTEBlock` abstraction. The engine stacks N blocks, each with independent state. The thought flows through all of them as a residual stream. This is the path to 1B+ parameters:

| Config | d_model | blocks | experts/block | params |
|---|---|---|---|---|
| Palier 0 | 128 | 1 | 4 | 6.6M |
| Palier 1 | 256 | 2 | 8 | ~25M |
| Palier 2 | 512 | 4 | 16 | ~120M |
| Palier 3 | 768 | 8 | 32 | ~350M |
| Palier 4 (1B) | 1280 | 16 | 128 | ~1B |

### 2.4 Persistent Memory

The engine maintains a bank of memory vectors (d_model-dimensional, with context labels and importance scores) that survives across sessions. Memories are recalled via cosine similarity and injected into the thought state at 5% blend (continuous injection).

A **salience head** (Linear(d_model β†’ 1)) learns to predict how much a memory injection will perturb the thought state β€” an intrinsic signal, not an external label. The system discovers its own sensitivity to memories.

### 2.5 Cognitive Modes

Cognitive modes emerge from the Kuramoto phase dynamics via unsupervised k-means clustering on phase features (synchronization degree r, mean phase, variance, per-oscillator sin/cos). No external labels β€” the modes are discovered from the structure of the phase space.

### 2.6 Self-Modification

Fractus is the only model that grows new capacity while it runs:

```python
engine.maybe_grow()
# β†’ "[Fractus] Self-modified: grew expert in all 16 blocks (now 129 experts)"
```

The new expert is zero-initialized (scale=0 β†’ output=0 β†’ no gradient spike), placed near the dominant expert (captures overflow traffic), and warms up gradually via backprop. Validated: new expert receives 50% of traffic, loss stable post-grow.

### 2.7 Progressive Growth

Instead of training 1B from scratch (months on GPU), Fractus grows palier by palier. Each palier:
1. Inherits the previous model's weights via zero-padding (`fractus/grow.py`)
2. Old knowledge preserved (top-left block of every matrix)
3. New capacity starts neutral (zeros for weights, ones for LayerNorm gamma)
4. Trains briefly to adapt the new dimensions

This is how a brain develops: small at first, growing new capacity on top of existing knowledge.

---

## 3. Training

### 3.1 Online Training

The CTE trains online: one chunk (32 tokens) per forward, one backward per chunk. The thought state carries forward (detached β€” no BPTT). Each chunk's attention state (S,z) starts from the previous chunk's accumulated state β€” continuous thought.

### 3.2 Training Optimizations (Measured)

| Optimization | What it does | Impact |
|---|---|---|
| Tied head | `output_head.weight = observe.weight` | Halves vocab params |
| Head-partial (`tick_chunk_train`) | Head on 1 position instead of C | 32x less head FLOPs |
| Sparse MoE low-rank | Only top-k experts computed (gather-first) | 64x at 128 experts |
| Gradient accumulation (accum=8) | 8x fewer optimizer steps | 16x fewer AdamW calls |
| Chunk_len=32 | Better Python amortization | ~1.1x |
| Detach Kuramoto | Phase computation in no_grad | Removes backward through clock |
| bf16 AMP (GPU) | 2x on all matmuls | GPU only |

**Combined measured result**: 4 tok/s β†’ **707 tok/s** on CPU (palier 0, d=128).

### 3.3 Profile Breakdown

Per-iteration cost at d=128, chunk_len=32:

| Component | Time (ms) | % of forward |
|---|---|---|
| Attention (QKV + causal vectorized) | 21.8 | 32% |
| Kuramoto (detached) | 0.0 | 0% |
| MoE (4 experts, dense) | 16.6 | 25% |
| Output head (1 position, tied) | 12.6 | 19% |
| Embedding + norms | 16.3 | 24% |

At 128 experts, the sparse MoE path reduces MoE cost by 64x, making attention the dominant cost β€” as it should be for a reasoning architecture.

---

## 4. Negative Results (Honest)

### 4.1 EDT (Expert Decoupled Training) β€” Refuted

EDT claimed 189x training speedup by pre-training experts independently. I tested 5 variants (vanilla, denoise, identity, residual objectives + routing filter) on the 13M CTE:

| Variant | Hold-out PPL | vs From-scratch |
|---|---|---|
| From-scratch | 1309.7 | β€” |
| EDT next_hidden | 1562.1 | +19.3% worse |
| EDT denoise + routing filter | 1555.1 | +18.7% worse |
| EDT identity + routing filter | 1556.8 | +18.9% worse |
| EDT residual + routing filter | 1573.7 | +20.2% worse |

**Root cause**: two structural defects. (1) The Phase-1 MSE objective (predict next hidden state) is misaligned with the Phase-3 CE objective (next-token prediction) β€” Pearson correlation never positive across 12 configs. (2) The Kuramoto router concentrates traffic on 2/4 experts, so half of pre-trained experts are never routed.

### 4.2 Forward-Forward (Hinton 2022) β€” Refuted for CTE

The Forward-Forward algorithm (local goodness signal, no global backprop) was adapted to the CTE. Result: NLL went UP (124 β†’ 221). The goodness signal (sum of squared activations) is not aligned with cross-entropy. Local learning objectives cannot replace global backprop for this architecture.

### 4.3 What Works

Only standard gradient descent (CE + backprop) produces a model that learns. The optimizations described in Β§3.2 are the path to making this feasible.

---

## 5. Experimental Results

### 5.1 Progressive Growth (CPU)

| Palier | d_model | blocks | params | tokens | loss | tok/s | time |
|---|---|---|---|---|---|---|---|
| 0 | 128 | 1 | 6.6M | 2M | 32.5 | 725 | 46 min |
| 1 | 256 | 2 | 25M | 1.5M | 27.1 | 395 | 63 min |
| 2 | 512 | 4 | 120M | 1M | 27.1 | 192 | 87 min |
| 3 | 768 | 8 | 350M | 500k | 23.0 | 122 | 68 min |

Each palier starts warm (inherited weights) and converges. The corpus includes Fractus's own source code (palimpseste principle: the model contains its own description).

### 5.2 Self-Modification Stability

After runtime `add_expert()` at tick 500 (4β†’5 experts):
- New expert receives 50% of routing traffic (placed near dominant)
- Gradient norm stable (zero-init β†’ no spike)
- Loss post-grow: +24.6% (vs control no-grow: +67.6%) β€” growth helps stability

### 5.3 Continuous Thought Verification

Attention state (S,z) verified to grow monotonically across chunk boundaries in all three paths (tick, tick_chunk, tick_chunk_train). The thought is truly continuous.

### 5.4 Cross-Session Memory

Session 1: 4 memories captured and saved to disk.
Session 2 (fresh engine): 4 memories loaded, thought state displaced by 100.16 units on first tick.
Cross-session persistence verified.

---

## 6. Comparison with GPT and Claude

| Property | GPT-4 / Claude | Fractus v2.0 |
|---|---|---|
| Processing | Static (1 forward) | Continuous (ticks through N blocks) |
| Memory | Context window | Persistent bank + salience-gated injection |
| Skills | Generic monolith | Specialized MoE experts (128 per block) |
| Mental state | Stateless | Cognitive modes (unsupervised) |
| Generation | Token-by-token | Plan then fill + adaptive depth |
| Training | Datacenter GPUs | Consumer CPU (progressive growth) + GPU for 1B |
| Deployment | Cloud API | Local device |
| User data | Sent to server | Stays local |
| Self-modification | None | Runtime expert growth |
| Depth scaling | Retrain from scratch | Progressive growth (warm start) |

---

## 7. Limitations and Future Work

1. **Model quality**: The trained model at palier 3 (350M, 500k tokens) produces repetitive text. More data and training are needed for coherent generation. Chinchilla-optimal (940M tokens) requires ~5 days on GPU.

2. **Multi-block training**: The multi-block architecture is validated (gradient flows through all blocks, continuous thought works) but has not yet been trained at scale. Palier 4 (16 blocks, 128 experts, 1B params) awaits GPU compute.

3. **EDT and Forward-Forward**: Both refuted. Alternative training acceleration methods must align their objective with the final CE loss.

4. **Vocabulary**: The GPT-2 BPE vocab (50257) dominates parameters (81% at d=768). A reduced vocab (8k-16k) would cut head FLOPs by 3-6x.

5. **Corpus**: The quality corpus (20.5M tokens) includes Fractus's own source code but is far below Chinchilla scale for the larger paliers.

---

## 8. Conclusion

Fractus v2.0 demonstrates that a continuous, multi-block, self-modifying cognitive agent can be constructed and progressively trained on consumer hardware. The architecture β€” CTEBlock stack with per-block continuous attention state, PhaseRoutedMoE with sparse low-rank experts, progressive growth via zero-padding, and runtime self-modification β€” is validated by 28 tests and measured benchmarks.

The negative results (EDT, Forward-Forward) are reported honestly. They do not weaken the architecture; they clarify what works (global backprop + architectural optimizations) and what does not (decoupled/local training).

The implications extend beyond performance metrics. If AI can be trained and deployed on any laptop, the centralization of intelligence by a handful of corporations is not inevitable. Fractus is a proof of concept for decentralized AI: intelligence that belongs to the user, runs on their hardware, remembers them, and grows.

This work is released as open source under the MIT license. All code, training scripts, datasets, and measured results are available at github.com/AFKmoney/fractus-test.

---

## References

[1] Katharopoulos et al. (2020). Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention. ICML.
[2] Sitzmann et al. (2020). Implicit Neural Representations with Periodic Activation Functions (SIREN). NeurIPS.
[3] Hinton, G. (2022). The Forward-Forward Algorithm: Some Preliminary Investigations.
[4] Kuramoto, Y. (1984). Chemical Oscillations, Waves, and Turbulence. Springer.
[5] Hinton, G. (2022). The Forward-Forward Algorithm.
[6] Rahimi & Recht (2007). Random Features for Large-Scale Kernel Machines. NeurIPS.
[7] Graves, A. (2016). Adaptive Computation Time for Recurrent Neural Networks. arXiv.
[8] Hu et al. (2021). LoRA: Low-Rank Adaptation of Large Language Models. arXiv.

---

*Β© 2026 Philippe-Antoine Robert. MIT License. Contact: rpa.tu@proton.me*