thefinalboss commited on
Commit
a80f736
Β·
verified Β·
1 Parent(s): 4e4b252

Upload Fractus_White_Paper_v2.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. Fractus_White_Paper_v2.md +274 -0
Fractus_White_Paper_v2.md ADDED
@@ -0,0 +1,274 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Fractus White Paper v2.0
2
+
3
+ **A Continuous Thought Engine with Multi-Block Depth, Self-Modification, and Progressive Growth**
4
+
5
+ ---
6
+
7
+ **Author:** Philippe-Antoine Robert
8
+ **Contact:** rpa.tu@proton.me
9
+ **Date:** August 6, 2026
10
+ **Version:** 2.0
11
+ **Repository:** github.com/AFKmoney/fractus-test
12
+ **Model Hub:** huggingface.co/thefinalboss/Fractus-1B
13
+ **License:** MIT
14
+
15
+ ---
16
+
17
+ ## Abstract
18
+
19
+ I present Fractus v2.0 β€” a continuous cognitive agent architecture that departs fundamentally from the transformer paradigm. Unlike static models that map input to output in a single forward pass, Fractus is a **dynamical system** that maintains a persistent thought state, advances it tick by tick through a multi-block residual stack, and emits output only when it has something confident to say.
20
+
21
+ This version introduces three structural advances over the original:
22
+
23
+ 1. **Multi-block depth** β€” the Continuous Thought Engine (CTE) now stacks N blocks, each with its own attention state (S,z), Kuramoto oscillator phases, and PhaseRoutedMoE. The thought flows through the stack as a residual stream, with per-block state carried continuously across chunk boundaries.
24
+
25
+ 2. **Progressive growth** β€” instead of training a large model from scratch, Fractus grows palier by palier (width + depth + experts), inheriting previous weights via zero-padding. Each palier starts warm and converges faster.
26
+
27
+ 3. **Runtime self-modification** β€” the model detects routing imbalance and grows new experts while it runs, with zero-init stability validated.
28
+
29
+ I also report **negative results** honestly: Expert Decoupled Training (EDT) and the Forward-Forward algorithm were both tested and refuted β€” their objectives are misaligned with the final cross-entropy loss. The only training method that works is standard gradient descent, but the architectural optimizations I describe (sparse low-rank MoE, head-partial training, gradient accumulation, Kuramoto detachment) achieve **707 tokens/second on a consumer CPU** β€” a 177x improvement over the baseline.
30
+
31
+ ---
32
+
33
+ ## 1. Introduction
34
+
35
+ Contemporary large language models (GPT-4, Claude, Llama) share four fundamental limitations: they are static functions (one forward pass per output), stateless (no memory between conversations), generic (one monolithic network for all tasks), and centralized (training requires datacenter GPUs).
36
+
37
+ Fractus challenges each of these assumptions. The Continuous Thought Engine replaces the static function with a dynamical system. Persistent Memory gives the engine a cross-session memory bank. Expert Specialization forces each MoE expert to own a distinct skill domain. And the LazyStructuredSiren compression combined with progressive growth enables training on consumer hardware.
38
+
39
+ The question is not whether Fractus matches GPT-4 on benchmarks. It does not. The question is whether the paradigm of continuous, personal, decentralized AI is viable. This work demonstrates that it is β€” with measured, reproducible results.
40
+
41
+ ---
42
+
43
+ ## 2. Architecture
44
+
45
+ ### 2.1 The Continuous Thought Engine (CTE)
46
+
47
+ The CTE is a dynamical system that maintains a persistent thought state `h ∈ R^{d_model}` and advances it tick by tick. The engine stacks `n_layers` blocks, each refining the thought:
48
+
49
+ ```
50
+ h β†’ [Block 0: norm β†’ attn β†’ +residual β†’ norm β†’ kuramoto β†’ phases β†’ norm β†’ moe β†’ +residual]
51
+ β†’ [Block 1: ... ]
52
+ β†’ ...
53
+ β†’ [Block N: ... ] β†’ thought_state = h_final
54
+ ```
55
+
56
+ Each `CTEBlock` owns:
57
+ - **FractalLinearAttention** (Katharopoulos 2020) β€” multi-level causal linear attention with a persistent state (S, z) that accumulates across ticks and chunk boundaries.
58
+ - **Kuramoto oscillators** β€” a coupled dynamical system (low-rank RK4) that acts as a "consciousness clock," producing phase vectors that route MoE experts.
59
+ - **PhaseRoutedMoE** β€” a sparse mixture-of-experts with von Mises gate on Farey-distributed expert phases.
60
+
61
+ The thought state `h` is a **residual stream** β€” each block adds its transformation. The attention state (S, z) is **per-block** and **continuous across chunk boundaries** (verified: S grows monotonically across chunks, never reset).
62
+
63
+ ### 2.2 PhaseRoutedMoE (Sparse, Low-Rank, Differentiable)
64
+
65
+ The MoE routes tokens via Kuramoto oscillator phases through a von Mises gate:
66
+
67
+ ```
68
+ g_e = exp(ΞΊ Β· cos(ΞΈ_token βˆ’ ΞΈ_expert)) / Ξ£_e' g_e'
69
+ ```
70
+
71
+ Expert phases are drawn from the Farey sequence F_{2E}, providing E angles in [0, 2Ο€) that are dense, non-collapsing, and deterministic. Only `top_k=2` experts are computed per token (gather-first sparse dispatch).
72
+
73
+ **Low-rank experts**: each expert weight matrix W is decomposed as `W = scale Β· U@Vα΅€` (rank r=64). The forward pass is two cheap matmuls that never materialize the full matrix. The sparse path gathers only the top-k experts' U/V factors, computing K experts instead of E. At 128 experts with top-k=2, this is 64x less compute.
74
+
75
+ **Self-modification**: `add_expert()` grows a new expert at runtime, placed near the dominant expert's phase (to capture overflow traffic), with zero-init (no forward perturbation). `maybe_grow()` triggers automatically when routing imbalance exceeds a threshold.
76
+
77
+ ### 2.3 Multi-Block Depth
78
+
79
+ The original CTE (v1.0) had a single attention + Kuramoto + MoE block. This was the deepest limitation: depth 1 means the thought traverses one refinement then exits.
80
+
81
+ v2.0 introduces the `CTEBlock` abstraction. The engine stacks N blocks, each with independent state. The thought flows through all of them as a residual stream. This is the path to 1B+ parameters:
82
+
83
+ | Config | d_model | blocks | experts/block | params |
84
+ |---|---|---|---|---|
85
+ | Palier 0 | 128 | 1 | 4 | 6.6M |
86
+ | Palier 1 | 256 | 2 | 8 | ~25M |
87
+ | Palier 2 | 512 | 4 | 16 | ~120M |
88
+ | Palier 3 | 768 | 8 | 32 | ~350M |
89
+ | Palier 4 (1B) | 1280 | 16 | 128 | ~1B |
90
+
91
+ ### 2.4 Persistent Memory
92
+
93
+ The engine maintains a bank of memory vectors (d_model-dimensional, with context labels and importance scores) that survives across sessions. Memories are recalled via cosine similarity and injected into the thought state at 5% blend (continuous injection).
94
+
95
+ A **salience head** (Linear(d_model β†’ 1)) learns to predict how much a memory injection will perturb the thought state β€” an intrinsic signal, not an external label. The system discovers its own sensitivity to memories.
96
+
97
+ ### 2.5 Cognitive Modes
98
+
99
+ Cognitive modes emerge from the Kuramoto phase dynamics via unsupervised k-means clustering on phase features (synchronization degree r, mean phase, variance, per-oscillator sin/cos). No external labels β€” the modes are discovered from the structure of the phase space.
100
+
101
+ ### 2.6 Self-Modification
102
+
103
+ Fractus is the only model that grows new capacity while it runs:
104
+
105
+ ```python
106
+ engine.maybe_grow()
107
+ # β†’ "[Fractus] Self-modified: grew expert in all 16 blocks (now 129 experts)"
108
+ ```
109
+
110
+ The new expert is zero-initialized (scale=0 β†’ output=0 β†’ no gradient spike), placed near the dominant expert (captures overflow traffic), and warms up gradually via backprop. Validated: new expert receives 50% of traffic, loss stable post-grow.
111
+
112
+ ### 2.7 Progressive Growth
113
+
114
+ Instead of training 1B from scratch (months on GPU), Fractus grows palier by palier. Each palier:
115
+ 1. Inherits the previous model's weights via zero-padding (`fractus/grow.py`)
116
+ 2. Old knowledge preserved (top-left block of every matrix)
117
+ 3. New capacity starts neutral (zeros for weights, ones for LayerNorm gamma)
118
+ 4. Trains briefly to adapt the new dimensions
119
+
120
+ This is how a brain develops: small at first, growing new capacity on top of existing knowledge.
121
+
122
+ ---
123
+
124
+ ## 3. Training
125
+
126
+ ### 3.1 Online Training
127
+
128
+ The CTE trains online: one chunk (32 tokens) per forward, one backward per chunk. The thought state carries forward (detached β€” no BPTT). Each chunk's attention state (S,z) starts from the previous chunk's accumulated state β€” continuous thought.
129
+
130
+ ### 3.2 Training Optimizations (Measured)
131
+
132
+ | Optimization | What it does | Impact |
133
+ |---|---|---|
134
+ | Tied head | `output_head.weight = observe.weight` | Halves vocab params |
135
+ | Head-partial (`tick_chunk_train`) | Head on 1 position instead of C | 32x less head FLOPs |
136
+ | Sparse MoE low-rank | Only top-k experts computed (gather-first) | 64x at 128 experts |
137
+ | Gradient accumulation (accum=8) | 8x fewer optimizer steps | 16x fewer AdamW calls |
138
+ | Chunk_len=32 | Better Python amortization | ~1.1x |
139
+ | Detach Kuramoto | Phase computation in no_grad | Removes backward through clock |
140
+ | bf16 AMP (GPU) | 2x on all matmuls | GPU only |
141
+
142
+ **Combined measured result**: 4 tok/s β†’ **707 tok/s** on CPU (palier 0, d=128).
143
+
144
+ ### 3.3 Profile Breakdown
145
+
146
+ Per-iteration cost at d=128, chunk_len=32:
147
+
148
+ | Component | Time (ms) | % of forward |
149
+ |---|---|---|
150
+ | Attention (QKV + causal vectorized) | 21.8 | 32% |
151
+ | Kuramoto (detached) | 0.0 | 0% |
152
+ | MoE (4 experts, dense) | 16.6 | 25% |
153
+ | Output head (1 position, tied) | 12.6 | 19% |
154
+ | Embedding + norms | 16.3 | 24% |
155
+
156
+ At 128 experts, the sparse MoE path reduces MoE cost by 64x, making attention the dominant cost β€” as it should be for a reasoning architecture.
157
+
158
+ ---
159
+
160
+ ## 4. Negative Results (Honest)
161
+
162
+ ### 4.1 EDT (Expert Decoupled Training) β€” Refuted
163
+
164
+ EDT claimed 189x training speedup by pre-training experts independently. I tested 5 variants (vanilla, denoise, identity, residual objectives + routing filter) on the 13M CTE:
165
+
166
+ | Variant | Hold-out PPL | vs From-scratch |
167
+ |---|---|---|
168
+ | From-scratch | 1309.7 | β€” |
169
+ | EDT next_hidden | 1562.1 | +19.3% worse |
170
+ | EDT denoise + routing filter | 1555.1 | +18.7% worse |
171
+ | EDT identity + routing filter | 1556.8 | +18.9% worse |
172
+ | EDT residual + routing filter | 1573.7 | +20.2% worse |
173
+
174
+ **Root cause**: two structural defects. (1) The Phase-1 MSE objective (predict next hidden state) is misaligned with the Phase-3 CE objective (next-token prediction) β€” Pearson correlation never positive across 12 configs. (2) The Kuramoto router concentrates traffic on 2/4 experts, so half of pre-trained experts are never routed.
175
+
176
+ ### 4.2 Forward-Forward (Hinton 2022) β€” Refuted for CTE
177
+
178
+ The Forward-Forward algorithm (local goodness signal, no global backprop) was adapted to the CTE. Result: NLL went UP (124 β†’ 221). The goodness signal (sum of squared activations) is not aligned with cross-entropy. Local learning objectives cannot replace global backprop for this architecture.
179
+
180
+ ### 4.3 What Works
181
+
182
+ Only standard gradient descent (CE + backprop) produces a model that learns. The optimizations described in Β§3.2 are the path to making this feasible.
183
+
184
+ ---
185
+
186
+ ## 5. Experimental Results
187
+
188
+ ### 5.1 Progressive Growth (CPU)
189
+
190
+ | Palier | d_model | blocks | params | tokens | loss | tok/s | time |
191
+ |---|---|---|---|---|---|---|---|
192
+ | 0 | 128 | 1 | 6.6M | 2M | 32.5 | 725 | 46 min |
193
+ | 1 | 256 | 2 | 25M | 1.5M | 27.1 | 395 | 63 min |
194
+ | 2 | 512 | 4 | 120M | 1M | 27.1 | 192 | 87 min |
195
+ | 3 | 768 | 8 | 350M | 500k | 23.0 | 122 | 68 min |
196
+
197
+ Each palier starts warm (inherited weights) and converges. The corpus includes Fractus's own source code (palimpseste principle: the model contains its own description).
198
+
199
+ ### 5.2 Self-Modification Stability
200
+
201
+ After runtime `add_expert()` at tick 500 (4β†’5 experts):
202
+ - New expert receives 50% of routing traffic (placed near dominant)
203
+ - Gradient norm stable (zero-init β†’ no spike)
204
+ - Loss post-grow: +24.6% (vs control no-grow: +67.6%) β€” growth helps stability
205
+
206
+ ### 5.3 Continuous Thought Verification
207
+
208
+ Attention state (S,z) verified to grow monotonically across chunk boundaries in all three paths (tick, tick_chunk, tick_chunk_train). The thought is truly continuous.
209
+
210
+ ### 5.4 Cross-Session Memory
211
+
212
+ Session 1: 4 memories captured and saved to disk.
213
+ Session 2 (fresh engine): 4 memories loaded, thought state displaced by 100.16 units on first tick.
214
+ Cross-session persistence verified.
215
+
216
+ ---
217
+
218
+ ## 6. Comparison with GPT and Claude
219
+
220
+ | Property | GPT-4 / Claude | Fractus v2.0 |
221
+ |---|---|---|
222
+ | Processing | Static (1 forward) | Continuous (ticks through N blocks) |
223
+ | Memory | Context window | Persistent bank + salience-gated injection |
224
+ | Skills | Generic monolith | Specialized MoE experts (128 per block) |
225
+ | Mental state | Stateless | Cognitive modes (unsupervised) |
226
+ | Generation | Token-by-token | Plan then fill + adaptive depth |
227
+ | Training | Datacenter GPUs | Consumer CPU (progressive growth) + GPU for 1B |
228
+ | Deployment | Cloud API | Local device |
229
+ | User data | Sent to server | Stays local |
230
+ | Self-modification | None | Runtime expert growth |
231
+ | Depth scaling | Retrain from scratch | Progressive growth (warm start) |
232
+
233
+ ---
234
+
235
+ ## 7. Limitations and Future Work
236
+
237
+ 1. **Model quality**: The trained model at palier 3 (350M, 500k tokens) produces repetitive text. More data and training are needed for coherent generation. Chinchilla-optimal (940M tokens) requires ~5 days on GPU.
238
+
239
+ 2. **Multi-block training**: The multi-block architecture is validated (gradient flows through all blocks, continuous thought works) but has not yet been trained at scale. Palier 4 (16 blocks, 128 experts, 1B params) awaits GPU compute.
240
+
241
+ 3. **EDT and Forward-Forward**: Both refuted. Alternative training acceleration methods must align their objective with the final CE loss.
242
+
243
+ 4. **Vocabulary**: The GPT-2 BPE vocab (50257) dominates parameters (81% at d=768). A reduced vocab (8k-16k) would cut head FLOPs by 3-6x.
244
+
245
+ 5. **Corpus**: The quality corpus (20.5M tokens) includes Fractus's own source code but is far below Chinchilla scale for the larger paliers.
246
+
247
+ ---
248
+
249
+ ## 8. Conclusion
250
+
251
+ Fractus v2.0 demonstrates that a continuous, multi-block, self-modifying cognitive agent can be constructed and progressively trained on consumer hardware. The architecture β€” CTEBlock stack with per-block continuous attention state, PhaseRoutedMoE with sparse low-rank experts, progressive growth via zero-padding, and runtime self-modification β€” is validated by 28 tests and measured benchmarks.
252
+
253
+ The negative results (EDT, Forward-Forward) are reported honestly. They do not weaken the architecture; they clarify what works (global backprop + architectural optimizations) and what does not (decoupled/local training).
254
+
255
+ The implications extend beyond performance metrics. If AI can be trained and deployed on any laptop, the centralization of intelligence by a handful of corporations is not inevitable. Fractus is a proof of concept for decentralized AI: intelligence that belongs to the user, runs on their hardware, remembers them, and grows.
256
+
257
+ This work is released as open source under the MIT license. All code, training scripts, datasets, and measured results are available at github.com/AFKmoney/fractus-test.
258
+
259
+ ---
260
+
261
+ ## References
262
+
263
+ [1] Katharopoulos et al. (2020). Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention. ICML.
264
+ [2] Sitzmann et al. (2020). Implicit Neural Representations with Periodic Activation Functions (SIREN). NeurIPS.
265
+ [3] Hinton, G. (2022). The Forward-Forward Algorithm: Some Preliminary Investigations.
266
+ [4] Kuramoto, Y. (1984). Chemical Oscillations, Waves, and Turbulence. Springer.
267
+ [5] Hinton, G. (2022). The Forward-Forward Algorithm.
268
+ [6] Rahimi & Recht (2007). Random Features for Large-Scale Kernel Machines. NeurIPS.
269
+ [7] Graves, A. (2016). Adaptive Computation Time for Recurrent Neural Networks. arXiv.
270
+ [8] Hu et al. (2021). LoRA: Low-Rank Adaptation of Large Language Models. arXiv.
271
+
272
+ ---
273
+
274
+ *Β© 2026 Philippe-Antoine Robert. MIT License. Contact: rpa.tu@proton.me*