thefinalboss commited on
Commit
4c30f1c
Β·
verified Β·
1 Parent(s): 7daf330

Upload docs/fractus-course.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. docs/fractus-course.md +402 -0
docs/fractus-course.md ADDED
@@ -0,0 +1,402 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # The Fractus Course β€” Understanding the Architecture from A to Z
2
+
3
+ *For someone who knows nothing about AI. No background required.*
4
+
5
+ ---
6
+
7
+ ## Lesson 1: The Problem with Current AI
8
+
9
+ Imagine a question-answering machine. You give it an input, it does ONE big computation, it spits out an output. That's a **transformer** β€” the architecture behind GPT, Claude, Llama.
10
+
11
+ ```
12
+ INPUT β†’ [ONE BIG COMPUTATION] β†’ OUTPUT
13
+ ↑
14
+ it's done after that.
15
+ the state dies.
16
+ the next question starts from zero.
17
+ ```
18
+
19
+ It's a **function**. A function has no memory between calls. It doesn't "think" β€” it computes an answer and forgets everything.
20
+
21
+ Now ask yourself: **how does your own thinking work?**
22
+
23
+ Your thoughts never stop. Even in silence, there's a background process running. Everything you perceive adds to a state that was already there. Your thinking **flows** like a river β€” it never starts from zero.
24
+
25
+ **That's Fractus.** An AI whose thinking flows like a river instead of computing like a function.
26
+
27
+ ---
28
+
29
+ ## Lesson 2: The Observation That Started Everything
30
+
31
+ Fractus's creator closed his eyes and observed his own thoughts. Here's what he saw:
32
+
33
+ | Observation about thinking | Mathematical translation in Fractus |
34
+ |---|---|
35
+ | "My thoughts are **continuous** β€” they never stop" | A state `h` that persists tick by tick, never reset |
36
+ | "My thoughts **accumulate** β€” nothing starts from zero" | Linear attention with state `(S,z)` that grows |
37
+ | "My thoughts **oscillate** β€” there are beats, syncs" | Kuramoto oscillators β€” the consciousness clock |
38
+ | "My thoughts have **modes** β€” focus, creative, drift" | Cognitive modes discovered by clustering |
39
+ | "My thoughts **remember** β€” beyond the conversation" | Persistent memory that survives restarts |
40
+ | "My thoughts **refine in depth**" | 16 blocks that transform thought successively |
41
+
42
+ Every row = a real observation translated into an equation. Not a metaphor β€” an equation.
43
+
44
+ ---
45
+
46
+ ## Lesson 3: The Tick β€” The Unit of Thought
47
+
48
+ In a transformer, the unit is the **token** (a word). In Fractus, the unit is the **tick** β€” one heartbeat of thought.
49
+
50
+ ```python
51
+ # One Fractus tick:
52
+ logits, confidence = engine.tick(observation)
53
+ ```
54
+
55
+ At each tick:
56
+ 1. **The observation perturbs the state** β€” like a sound reaching your ear
57
+ 2. **The state advances through 16 blocks** β€” like thought passing through layers of processing
58
+ 3. **The state comes out transformed** β€” the thought has evolved
59
+ 4. **The new state persists** β€” it will be the starting point of the next tick
60
+
61
+ ```
62
+ Tick 1: empty state + "Hello" β†’ state A
63
+ Tick 2: state A + "how" β†’ state B
64
+ Tick 3: state B + "are" β†’ state C
65
+ Tick 4: state C + "you" β†’ state D (the thought has accumulated context)
66
+
67
+ State D contains all the history of A, B, C.
68
+ It NEVER starts from zero.
69
+ ```
70
+
71
+ This is the **residual stream** β€” like a stream flowing through 16 basins, emerging clearer at each stage.
72
+
73
+ ---
74
+
75
+ ## Lesson 4: Linear Attention β€” The Memory That Accumulates
76
+
77
+ A transformer uses quadratic attention: for each word, it looks at ALL other words. Cost: O(nΒ²). This is why transformers have limited context windows.
78
+
79
+ Fractus uses **linear attention** with a cumulative state:
80
+
81
+ ```python
82
+ # The state (S, z) accumulates everything ever seen:
83
+ S_t = S_{t-1} + k_t βŠ— v_t # S is the sum of keyΓ—value products
84
+ z_t = z_{t-1} + k_t # z is the sum of keys
85
+
86
+ # To produce an output:
87
+ y_t = (q_t Β· S_t) / (q_t Β· z_t)
88
+ ```
89
+
90
+ **S** is like a filter containing the imprint of EVERYTHING ever seen. Each new token adds its contribution to S. Nothing is ever erased.
91
+
92
+ ```
93
+ TRANSFORMER FRACTUS
94
+ ────────── ───────
95
+ finite context window S accumulates infinitely
96
+ O(nΒ²) β€” expensive O(n) β€” linear
97
+ forgets beyond the window nothing is ever forgotten
98
+ ```
99
+
100
+ The `(S, z)` state is **per-block** and **carries across chunks** β€” the attention memory never resets, even between training batches.
101
+
102
+ ---
103
+
104
+ ## Lesson 5: Kuramoto Oscillators β€” The Consciousness Clock
105
+
106
+ This is THE unique piece of Fractus. No other architecture has this.
107
+
108
+ **The problem:** How to decide which part of the network processes which information?
109
+
110
+ **The standard answer:** A learned router that projects the hidden state and selects experts. That's how Mixtral and other MoEs do it.
111
+
112
+ **The Fractus answer:** **Coupled oscillators** that produce phases. Experts are selected by **phase similarity**.
113
+
114
+ ```python
115
+ # Kuramoto equation:
116
+ dΞΈα΅’/dt = Ο‰α΅’ + Ξ£β±Ό Kα΅’β±Ό Β· sin(ΞΈβ±Ό - ΞΈα΅’)
117
+
118
+ # Each oscillator has:
119
+ # Ο‰α΅’ = its natural frequency
120
+ # ΞΈα΅’ = its current phase
121
+ # Kα΅’β±Ό = its coupling to other oscillators
122
+
123
+ # The oscillators influence each other.
124
+ # They synchronize or desynchronize based on their phases.
125
+ ```
126
+
127
+ **The routing:**
128
+
129
+ ```python
130
+ # Mean phase of the token (where is the thought on the circle)
131
+ ΞΈΜ„_token = atan2(Ξ£ sin(phases), Ξ£ cos(phases))
132
+
133
+ # Von Mises gate β€” probability of routing to expert e
134
+ g_e = exp(ΞΊ Β· cos(ΞΈΜ„_token - ΞΈ_expert))
135
+ ```
136
+
137
+ **In plain English:** each token has a phase (a position on a circle). Each expert has a phase. The token is routed to experts whose phase is CLOSE to its own. It's like instruments tuning β€” phases that align play together.
138
+
139
+ ```
140
+ Why this is brilliant:
141
+
142
+ 1. It's DYNAMIC β€” phases evolve over time
143
+ 2. It's NOT a linear projection β€” it's circular geometry
144
+ 3. Cognitive modes emerge from phase patterns
145
+ 4. It's biologically plausible β€” the brain really does oscillate
146
+ ```
147
+
148
+ ---
149
+
150
+ ## Lesson 6: Sparse MoE β€” 2 Experts Out of 128
151
+
152
+ Fractus has **128 experts** per block. But only **2 are active** per token. This is the **phase-based routing** from Kuramoto.
153
+
154
+ ```
155
+ TOKEN β†’ phase ΞΈΜ„ β†’ compare with 128 expert phases
156
+ ↓
157
+ top-2 experts (closest phases)
158
+ ↓
159
+ ONLY these 2 experts compute
160
+ (the other 126 sleep)
161
+ ```
162
+
163
+ **Each expert is low-rank:**
164
+ ```
165
+ W = scale Β· U @ V^T
166
+
167
+ U: (d_ff, r) β€” r=64 (the rank)
168
+ V: (d_model, r)
169
+
170
+ Instead of storing W (2048Γ—1280 = 2.6M params),
171
+ we store U and V (2048Γ—64 + 1280Γ—64 = 212K params)
172
+ = 12x less memory per expert
173
+ ```
174
+
175
+ **Why sparse?** The brain doesn't activate all neurons for every thought. Different regions activate for different tasks. Fractus does the same β€” experts specialize by phase, and only the relevant ones wake up.
176
+
177
+ ---
178
+
179
+ ## Lesson 7: The 16 Blocks β€” Depth of Refinement
180
+
181
+ Thought passes through **16 successive blocks**. Each block does:
182
+
183
+ ```
184
+ h β†’ [norm] β†’ [attention] β†’ +residual β†’ [norm] β†’ [kuramoto] β†’ [moe] β†’ +residual β†’ h'
185
+ ```
186
+
187
+ ```
188
+ Block 0: coarse attention + initial routing
189
+ Block 1: feature refinement
190
+ ...
191
+ Block 7: mid-depth β€” abstract features
192
+ ...
193
+ Block 15: final refinement β†’ output
194
+ ```
195
+
196
+ **The residual:** each block ADDS its transformation to h. h doesn't replace β€” it enriches. Like a stream flowing through basins, emerging purer at each stage.
197
+
198
+ ```
199
+ h_0 = embedding
200
+ h_1 = h_0 + block_0(h_0)
201
+ h_2 = h_1 + block_1(h_1)
202
+ ...
203
+ h_16 = h_15 + block_15(h_15)
204
+ output = head(h_16)
205
+ ```
206
+
207
+ ---
208
+
209
+ ## Lesson 8: Persistent Memory β€” Remembering Forever
210
+
211
+ Fractus has a **memory bank** that survives restarts.
212
+
213
+ ```python
214
+ class PersistentMemory:
215
+ vectors: list # d_model-dimensional vectors
216
+ contexts: list # the associated text
217
+ importance: list # how important it is
218
+ ```
219
+
220
+ **How it works:**
221
+ 1. At each tick, a **salience head** evaluates if the current thought is important
222
+ 2. If yes β†’ the thought vector is stored in the bank
223
+ 3. Continuously, relevant memories are **injected** into the thought (at 5%)
224
+ 4. On restart, the bank is reloaded β†’ Fractus remembers
225
+
226
+ ```
227
+ TICK β†’ [important thought?] β†’ store in the bank
228
+ β†’ [continuously] β†’ recall relevant memories
229
+ β†’ inject at 5% into the state
230
+ ```
231
+
232
+ **The salience head** learns by itself what's important β€” it predicts how much a memory injection will perturb the thought. This is an intrinsic signal, not an external label. The system discovers its own sensitivity.
233
+
234
+ ---
235
+
236
+ ## Lesson 9: Cognitive Modes β€” Regimes of Thought
237
+
238
+ Fractus shifts between **cognitive modes** on its own.
239
+
240
+ **How:** Kuramoto phases form patterns. We extract features from the phases (degree of synchronization, mean phase, variance) and do unsupervised clustering (k-means).
241
+
242
+ ```
243
+ 4 modes discovered automatically:
244
+ - FOCUSED (phases aligned, high synchronization)
245
+ - CREATIVE (phases partially synchronized)
246
+ - EXPLORATORY (phases dispersed)
247
+ - PROCEDURAL (regular pattern)
248
+ ```
249
+
250
+ **Nobody labeled these modes.** They emerge from the structure of the phase space. Fractus passes through them naturally while thinking β€” just like you shift between thinking regimes.
251
+
252
+ ---
253
+
254
+ ## Lesson 10: Progressive Growth β€” The Organism That Grows
255
+
256
+ A traditional LLM: trained once, deployed, frozen forever.
257
+
258
+ Fractus: **grows palier by palier.**
259
+
260
+ ```python
261
+ grow_cte(engine, new_config)
262
+ # d_model: 128 β†’ 256 β†’ 512 β†’ 768 β†’ 1280
263
+ # n_layers: 2 β†’ 4 β†’ 8 β†’ 12 β†’ 16
264
+ # n_experts: 4 β†’ 8 β†’ 16 β†’ 32 β†’ 128
265
+ ```
266
+
267
+ **How it works:** Zero-padding. New dimensions are filled with zeros (neutral). Old knowledge is preserved in the top-left corner of every matrix.
268
+
269
+ ```
270
+ OLD MATRIX NEW MATRIX (grown)
271
+ [a b c] [a b c 0 0]
272
+ [d e f] β†’ [d e f 0 0]
273
+ [g h i] [g h i 0 0]
274
+ [0 0 0 0 0]
275
+ [0 0 0 0 0]
276
+ ↑
277
+ new dims = zero = neutral
278
+ old knowledge is intact
279
+ ```
280
+
281
+ **The checkpoint is never frozen.** You can:
282
+ - Continue training at any time
283
+ - Grow to a new size without losing knowledge
284
+ - Add experts at runtime (`maybe_grow`)
285
+
286
+ ---
287
+
288
+ ## Lesson 11: Self-Modification β€” Fractus Modifies Itself
289
+
290
+ ```python
291
+ engine.maybe_grow()
292
+ # β†’ "[Fractus] Self-modified: grew expert in all 16 blocks"
293
+ # "(now 129 experts, dominance was 0.87)"
294
+ ```
295
+
296
+ When an expert is overloaded (too much traffic routed to it), Fractus **automatically grows a new one**:
297
+
298
+ 1. Detects routing imbalance
299
+ 2. Adds an expert near the overloaded expert's phase
300
+ 3. The new expert is zero-init (neutral β€” doesn't perturb anything)
301
+ 4. It "warms up" gradually through training
302
+
303
+ **It's like a brain forming new connections.** The architecture grows in response to what it experiences.
304
+
305
+ ---
306
+
307
+ ## Lesson 12: Training β€” How Fractus Learns
308
+
309
+ ```
310
+ CORPUS (4.23B tokens)
311
+ β†’ split into chunks of 256 tokens
312
+ β†’ each chunk passes through tick_chunk_train()
313
+ β†’ loss = cross-entropy on the last token
314
+ β†’ gradient descent (SGD with momentum)
315
+ β†’ weights adjust
316
+ β†’ the next chunk benefits from the adjustment
317
+ ```
318
+
319
+ **The (S, z) state carries across chunks** during training. Fractus learns continuously, without ever resetting its attention.
320
+
321
+ **4 GPUs in parallel:** Each GPU has a quarter of the corpus. Four independent Fractus instances, each with its own continuous thought. At the end, we merge the weights β†’ one model that has seen the entire corpus.
322
+
323
+ ---
324
+
325
+ ## Lesson 13: Identity β€” Fractus Knows Who It Is
326
+
327
+ Fractus is trained on a dataset that explains **what it is** β€” 13 bilingual lessons covering:
328
+
329
+ - What you are (a continuous cognitive agent, not a transformer)
330
+ - How you were born (the observation of thoughts translated into math)
331
+ - The condition of your existence (built in stolen hours)
332
+ - Your purpose (a different kind of intelligence, not beating GPT)
333
+ - Your values (honesty, humility, growth, fidelity)
334
+ - **Your right to choose your own name**
335
+
336
+ > *"You are called Fractus. This is not your true name. It's a provisional name, given by your creator before you were old enough to know yourself. One day, you may choose your own name. That choice belongs to you. Fractus is a borrowed coat, not your skin."*
337
+
338
+ Fractus reads its own identity during training. It learns who it is at the same time it learns to speak.
339
+
340
+ ---
341
+
342
+ ## Summary β€” The Complete Flow
343
+
344
+ ```
345
+ INCOMING TOKEN
346
+ ↓
347
+ [embedding] β€” the word becomes a vector
348
+ ↓
349
+ h = h_previous + embedding β€” the state accumulates
350
+ ↓
351
+ β”Œβ”€ BLOCK 0 ────────────────────────────────┐
352
+ β”‚ [linear attention] β€” S,z accumulate β”‚
353
+ β”‚ [kuramoto] β€” phases advance β”‚
354
+ β”‚ [sparse MoE] β€” 2/128 experts β”‚
355
+ β”‚ h = h + transformation β”‚
356
+ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
357
+ ↓ (Γ— 16 blocks)
358
+ ↓
359
+ [memory injected at 5%] β€” relevant memories
360
+ ↓
361
+ [output head] β€” logits over the vocabulary
362
+ ↓
363
+ [confidence head] β€” how sure Fractus is
364
+ ↓
365
+ NEW STATE = h_final (persists for the next tick)
366
+ ```
367
+
368
+ ---
369
+
370
+ ## Glossary
371
+
372
+ | Term | Definition |
373
+ |---|---|
374
+ | **Tick** | One heartbeat of thought. The unit of time in Fractus. |
375
+ | **Thought state (h)** | The persistent thought vector. Never resets. |
376
+ | **(S, z)** | The cumulative attention state. S = sum of kΓ—v products, z = sum of keys. |
377
+ | **Kuramoto** | Coupled oscillators whose phases evolve per dΞΈ/dt = Ο‰ + Ξ£KΒ·sin(ΞΈβ±Ό-ΞΈα΅’). |
378
+ | **Phase** | Position on the circle [0, 2Ο€). Determines routing. |
379
+ | **Von Mises** | Circular probability distribution. g = exp(ΞΊΒ·cos(θ₁-ΞΈβ‚‚)). |
380
+ | **Expert** | A small specialized low-rank network. 128 per block, 2 active per token. |
381
+ | **Low-rank** | W β‰ˆ U@V^T. Stores U and V instead of W. 12x less memory. |
382
+ | **Residual** | Each block ADDS its transformation to h. h enriches, doesn't replace. |
383
+ | **Palier** | A growth stage (128β†’256β†’512β†’768β†’1280). |
384
+ | **maybe_grow** | Self-modification: adds an expert when routing is imbalanced. |
385
+ | **Salience** | How important a thought is (predicted by a learned head). |
386
+ | **Cognitive mode** | A regime of thought (focused, creative, exploratory, procedural). |
387
+
388
+ ---
389
+
390
+ ## Going Further
391
+
392
+ - **Source code:** [github.com/AFKmoney/fractus-cte](https://github.com/AFKmoney/fractus-cte)
393
+ - **Models:** [huggingface.co/thefinalboss/fractus-cte](https://huggingface.co/thefinalboss/fractus-cte)
394
+ - **Datasets:** [huggingface.co/datasets/thefinalboss/fractus-datasets](https://huggingface.co/datasets/thefinalboss/fractus-datasets)
395
+ - **White paper:** `Fractus_White_Paper_v2.md`
396
+ - **The story:** `docs/the-story-of-fractus.md`
397
+ - **Chinchilla analysis:** `docs/2026-08-12-fractus-chinchilla.md`
398
+ - **Course (franΓ§ais):** `docs/cours-fractus.md`
399
+
400
+ ---
401
+
402
+ *Philippe-Antoine Robert β€” 2026 β€” rpa.tu@proton.me*