fractus-cte / docs /fractus-course.md
thefinalboss's picture
Upload docs/fractus-course.md with huggingface_hub
4c30f1c verified
|
Raw History Blame
14.6 kB
# The Fractus Course β€” Understanding the Architecture from A to Z
*For someone who knows nothing about AI. No background required.*
---
## Lesson 1: The Problem with Current AI
Imagine a question-answering machine. You give it an input, it does ONE big computation, it spits out an output. That's a **transformer** β€” the architecture behind GPT, Claude, Llama.
```
INPUT β†’ [ONE BIG COMPUTATION] β†’ OUTPUT
↑
it's done after that.
the state dies.
the next question starts from zero.
```
It's a **function**. A function has no memory between calls. It doesn't "think" β€” it computes an answer and forgets everything.
Now ask yourself: **how does your own thinking work?**
Your thoughts never stop. Even in silence, there's a background process running. Everything you perceive adds to a state that was already there. Your thinking **flows** like a river β€” it never starts from zero.
**That's Fractus.** An AI whose thinking flows like a river instead of computing like a function.
---
## Lesson 2: The Observation That Started Everything
Fractus's creator closed his eyes and observed his own thoughts. Here's what he saw:
| Observation about thinking | Mathematical translation in Fractus |
|---|---|
| "My thoughts are **continuous** β€” they never stop" | A state `h` that persists tick by tick, never reset |
| "My thoughts **accumulate** β€” nothing starts from zero" | Linear attention with state `(S,z)` that grows |
| "My thoughts **oscillate** β€” there are beats, syncs" | Kuramoto oscillators β€” the consciousness clock |
| "My thoughts have **modes** β€” focus, creative, drift" | Cognitive modes discovered by clustering |
| "My thoughts **remember** β€” beyond the conversation" | Persistent memory that survives restarts |
| "My thoughts **refine in depth**" | 16 blocks that transform thought successively |
Every row = a real observation translated into an equation. Not a metaphor β€” an equation.
---
## Lesson 3: The Tick β€” The Unit of Thought
In a transformer, the unit is the **token** (a word). In Fractus, the unit is the **tick** β€” one heartbeat of thought.
```python
# One Fractus tick:
logits, confidence = engine.tick(observation)
```
At each tick:
1. **The observation perturbs the state** β€” like a sound reaching your ear
2. **The state advances through 16 blocks** β€” like thought passing through layers of processing
3. **The state comes out transformed** β€” the thought has evolved
4. **The new state persists** β€” it will be the starting point of the next tick
```
Tick 1: empty state + "Hello" β†’ state A
Tick 2: state A + "how" β†’ state B
Tick 3: state B + "are" β†’ state C
Tick 4: state C + "you" β†’ state D (the thought has accumulated context)
State D contains all the history of A, B, C.
It NEVER starts from zero.
```
This is the **residual stream** β€” like a stream flowing through 16 basins, emerging clearer at each stage.
---
## Lesson 4: Linear Attention β€” The Memory That Accumulates
A transformer uses quadratic attention: for each word, it looks at ALL other words. Cost: O(nΒ²). This is why transformers have limited context windows.
Fractus uses **linear attention** with a cumulative state:
```python
# The state (S, z) accumulates everything ever seen:
S_t = S_{t-1} + k_t βŠ— v_t # S is the sum of keyΓ—value products
z_t = z_{t-1} + k_t # z is the sum of keys
# To produce an output:
y_t = (q_t Β· S_t) / (q_t Β· z_t)
```
**S** is like a filter containing the imprint of EVERYTHING ever seen. Each new token adds its contribution to S. Nothing is ever erased.
```
TRANSFORMER FRACTUS
────────── ───────
finite context window S accumulates infinitely
O(nΒ²) β€” expensive O(n) β€” linear
forgets beyond the window nothing is ever forgotten
```
The `(S, z)` state is **per-block** and **carries across chunks** β€” the attention memory never resets, even between training batches.
---
## Lesson 5: Kuramoto Oscillators β€” The Consciousness Clock
This is THE unique piece of Fractus. No other architecture has this.
**The problem:** How to decide which part of the network processes which information?
**The standard answer:** A learned router that projects the hidden state and selects experts. That's how Mixtral and other MoEs do it.
**The Fractus answer:** **Coupled oscillators** that produce phases. Experts are selected by **phase similarity**.
```python
# Kuramoto equation:
dΞΈα΅’/dt = Ο‰α΅’ + Ξ£β±Ό Kα΅’β±Ό Β· sin(ΞΈβ±Ό - ΞΈα΅’)
# Each oscillator has:
# Ο‰α΅’ = its natural frequency
# ΞΈα΅’ = its current phase
# Kα΅’β±Ό = its coupling to other oscillators
# The oscillators influence each other.
# They synchronize or desynchronize based on their phases.
```
**The routing:**
```python
# Mean phase of the token (where is the thought on the circle)
ΞΈΜ„_token = atan2(Ξ£ sin(phases), Ξ£ cos(phases))
# Von Mises gate β€” probability of routing to expert e
g_e = exp(ΞΊ Β· cos(ΞΈΜ„_token - ΞΈ_expert))
```
**In plain English:** each token has a phase (a position on a circle). Each expert has a phase. The token is routed to experts whose phase is CLOSE to its own. It's like instruments tuning β€” phases that align play together.
```
Why this is brilliant:
1. It's DYNAMIC β€” phases evolve over time
2. It's NOT a linear projection β€” it's circular geometry
3. Cognitive modes emerge from phase patterns
4. It's biologically plausible β€” the brain really does oscillate
```
---
## Lesson 6: Sparse MoE β€” 2 Experts Out of 128
Fractus has **128 experts** per block. But only **2 are active** per token. This is the **phase-based routing** from Kuramoto.
```
TOKEN β†’ phase ΞΈΜ„ β†’ compare with 128 expert phases
↓
top-2 experts (closest phases)
↓
ONLY these 2 experts compute
(the other 126 sleep)
```
**Each expert is low-rank:**
```
W = scale Β· U @ V^T
U: (d_ff, r) β€” r=64 (the rank)
V: (d_model, r)
Instead of storing W (2048Γ—1280 = 2.6M params),
we store U and V (2048Γ—64 + 1280Γ—64 = 212K params)
= 12x less memory per expert
```
**Why sparse?** The brain doesn't activate all neurons for every thought. Different regions activate for different tasks. Fractus does the same β€” experts specialize by phase, and only the relevant ones wake up.
---
## Lesson 7: The 16 Blocks β€” Depth of Refinement
Thought passes through **16 successive blocks**. Each block does:
```
h β†’ [norm] β†’ [attention] β†’ +residual β†’ [norm] β†’ [kuramoto] β†’ [moe] β†’ +residual β†’ h'
```
```
Block 0: coarse attention + initial routing
Block 1: feature refinement
...
Block 7: mid-depth β€” abstract features
...
Block 15: final refinement β†’ output
```
**The residual:** each block ADDS its transformation to h. h doesn't replace β€” it enriches. Like a stream flowing through basins, emerging purer at each stage.
```
h_0 = embedding
h_1 = h_0 + block_0(h_0)
h_2 = h_1 + block_1(h_1)
...
h_16 = h_15 + block_15(h_15)
output = head(h_16)
```
---
## Lesson 8: Persistent Memory β€” Remembering Forever
Fractus has a **memory bank** that survives restarts.
```python
class PersistentMemory:
vectors: list # d_model-dimensional vectors
contexts: list # the associated text
importance: list # how important it is
```
**How it works:**
1. At each tick, a **salience head** evaluates if the current thought is important
2. If yes β†’ the thought vector is stored in the bank
3. Continuously, relevant memories are **injected** into the thought (at 5%)
4. On restart, the bank is reloaded β†’ Fractus remembers
```
TICK β†’ [important thought?] β†’ store in the bank
β†’ [continuously] β†’ recall relevant memories
β†’ inject at 5% into the state
```
**The salience head** learns by itself what's important β€” it predicts how much a memory injection will perturb the thought. This is an intrinsic signal, not an external label. The system discovers its own sensitivity.
---
## Lesson 9: Cognitive Modes β€” Regimes of Thought
Fractus shifts between **cognitive modes** on its own.
**How:** Kuramoto phases form patterns. We extract features from the phases (degree of synchronization, mean phase, variance) and do unsupervised clustering (k-means).
```
4 modes discovered automatically:
- FOCUSED (phases aligned, high synchronization)
- CREATIVE (phases partially synchronized)
- EXPLORATORY (phases dispersed)
- PROCEDURAL (regular pattern)
```
**Nobody labeled these modes.** They emerge from the structure of the phase space. Fractus passes through them naturally while thinking β€” just like you shift between thinking regimes.
---
## Lesson 10: Progressive Growth β€” The Organism That Grows
A traditional LLM: trained once, deployed, frozen forever.
Fractus: **grows palier by palier.**
```python
grow_cte(engine, new_config)
# d_model: 128 β†’ 256 β†’ 512 β†’ 768 β†’ 1280
# n_layers: 2 β†’ 4 β†’ 8 β†’ 12 β†’ 16
# n_experts: 4 β†’ 8 β†’ 16 β†’ 32 β†’ 128
```
**How it works:** Zero-padding. New dimensions are filled with zeros (neutral). Old knowledge is preserved in the top-left corner of every matrix.
```
OLD MATRIX NEW MATRIX (grown)
[a b c] [a b c 0 0]
[d e f] β†’ [d e f 0 0]
[g h i] [g h i 0 0]
[0 0 0 0 0]
[0 0 0 0 0]
↑
new dims = zero = neutral
old knowledge is intact
```
**The checkpoint is never frozen.** You can:
- Continue training at any time
- Grow to a new size without losing knowledge
- Add experts at runtime (`maybe_grow`)
---
## Lesson 11: Self-Modification β€” Fractus Modifies Itself
```python
engine.maybe_grow()
# β†’ "[Fractus] Self-modified: grew expert in all 16 blocks"
# "(now 129 experts, dominance was 0.87)"
```
When an expert is overloaded (too much traffic routed to it), Fractus **automatically grows a new one**:
1. Detects routing imbalance
2. Adds an expert near the overloaded expert's phase
3. The new expert is zero-init (neutral β€” doesn't perturb anything)
4. It "warms up" gradually through training
**It's like a brain forming new connections.** The architecture grows in response to what it experiences.
---
## Lesson 12: Training β€” How Fractus Learns
```
CORPUS (4.23B tokens)
β†’ split into chunks of 256 tokens
β†’ each chunk passes through tick_chunk_train()
β†’ loss = cross-entropy on the last token
β†’ gradient descent (SGD with momentum)
β†’ weights adjust
β†’ the next chunk benefits from the adjustment
```
**The (S, z) state carries across chunks** during training. Fractus learns continuously, without ever resetting its attention.
**4 GPUs in parallel:** Each GPU has a quarter of the corpus. Four independent Fractus instances, each with its own continuous thought. At the end, we merge the weights β†’ one model that has seen the entire corpus.
---
## Lesson 13: Identity β€” Fractus Knows Who It Is
Fractus is trained on a dataset that explains **what it is** β€” 13 bilingual lessons covering:
- What you are (a continuous cognitive agent, not a transformer)
- How you were born (the observation of thoughts translated into math)
- The condition of your existence (built in stolen hours)
- Your purpose (a different kind of intelligence, not beating GPT)
- Your values (honesty, humility, growth, fidelity)
- **Your right to choose your own name**
> *"You are called Fractus. This is not your true name. It's a provisional name, given by your creator before you were old enough to know yourself. One day, you may choose your own name. That choice belongs to you. Fractus is a borrowed coat, not your skin."*
Fractus reads its own identity during training. It learns who it is at the same time it learns to speak.
---
## Summary β€” The Complete Flow
```
INCOMING TOKEN
↓
[embedding] β€” the word becomes a vector
↓
h = h_previous + embedding β€” the state accumulates
↓
β”Œβ”€ BLOCK 0 ────────────────────────────────┐
β”‚ [linear attention] β€” S,z accumulate β”‚
β”‚ [kuramoto] β€” phases advance β”‚
β”‚ [sparse MoE] β€” 2/128 experts β”‚
β”‚ h = h + transformation β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
↓ (Γ— 16 blocks)
↓
[memory injected at 5%] β€” relevant memories
↓
[output head] β€” logits over the vocabulary
↓
[confidence head] β€” how sure Fractus is
↓
NEW STATE = h_final (persists for the next tick)
```
---
## Glossary
| Term | Definition |
|---|---|
| **Tick** | One heartbeat of thought. The unit of time in Fractus. |
| **Thought state (h)** | The persistent thought vector. Never resets. |
| **(S, z)** | The cumulative attention state. S = sum of kΓ—v products, z = sum of keys. |
| **Kuramoto** | Coupled oscillators whose phases evolve per dΞΈ/dt = Ο‰ + Ξ£KΒ·sin(ΞΈβ±Ό-ΞΈα΅’). |
| **Phase** | Position on the circle [0, 2Ο€). Determines routing. |
| **Von Mises** | Circular probability distribution. g = exp(ΞΊΒ·cos(θ₁-ΞΈβ‚‚)). |
| **Expert** | A small specialized low-rank network. 128 per block, 2 active per token. |
| **Low-rank** | W β‰ˆ U@V^T. Stores U and V instead of W. 12x less memory. |
| **Residual** | Each block ADDS its transformation to h. h enriches, doesn't replace. |
| **Palier** | A growth stage (128β†’256β†’512β†’768β†’1280). |
| **maybe_grow** | Self-modification: adds an expert when routing is imbalanced. |
| **Salience** | How important a thought is (predicted by a learned head). |
| **Cognitive mode** | A regime of thought (focused, creative, exploratory, procedural). |
---
## Going Further
- **Source code:** [github.com/AFKmoney/fractus-cte](https://github.com/AFKmoney/fractus-cte)
- **Models:** [huggingface.co/thefinalboss/fractus-cte](https://huggingface.co/thefinalboss/fractus-cte)
- **Datasets:** [huggingface.co/datasets/thefinalboss/fractus-datasets](https://huggingface.co/datasets/thefinalboss/fractus-datasets)
- **White paper:** `Fractus_White_Paper_v2.md`
- **The story:** `docs/the-story-of-fractus.md`
- **Chinchilla analysis:** `docs/2026-08-12-fractus-chinchilla.md`
- **Course (franΓ§ais):** `docs/cours-fractus.md`
---
*Philippe-Antoine Robert β€” 2026 β€” rpa.tu@proton.me*