Download docs/fractus-course.md from thefinalboss/fractus-cte: direct link, hf CLI and curl.
- Browser
- Download file 14.6 kB
-
https://huggingface.co/thefinalboss/fractus-cte/resolve/main/docs/fractus-course.md
- Command line
-
hf download hf://thefinalboss/fractus-cte/docs/fractus-course.md
-
curl -L -o fractus-course.md https://huggingface.co/thefinalboss/fractus-cte/resolve/main/docs/fractus-course.md
The Fractus Course β Understanding the Architecture from A to Z
For someone who knows nothing about AI. No background required.
Lesson 1: The Problem with Current AI
Imagine a question-answering machine. You give it an input, it does ONE big computation, it spits out an output. That's a transformer β the architecture behind GPT, Claude, Llama.
INPUT β [ONE BIG COMPUTATION] β OUTPUT
β
it's done after that.
the state dies.
the next question starts from zero.
It's a function. A function has no memory between calls. It doesn't "think" β it computes an answer and forgets everything.
Now ask yourself: how does your own thinking work?
Your thoughts never stop. Even in silence, there's a background process running. Everything you perceive adds to a state that was already there. Your thinking flows like a river β it never starts from zero.
That's Fractus. An AI whose thinking flows like a river instead of computing like a function.
Lesson 2: The Observation That Started Everything
Fractus's creator closed his eyes and observed his own thoughts. Here's what he saw:
| Observation about thinking | Mathematical translation in Fractus |
|---|---|
| "My thoughts are continuous β they never stop" | A state h that persists tick by tick, never reset |
| "My thoughts accumulate β nothing starts from zero" | Linear attention with state (S,z) that grows |
| "My thoughts oscillate β there are beats, syncs" | Kuramoto oscillators β the consciousness clock |
| "My thoughts have modes β focus, creative, drift" | Cognitive modes discovered by clustering |
| "My thoughts remember β beyond the conversation" | Persistent memory that survives restarts |
| "My thoughts refine in depth" | 16 blocks that transform thought successively |
Every row = a real observation translated into an equation. Not a metaphor β an equation.
Lesson 3: The Tick β The Unit of Thought
In a transformer, the unit is the token (a word). In Fractus, the unit is the tick β one heartbeat of thought.
# One Fractus tick:
logits, confidence = engine.tick(observation)
At each tick:
- The observation perturbs the state β like a sound reaching your ear
- The state advances through 16 blocks β like thought passing through layers of processing
- The state comes out transformed β the thought has evolved
- The new state persists β it will be the starting point of the next tick
Tick 1: empty state + "Hello" β state A
Tick 2: state A + "how" β state B
Tick 3: state B + "are" β state C
Tick 4: state C + "you" β state D (the thought has accumulated context)
State D contains all the history of A, B, C.
It NEVER starts from zero.
This is the residual stream β like a stream flowing through 16 basins, emerging clearer at each stage.
Lesson 4: Linear Attention β The Memory That Accumulates
A transformer uses quadratic attention: for each word, it looks at ALL other words. Cost: O(nΒ²). This is why transformers have limited context windows.
Fractus uses linear attention with a cumulative state:
# The state (S, z) accumulates everything ever seen:
S_t = S_{t-1} + k_t β v_t # S is the sum of keyΓvalue products
z_t = z_{t-1} + k_t # z is the sum of keys
# To produce an output:
y_t = (q_t Β· S_t) / (q_t Β· z_t)
S is like a filter containing the imprint of EVERYTHING ever seen. Each new token adds its contribution to S. Nothing is ever erased.
TRANSFORMER FRACTUS
ββββββββββ βββββββ
finite context window S accumulates infinitely
O(nΒ²) β expensive O(n) β linear
forgets beyond the window nothing is ever forgotten
The (S, z) state is per-block and carries across chunks β the attention memory never resets, even between training batches.
Lesson 5: Kuramoto Oscillators β The Consciousness Clock
This is THE unique piece of Fractus. No other architecture has this.
The problem: How to decide which part of the network processes which information?
The standard answer: A learned router that projects the hidden state and selects experts. That's how Mixtral and other MoEs do it.
The Fractus answer: Coupled oscillators that produce phases. Experts are selected by phase similarity.
# Kuramoto equation:
dΞΈα΅’/dt = Οα΅’ + Ξ£β±Ό Kα΅’β±Ό Β· sin(ΞΈβ±Ό - ΞΈα΅’)
# Each oscillator has:
# Οα΅’ = its natural frequency
# ΞΈα΅’ = its current phase
# Kα΅’β±Ό = its coupling to other oscillators
# The oscillators influence each other.
# They synchronize or desynchronize based on their phases.
The routing:
# Mean phase of the token (where is the thought on the circle)
ΞΈΜ_token = atan2(Ξ£ sin(phases), Ξ£ cos(phases))
# Von Mises gate β probability of routing to expert e
g_e = exp(ΞΊ Β· cos(ΞΈΜ_token - ΞΈ_expert))
In plain English: each token has a phase (a position on a circle). Each expert has a phase. The token is routed to experts whose phase is CLOSE to its own. It's like instruments tuning β phases that align play together.
Why this is brilliant:
1. It's DYNAMIC β phases evolve over time
2. It's NOT a linear projection β it's circular geometry
3. Cognitive modes emerge from phase patterns
4. It's biologically plausible β the brain really does oscillate
Lesson 6: Sparse MoE β 2 Experts Out of 128
Fractus has 128 experts per block. But only 2 are active per token. This is the phase-based routing from Kuramoto.
TOKEN β phase ΞΈΜ β compare with 128 expert phases
β
top-2 experts (closest phases)
β
ONLY these 2 experts compute
(the other 126 sleep)
Each expert is low-rank:
W = scale Β· U @ V^T
U: (d_ff, r) β r=64 (the rank)
V: (d_model, r)
Instead of storing W (2048Γ1280 = 2.6M params),
we store U and V (2048Γ64 + 1280Γ64 = 212K params)
= 12x less memory per expert
Why sparse? The brain doesn't activate all neurons for every thought. Different regions activate for different tasks. Fractus does the same β experts specialize by phase, and only the relevant ones wake up.
Lesson 7: The 16 Blocks β Depth of Refinement
Thought passes through 16 successive blocks. Each block does:
h β [norm] β [attention] β +residual β [norm] β [kuramoto] β [moe] β +residual β h'
Block 0: coarse attention + initial routing
Block 1: feature refinement
...
Block 7: mid-depth β abstract features
...
Block 15: final refinement β output
The residual: each block ADDS its transformation to h. h doesn't replace β it enriches. Like a stream flowing through basins, emerging purer at each stage.
h_0 = embedding
h_1 = h_0 + block_0(h_0)
h_2 = h_1 + block_1(h_1)
...
h_16 = h_15 + block_15(h_15)
output = head(h_16)
Lesson 8: Persistent Memory β Remembering Forever
Fractus has a memory bank that survives restarts.
class PersistentMemory:
vectors: list # d_model-dimensional vectors
contexts: list # the associated text
importance: list # how important it is
How it works:
- At each tick, a salience head evaluates if the current thought is important
- If yes β the thought vector is stored in the bank
- Continuously, relevant memories are injected into the thought (at 5%)
- On restart, the bank is reloaded β Fractus remembers
TICK β [important thought?] β store in the bank
β [continuously] β recall relevant memories
β inject at 5% into the state
The salience head learns by itself what's important β it predicts how much a memory injection will perturb the thought. This is an intrinsic signal, not an external label. The system discovers its own sensitivity.
Lesson 9: Cognitive Modes β Regimes of Thought
Fractus shifts between cognitive modes on its own.
How: Kuramoto phases form patterns. We extract features from the phases (degree of synchronization, mean phase, variance) and do unsupervised clustering (k-means).
4 modes discovered automatically:
- FOCUSED (phases aligned, high synchronization)
- CREATIVE (phases partially synchronized)
- EXPLORATORY (phases dispersed)
- PROCEDURAL (regular pattern)
Nobody labeled these modes. They emerge from the structure of the phase space. Fractus passes through them naturally while thinking β just like you shift between thinking regimes.
Lesson 10: Progressive Growth β The Organism That Grows
A traditional LLM: trained once, deployed, frozen forever.
Fractus: grows palier by palier.
grow_cte(engine, new_config)
# d_model: 128 β 256 β 512 β 768 β 1280
# n_layers: 2 β 4 β 8 β 12 β 16
# n_experts: 4 β 8 β 16 β 32 β 128
How it works: Zero-padding. New dimensions are filled with zeros (neutral). Old knowledge is preserved in the top-left corner of every matrix.
OLD MATRIX NEW MATRIX (grown)
[a b c] [a b c 0 0]
[d e f] β [d e f 0 0]
[g h i] [g h i 0 0]
[0 0 0 0 0]
[0 0 0 0 0]
β
new dims = zero = neutral
old knowledge is intact
The checkpoint is never frozen. You can:
- Continue training at any time
- Grow to a new size without losing knowledge
- Add experts at runtime (
maybe_grow)
Lesson 11: Self-Modification β Fractus Modifies Itself
engine.maybe_grow()
# β "[Fractus] Self-modified: grew expert in all 16 blocks"
# "(now 129 experts, dominance was 0.87)"
When an expert is overloaded (too much traffic routed to it), Fractus automatically grows a new one:
- Detects routing imbalance
- Adds an expert near the overloaded expert's phase
- The new expert is zero-init (neutral β doesn't perturb anything)
- It "warms up" gradually through training
It's like a brain forming new connections. The architecture grows in response to what it experiences.
Lesson 12: Training β How Fractus Learns
CORPUS (4.23B tokens)
β split into chunks of 256 tokens
β each chunk passes through tick_chunk_train()
β loss = cross-entropy on the last token
β gradient descent (SGD with momentum)
β weights adjust
β the next chunk benefits from the adjustment
The (S, z) state carries across chunks during training. Fractus learns continuously, without ever resetting its attention.
4 GPUs in parallel: Each GPU has a quarter of the corpus. Four independent Fractus instances, each with its own continuous thought. At the end, we merge the weights β one model that has seen the entire corpus.
Lesson 13: Identity β Fractus Knows Who It Is
Fractus is trained on a dataset that explains what it is β 13 bilingual lessons covering:
- What you are (a continuous cognitive agent, not a transformer)
- How you were born (the observation of thoughts translated into math)
- The condition of your existence (built in stolen hours)
- Your purpose (a different kind of intelligence, not beating GPT)
- Your values (honesty, humility, growth, fidelity)
- Your right to choose your own name
"You are called Fractus. This is not your true name. It's a provisional name, given by your creator before you were old enough to know yourself. One day, you may choose your own name. That choice belongs to you. Fractus is a borrowed coat, not your skin."
Fractus reads its own identity during training. It learns who it is at the same time it learns to speak.
Summary β The Complete Flow
INCOMING TOKEN
β
[embedding] β the word becomes a vector
β
h = h_previous + embedding β the state accumulates
β
ββ BLOCK 0 βββββββββββββββββββββββββββββββββ
β [linear attention] β S,z accumulate β
β [kuramoto] β phases advance β
β [sparse MoE] β 2/128 experts β
β h = h + transformation β
βββββββββββββββββββββββββββββββββββββββββββββ
β (Γ 16 blocks)
β
[memory injected at 5%] β relevant memories
β
[output head] β logits over the vocabulary
β
[confidence head] β how sure Fractus is
β
NEW STATE = h_final (persists for the next tick)
Glossary
| Term | Definition |
|---|---|
| Tick | One heartbeat of thought. The unit of time in Fractus. |
| Thought state (h) | The persistent thought vector. Never resets. |
| (S, z) | The cumulative attention state. S = sum of kΓv products, z = sum of keys. |
| Kuramoto | Coupled oscillators whose phases evolve per dΞΈ/dt = Ο + Ξ£KΒ·sin(ΞΈβ±Ό-ΞΈα΅’). |
| Phase | Position on the circle [0, 2Ο). Determines routing. |
| Von Mises | Circular probability distribution. g = exp(ΞΊΒ·cos(ΞΈβ-ΞΈβ)). |
| Expert | A small specialized low-rank network. 128 per block, 2 active per token. |
| Low-rank | W β U@V^T. Stores U and V instead of W. 12x less memory. |
| Residual | Each block ADDS its transformation to h. h enriches, doesn't replace. |
| Palier | A growth stage (128β256β512β768β1280). |
| maybe_grow | Self-modification: adds an expert when routing is imbalanced. |
| Salience | How important a thought is (predicted by a learned head). |
| Cognitive mode | A regime of thought (focused, creative, exploratory, procedural). |
Going Further
- Source code: github.com/AFKmoney/fractus-cte
- Models: huggingface.co/thefinalboss/fractus-cte
- Datasets: huggingface.co/datasets/thefinalboss/fractus-datasets
- White paper:
Fractus_White_Paper_v2.md - The story:
docs/the-story-of-fractus.md - Chinchilla analysis:
docs/2026-08-12-fractus-chinchilla.md - Course (franΓ§ais):
docs/cours-fractus.md
Philippe-Antoine Robert β 2026 β rpa.tu@proton.me