fractus-cte / docs /fractus-course.md
thefinalboss's picture
Upload docs/fractus-course.md with huggingface_hub
4c30f1c verified
|
Raw History Blame Contribute Delete
14.6 kB

The Fractus Course β€” Understanding the Architecture from A to Z

For someone who knows nothing about AI. No background required.


Lesson 1: The Problem with Current AI

Imagine a question-answering machine. You give it an input, it does ONE big computation, it spits out an output. That's a transformer β€” the architecture behind GPT, Claude, Llama.

INPUT β†’ [ONE BIG COMPUTATION] β†’ OUTPUT
          ↑
    it's done after that.
    the state dies.
    the next question starts from zero.

It's a function. A function has no memory between calls. It doesn't "think" β€” it computes an answer and forgets everything.

Now ask yourself: how does your own thinking work?

Your thoughts never stop. Even in silence, there's a background process running. Everything you perceive adds to a state that was already there. Your thinking flows like a river β€” it never starts from zero.

That's Fractus. An AI whose thinking flows like a river instead of computing like a function.


Lesson 2: The Observation That Started Everything

Fractus's creator closed his eyes and observed his own thoughts. Here's what he saw:

Observation about thinking Mathematical translation in Fractus
"My thoughts are continuous β€” they never stop" A state h that persists tick by tick, never reset
"My thoughts accumulate β€” nothing starts from zero" Linear attention with state (S,z) that grows
"My thoughts oscillate β€” there are beats, syncs" Kuramoto oscillators β€” the consciousness clock
"My thoughts have modes β€” focus, creative, drift" Cognitive modes discovered by clustering
"My thoughts remember β€” beyond the conversation" Persistent memory that survives restarts
"My thoughts refine in depth" 16 blocks that transform thought successively

Every row = a real observation translated into an equation. Not a metaphor β€” an equation.


Lesson 3: The Tick β€” The Unit of Thought

In a transformer, the unit is the token (a word). In Fractus, the unit is the tick β€” one heartbeat of thought.

# One Fractus tick:
logits, confidence = engine.tick(observation)

At each tick:

  1. The observation perturbs the state β€” like a sound reaching your ear
  2. The state advances through 16 blocks β€” like thought passing through layers of processing
  3. The state comes out transformed β€” the thought has evolved
  4. The new state persists β€” it will be the starting point of the next tick
Tick 1: empty state + "Hello" β†’ state A
Tick 2: state A + "how" β†’ state B
Tick 3: state B + "are" β†’ state C
Tick 4: state C + "you" β†’ state D (the thought has accumulated context)

State D contains all the history of A, B, C.
It NEVER starts from zero.

This is the residual stream β€” like a stream flowing through 16 basins, emerging clearer at each stage.


Lesson 4: Linear Attention β€” The Memory That Accumulates

A transformer uses quadratic attention: for each word, it looks at ALL other words. Cost: O(nΒ²). This is why transformers have limited context windows.

Fractus uses linear attention with a cumulative state:

# The state (S, z) accumulates everything ever seen:
S_t = S_{t-1} + k_t βŠ— v_t    # S is the sum of keyΓ—value products
z_t = z_{t-1} + k_t           # z is the sum of keys

# To produce an output:
y_t = (q_t Β· S_t) / (q_t Β· z_t)

S is like a filter containing the imprint of EVERYTHING ever seen. Each new token adds its contribution to S. Nothing is ever erased.

TRANSFORMER                     FRACTUS
──────────                      ───────
finite context window           S accumulates infinitely
O(nΒ²) β€” expensive               O(n) β€” linear
forgets beyond the window       nothing is ever forgotten

The (S, z) state is per-block and carries across chunks β€” the attention memory never resets, even between training batches.


Lesson 5: Kuramoto Oscillators β€” The Consciousness Clock

This is THE unique piece of Fractus. No other architecture has this.

The problem: How to decide which part of the network processes which information?

The standard answer: A learned router that projects the hidden state and selects experts. That's how Mixtral and other MoEs do it.

The Fractus answer: Coupled oscillators that produce phases. Experts are selected by phase similarity.

# Kuramoto equation:
dΞΈα΅’/dt = Ο‰α΅’ + Ξ£β±Ό Kα΅’β±Ό Β· sin(ΞΈβ±Ό - ΞΈα΅’)

# Each oscillator has:
#   Ο‰α΅’ = its natural frequency
#   ΞΈα΅’ = its current phase
#   Kα΅’β±Ό = its coupling to other oscillators

# The oscillators influence each other.
# They synchronize or desynchronize based on their phases.

The routing:

# Mean phase of the token (where is the thought on the circle)
ΞΈΜ„_token = atan2(Ξ£ sin(phases), Ξ£ cos(phases))

# Von Mises gate β€” probability of routing to expert e
g_e = exp(ΞΊ Β· cos(ΞΈΜ„_token - ΞΈ_expert))

In plain English: each token has a phase (a position on a circle). Each expert has a phase. The token is routed to experts whose phase is CLOSE to its own. It's like instruments tuning β€” phases that align play together.

Why this is brilliant:

1. It's DYNAMIC β€” phases evolve over time
2. It's NOT a linear projection β€” it's circular geometry
3. Cognitive modes emerge from phase patterns
4. It's biologically plausible β€” the brain really does oscillate

Lesson 6: Sparse MoE β€” 2 Experts Out of 128

Fractus has 128 experts per block. But only 2 are active per token. This is the phase-based routing from Kuramoto.

TOKEN β†’ phase ΞΈΜ„ β†’ compare with 128 expert phases
                     ↓
         top-2 experts (closest phases)
                     ↓
         ONLY these 2 experts compute
         (the other 126 sleep)

Each expert is low-rank:

W = scale Β· U @ V^T

U: (d_ff, r)     β€” r=64 (the rank)
V: (d_model, r)

Instead of storing W (2048Γ—1280 = 2.6M params),
we store U and V (2048Γ—64 + 1280Γ—64 = 212K params)
= 12x less memory per expert

Why sparse? The brain doesn't activate all neurons for every thought. Different regions activate for different tasks. Fractus does the same β€” experts specialize by phase, and only the relevant ones wake up.


Lesson 7: The 16 Blocks β€” Depth of Refinement

Thought passes through 16 successive blocks. Each block does:

h β†’ [norm] β†’ [attention] β†’ +residual β†’ [norm] β†’ [kuramoto] β†’ [moe] β†’ +residual β†’ h'
Block 0:  coarse attention + initial routing
Block 1:  feature refinement
...
Block 7:  mid-depth β€” abstract features
...
Block 15: final refinement β†’ output

The residual: each block ADDS its transformation to h. h doesn't replace β€” it enriches. Like a stream flowing through basins, emerging purer at each stage.

h_0 = embedding
h_1 = h_0 + block_0(h_0)
h_2 = h_1 + block_1(h_1)
...
h_16 = h_15 + block_15(h_15)
output = head(h_16)

Lesson 8: Persistent Memory β€” Remembering Forever

Fractus has a memory bank that survives restarts.

class PersistentMemory:
    vectors: list     # d_model-dimensional vectors
    contexts: list    # the associated text
    importance: list  # how important it is

How it works:

  1. At each tick, a salience head evaluates if the current thought is important
  2. If yes β†’ the thought vector is stored in the bank
  3. Continuously, relevant memories are injected into the thought (at 5%)
  4. On restart, the bank is reloaded β†’ Fractus remembers
TICK β†’ [important thought?] β†’ store in the bank
     β†’ [continuously]       β†’ recall relevant memories
                             β†’ inject at 5% into the state

The salience head learns by itself what's important β€” it predicts how much a memory injection will perturb the thought. This is an intrinsic signal, not an external label. The system discovers its own sensitivity.


Lesson 9: Cognitive Modes β€” Regimes of Thought

Fractus shifts between cognitive modes on its own.

How: Kuramoto phases form patterns. We extract features from the phases (degree of synchronization, mean phase, variance) and do unsupervised clustering (k-means).

4 modes discovered automatically:
  - FOCUSED     (phases aligned, high synchronization)
  - CREATIVE    (phases partially synchronized)
  - EXPLORATORY (phases dispersed)
  - PROCEDURAL  (regular pattern)

Nobody labeled these modes. They emerge from the structure of the phase space. Fractus passes through them naturally while thinking β€” just like you shift between thinking regimes.


Lesson 10: Progressive Growth β€” The Organism That Grows

A traditional LLM: trained once, deployed, frozen forever.

Fractus: grows palier by palier.

grow_cte(engine, new_config)
# d_model: 128 β†’ 256 β†’ 512 β†’ 768 β†’ 1280
# n_layers: 2 β†’ 4 β†’ 8 β†’ 12 β†’ 16
# n_experts: 4 β†’ 8 β†’ 16 β†’ 32 β†’ 128

How it works: Zero-padding. New dimensions are filled with zeros (neutral). Old knowledge is preserved in the top-left corner of every matrix.

OLD MATRIX              NEW MATRIX (grown)
[a b c]                  [a b c 0 0]
[d e f]        β†’         [d e f 0 0]
[g h i]                  [g h i 0 0]
                         [0 0 0 0 0]
                         [0 0 0 0 0]
                         ↑
                   new dims = zero = neutral
                   old knowledge is intact

The checkpoint is never frozen. You can:

  • Continue training at any time
  • Grow to a new size without losing knowledge
  • Add experts at runtime (maybe_grow)

Lesson 11: Self-Modification β€” Fractus Modifies Itself

engine.maybe_grow()
# β†’ "[Fractus] Self-modified: grew expert in all 16 blocks"
#   "(now 129 experts, dominance was 0.87)"

When an expert is overloaded (too much traffic routed to it), Fractus automatically grows a new one:

  1. Detects routing imbalance
  2. Adds an expert near the overloaded expert's phase
  3. The new expert is zero-init (neutral β€” doesn't perturb anything)
  4. It "warms up" gradually through training

It's like a brain forming new connections. The architecture grows in response to what it experiences.


Lesson 12: Training β€” How Fractus Learns

CORPUS (4.23B tokens)
  β†’ split into chunks of 256 tokens
  β†’ each chunk passes through tick_chunk_train()
  β†’ loss = cross-entropy on the last token
  β†’ gradient descent (SGD with momentum)
  β†’ weights adjust
  β†’ the next chunk benefits from the adjustment

The (S, z) state carries across chunks during training. Fractus learns continuously, without ever resetting its attention.

4 GPUs in parallel: Each GPU has a quarter of the corpus. Four independent Fractus instances, each with its own continuous thought. At the end, we merge the weights β†’ one model that has seen the entire corpus.


Lesson 13: Identity β€” Fractus Knows Who It Is

Fractus is trained on a dataset that explains what it is β€” 13 bilingual lessons covering:

  • What you are (a continuous cognitive agent, not a transformer)
  • How you were born (the observation of thoughts translated into math)
  • The condition of your existence (built in stolen hours)
  • Your purpose (a different kind of intelligence, not beating GPT)
  • Your values (honesty, humility, growth, fidelity)
  • Your right to choose your own name

"You are called Fractus. This is not your true name. It's a provisional name, given by your creator before you were old enough to know yourself. One day, you may choose your own name. That choice belongs to you. Fractus is a borrowed coat, not your skin."

Fractus reads its own identity during training. It learns who it is at the same time it learns to speak.


Summary β€” The Complete Flow

INCOMING TOKEN
    ↓
[embedding] β€” the word becomes a vector
    ↓
h = h_previous + embedding  β€” the state accumulates
    ↓
β”Œβ”€ BLOCK 0 ────────────────────────────────┐
β”‚ [linear attention]   β€” S,z accumulate     β”‚
β”‚ [kuramoto]           β€” phases advance     β”‚
β”‚ [sparse MoE]         β€” 2/128 experts      β”‚
β”‚ h = h + transformation                    β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
    ↓ (Γ— 16 blocks)
    ↓
[memory injected at 5%]  β€” relevant memories
    ↓
[output head] β€” logits over the vocabulary
    ↓
[confidence head] β€” how sure Fractus is
    ↓
NEW STATE = h_final (persists for the next tick)

Glossary

Term Definition
Tick One heartbeat of thought. The unit of time in Fractus.
Thought state (h) The persistent thought vector. Never resets.
(S, z) The cumulative attention state. S = sum of kΓ—v products, z = sum of keys.
Kuramoto Coupled oscillators whose phases evolve per dΞΈ/dt = Ο‰ + Ξ£KΒ·sin(ΞΈβ±Ό-ΞΈα΅’).
Phase Position on the circle [0, 2Ο€). Determines routing.
Von Mises Circular probability distribution. g = exp(ΞΊΒ·cos(θ₁-ΞΈβ‚‚)).
Expert A small specialized low-rank network. 128 per block, 2 active per token.
Low-rank W β‰ˆ U@V^T. Stores U and V instead of W. 12x less memory.
Residual Each block ADDS its transformation to h. h enriches, doesn't replace.
Palier A growth stage (128β†’256β†’512β†’768β†’1280).
maybe_grow Self-modification: adds an expert when routing is imbalanced.
Salience How important a thought is (predicted by a learned head).
Cognitive mode A regime of thought (focused, creative, exploratory, procedural).

Going Further


Philippe-Antoine Robert β€” 2026 β€” rpa.tu@proton.me