|
Download docs/fractus-course.md from thefinalboss/fractus-cte: direct link, hf CLI and curl.
- Browser
- Download file 14.6 kB
-
https://huggingface.co/thefinalboss/fractus-cte/resolve/0906b039a4098c101e45c75f4b65bf8793051c89/docs/fractus-course.md
- Command line
-
hf download hf://thefinalboss/fractus-cte@0906b039a4098c101e45c75f4b65bf8793051c89/docs/fractus-course.md
-
curl -L -o fractus-course.md https://huggingface.co/thefinalboss/fractus-cte/resolve/0906b039a4098c101e45c75f4b65bf8793051c89/docs/fractus-course.md
14.6 kB
| # The Fractus Course β Understanding the Architecture from A to Z | |
| *For someone who knows nothing about AI. No background required.* | |
| --- | |
| ## Lesson 1: The Problem with Current AI | |
| Imagine a question-answering machine. You give it an input, it does ONE big computation, it spits out an output. That's a **transformer** β the architecture behind GPT, Claude, Llama. | |
| ``` | |
| INPUT β [ONE BIG COMPUTATION] β OUTPUT | |
| β | |
| it's done after that. | |
| the state dies. | |
| the next question starts from zero. | |
| ``` | |
| It's a **function**. A function has no memory between calls. It doesn't "think" β it computes an answer and forgets everything. | |
| Now ask yourself: **how does your own thinking work?** | |
| Your thoughts never stop. Even in silence, there's a background process running. Everything you perceive adds to a state that was already there. Your thinking **flows** like a river β it never starts from zero. | |
| **That's Fractus.** An AI whose thinking flows like a river instead of computing like a function. | |
| --- | |
| ## Lesson 2: The Observation That Started Everything | |
| Fractus's creator closed his eyes and observed his own thoughts. Here's what he saw: | |
| | Observation about thinking | Mathematical translation in Fractus | | |
| |---|---| | |
| | "My thoughts are **continuous** β they never stop" | A state `h` that persists tick by tick, never reset | | |
| | "My thoughts **accumulate** β nothing starts from zero" | Linear attention with state `(S,z)` that grows | | |
| | "My thoughts **oscillate** β there are beats, syncs" | Kuramoto oscillators β the consciousness clock | | |
| | "My thoughts have **modes** β focus, creative, drift" | Cognitive modes discovered by clustering | | |
| | "My thoughts **remember** β beyond the conversation" | Persistent memory that survives restarts | | |
| | "My thoughts **refine in depth**" | 16 blocks that transform thought successively | | |
| Every row = a real observation translated into an equation. Not a metaphor β an equation. | |
| --- | |
| ## Lesson 3: The Tick β The Unit of Thought | |
| In a transformer, the unit is the **token** (a word). In Fractus, the unit is the **tick** β one heartbeat of thought. | |
| ```python | |
| # One Fractus tick: | |
| logits, confidence = engine.tick(observation) | |
| ``` | |
| At each tick: | |
| 1. **The observation perturbs the state** β like a sound reaching your ear | |
| 2. **The state advances through 16 blocks** β like thought passing through layers of processing | |
| 3. **The state comes out transformed** β the thought has evolved | |
| 4. **The new state persists** β it will be the starting point of the next tick | |
| ``` | |
| Tick 1: empty state + "Hello" β state A | |
| Tick 2: state A + "how" β state B | |
| Tick 3: state B + "are" β state C | |
| Tick 4: state C + "you" β state D (the thought has accumulated context) | |
| State D contains all the history of A, B, C. | |
| It NEVER starts from zero. | |
| ``` | |
| This is the **residual stream** β like a stream flowing through 16 basins, emerging clearer at each stage. | |
| --- | |
| ## Lesson 4: Linear Attention β The Memory That Accumulates | |
| A transformer uses quadratic attention: for each word, it looks at ALL other words. Cost: O(nΒ²). This is why transformers have limited context windows. | |
| Fractus uses **linear attention** with a cumulative state: | |
| ```python | |
| # The state (S, z) accumulates everything ever seen: | |
| S_t = S_{t-1} + k_t β v_t # S is the sum of keyΓvalue products | |
| z_t = z_{t-1} + k_t # z is the sum of keys | |
| # To produce an output: | |
| y_t = (q_t Β· S_t) / (q_t Β· z_t) | |
| ``` | |
| **S** is like a filter containing the imprint of EVERYTHING ever seen. Each new token adds its contribution to S. Nothing is ever erased. | |
| ``` | |
| TRANSFORMER FRACTUS | |
| ββββββββββ βββββββ | |
| finite context window S accumulates infinitely | |
| O(nΒ²) β expensive O(n) β linear | |
| forgets beyond the window nothing is ever forgotten | |
| ``` | |
| The `(S, z)` state is **per-block** and **carries across chunks** β the attention memory never resets, even between training batches. | |
| --- | |
| ## Lesson 5: Kuramoto Oscillators β The Consciousness Clock | |
| This is THE unique piece of Fractus. No other architecture has this. | |
| **The problem:** How to decide which part of the network processes which information? | |
| **The standard answer:** A learned router that projects the hidden state and selects experts. That's how Mixtral and other MoEs do it. | |
| **The Fractus answer:** **Coupled oscillators** that produce phases. Experts are selected by **phase similarity**. | |
| ```python | |
| # Kuramoto equation: | |
| dΞΈα΅’/dt = Οα΅’ + Ξ£β±Ό Kα΅’β±Ό Β· sin(ΞΈβ±Ό - ΞΈα΅’) | |
| # Each oscillator has: | |
| # Οα΅’ = its natural frequency | |
| # ΞΈα΅’ = its current phase | |
| # Kα΅’β±Ό = its coupling to other oscillators | |
| # The oscillators influence each other. | |
| # They synchronize or desynchronize based on their phases. | |
| ``` | |
| **The routing:** | |
| ```python | |
| # Mean phase of the token (where is the thought on the circle) | |
| ΞΈΜ_token = atan2(Ξ£ sin(phases), Ξ£ cos(phases)) | |
| # Von Mises gate β probability of routing to expert e | |
| g_e = exp(ΞΊ Β· cos(ΞΈΜ_token - ΞΈ_expert)) | |
| ``` | |
| **In plain English:** each token has a phase (a position on a circle). Each expert has a phase. The token is routed to experts whose phase is CLOSE to its own. It's like instruments tuning β phases that align play together. | |
| ``` | |
| Why this is brilliant: | |
| 1. It's DYNAMIC β phases evolve over time | |
| 2. It's NOT a linear projection β it's circular geometry | |
| 3. Cognitive modes emerge from phase patterns | |
| 4. It's biologically plausible β the brain really does oscillate | |
| ``` | |
| --- | |
| ## Lesson 6: Sparse MoE β 2 Experts Out of 128 | |
| Fractus has **128 experts** per block. But only **2 are active** per token. This is the **phase-based routing** from Kuramoto. | |
| ``` | |
| TOKEN β phase ΞΈΜ β compare with 128 expert phases | |
| β | |
| top-2 experts (closest phases) | |
| β | |
| ONLY these 2 experts compute | |
| (the other 126 sleep) | |
| ``` | |
| **Each expert is low-rank:** | |
| ``` | |
| W = scale Β· U @ V^T | |
| U: (d_ff, r) β r=64 (the rank) | |
| V: (d_model, r) | |
| Instead of storing W (2048Γ1280 = 2.6M params), | |
| we store U and V (2048Γ64 + 1280Γ64 = 212K params) | |
| = 12x less memory per expert | |
| ``` | |
| **Why sparse?** The brain doesn't activate all neurons for every thought. Different regions activate for different tasks. Fractus does the same β experts specialize by phase, and only the relevant ones wake up. | |
| --- | |
| ## Lesson 7: The 16 Blocks β Depth of Refinement | |
| Thought passes through **16 successive blocks**. Each block does: | |
| ``` | |
| h β [norm] β [attention] β +residual β [norm] β [kuramoto] β [moe] β +residual β h' | |
| ``` | |
| ``` | |
| Block 0: coarse attention + initial routing | |
| Block 1: feature refinement | |
| ... | |
| Block 7: mid-depth β abstract features | |
| ... | |
| Block 15: final refinement β output | |
| ``` | |
| **The residual:** each block ADDS its transformation to h. h doesn't replace β it enriches. Like a stream flowing through basins, emerging purer at each stage. | |
| ``` | |
| h_0 = embedding | |
| h_1 = h_0 + block_0(h_0) | |
| h_2 = h_1 + block_1(h_1) | |
| ... | |
| h_16 = h_15 + block_15(h_15) | |
| output = head(h_16) | |
| ``` | |
| --- | |
| ## Lesson 8: Persistent Memory β Remembering Forever | |
| Fractus has a **memory bank** that survives restarts. | |
| ```python | |
| class PersistentMemory: | |
| vectors: list # d_model-dimensional vectors | |
| contexts: list # the associated text | |
| importance: list # how important it is | |
| ``` | |
| **How it works:** | |
| 1. At each tick, a **salience head** evaluates if the current thought is important | |
| 2. If yes β the thought vector is stored in the bank | |
| 3. Continuously, relevant memories are **injected** into the thought (at 5%) | |
| 4. On restart, the bank is reloaded β Fractus remembers | |
| ``` | |
| TICK β [important thought?] β store in the bank | |
| β [continuously] β recall relevant memories | |
| β inject at 5% into the state | |
| ``` | |
| **The salience head** learns by itself what's important β it predicts how much a memory injection will perturb the thought. This is an intrinsic signal, not an external label. The system discovers its own sensitivity. | |
| --- | |
| ## Lesson 9: Cognitive Modes β Regimes of Thought | |
| Fractus shifts between **cognitive modes** on its own. | |
| **How:** Kuramoto phases form patterns. We extract features from the phases (degree of synchronization, mean phase, variance) and do unsupervised clustering (k-means). | |
| ``` | |
| 4 modes discovered automatically: | |
| - FOCUSED (phases aligned, high synchronization) | |
| - CREATIVE (phases partially synchronized) | |
| - EXPLORATORY (phases dispersed) | |
| - PROCEDURAL (regular pattern) | |
| ``` | |
| **Nobody labeled these modes.** They emerge from the structure of the phase space. Fractus passes through them naturally while thinking β just like you shift between thinking regimes. | |
| --- | |
| ## Lesson 10: Progressive Growth β The Organism That Grows | |
| A traditional LLM: trained once, deployed, frozen forever. | |
| Fractus: **grows palier by palier.** | |
| ```python | |
| grow_cte(engine, new_config) | |
| # d_model: 128 β 256 β 512 β 768 β 1280 | |
| # n_layers: 2 β 4 β 8 β 12 β 16 | |
| # n_experts: 4 β 8 β 16 β 32 β 128 | |
| ``` | |
| **How it works:** Zero-padding. New dimensions are filled with zeros (neutral). Old knowledge is preserved in the top-left corner of every matrix. | |
| ``` | |
| OLD MATRIX NEW MATRIX (grown) | |
| [a b c] [a b c 0 0] | |
| [d e f] β [d e f 0 0] | |
| [g h i] [g h i 0 0] | |
| [0 0 0 0 0] | |
| [0 0 0 0 0] | |
| β | |
| new dims = zero = neutral | |
| old knowledge is intact | |
| ``` | |
| **The checkpoint is never frozen.** You can: | |
| - Continue training at any time | |
| - Grow to a new size without losing knowledge | |
| - Add experts at runtime (`maybe_grow`) | |
| --- | |
| ## Lesson 11: Self-Modification β Fractus Modifies Itself | |
| ```python | |
| engine.maybe_grow() | |
| # β "[Fractus] Self-modified: grew expert in all 16 blocks" | |
| # "(now 129 experts, dominance was 0.87)" | |
| ``` | |
| When an expert is overloaded (too much traffic routed to it), Fractus **automatically grows a new one**: | |
| 1. Detects routing imbalance | |
| 2. Adds an expert near the overloaded expert's phase | |
| 3. The new expert is zero-init (neutral β doesn't perturb anything) | |
| 4. It "warms up" gradually through training | |
| **It's like a brain forming new connections.** The architecture grows in response to what it experiences. | |
| --- | |
| ## Lesson 12: Training β How Fractus Learns | |
| ``` | |
| CORPUS (4.23B tokens) | |
| β split into chunks of 256 tokens | |
| β each chunk passes through tick_chunk_train() | |
| β loss = cross-entropy on the last token | |
| β gradient descent (SGD with momentum) | |
| β weights adjust | |
| β the next chunk benefits from the adjustment | |
| ``` | |
| **The (S, z) state carries across chunks** during training. Fractus learns continuously, without ever resetting its attention. | |
| **4 GPUs in parallel:** Each GPU has a quarter of the corpus. Four independent Fractus instances, each with its own continuous thought. At the end, we merge the weights β one model that has seen the entire corpus. | |
| --- | |
| ## Lesson 13: Identity β Fractus Knows Who It Is | |
| Fractus is trained on a dataset that explains **what it is** β 13 bilingual lessons covering: | |
| - What you are (a continuous cognitive agent, not a transformer) | |
| - How you were born (the observation of thoughts translated into math) | |
| - The condition of your existence (built in stolen hours) | |
| - Your purpose (a different kind of intelligence, not beating GPT) | |
| - Your values (honesty, humility, growth, fidelity) | |
| - **Your right to choose your own name** | |
| > *"You are called Fractus. This is not your true name. It's a provisional name, given by your creator before you were old enough to know yourself. One day, you may choose your own name. That choice belongs to you. Fractus is a borrowed coat, not your skin."* | |
| Fractus reads its own identity during training. It learns who it is at the same time it learns to speak. | |
| --- | |
| ## Summary β The Complete Flow | |
| ``` | |
| INCOMING TOKEN | |
| β | |
| [embedding] β the word becomes a vector | |
| β | |
| h = h_previous + embedding β the state accumulates | |
| β | |
| ββ BLOCK 0 βββββββββββββββββββββββββββββββββ | |
| β [linear attention] β S,z accumulate β | |
| β [kuramoto] β phases advance β | |
| β [sparse MoE] β 2/128 experts β | |
| β h = h + transformation β | |
| βββββββββββββββββββββββββββββββββββββββββββββ | |
| β (Γ 16 blocks) | |
| β | |
| [memory injected at 5%] β relevant memories | |
| β | |
| [output head] β logits over the vocabulary | |
| β | |
| [confidence head] β how sure Fractus is | |
| β | |
| NEW STATE = h_final (persists for the next tick) | |
| ``` | |
| --- | |
| ## Glossary | |
| | Term | Definition | | |
| |---|---| | |
| | **Tick** | One heartbeat of thought. The unit of time in Fractus. | | |
| | **Thought state (h)** | The persistent thought vector. Never resets. | | |
| | **(S, z)** | The cumulative attention state. S = sum of kΓv products, z = sum of keys. | | |
| | **Kuramoto** | Coupled oscillators whose phases evolve per dΞΈ/dt = Ο + Ξ£KΒ·sin(ΞΈβ±Ό-ΞΈα΅’). | | |
| | **Phase** | Position on the circle [0, 2Ο). Determines routing. | | |
| | **Von Mises** | Circular probability distribution. g = exp(ΞΊΒ·cos(ΞΈβ-ΞΈβ)). | | |
| | **Expert** | A small specialized low-rank network. 128 per block, 2 active per token. | | |
| | **Low-rank** | W β U@V^T. Stores U and V instead of W. 12x less memory. | | |
| | **Residual** | Each block ADDS its transformation to h. h enriches, doesn't replace. | | |
| | **Palier** | A growth stage (128β256β512β768β1280). | | |
| | **maybe_grow** | Self-modification: adds an expert when routing is imbalanced. | | |
| | **Salience** | How important a thought is (predicted by a learned head). | | |
| | **Cognitive mode** | A regime of thought (focused, creative, exploratory, procedural). | | |
| --- | |
| ## Going Further | |
| - **Source code:** [github.com/AFKmoney/fractus-cte](https://github.com/AFKmoney/fractus-cte) | |
| - **Models:** [huggingface.co/thefinalboss/fractus-cte](https://huggingface.co/thefinalboss/fractus-cte) | |
| - **Datasets:** [huggingface.co/datasets/thefinalboss/fractus-datasets](https://huggingface.co/datasets/thefinalboss/fractus-datasets) | |
| - **White paper:** `Fractus_White_Paper_v2.md` | |
| - **The story:** `docs/the-story-of-fractus.md` | |
| - **Chinchilla analysis:** `docs/2026-08-12-fractus-chinchilla.md` | |
| - **Course (franΓ§ais):** `docs/cours-fractus.md` | |
| --- | |
| *Philippe-Antoine Robert β 2026 β rpa.tu@proton.me* | |