docs: add Vision section — tick_vec multimodal path + CIFAR-10 eyes prototype
Browse files
README.md
CHANGED
|
@@ -111,7 +111,7 @@ ContinuousThoughtEngine
|
|
| 111 |
| **PhaseRoutedMoE** | Sparse experts routed by oscillator phases |
|
| 112 |
| **Kuramoto Clock** | A dynamical system that drives routing decisions |
|
| 113 |
| **Online Trainer** | Learns continuously, one chunk at a time |
|
| 114 |
-
| **
|
| 115 |
|
| 116 |
---
|
| 117 |
|
|
@@ -131,6 +131,26 @@ ContinuousThoughtEngine
|
|
| 131 |
|
| 132 |
**Honest generation state:** teacher-forced loss keeps falling (ema_tf ≈ 1.3–1.8 on the best GPUs), but free-run generation is still non-linguistic — diverse word-salad with repetition loops, better with window-based decoding (~25–28 unique tokens vs ~10–12 chunk-based). This is the documented exposure-bias gap, being closed by scheduled sampling + aligned decoding. Low TF loss does **not** mean the model can speak yet. See [Trusted Loss](#trusted-loss-reading-the-numbers).
|
| 133 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 134 |
### Adapted Chinchilla target
|
| 135 |
|
| 136 |
Standard Chinchilla (20 tokens/param) assumes dense transformers where all parameters train on every token. Fractus violates that: only ~119M of 1.05B params are active per token, and the Kuramoto clock trains end-to-end now. The adapted target is **20 × active params ≈ 2.4–3.7B tokens**; with a warm-started checkpoint, the practical target is **~2.5–3B tokens** — which is why the phase-2 corpus is 3.44B. Since Fractus grows continuously (`maybe_grow`), this is an instantaneous cumulative-exposure requirement that grows with the model, not a final-stop condition. Details: [`docs/2026-08-12-fractus-chinchilla.md`](docs/2026-08-12-fractus-chinchilla.md).
|
|
@@ -343,6 +363,10 @@ engine.attach_memory(memory)
|
|
| 343 |
engine.reset_thought(batch_size=1)
|
| 344 |
logits, confidence = engine.tick(torch.tensor([42]))
|
| 345 |
print(f"Confidence: {confidence.item():.2f}")
|
|
|
|
|
|
|
|
|
|
|
|
|
| 346 |
```
|
| 347 |
|
| 348 |
---
|
|
@@ -402,6 +426,8 @@ fractus-cte/
|
|
| 402 |
|
| 403 |
**Kuramoto clock**: coupled oscillators producing phase vectors that route tokens to experts — learned end-to-end since the routing surgery.
|
| 404 |
|
|
|
|
|
|
|
| 405 |
**Surgery**: patching code around a preserved checkpoint mid-training, resuming at the exact token offset. Weights are never wiped for a routing, objective, or decode bug.
|
| 406 |
|
| 407 |
---
|
|
|
|
| 111 |
| **PhaseRoutedMoE** | Sparse experts routed by oscillator phases |
|
| 112 |
| **Kuramoto Clock** | A dynamical system that drives routing decisions |
|
| 113 |
| **Online Trainer** | Learns continuously, one chunk at a time |
|
| 114 |
+
| **Vision (`tick_vec`)** | Multimodal input path — image patches drive the engine directly, bypassing the token embedding (CIFAR "eyes" prototype trained while the text run continued) |
|
| 115 |
|
| 116 |
---
|
| 117 |
|
|
|
|
| 131 |
|
| 132 |
**Honest generation state:** teacher-forced loss keeps falling (ema_tf ≈ 1.3–1.8 on the best GPUs), but free-run generation is still non-linguistic — diverse word-salad with repetition loops, better with window-based decoding (~25–28 unique tokens vs ~10–12 chunk-based). This is the documented exposure-bias gap, being closed by scheduled sampling + aligned decoding. Low TF loss does **not** mean the model can speak yet. See [Trusted Loss](#trusted-loss-reading-the-numbers).
|
| 133 |
|
| 134 |
+
## Vision — the CIFAR "eyes" prototype
|
| 135 |
+
|
| 136 |
+
Fractus is not text-only. The engine exposes **`tick_vec(obs_vec)`**: a multimodal entry point that accepts a precomputed `(B, d_model)` vector and injects it directly into the residual thought state — **bypassing the token embedding entirely**. Any modality that can be embedded into `d_model` dimensions can drive continuous thought: images, audio, sensor streams.
|
| 137 |
+
|
| 138 |
+
```
|
| 139 |
+
image → PatchEmbed → (B, d_model) patch vectors
|
| 140 |
+
│
|
| 141 |
+
▼
|
| 142 |
+
engine.tick_vec(patch) ← no tokenizer involved
|
| 143 |
+
│
|
| 144 |
+
thought state advances through all 16 blocks
|
| 145 |
+
(same Kuramoto routing + MoE as text)
|
| 146 |
+
```
|
| 147 |
+
|
| 148 |
+
**Proof of concept — CIFAR-10 eyes:** a small CTE + PatchEmbed stack (`fractus/nn/vision.py`) was trained on CPU on real CIFAR-10 images (`fractus_eyes_cifar_final.pt`), demonstrating that image patches can drive the continuous-thought loop. Crucially, this ran as a **parallel track**: the 8-GPU text run was never interrupted for eyes work — GPU text digestion and CPU vision learning proceed simultaneously on the same living system.
|
| 149 |
+
|
| 150 |
+
This mirrors the biological premise: eyes evolved as a peripheral sense feeding a central dynamical brain, not as a core feature of it. The text 1B is the brain; vision is a front-end that plugs into `tick_vec`.
|
| 151 |
+
|
| 152 |
+
---
|
| 153 |
+
|
| 154 |
### Adapted Chinchilla target
|
| 155 |
|
| 156 |
Standard Chinchilla (20 tokens/param) assumes dense transformers where all parameters train on every token. Fractus violates that: only ~119M of 1.05B params are active per token, and the Kuramoto clock trains end-to-end now. The adapted target is **20 × active params ≈ 2.4–3.7B tokens**; with a warm-started checkpoint, the practical target is **~2.5–3B tokens** — which is why the phase-2 corpus is 3.44B. Since Fractus grows continuously (`maybe_grow`), this is an instantaneous cumulative-exposure requirement that grows with the model, not a final-stop condition. Details: [`docs/2026-08-12-fractus-chinchilla.md`](docs/2026-08-12-fractus-chinchilla.md).
|
|
|
|
| 363 |
engine.reset_thought(batch_size=1)
|
| 364 |
logits, confidence = engine.tick(torch.tensor([42]))
|
| 365 |
print(f"Confidence: {confidence.item():.2f}")
|
| 366 |
+
|
| 367 |
+
# Vision: drive thought with an image patch vector (no tokenizer)
|
| 368 |
+
patch_vec = torch.randn(1, 128) # any (B, d_model) embedding
|
| 369 |
+
logits, confidence = engine.tick_vec(patch_vec)
|
| 370 |
```
|
| 371 |
|
| 372 |
---
|
|
|
|
| 426 |
|
| 427 |
**Kuramoto clock**: coupled oscillators producing phase vectors that route tokens to experts — learned end-to-end since the routing surgery.
|
| 428 |
|
| 429 |
+
**`tick_vec`**: the multimodal tick — feed any precomputed `(B, d_model)` vector (image patches, embeddings from any encoder) straight into the thought state without tokens. This is how vision plugs in.
|
| 430 |
+
|
| 431 |
**Surgery**: patching code around a preserved checkpoint mid-training, resuming at the exact token offset. Weights are never wiped for a routing, objective, or decode bug.
|
| 432 |
|
| 433 |
---
|