thefinalboss commited on
Commit
8d13dac
·
verified ·
1 Parent(s): 7a89bec

docs: add Vision section — tick_vec multimodal path + CIFAR-10 eyes prototype

Browse files
Files changed (1) hide show
  1. README.md +27 -1
README.md CHANGED
@@ -111,7 +111,7 @@ ContinuousThoughtEngine
111
  | **PhaseRoutedMoE** | Sparse experts routed by oscillator phases |
112
  | **Kuramoto Clock** | A dynamical system that drives routing decisions |
113
  | **Online Trainer** | Learns continuously, one chunk at a time |
114
- | **Multimodal `tick_vec`** | Vision patches bypass token embedding (CIFAR "eyes" prototype trained alongside text) |
115
 
116
  ---
117
 
@@ -131,6 +131,26 @@ ContinuousThoughtEngine
131
 
132
  **Honest generation state:** teacher-forced loss keeps falling (ema_tf ≈ 1.3–1.8 on the best GPUs), but free-run generation is still non-linguistic — diverse word-salad with repetition loops, better with window-based decoding (~25–28 unique tokens vs ~10–12 chunk-based). This is the documented exposure-bias gap, being closed by scheduled sampling + aligned decoding. Low TF loss does **not** mean the model can speak yet. See [Trusted Loss](#trusted-loss-reading-the-numbers).
133
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
134
  ### Adapted Chinchilla target
135
 
136
  Standard Chinchilla (20 tokens/param) assumes dense transformers where all parameters train on every token. Fractus violates that: only ~119M of 1.05B params are active per token, and the Kuramoto clock trains end-to-end now. The adapted target is **20 × active params ≈ 2.4–3.7B tokens**; with a warm-started checkpoint, the practical target is **~2.5–3B tokens** — which is why the phase-2 corpus is 3.44B. Since Fractus grows continuously (`maybe_grow`), this is an instantaneous cumulative-exposure requirement that grows with the model, not a final-stop condition. Details: [`docs/2026-08-12-fractus-chinchilla.md`](docs/2026-08-12-fractus-chinchilla.md).
@@ -343,6 +363,10 @@ engine.attach_memory(memory)
343
  engine.reset_thought(batch_size=1)
344
  logits, confidence = engine.tick(torch.tensor([42]))
345
  print(f"Confidence: {confidence.item():.2f}")
 
 
 
 
346
  ```
347
 
348
  ---
@@ -402,6 +426,8 @@ fractus-cte/
402
 
403
  **Kuramoto clock**: coupled oscillators producing phase vectors that route tokens to experts — learned end-to-end since the routing surgery.
404
 
 
 
405
  **Surgery**: patching code around a preserved checkpoint mid-training, resuming at the exact token offset. Weights are never wiped for a routing, objective, or decode bug.
406
 
407
  ---
 
111
  | **PhaseRoutedMoE** | Sparse experts routed by oscillator phases |
112
  | **Kuramoto Clock** | A dynamical system that drives routing decisions |
113
  | **Online Trainer** | Learns continuously, one chunk at a time |
114
+ | **Vision (`tick_vec`)** | Multimodal input path — image patches drive the engine directly, bypassing the token embedding (CIFAR "eyes" prototype trained while the text run continued) |
115
 
116
  ---
117
 
 
131
 
132
  **Honest generation state:** teacher-forced loss keeps falling (ema_tf ≈ 1.3–1.8 on the best GPUs), but free-run generation is still non-linguistic — diverse word-salad with repetition loops, better with window-based decoding (~25–28 unique tokens vs ~10–12 chunk-based). This is the documented exposure-bias gap, being closed by scheduled sampling + aligned decoding. Low TF loss does **not** mean the model can speak yet. See [Trusted Loss](#trusted-loss-reading-the-numbers).
133
 
134
+ ## Vision — the CIFAR "eyes" prototype
135
+
136
+ Fractus is not text-only. The engine exposes **`tick_vec(obs_vec)`**: a multimodal entry point that accepts a precomputed `(B, d_model)` vector and injects it directly into the residual thought state — **bypassing the token embedding entirely**. Any modality that can be embedded into `d_model` dimensions can drive continuous thought: images, audio, sensor streams.
137
+
138
+ ```
139
+ image → PatchEmbed → (B, d_model) patch vectors
140
+ │
141
+ ▼
142
+ engine.tick_vec(patch) ← no tokenizer involved
143
+ │
144
+ thought state advances through all 16 blocks
145
+ (same Kuramoto routing + MoE as text)
146
+ ```
147
+
148
+ **Proof of concept — CIFAR-10 eyes:** a small CTE + PatchEmbed stack (`fractus/nn/vision.py`) was trained on CPU on real CIFAR-10 images (`fractus_eyes_cifar_final.pt`), demonstrating that image patches can drive the continuous-thought loop. Crucially, this ran as a **parallel track**: the 8-GPU text run was never interrupted for eyes work — GPU text digestion and CPU vision learning proceed simultaneously on the same living system.
149
+
150
+ This mirrors the biological premise: eyes evolved as a peripheral sense feeding a central dynamical brain, not as a core feature of it. The text 1B is the brain; vision is a front-end that plugs into `tick_vec`.
151
+
152
+ ---
153
+
154
  ### Adapted Chinchilla target
155
 
156
  Standard Chinchilla (20 tokens/param) assumes dense transformers where all parameters train on every token. Fractus violates that: only ~119M of 1.05B params are active per token, and the Kuramoto clock trains end-to-end now. The adapted target is **20 × active params ≈ 2.4–3.7B tokens**; with a warm-started checkpoint, the practical target is **~2.5–3B tokens** — which is why the phase-2 corpus is 3.44B. Since Fractus grows continuously (`maybe_grow`), this is an instantaneous cumulative-exposure requirement that grows with the model, not a final-stop condition. Details: [`docs/2026-08-12-fractus-chinchilla.md`](docs/2026-08-12-fractus-chinchilla.md).
 
363
  engine.reset_thought(batch_size=1)
364
  logits, confidence = engine.tick(torch.tensor([42]))
365
  print(f"Confidence: {confidence.item():.2f}")
366
+
367
+ # Vision: drive thought with an image patch vector (no tokenizer)
368
+ patch_vec = torch.randn(1, 128) # any (B, d_model) embedding
369
+ logits, confidence = engine.tick_vec(patch_vec)
370
  ```
371
 
372
  ---
 
426
 
427
  **Kuramoto clock**: coupled oscillators producing phase vectors that route tokens to experts — learned end-to-end since the routing surgery.
428
 
429
+ **`tick_vec`**: the multimodal tick — feed any precomputed `(B, d_model)` vector (image patches, embeddings from any encoder) straight into the thought state without tokens. This is how vision plugs in.
430
+
431
  **Surgery**: patching code around a preserved checkpoint mid-training, resuming at the exact token offset. Weights are never wiped for a routing, objective, or decode bug.
432
 
433
  ---