fractus-cte / docs /DISCOVERY_LOG.md
Philippe-Antoine Robert
docs: X8 unified checkpoint — sandbox live operation (2026-09-02)
5326f88
|
Raw History Blame Contribute Delete
9.93 kB

Fractus Discovery Log — Bugs, Optimizations, Emergent Features

Updated: 2026-09-02 (adds sandbox live-operation discoveries on the X8 unified checkpoint)

This document records what we found while running Fractus-1B — things that go beyond the original design notes. Empirical, not marketing.


1. Critical training bugs found (and fixed)

1.1 Kuramoto was not learning during training

Symptom: omega stayed near init (±0.05), phase order parameter r ≈ 0.01–0.03, soft dynamics.

Root cause: In CTEBlock.tick_chunk_core, Kuramoto integration ran under torch.no_grad():

Comment in code even said "clock, not learned". So CE loss never reached omega / coupling.

Fix: Remove no_grad around Kuramoto; keep phase state detached for carry, keep differentiable theta for MoE routing so parameters receive gradients.

Status: Fixed in fractus/continuous_engine.py (surgery 2026-08-16).

1.2 Load-balance loss was computed then thrown away

Symptom: Offline probe showed ~70% experts dead (90+/128), only 2–3 experts active per block.

Root causes:

  1. tick_chunk_train did total_lb + lb.detach() — no gradient through LB
  2. fast4gpu.py optimized only cross-entropy — never added lb_loss to the loss

Fix:

  • Keep LB in the graph
  • tick_chunk_train returns (logits, lb_loss)
  • Surgery trainer: loss = CE + 0.02 * lb

Status: Fixed; live lb≈14 on all GPUs after resume.

1.3 Probe false alarm: "all phases are zero"

Symptom: A probe script reported all Kuramoto phases at 0.0.

Root cause: Script called reset_thought() before measuring — which zeros state buffers.

Reality in checkpoints: phases nonzero, std ≈ 1.8 across all 16 blocks × 4 GPUs.

Lesson: Never diagnose dynamical state after an intentional reset.


2. Optimizations discovered in production

Optimization What we learned
4-GPU independent shards + mean-merge Works; unified .pt generates; different lexical attractors than single shard
Resume from exact token offset Manifest-driven start_token preserves progress across pod reboot
Gate temperature ↑ (1.0 → 2.5) Softens von Mises routing; more experts can enter the top-k mix
Omega scale ×4 on resume Restores phase-rate diversity without wiping weights
tick_vec multimodal path Vision patches can drive CTE without touching token embedding
CPU eyes prototype (CIFAR) Small CTE+PatchEmbed learns real images offline while 1B trains on GPU

3. Features that emerged beyond the original plan

3.1 Operable / open-heart model

Weights (.pt) + body (fractus/ code) are separable. We can:

  • merge brains
  • change routing temperature
  • inject LB pressure
  • add vision front-end without a full retrain from zero.

3.2 Infinite-ish checkpoint fusion (same architecture)

Compatible checkpoints can be mean-merged and trained again:

train → merge → train → merge → ...

Constraint: same shapes (d_model, layers, experts, etc.). Divergent merges can soup skills; shard-merge is the proven path.

3.3 Mid-training generation behavior

At loss ~33→25, generation pipeline works but outputs word-level repetition collapse (real tokens stuck in loops: "Colorado", "Population", "Fate", "thinks"…). Documents that Fractus emits lexical tokens before coherent sentences.

3.4 Parallel modality track

Text 1B digestion and vision prototype can run in parallel (GPU text + CPU eyes) without stopping the main run.

3.5 Routing pathology as first-class debug target

Expert-hit histograms + phase order parameter r are necessary metrics. Loss alone hides "model learns with 3 experts".


4. Live metrics after routing surgery (resume)

Resume offsets preserved from pre-crash run (~176–187M tokens/GPU).

GPU Resume start Snapshot tokens CE loss lb
0 176,281,600 176,614,400 36.5 14.026
1 186,137,600 186,444,800 12.4 14.024
2 186,854,400 187,212,800 34.7 14.026
3 183,833,600 184,166,400 41.0 14.027

Post-surgery CE can spike briefly then fall (routing distribution shift). GPU1 recovered fastest into the teens.


5. What this means for the Fractus thesis

Fractus is not only "another 1B trained on shards". The run forced discovery of:

  1. Silent non-learning of the phase clock under no_grad
  2. Silent expert death without LB in the loss
  3. Composable checkpoints as a workflow
  4. Operability (surgery without discarding digestion)
  5. A body that speaks before the brain converges — decode-level surgery on the same weights turns 2.14 unique/48 (greedy carry) into 32.86 unique/48 fluid, repeat-free speech. The collapse is a sampling-regime attractor, not a weights defect: the thesis holds above the training gate (see §7).

The architecture does more than the first README described — because production training exposed the dynamical bottlenecks.


6. Discoveries from the optimization session (2026-08-22/23)

6.1 The "attention matmul" was a disguised cumsum

The production kernel computed causal sums via einsum("tj,bjpq->btpq", tril_mask, outer) — an O(C²·dH²) masked contraction that IS mathematically a prefix sum. Replaced by torch.cumsum(outer, dim=1) (O(C·dH²)), proven bit-close to the kept einsum reference (forward, gradients, carried S/z). Lesson: profile semantics, not just FLOPs — the mask+matmul form hid a linear-scan structure for months.

6.2 Discrete routing makes strict equivalence claims wrong at block level

topk over von Mises gates is DISCRETE: a float32 rounding difference near a gate tie flips expert choice for isolated tokens (measure-zero boundary), and depth amplifies it. Honest framing adopted: continuous stages are proven bit-close; full-block equality is asserted statistically (median at rounding scale, mismatch fraction bounded). Tests encode exactly this.

6.3 Two constructions under one seed have DIFFERENT weights

manual_seed(s); A = Model(); B = Model() gives two different networks — the RNG stream continues across constructions. Any A/B equivalence test must clone weights (load_state_dict) instead of re-seeding. This masqueraded as a "routing divergence" for half a session before being identified.

6.4 CPU commit-limit segfaults are data, not noise

At 1B shapes the reference kernels materialize (G, C, dH, dH) and hit the native commit limit on a RAM-constrained host (segfault, not MemoryError). The benchmark harness now runs each cell in its own subprocess so one crash documents itself without killing the table — and the crash boundary itself measures the memory win of the chunked kernel.


7. Sandbox live operation on the X8 unified checkpoint (2026-09-02)

External CPU sandbox (fp32) loaded FRACTUS_1B_X8_MERGED.pt and ran the full README gate protocol + a live speech operation on the continuous state. Empirical findings:

7.1 The generation collapse is a greedy-regime attractor, not a distribution property

Symptom: greedy carry (native continuous mode) = 48 consecutive identical tokens (2.14 unique/48). Same weights, top-150 sampling @ T 1.15 with ban 20: 32.86 unique/48, zero repeats, no 2-cycles.

Root cause: von Mises soft gates (architecture) + anti-copy/LB training spread mass across the top-150 support; the argmax landscape still holds short repetition attractors. The gate's greedy unique@40 measures the most pathological regime of the model.

Status: Measured. Action: add a second birth metric — sampled unique@48 (top-k 150, T 1.15, ban 20, continuous state). Gate v2 = greedy (formal) + sampled (the "body") + state stability. See docs/X8_CIEL_OUVERT.md §5.

7.2 The continuous (S,z) state is a memory of degradation, not a memory

Symptom: in one unreset thought stream, vocabulary overlap with earlier segments: 91 % (segment 2), 100 % (segment 3). 192-token run = a ~12-token motif macro-cycle (ban 20 breaks <20 loops, not 20+ motifs).

Root cause: attn_S += outer.detach() is unbounded, no decay — the readout q·S / q·z becomes a function of the whole history and the effective vocabulary shrinks with state age. A multiplicative leak on (S,z) (λ 0.99, half-life ~69 ticks) makes continuity sustainable but does not invert the shrinkage.

Status: Measured. Action: active forgetting (selective/importance-weighted decay, or normalize S/z to code a mean). Continuity metric: overlap < 0.7 AND unique ≥ 20 over 3+ segments.

7.3 Decode-surgery phase noise is a no-op on the tick path

Symptom: r ≈ 0.0028 in every configuration; perturbing blk.kuramoto_phases never changes outputs.

Root cause: CTEBlock.tick_single computes a fresh θ from the hidden state each tick and feeds it to the MoE; the stored kuramoto_phases buffer is read only by the expert-hit counter. decode_surgery.generate_with_surgery mutates the dead buffer.

Status: Verified in code. Action: apply phase noise to the θ the MoE actually reads (one line) or delete the dead code. Kuramoto order must be raised in training (LB under gradient), not in decode.

7.4 Publish gap: the load fix is not wired into the public path

Symptom: a fresh load of the public checkpoint runs at moe.temperature = 1.0 (code default) — the RAW config — while QuickPod runs at 2.5 after apply_kuramoto_routing_fix. Nothing in the public code calls the fix; memory.py docstring promises an engine.inject_memory() that does not exist.

Status: Measured (probe RAW = 1.0 vs FIX = 2.5). Action: README line — "apply apply_kuramoto_routing_fix at load"; implement or de-document inject_memory.



Keep this log updated when new probes, merges, or surgeries land.