fractus-cte / docs /MASTER_RUN_LOG.md
Philippe-Antoine Robert
docs: X8 unified checkpoint β€” sandbox live operation (2026-09-02)
5326f88
|
Raw History Blame Contribute Delete
6.04 kB

Fractus-1B Production Run Log (Complete)

Updated: 2026-09-02 23:55 UTC

English master log of the Fractus continuous-thought 1B training run: discoveries, bugs, surgeries, metrics, and generation probes.


1. Architecture

Fractus is not a standard decoder-only transformer training loop.

  • ContinuousThoughtEngine (CTE): residual thought state across ticks
  • Per-block linear attention with continuous carry (S, z)
  • Kuramoto phase oscillators for routing dynamics
  • PhaseRoutedMoE (von Mises gates, top-k experts, load-balance loss)
  • Tied embedding / output head
  • Target: d_model=1280, 16 layers, 128 experts, top-k=2

Checkpoints store weights + dynamic state; package fractus/ is the body.


2. Hardware and data

  • 4x NVIDIA RTX 5090
  • Independent shard per GPU (~1.057B tokens each, GPT-2 BPE)
  • Batch B=2, seq_len=128 (256 tokens/step)
  • Optimizer: SGD momentum 0.9
  • Precision: bfloat16 autocast

3. Chronology of interventions

Phase A β€” Initial distributed training

  • Four independent GPU runs
  • Early objective: last-position CE only
  • Stable throughput, descending loss
  • Generation collapsed to single-token loops

Phase B β€” Routing surgery (weights kept)

Bugs found:

  1. Kuramoto under torch.no_grad in tick_chunk_core (omega never trained)
  2. lb_loss detached and never in training loss (dead experts)
  3. Hard gate temperature

Fixes: Kuramoto gradients, CE+0.02*lb, gate temp 2.5, omega diversity, exact token resume

Phase C β€” Stage2 dense CE

  • tick_chunk_train returns logits for all positions
  • Dense next-token CE over full chunk
  • CE dropped hard (GPU1 into ~2 range)

Phase D β€” Decode surgery (weights unchanged)

  • Phase/thought noise, frequency penalty, cycle bans, forced escape tokens
  • Broke single-token decode lock
  • Did not produce coherent English by itself

Phase E β€” Train/gen path mismatch

  • Train: tick_chunk_core (causal attn + RK4 Kuramoto)
  • Default gen: tick_single (simpler attn + Euler Kuramoto)
  • Fix: align tick_single to RK4; prefer 100% tick_chunk generation
  • Module: fractus/generate_aligned.py (generate_chunk, generate_window)

Phase F β€” Fine phase loss recalibration

  • Cumulative average CE misleading near 2.0
  • Batch CE + EMA; real session tok/s
  • LR 1e-3 -> 5e-4

Phase G β€” Scheduled sampling (loss vs gen gap)

  • Warm TF CE ~1.3-1.5 can coexist with broken free-run generation
  • Free-run AR CE vs ground truth is catastrophic (exposure bias)
  • Fix: sequential scheduled sampling
    • Pass A: teacher-forced CE + LB
    • Pass B (~30% steps): mix ~20% model samples into inputs
  • Logs: tf/ema_tf and ss/ema_ss

4. Live SS metrics (at doc time)

  • GPU 0: GPU 0: 190,809,600 tf=7.760 ema_tf=7.689 lb=14.027 670 tok/s [ss]
  • GPU 1: GPU 1: 200,499,200 tf=1.292 ema_tf=1.334 lb=14.027 670 tok/s [ss]
  • GPU 2: GPU 2: 201,446,400 tf=1.766 ema_tf=1.789 lb=14.026 656 tok/s [ss]
  • GPU 3: GPU 3: 198,656,000 tf=5.662 ema_tf=5.998 lb=14.027 670 tok/s [ss]

5. Generation summary

Stage Behavior
Early mid-train Single-token loops
After decode surgery Multi-token diversity, non-sentences
After path alignment WINDOW more diverse than CHUNK
Current Lexical noise / short cycles; not coherent prose

Expected while SS is still teaching free-run and only a fraction of each shard is consumed.


6. Operability principle

  1. Keep .pt weights
  2. Patch body (engine / trainer / decode)
  3. Resume at recorded token offset
  4. Document the surgery

Do not discard multi-day digestion for routing, objective, or decode bugs.


7. Key files

  • fractus/continuous_engine.py β€” CTE body
  • fractus/generate_aligned.py β€” train-aligned decode
  • fractus/decode_surgery.py β€” optional decode stack
  • scripts/fast4gpu_stage2_ss.py β€” current trainer
  • scripts/fast4gpu_stage2_fine.py β€” fine phase trainer
  • RESUME_MANIFEST_*.json β€” exact offsets
  • checkpoints/fractus_1b_gpu0-3.pt β€” per-GPU brains

8. Related docs

  • docs/COMPOSABILITY_AND_SURGERY.md β€” in-training surgery + multi-pt merge/grow

  • docs/DISCOVERY_LOG.md

  • docs/OPERABILITY_MIDTRAIN.md

  • docs/DECODE_SURGERY.md

  • docs/TRAIN_GEN_MISMATCH.md

  • docs/GENERATE_ALIGNED.md

  • docs/FINE_PHASE.md

  • docs/LOSS_VS_GEN.md

  • docs/TRUSTED_LOSS.md β€” which loss to trust for train vs gen

  • docs/STAGE2_SS_GEN_PROBE.md (generation probe with outputs)


9. Conclusion

  1. Learning is real under teacher-forced dense CE.
  2. Generation coherence is not yet achieved.
  3. Low TF loss does not imply clean free-run text.
  4. Active remediation: scheduled sampling + aligned tick_chunk decode.
  5. Keep rolling; re-probe generation after more SS tokens.

10. 2026-08-28 23:35 UTC β€” Decode surgery I (window + anti-copy)

  • Space showed RetailΓ—32 / unique@32=1 on greedy length-1 tick_chunk + carry.
  • Replaced that path in fractus/generate_aligned.py: sliding causal window 64 + mask previous token.
  • Weights / trainers unchanged. QuickPod still on boost_v4 ~89M/430M GPU0.
  • Post-op unique@32: 16 / 14 / 27 / 12 vs 1 / 1 / 1 / 2.
  • Short cycles remain. Speech not claimed.
  • Full note: docs/2026-08-28-DECODE-WINDOW.md

11. 2026-09-02 22:56–23:40 UTC β€” X8 unified checkpoint, sandbox live operation

  • 8 QuickPod brains merged (mean_float_params) β†’ FRACTUS_1B_X8_MERGED.pt
  • Full README gate protocol reproduced in an external CPU sandbox
  • Live operation on the same weights: greedy carry collapses to single-token
  • Continuous state without forgetting shrinks: vocabulary overlap 91 % β†’ 100 %
  • Kuramoto order r β‰ˆ 0.003 everywhere; decode-surgery phase noise is a no-op
  • Gate v2 proposed: greedy unique@40 (formal, unchanged) + sampled unique@48 (the "body" β€” GO today at 32.86) + continuous-state stability (overlap < 0.7, unique β‰₯ 20 over 3 segments).
  • Full journal + repro: docs/X8_CIEL_OUVERT.md Β· results: docs/X8_UNIFIED_PROBE.json, docs/OPERATION_CIEL_OUVERT_RESULTS.json Β· harness: scripts/probe_ciel_ouvert.py.