thefinalboss commited on
Commit
dbc0639
·
verified ·
1 Parent(s): 1d36313

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +100 -71
README.md CHANGED
@@ -1,98 +1,127 @@
1
- # fractus-opt
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2
 
3
- Training-optimization workshop for
4
- [fractus-cte](https://huggingface.co/thefinalboss/fractus-cte)
5
- (Continuous Thought Engine — a dynamical system, not a transformer).
6
 
7
- **Mission (2026-08-22): make phase-2 training (8×RTX 5090) faster and more
8
- efficient — without ever breaking a living checkpoint.**
9
 
10
- ## The open-heart rule
 
 
 
 
 
 
11
 
12
- Fractus is operable "open-heart": organs of the code are replaced while the
13
- patient (the live training) keeps going. The rule is strict:
 
14
 
15
- 1. **Every change to the production path must be PROVEN mathematically
16
- equivalent to the reference** — forward, gradients, carried states
17
- (`tests/test_attention_equivalence.py`, 14 proofs).
18
- 2. The old implementation stays in the code as an executable reference
19
- (`_linear_attention_causal_einsum`).
20
- 3. The `.pt` checkpoints never change format; resumption happens at exact
21
- token offsets (`RESUME_MANIFEST`), zero progress thrown away.
22
- 4. Honest documented caveat: topk over von Mises gates is discrete — fp32
23
- rounding near a gate tie can flip the expert choice for isolated tokens
24
- (zero net effect). The engine tests bound this behavior instead of
25
- denying it.
26
 
27
- ## Contents
28
 
29
- | Path | Role |
30
  |---|---|
31
- | `fractus/nn/attention.py` | Causal linear-attention kernel: einsum (reference O(C²·dH²)), **cumsum** (O(C·dH²)), **chunked** (flat memory, block²+dH²). Dispatcher `_linear_attention_causal_vectorized` + env `FRACTUS_ATTN_IMPL`. |
32
- | `fractus/nn/ce.py` | `chunked_cross_entropy` (per-chunk checkpointing → logits never held for backward) + `sample_tokens_chunked` (memory-flat sampling for scheduled sampling). |
33
- | `fractus/continuous_engine.py` | + method `tick_chunk_train_ce(obs, targets, ce_chunk, return_hidden)` — same forward, CE without dense logits; + pure-core block checkpointing (`_tick_chunk_core_pure`, `BLOCK_CKPT`). Nothing removed. |
34
- | `scripts/fast4gpu_boost_v3.py` | Trainer v3 = consolidation of everything since v2: exact v1 semantics at ACCUM=1, zero-copy int32 memmap fetch (~27 GB RAM saved/pod), `CE_CHUNK`, `ACCUM`, `COMPILE`, `BLOCK_CKPT`, atomic checkpoint writes (tmp + rename), periodic resume manifest. |
35
- | `scripts/fast4gpu_boost_v2.py` | Trainer v2 (kept for provenance; v3 supersedes it). |
36
- | `scripts/pod_deploy_x8.sh` | One-shot detached pod deployer: env → assets from HF → smoke → launch → verify; stage markers make re-runs cheap. |
37
- | `tests/` | 46 tests: repo suite (28) + equivalences (14) + v2 loop smoke (2) + block-ckpt (2). **46/46 passing** (2026-08-24). |
38
- | `benchmarks/` | Attention micro-bench at real 1B shapes (subprocess-isolated cells) + end-to-end tok/s bench. |
39
- | `docs/OPTIMIZATION_2026-08-22.md` | **THE document**: bottleneck diagnosis, mathematical proofs, step-by-step pod deployment guide, measured results (CPU / 6×5060 Ti / production 8×5090). |
40
- | `docs/…` | Copies of the project documents (DISCOVERY_LOG, TRUSTED_LOSS, HOW_FRACTUS_IS_TRAINED, …). |
41
 
42
- ## Key results
43
 
44
- CPU local (torch 2.9.1+cpu) — attention micro-bench, real 1B shapes
45
- (G=B·40, C=128, dH=64), fwd+bwd:
46
 
47
- | Impl | B=2 | B=4 | B=8 |
48
- |---|---|---|---|
49
- | einsum (ref) | 572 ms | 1127 ms | crash (memory) |
50
- | cumsum | 799 ms | crash (memory) | crash |
51
- | **chunked** | **32 ms** | **72 ms** | **150 ms** |
52
 
53
- - chunked = **×15–25 vs reference**, flat memory ((G, block²+dH²)) where
54
- einsum/cumsum materialize (G, C, dH, dH) and explode.
55
- - End-to-end CPU: 703 → **764 tok/s** (+8.7 %).
56
- - Absolute CPU timings do not transfer to CUDA; relative ordering and memory
57
- profile do.
 
 
58
 
59
- GPU validation (6× RTX 5060 Ti 16 GB, torch 2.11+cu128):
 
 
 
 
 
 
60
 
61
- - chunked beats cumsum ×1.27 on GPU; full trainer sustained ~365 tok/s/GPU
62
- at B=3; without `BLOCK_CKPT` the 1B config does not fit in 16 GB.
63
 
64
- Production run (8× RTX 5090 32 GB, 2026-08-26, live):
65
 
66
- - Smoke B=8: **1,527 tok/s/GPU** @ 19.85 GB; sustained all-8:
67
- **~1,554–1,604 tok/s/GPU ≈ 12,600 aggregate**, lb = 14.028 stable.
68
- - Phase-2 finish projected in **≈ 3.1 days** (~3,420M tokens remaining).
69
 
70
- ## Usage
 
 
 
 
71
 
72
- ```bash
73
- # tests (no GPU required)
74
- py -m pytest tests/ -q # 46 passed
75
 
76
- # benchmarks
77
- py benchmarks/bench_attention.py --iters 8 # per-kernel table
78
- py benchmarks/bench_engine.py # end-to-end tok/s
79
 
80
- # trainer v3, one process per GPU (details: docs/OPTIMIZATION_2026-08-22.md §3)
81
- CUDA_VISIBLE_DEVICES=$i GPU_ID=$i START_TOKEN=<manifest offset> \
82
- BATCH=8 CE_CHUNK=2048 FRACTUS_ATTN_IMPL=chunked BLOCK_CKPT=1 \
 
 
 
83
  python -u scripts/fast4gpu_boost_v3.py
84
  ```
85
 
86
- Recommended escalation on a pod (validate ema_tf at every step):
87
- code swap with same settings → `CE_CHUNK=2048` → `FRACTUS_ATTN_IMPL=chunked`
88
- → `BLOCK_CKPT=1` → `COMPILE=1` + higher `BATCH`.
89
 
90
- ## Local Windows environment
91
 
92
- Python via the `py` launcher (torch 2.9.1+cpu); bare `python` = msys64
93
- without torch. cp1252 console: no non-ASCII characters in prints.
 
 
 
 
 
 
 
 
 
94
 
95
  ---
96
- *Work done 2026-08-22/23 with Claude Code; GPU validation 2026-08-24;
97
- production x8 launch 2026-08-26. Model and corpus:
98
- HF [thefinalboss](https://huggingface.co/thefinalboss).*
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: mit
3
+ language:
4
+ - en
5
+ - fr
6
+ tags:
7
+ - continuous-thought
8
+ - linear-attention
9
+ - kuramoto
10
+ - moe
11
+ - rnn
12
+ library_name: pytorch
13
+ ---
14
+
15
+ # Fractus-CTE
16
 
17
+ **Continuous Thought Engine.** A dynamical system with a fixed-size recurrent state, not a transformer.
 
 
18
 
19
+ Fractus belongs to the linear-attention RNN family (Katharopoulos 2020, RetNet, RWKV, GLA, DeltaNet, Mamba/S6) **plus** phase-routed MoE and a persistent thought state across chunks. Depth is time. The checkpoint is a living state (weights + phases + attention memory), not a frozen function.
 
20
 
21
+ | | Transformer | Fractus |
22
+ |---|---|---|
23
+ | Computation | one forward per prompt | continuous ticks / chunks |
24
+ | Memory | KV cache grows with length | fixed-size \(S, z\), thought state |
25
+ | Routing | (optional) learned softmax MoE | Kuramoto phases on a circle |
26
+ | Knowledge after train | fine-tune / RAG | **Vorax** organs, append-only `.kn` |
27
+ | Operability | replace the blob | open-heart: code changes, `.pt` shapes stay |
28
 
29
+ **Author:** Philippe-Antoine Robert ([thefinalboss](https://huggingface.co/thefinalboss))
30
+ **License:** MIT
31
+ **Sibling repos:** [fractus-p0](https://huggingface.co/thefinalboss/fractus-p0) (routing + v4 + DiffusionBlocks body) · [fractus-vorax](https://huggingface.co/thefinalboss/fractus-vorax) (sealed brain + ingest) · [fractus-datasets](https://huggingface.co/datasets/thefinalboss/fractus-datasets)
32
 
33
+ Full chronology: [`docs/EVOLUTION.md`](docs/EVOLUTION.md) · français [`docs/EVOLUTION.fr.md`](docs/EVOLUTION.fr.md)
34
+
35
+ ---
 
 
 
 
 
 
 
 
36
 
37
+ ## Architecture (1B production config)
38
 
39
+ | | |
40
  |---|---|
41
+ | Params | ~1.165B |
42
+ | `d_model` | 1280 |
43
+ | `n_layers` | 16 CTEBlocks |
44
+ | Attention | causal **linear** (cumsum / chunked). Env `FRACTUS_ATTN_IMPL` |
45
+ | Oscillators / block | 16, coupling rank 8 |
46
+ | Experts / block | 128, top-2, phase-gated |
47
+ | Vocab | 50257 GPT-2 BPE |
48
+ | Train objective | dense next-token CE on the chunk + Switch load-balance (+ optional SS / anti-repeat in v4) |
 
 
49
 
50
+ Each block: attention → Kuramoto → PhaseRoutedMoE. \(S,z\) and phases are **per-block** and carried across chunk boundaries.
51
 
52
+ ---
 
53
 
54
+ ## What is true right now (2026-08-27)
 
 
 
 
55
 
56
+ **Done**
57
+ - Multi-GPU data-parallel shards, exact `START_TOKEN` resume, hourly mean-merge of same-shape `.pt`
58
+ - Open-heart speed work (22–26 Aug): chunked linear attention, `chunked_cross_entropy`, `BLOCK_CKPT`, atomic ckpts — **shapes unchanged**
59
+ - Phase-2 corpus on HF; freeze then x8 resume (~35M tokens/GPU on ~430M-token shards at last pushed manifest)
60
+ - P0 routing body (atan2 encode, phase carry, per-token MoE, Switch LB) in [fractus-p0](https://huggingface.co/thefinalboss/fractus-p0)
61
+ - PREFIX vs CARRY decode hole **named and measured** on a CPU mini
62
+ - Fractus-native DiffusionBlocks prototype (one CTEBlock / step, 0 new params) — experiment, not default
63
 
64
+ **Open (do not skip)**
65
+ - Attention state \(S\) still looks like \(S_t = S_{t-1} + k\otimes v\) — **no learned decay** yet (RWKV/RetNet/Mamba all added one)
66
+ - Mean-merge of independent MoE shards ≠ DDP; expert #k is not aligned across GPUs
67
+ - Kuramoto decision is still a **circle** (need \(\mathrm{MI}(\bar\theta, \text{token})\) vs tick)
68
+ - No published **matched-compute PPL** vs a vanilla transformer on a clean held-out
69
+ - CARRY length-1 generation is not the trained graph (PREFIX is the language gate)
70
+ - Eval must be split by **source**, not by line (identity text in the corpus)
71
 
72
+ Default next run: **v4 + P0 body + resume offsets + 24GB+ identical GPUs (x4 is enough)**.
73
+ DiffusionBlocks and DDP are options. Finishing the pass without a baseline PPL is not a paper.
74
 
75
+ ---
76
 
77
+ ## Load
 
 
78
 
79
+ ```python
80
+ from fractus.continuous_engine import ContinuousThoughtEngine
81
+ eng = ContinuousThoughtEngine.from_pretrained("checkpoints/FRACTUS_1B_PHASE2_FROZEN_MERGED.pt")
82
+ # .pt = weights + live state. fractus/ = the body that ticks.
83
+ ```
84
 
85
+ Same-shape checkpoints can be mean-merged. Different `d_model` / layer / expert counts **cannot**. See `docs/CPU_MINI_MERGE_AND_DIMENSIONS.md`.
 
 
86
 
87
+ ---
 
 
88
 
89
+ ## Train (resume, never from token 0 unless you mean it)
90
+
91
+ ```bash
92
+ # offsets: checkpoints/X8_MANIFEST.json or FROZEN_RESUME_MANIFEST.json
93
+ CUDA_VISIBLE_DEVICES=0 GPU_ID=0 START_TOKEN=<manifest> \
94
+ BATCH=2 CE_CHUNK=2048 FRACTUS_ATTN_IMPL=chunked BLOCK_CKPT=1 \
95
  python -u scripts/fast4gpu_boost_v3.py
96
  ```
97
 
98
+ P0 + v4 launcher lives in [fractus-p0](https://huggingface.co/thefinalboss/fractus-p0) (`scripts/fast4gpu_boost_v4.py`).
99
+ PREFIX `unique@40` is the language gate. Do not raise `REPEAT_COEF` above 0.1.
 
100
 
101
+ ---
102
 
103
+ ## Docs worth reading first
104
+
105
+ | Doc | Why |
106
+ |---|---|
107
+ | `docs/EVOLUTION.md` | timestamped history |
108
+ | `docs/HOW_FRACTUS_IS_TRAINED.md` | phase-2 recipe |
109
+ | `docs/KURAMOTO_BOTTLENECK_AND_FIX.md` | dead experts / 25° arc |
110
+ | `docs/OPTIMIZATION_2026-08-22.md` | open-heart kernels |
111
+ | `docs/TRAIN_GEN_MISMATCH.md` / p0 `V4_AR.md` | PREFIX vs CARRY |
112
+ | `docs/TRUSTED_LOSS.md` | loss ≠ generation |
113
+ | `Fractus_White_Paper_v2.md` | architecture write-up |
114
 
115
  ---
116
+
117
+ ## Citation
118
+
119
+ ```
120
+ @misc{robert2026fractus,
121
+ title={Fractus: a Continuous Thought Engine},
122
+ author={Robert, Philippe-Antoine},
123
+ year={2026},
124
+ howpublished={https://huggingface.co/thefinalboss/fractus-cte},
125
+ note={MIT. Dynamical system; linear-attention RNN lineage + phase-routed MoE.}
126
+ }
127
+ ```