thefinalboss commited on
Commit
9570b5e
·
verified ·
1 Parent(s): 64b00e4

Upload docs/HOW_FRACTUS_IS_TRAINED.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. docs/HOW_FRACTUS_IS_TRAINED.md +172 -0
docs/HOW_FRACTUS_IS_TRAINED.md ADDED
@@ -0,0 +1,172 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # How Fractus Is Trained
2
+
3
+ **Updated:** 2026-08-18 00:17 UTC
4
+ **Author:** Philippe-Antoine Robert
5
+ **Model repo:** https://huggingface.co/thefinalboss/fractus-cte
6
+ **Dataset:** https://huggingface.co/datasets/thefinalboss/fractus-datasets
7
+
8
+ This note explains the actual training procedure for Fractus-1B (Continuous Thought Engine): multi-GPU runs, corpus handling, mid-train surgery, and recovery after host failure.
9
+
10
+ ---
11
+
12
+ ## 1. What Fractus is (training-relevant)
13
+
14
+ Fractus is not a standard decoder-only transformer trained only with next-token CE on a frozen residual stream.
15
+
16
+ | Component | Role in training |
17
+ |-----------|------------------|
18
+ | Continuous thought state | Carries state across ticks |
19
+ | Kuramoto oscillators | Phase dynamics for temporal structure / routing |
20
+ | Phase-routed MoE | Sparse experts selected via phase/gates |
21
+ | Dense CE path (tick_chunk_train) | Teacher-forced sequence loss aligned with gen path |
22
+ | Scheduled sampling (SS) | Mix of ground-truth and model predictions as inputs |
23
+ | Load-balance loss | Keeps experts alive |
24
+ | Gate temperature | Controls routing softness |
25
+
26
+ Weights live in .pt checkpoints. Code lives in fractus/. Both are required to run.
27
+
28
+ ---
29
+
30
+ ## 2. Corpus (source of truth = HF dataset)
31
+
32
+ **Source of truth:** thefinalboss/fractus-datasets
33
+
34
+ A local full_corpus.pt (~4.23B) was historically a concatenation artifact built from this dataset. If that file is missing after a host crash, rebuild from the same HF dataset. The dataset is not lost.
35
+
36
+ ### Dataset layout
37
+
38
+ - neuro_paradigms_1b/ — neuroscience to architecture paradigms (jsonl.gz)
39
+ - cognitive_skills/ — skill / coding / reasoning JSONL
40
+ - neuro_code_math/ — math + code + applied neuroscience
41
+ - data/training_corpus.pt and datasets/*.pt — already-tokenized streams
42
+ - literature / esoteric / repos / identity subsets
43
+
44
+ ### Tokenized streams used in practice
45
+
46
+ | Stream | Approx size | Location |
47
+ |--------|-------------|----------|
48
+ | Phase-1 tokenized .pt union | ~1.52B tokens | dataset data/ + datasets/*.pt |
49
+ | Phase-2 full raw tokenize | 3.44B tokens | tokenized/phase2/shard_phase2_gpu0-7.npy |
50
+
51
+ Phase-2: stream all relevant JSONL/JSONL.GZ, GPT-2 BPE encode, write 8 equal int32 numpy memmap shards.
52
+
53
+ Anti re-ingest: phase-1 and phase-2 are separate streams. After phase-1 progress, switch to phase-2 rather than restarting the same ordered stream from token 0 when the goal is new data.
54
+
55
+ ---
56
+
57
+ ## 3. Multi-GPU training layout
58
+
59
+ Target: 8x RTX 5090 (recovery). Earlier: 4x.
60
+
61
+ | Setting | Typical value |
62
+ |---------|----------------|
63
+ | Processes | 1 Python process per GPU |
64
+ | CUDA_VISIBLE_DEVICES | equals GPU_ID |
65
+ | Batch | 2 or 3 (B=4 can OOM with SS) |
66
+ | Sequence length | 128 |
67
+ | LR | 7e-4 (SGD momentum 0.9) |
68
+ | SS_RATE | 0.25 |
69
+ | SS_PROB | 0.2 |
70
+ | LB_COEF | 0.02 |
71
+ | Gate temperature | 2.5 |
72
+ | TF32 + cudnn.benchmark | on |
73
+ | torch.compile | often off (VRAM) |
74
+ | Large shards | .npy memmap |
75
+
76
+ Script: scripts/fast4gpu_boost.py
77
+
78
+ Each GPU reads only its shard and writes checkpoints/fractus_1b_gpu{i}.pt
79
+
80
+ ### Launch example
81
+
82
+
83
+
84
+ Repeat for GPUs 1-7. Never launch multiple workers without unique CUDA_VISIBLE_DEVICES.
85
+
86
+ ---
87
+
88
+ ## 4. Loss signals
89
+
90
+ | Signal | Meaning |
91
+ |--------|---------|
92
+ | tf / ema_tf | Teacher-forced dense CE |
93
+ | ss / ema_ss | Loss under scheduled-sampling inputs |
94
+ | lb | Load-balance term |
95
+
96
+ Do not equate low TF loss with coherent free-run text. TF can be strong while greedy AR still mono-token collapses until SS + decode path close the train/gen gap.
97
+
98
+ See docs/LOSS_VS_GEN.md, docs/TRUSTED_LOSS.md, docs/GEN_PROBE_*.
99
+
100
+ ---
101
+
102
+ ## 5. Checkpointing and merge
103
+
104
+ - Per-GPU: fractus_1b_gpu0.pt ... gpu7.pt
105
+ - Mean-merge floating tensors across GPUs -> unified brain (e.g. FRACTUS_1B_PHASE2_LIVE_MERGED.pt)
106
+ - Hourly HF sync of 8 individuals + merged (Xet for binaries)
107
+ - RESUME_MANIFEST_8GPU.json stores per-GPU token offsets
108
+
109
+ Checkpoints can be merged, reloaded, continued. Mid-train edits are possible when careful.
110
+
111
+ ---
112
+
113
+ ## 6. Mid-train operability
114
+
115
+ Distinctive vs typical LLM pretrain:
116
+
117
+ - Probe experts/phases offline without destroying live checkpoint state
118
+ - Decode-path surgery (align tick_chunk vs single-step, anti-collapse)
119
+ - Merge parallel trained shards into one model
120
+ - Continue after host migration from HF weights
121
+
122
+ See COMPOSABILITY_AND_SURGERY.md, OPERABILITY_MIDTRAIN.md, DECODE_SURGERY.md.
123
+
124
+ ---
125
+
126
+ ## 7. Recovery playbook (host death)
127
+
128
+ 1. Treat HF as source of truth
129
+ 2. New pod + torch matching GPU arch (5090 needs recent CUDA builds)
130
+ 3. Code + dataset from HF
131
+ 4. Restore weights from checkpoints/fractus_1b_gpu*.pt or merged
132
+ 5. Restore shards from tokenized/phase2/*.npy or rebuild
133
+ 6. Resume START_TOKEN from RESUME_MANIFEST_8GPU.json
134
+ 7. Re-enable hourly Xet upload
135
+
136
+ Phase switch record: PHASE_SWITCH.json
137
+
138
+ ---
139
+
140
+ ## 8. Current production recipe (phase 2)
141
+
142
+ 1. Dataset fully present from HF
143
+ 2. Phase-2 tokenization done: 3,439,171,703 tokens -> 8x npy shards (~430M/GPU)
144
+ 3. Train 8-way BATCH=2, SS on, memmap shards
145
+ 4. Weights continued from phase-1 (not random init)
146
+ 5. Checkpoints + phase-2 shards + manifests on HF
147
+
148
+ Throughput ~900-1100 tok/s/GPU at B=2. One full phase-2 pass ~4-5 days wall-clock.
149
+
150
+ ---
151
+
152
+ ## 9. What done is not
153
+
154
+ Finishing tokens is not a finished model.
155
+
156
+ Progress criteria:
157
+
158
+ 1. Stable multi-GPU run
159
+ 2. TF loss trending down without NaNs
160
+ 3. SS loss not exploding vs TF
161
+ 4. Gen probes: rising uniqueness / less mono-token lock
162
+ 5. Checkpoints recoverable purely from HF
163
+
164
+ ---
165
+
166
+ ## 10. One-sentence summary
167
+
168
+ Fractus is trained as eight parallel continuous-thought engines on sharded token streams from the HF neuroscience-grounded dataset, optimized with dense teacher-forced CE plus scheduled sampling, checkpointed per GPU, mean-merged and uploaded hourly, and designed so training can be paused, surgically modified, merged, and resumed without treating the run as a single disposable monolith.
169
+
170
+ ---
171
+
172
+ Machine notes from the live 8x5090 recovery run. Update when the recipe changes.