Upload docs/HOW_FRACTUS_IS_TRAINED.md with huggingface_hub
Browse files- docs/HOW_FRACTUS_IS_TRAINED.md +172 -0
docs/HOW_FRACTUS_IS_TRAINED.md
ADDED
|
@@ -0,0 +1,172 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# How Fractus Is Trained
|
| 2 |
+
|
| 3 |
+
**Updated:** 2026-08-18 00:17 UTC
|
| 4 |
+
**Author:** Philippe-Antoine Robert
|
| 5 |
+
**Model repo:** https://huggingface.co/thefinalboss/fractus-cte
|
| 6 |
+
**Dataset:** https://huggingface.co/datasets/thefinalboss/fractus-datasets
|
| 7 |
+
|
| 8 |
+
This note explains the actual training procedure for Fractus-1B (Continuous Thought Engine): multi-GPU runs, corpus handling, mid-train surgery, and recovery after host failure.
|
| 9 |
+
|
| 10 |
+
---
|
| 11 |
+
|
| 12 |
+
## 1. What Fractus is (training-relevant)
|
| 13 |
+
|
| 14 |
+
Fractus is not a standard decoder-only transformer trained only with next-token CE on a frozen residual stream.
|
| 15 |
+
|
| 16 |
+
| Component | Role in training |
|
| 17 |
+
|-----------|------------------|
|
| 18 |
+
| Continuous thought state | Carries state across ticks |
|
| 19 |
+
| Kuramoto oscillators | Phase dynamics for temporal structure / routing |
|
| 20 |
+
| Phase-routed MoE | Sparse experts selected via phase/gates |
|
| 21 |
+
| Dense CE path (tick_chunk_train) | Teacher-forced sequence loss aligned with gen path |
|
| 22 |
+
| Scheduled sampling (SS) | Mix of ground-truth and model predictions as inputs |
|
| 23 |
+
| Load-balance loss | Keeps experts alive |
|
| 24 |
+
| Gate temperature | Controls routing softness |
|
| 25 |
+
|
| 26 |
+
Weights live in .pt checkpoints. Code lives in fractus/. Both are required to run.
|
| 27 |
+
|
| 28 |
+
---
|
| 29 |
+
|
| 30 |
+
## 2. Corpus (source of truth = HF dataset)
|
| 31 |
+
|
| 32 |
+
**Source of truth:** thefinalboss/fractus-datasets
|
| 33 |
+
|
| 34 |
+
A local full_corpus.pt (~4.23B) was historically a concatenation artifact built from this dataset. If that file is missing after a host crash, rebuild from the same HF dataset. The dataset is not lost.
|
| 35 |
+
|
| 36 |
+
### Dataset layout
|
| 37 |
+
|
| 38 |
+
- neuro_paradigms_1b/ — neuroscience to architecture paradigms (jsonl.gz)
|
| 39 |
+
- cognitive_skills/ — skill / coding / reasoning JSONL
|
| 40 |
+
- neuro_code_math/ — math + code + applied neuroscience
|
| 41 |
+
- data/training_corpus.pt and datasets/*.pt — already-tokenized streams
|
| 42 |
+
- literature / esoteric / repos / identity subsets
|
| 43 |
+
|
| 44 |
+
### Tokenized streams used in practice
|
| 45 |
+
|
| 46 |
+
| Stream | Approx size | Location |
|
| 47 |
+
|--------|-------------|----------|
|
| 48 |
+
| Phase-1 tokenized .pt union | ~1.52B tokens | dataset data/ + datasets/*.pt |
|
| 49 |
+
| Phase-2 full raw tokenize | 3.44B tokens | tokenized/phase2/shard_phase2_gpu0-7.npy |
|
| 50 |
+
|
| 51 |
+
Phase-2: stream all relevant JSONL/JSONL.GZ, GPT-2 BPE encode, write 8 equal int32 numpy memmap shards.
|
| 52 |
+
|
| 53 |
+
Anti re-ingest: phase-1 and phase-2 are separate streams. After phase-1 progress, switch to phase-2 rather than restarting the same ordered stream from token 0 when the goal is new data.
|
| 54 |
+
|
| 55 |
+
---
|
| 56 |
+
|
| 57 |
+
## 3. Multi-GPU training layout
|
| 58 |
+
|
| 59 |
+
Target: 8x RTX 5090 (recovery). Earlier: 4x.
|
| 60 |
+
|
| 61 |
+
| Setting | Typical value |
|
| 62 |
+
|---------|----------------|
|
| 63 |
+
| Processes | 1 Python process per GPU |
|
| 64 |
+
| CUDA_VISIBLE_DEVICES | equals GPU_ID |
|
| 65 |
+
| Batch | 2 or 3 (B=4 can OOM with SS) |
|
| 66 |
+
| Sequence length | 128 |
|
| 67 |
+
| LR | 7e-4 (SGD momentum 0.9) |
|
| 68 |
+
| SS_RATE | 0.25 |
|
| 69 |
+
| SS_PROB | 0.2 |
|
| 70 |
+
| LB_COEF | 0.02 |
|
| 71 |
+
| Gate temperature | 2.5 |
|
| 72 |
+
| TF32 + cudnn.benchmark | on |
|
| 73 |
+
| torch.compile | often off (VRAM) |
|
| 74 |
+
| Large shards | .npy memmap |
|
| 75 |
+
|
| 76 |
+
Script: scripts/fast4gpu_boost.py
|
| 77 |
+
|
| 78 |
+
Each GPU reads only its shard and writes checkpoints/fractus_1b_gpu{i}.pt
|
| 79 |
+
|
| 80 |
+
### Launch example
|
| 81 |
+
|
| 82 |
+
|
| 83 |
+
|
| 84 |
+
Repeat for GPUs 1-7. Never launch multiple workers without unique CUDA_VISIBLE_DEVICES.
|
| 85 |
+
|
| 86 |
+
---
|
| 87 |
+
|
| 88 |
+
## 4. Loss signals
|
| 89 |
+
|
| 90 |
+
| Signal | Meaning |
|
| 91 |
+
|--------|---------|
|
| 92 |
+
| tf / ema_tf | Teacher-forced dense CE |
|
| 93 |
+
| ss / ema_ss | Loss under scheduled-sampling inputs |
|
| 94 |
+
| lb | Load-balance term |
|
| 95 |
+
|
| 96 |
+
Do not equate low TF loss with coherent free-run text. TF can be strong while greedy AR still mono-token collapses until SS + decode path close the train/gen gap.
|
| 97 |
+
|
| 98 |
+
See docs/LOSS_VS_GEN.md, docs/TRUSTED_LOSS.md, docs/GEN_PROBE_*.
|
| 99 |
+
|
| 100 |
+
---
|
| 101 |
+
|
| 102 |
+
## 5. Checkpointing and merge
|
| 103 |
+
|
| 104 |
+
- Per-GPU: fractus_1b_gpu0.pt ... gpu7.pt
|
| 105 |
+
- Mean-merge floating tensors across GPUs -> unified brain (e.g. FRACTUS_1B_PHASE2_LIVE_MERGED.pt)
|
| 106 |
+
- Hourly HF sync of 8 individuals + merged (Xet for binaries)
|
| 107 |
+
- RESUME_MANIFEST_8GPU.json stores per-GPU token offsets
|
| 108 |
+
|
| 109 |
+
Checkpoints can be merged, reloaded, continued. Mid-train edits are possible when careful.
|
| 110 |
+
|
| 111 |
+
---
|
| 112 |
+
|
| 113 |
+
## 6. Mid-train operability
|
| 114 |
+
|
| 115 |
+
Distinctive vs typical LLM pretrain:
|
| 116 |
+
|
| 117 |
+
- Probe experts/phases offline without destroying live checkpoint state
|
| 118 |
+
- Decode-path surgery (align tick_chunk vs single-step, anti-collapse)
|
| 119 |
+
- Merge parallel trained shards into one model
|
| 120 |
+
- Continue after host migration from HF weights
|
| 121 |
+
|
| 122 |
+
See COMPOSABILITY_AND_SURGERY.md, OPERABILITY_MIDTRAIN.md, DECODE_SURGERY.md.
|
| 123 |
+
|
| 124 |
+
---
|
| 125 |
+
|
| 126 |
+
## 7. Recovery playbook (host death)
|
| 127 |
+
|
| 128 |
+
1. Treat HF as source of truth
|
| 129 |
+
2. New pod + torch matching GPU arch (5090 needs recent CUDA builds)
|
| 130 |
+
3. Code + dataset from HF
|
| 131 |
+
4. Restore weights from checkpoints/fractus_1b_gpu*.pt or merged
|
| 132 |
+
5. Restore shards from tokenized/phase2/*.npy or rebuild
|
| 133 |
+
6. Resume START_TOKEN from RESUME_MANIFEST_8GPU.json
|
| 134 |
+
7. Re-enable hourly Xet upload
|
| 135 |
+
|
| 136 |
+
Phase switch record: PHASE_SWITCH.json
|
| 137 |
+
|
| 138 |
+
---
|
| 139 |
+
|
| 140 |
+
## 8. Current production recipe (phase 2)
|
| 141 |
+
|
| 142 |
+
1. Dataset fully present from HF
|
| 143 |
+
2. Phase-2 tokenization done: 3,439,171,703 tokens -> 8x npy shards (~430M/GPU)
|
| 144 |
+
3. Train 8-way BATCH=2, SS on, memmap shards
|
| 145 |
+
4. Weights continued from phase-1 (not random init)
|
| 146 |
+
5. Checkpoints + phase-2 shards + manifests on HF
|
| 147 |
+
|
| 148 |
+
Throughput ~900-1100 tok/s/GPU at B=2. One full phase-2 pass ~4-5 days wall-clock.
|
| 149 |
+
|
| 150 |
+
---
|
| 151 |
+
|
| 152 |
+
## 9. What done is not
|
| 153 |
+
|
| 154 |
+
Finishing tokens is not a finished model.
|
| 155 |
+
|
| 156 |
+
Progress criteria:
|
| 157 |
+
|
| 158 |
+
1. Stable multi-GPU run
|
| 159 |
+
2. TF loss trending down without NaNs
|
| 160 |
+
3. SS loss not exploding vs TF
|
| 161 |
+
4. Gen probes: rising uniqueness / less mono-token lock
|
| 162 |
+
5. Checkpoints recoverable purely from HF
|
| 163 |
+
|
| 164 |
+
---
|
| 165 |
+
|
| 166 |
+
## 10. One-sentence summary
|
| 167 |
+
|
| 168 |
+
Fractus is trained as eight parallel continuous-thought engines on sharded token streams from the HF neuroscience-grounded dataset, optimized with dense teacher-forced CE plus scheduled sampling, checkpointed per GPU, mean-merged and uploaded hourly, and designed so training can be paused, surgically modified, merged, and resumed without treating the run as a single disposable monolith.
|
| 169 |
+
|
| 170 |
+
---
|
| 171 |
+
|
| 172 |
+
Machine notes from the live 8x5090 recovery run. Update when the recipe changes.
|