# Training-data expansion, 2026-09-17 This document preserves the expansion experiment and its verification. The campaign has completed; see the [current handover](../operations/handover.md) and [final results](../../RESULTS.md). The dated launch and control details below remain as provenance. The user requested a broader training set. Version 2 adds official training data from HellaSwag, PIQA and CommonsenseQA to the complete version-1 replay set. These sources add plausible continuations, physical problem solving and everyday commonsense questions with two to five answer choices. They train the existing decision scorer, not a text-generation model. Data directory: `/home/andy/ai/opensysone/data/public-decisions-v2-20260917`. Build with `scripts/prepare_expanded_data.py`; raw downloads, transformed rows, weights and token caches stay outside Git. Source revisions, licenses, file hashes, sampling decisions and exclusions are recorded in the dataset manifest and `results/20260917-expanded-data/dataset-manifest.json`. | Source | Training decisions after 512-token filtering | License | | --- | ---: | --- | | HellaSwag | 16,000 | MIT | | PIQA | 14,360 | AFL-3.0, creator's README | | CommonsenseQA | 9,490 | MIT | | Original replay | 40,915 | Original per-source provenance | | Total | 80,765 | Mixed; no blanket relicensing | Only labeled official training rows are added. Exclusions use normalized source groups and exact normalized context text, including official evaluation metadata. Duplicate choices and invalid labels are rejected. Original training bytes form an unchanged prefix; the existing deterministic per-epoch shuffle samples the combined rows uniformly. Model-specific token filtering rejects oversize examples without truncation. Exact matching does not exclude semantic duplicates or contamination already present in the pretrained model. CPU tokenization completed in 60.96 seconds without loading model weights or initializing CUDA. Of 80,792 raw rows, 27 exceed the limit (26 original BoolQ and one new PIQA); the longest retained branch is 509 tokens. Original tokenized training replay and all four protected tokenized splits compare exactly with the verified version-1 cache. Data signature: `76183c642668602f42b7f3e71a3fe03bd5bd76f064fce4ba92351d8703396207`. The original validation, calibration, test and Social IQA holdout files remain byte-for-byte identical. Another 128 source groups per new family are withheld as separate diagnostics, excluded from training and from checkpoint selection. After length filtering these contain 383 decisions (one PIQA item is too long). The existing fixed 512-decision selection therefore measures retention on the original four tasks; gains on the new tasks require separate diagnostic evidence. No new-source accuracy or probability-calibration claim is made from expansion. The new candidate will warm-start from Spark B's frozen selected step 1,500: `/home/andy/ai/opensysone/runs/20260917T070415Z-expanded-parent/best.pt`. Its 512-decision validation accuracy is 94.7266%, crossfit NLL 0.170149713. Parent SHA256: `5f57ec38796d132edfa23638fbce66131fd4e7dfe87ceeadaba2b0e0d5c78024`. The original artifact and its data signature are preserved. An explicit `--warm-start --allow-train-data-change` operation verifies both real datasets, parent manifest lineage, pinned model, prompt, adapter layout and token policy, then restores only trained weights with fresh Adam and RNG. Exact resume remains strict about the full data signature. Fleet dataset overrides require unchanged reserved bytes and finalize against the selected candidate's verified dataset. GX10 is the replacement target: its selected step 2,500 survived three later non-improving validations by 07:00 UTC. Both improving Spark runs continue. Preserve GX10's stopped run and its fleet candidacy. The planned expanded pilot uses Qwen3-4B-Instruct-2507, FP32 rank 8 / alpha 16, branch batch 1, exact two-pass gradients, effective batch 4, seed 433, learning rates 2e-5, schedule 3,500 steps. The original 16 GiB allocation cap and absolute 16:00 UTC training / 18:16:10 UTC final deadlines remain in force. Startup, progress and process controls are recorded below after the real pilot and resume checks complete. ## Verification and execution Source revision `24b8ccf60d388f9cbb184e03a6ae260a1f5a8b86` is frozen in the clean detached worktree `/home/andy/ai/opensysone/source/expanded-24b8ccf`. Training uses the existing isolated Python at `/home/andy/ai/envs/opensysone/bin/python`. All 86 source tests passed. The independent postbuild audit reconstructed every added row against its pinned raw source, including the correct answer after permutation; all 13 downloaded source artifacts matched their hashes. The final training mix is 50.66% original replay and 49.34% additions. GX10's original campaign `20260916T193741Z-24h` stopped gracefully at step 4,380, with trainer and supervisor exit 0. The resumable checkpoint retains all 506 Adam states, and its selected step 2,500 is unchanged. Its smoke lock was released and the old trainer disappeared from the GPU process list before loading the pilot. The playground on port 7466 remains running. Expanded pilot: `/home/andy/ai/opensysone/runs/20260917T070758Z-train/artifacts`. It started at 07:07:59 UTC with approximately 100 GiB available memory and OOM adjustment 0. A frozen copy of the actual step-0 checkpoint verifies all 506 trainable tensors exactly equal the parent and the Adam state is empty; the new Python RNG matches seed 433. No parent weights or data signature were rewritten. Initial FP32 correctness passed, worst probability difference 2.65e-7 against the 1e-4 tolerance. The first 32 sampled training decisions cover all seven families. Small evidence is in `results/20260917-expanded-data/`; runtime verification proofs are also under `/home/andy/ai/opensysone/fleet-20260917/expanded-startup-verification/`. Captured initialization weights live outside Git under `/home/andy/ai/opensysone/runs/20260917T070758Z-expanded-startup-snapshots/`. The reusable helpers and their tests are also archived as `scripts/verify_expanded_startup.py` and `scripts/prepare_expansion_backup.py`. The pilot completed all eight updates and both correctness gates, **exit 0**. All 512 initial logits/probabilities/log-probabilities exactly match the frozen parent. Every Adam counter is eight; every loss, gradient and saved tensor is finite. Median update time is 8.90 seconds. Peak allocation including final checks is 15.426 GiB, under the 16 GiB cap; final correctness worst difference is 3.58e-7. Step-8 validation crossfit NLL is 0.170108 and accuracy 94.7266%; its improvement is below the fixed 0.001 threshold, so branch step 0 remains selected at 0.170150. Eight updates establish startup, not new-task generalization. ## Active campaign and controls Expanded campaign: `/home/andy/ai/opensysone/runs/20260917T072142Z-24h`, launched 07:21:42 UTC. Supervisor **1630617**, trainer **1630638**, source **24b8ccf**. It resumes the complete pilot step-8 weights, Adam and RNG. A frozen copy of the new campaign's initial checkpoint matches every trainable tensor, every Adam state and Python/Torch/CUDA RNG exactly; all launch processes have OOM adjustment 0. Final training exit is pending. Check live state rather than relying on PIDs. Full startup verification passed at **07:28:29 UTC**: all 512 resumed logits, probabilities and log-probabilities exactly reproduce pilot step 8; the captured checkpoint preserves every model/Adam/Python/Torch/CUDA state. Updates 9–12 are finite, with peak allocation 15.38 GiB. At **07:28:58 UTC**, GX10 had reached step 15, Spark A step 4,599 and Spark B refinement step 1,876. The five-candidate coordinator was waiting normally, the GUI on 7466 returned HTTP 200, and final publication watcher PID 1427060 was waiting for completion. Active job exits remain pending; these startup checks do not claim bitwise equality of future training trajectories or performance improvements on new tasks. ```bash python3 scripts/fleet_status.py --json tail -n 5 ~/ai/opensysone/runs/20260917T072142Z-24h/training.log ``` To stop only this training candidate gracefully: ```bash ~/ai/envs/opensysone/bin/python \ ~/ai/opensysone/source/expanded-24b8ccf/scripts/campaign_status.py \ --campaign ~/ai/opensysone/runs/20260917T072142Z-24h --stop ``` To recover after a verified stop, inspect the last saved step and free memory, then create a fresh campaign from its training directory with the same deadlines: ```bash ~/ai/envs/opensysone/bin/python \ ~/ai/opensysone/source/expanded-24b8ccf/scripts/launch_24h.py \ --pilot ~/ai/opensysone/runs/20260917T072142Z-24h/training \ --train-only --training-deadline 2026-09-17T16:00:00Z \ --deadline 2026-09-17T18:16:10Z --inference-max-tokens 1024 \ --selection-metric crossfit_temperature_nll_v1 ``` Before restarting a recovered fleet candidate, stop the waiting coordinator, update that candidate's campaign/training/evidence paths in its backed-up plan, retain its explicit dataset override and frozen source project path, and resume the coordinator. Preserve the original stopped candidate and previous artifacts. The coordinator is `/home/andy/ai/opensysone/runs/20260916T194403396250Z-fleet`, resumed as PID **1630841** after registering `gx10-4b-expanded` as candidate five. Use `scripts/fleet_campaign.py --campaign PATH --stop` and then `--resume` from the main checkout. It will finalize using the winner's verified dataset. Its absolute cutoffs remain unchanged; never reset them on recovery. ## Hugging Face backup The immutable expanded snapshot is `/home/andy/ai/opensysone/exports/20260917-expanded-pilot-hf-backup`: 41 files, 390,353,856 bytes. It includes the frozen parent, pilot selected step 0 and resumable step 8, complete transformed v2 data, diagnostic exclusions, original lineage, source notices, fixed validation reference and captured five-candidate fleet plan. Source archives preserve both training revisions and current source. The first expanded publication completed successfully, exit 0, at 07:24:36 UTC to the existing private [andyshu/opensysone](https://huggingface.co/andyshu/opensysone) repository. Payload commit `1a0a7fb679b88043c84325eadbbfc37c83345a31`, verified pointer commit `8f7cfb622e5dc03ccf2c480a7f7a03b6cb171021`, source `fe7e48f`. The repository remains private; no new blanket license is assigned. Later source refreshes use the same verified snapshot, and `CURRENT_SNAPSHOT.json` identifies the latest archived source revision. Publication run receipts remain under `~/ai/opensysone/runs/`; `LAST_EXPANDED_HF_PUBLISH` locates the latest attempt. The existing final-publication watcher still awaits the fleet's separately calibrated, evaluated winner and will publish `FINAL_MODEL.json` only after its completion gates pass.