|
Download source/EXPANDED_DATA.md from andyshu/opensysone: direct link, hf CLI and curl.
- Browser
- Download file 10.6 kB
-
https://huggingface.co/andyshu/opensysone/resolve/58f289696f58962a8ec98293d7b1abf9fd0c6b8b/source/EXPANDED_DATA.md
- Command line
-
hf download hf://andyshu/opensysone@58f289696f58962a8ec98293d7b1abf9fd0c6b8b/source/EXPANDED_DATA.md
-
curl -L -o EXPANDED_DATA.md https://huggingface.co/andyshu/opensysone/resolve/58f289696f58962a8ec98293d7b1abf9fd0c6b8b/source/EXPANDED_DATA.md
10.6 kB
| # Training-data expansion, 2026-09-17 | |
| The user requested a broader training set. Version 2 adds official training data | |
| from HellaSwag, PIQA and CommonsenseQA to the complete version-1 replay set. | |
| These sources add plausible continuations, physical problem solving and everyday | |
| commonsense questions with two to five answer choices. They train the existing | |
| decision scorer, not a text-generation model. | |
| Data directory: `/home/andy/ai/opensysone/data/public-decisions-v2-20260917`. | |
| Build with `scripts/prepare_expanded_data.py`; raw downloads, transformed rows, | |
| weights and token caches stay outside Git. Source revisions, licenses, file hashes, | |
| sampling decisions and exclusions are recorded in the dataset manifest and | |
| `results/20260917-expanded-data/dataset-manifest.json`. | |
| | Source | Training decisions after 512-token filtering | License | | |
| | --- | ---: | --- | | |
| | HellaSwag | 16,000 | MIT | | |
| | PIQA | 14,360 | AFL-3.0, creator's README | | |
| | CommonsenseQA | 9,490 | MIT | | |
| | Original replay | 40,915 | Original per-source provenance | | |
| | Total | 80,765 | Mixed; no blanket relicensing | | |
| Only labeled official training rows are added. Exclusions use normalized source | |
| groups and exact normalized context text, including official evaluation metadata. | |
| Duplicate choices and invalid labels are rejected. Original training bytes form | |
| an unchanged prefix; the existing deterministic per-epoch shuffle samples the | |
| combined rows uniformly. Model-specific token filtering rejects oversize examples | |
| without truncation. Exact matching does not exclude semantic duplicates or | |
| contamination already present in the pretrained model. | |
| CPU tokenization completed in 60.96 seconds without loading model weights or | |
| initializing CUDA. Of 80,792 raw rows, 27 exceed the limit (26 original BoolQ and | |
| one new PIQA); the longest retained branch is 509 tokens. Original tokenized | |
| training replay and all four protected tokenized splits compare exactly with the | |
| verified version-1 cache. Data signature: | |
| `76183c642668602f42b7f3e71a3fe03bd5bd76f064fce4ba92351d8703396207`. | |
| The original validation, calibration, test and Social IQA holdout files remain | |
| byte-for-byte identical. Another 128 source groups per new family are withheld | |
| as separate diagnostics, excluded from training and from checkpoint selection. | |
| After length filtering these contain 383 decisions (one PIQA item is too long). | |
| The existing fixed 512-decision selection therefore measures retention on the | |
| original four tasks; gains on the new tasks require separate diagnostic evidence. | |
| No new-source accuracy or probability-calibration claim is made from expansion. | |
| The new candidate will warm-start from Spark B's frozen selected step 1,500: | |
| `/home/andy/ai/opensysone/runs/20260917T070415Z-expanded-parent/best.pt`. | |
| Its 512-decision validation accuracy is 94.7266%, crossfit NLL 0.170149713. | |
| Parent SHA256: `5f57ec38796d132edfa23638fbce66131fd4e7dfe87ceeadaba2b0e0d5c78024`. | |
| The original artifact and its data signature are preserved. An explicit | |
| `--warm-start --allow-train-data-change` operation verifies both real datasets, | |
| parent manifest lineage, pinned model, prompt, adapter layout and token policy, | |
| then restores only trained weights with fresh Adam and RNG. Exact resume remains | |
| strict about the full data signature. Fleet dataset overrides require unchanged | |
| reserved bytes and finalize against the selected candidate's verified dataset. | |
| GX10 is the replacement target: its selected step 2,500 survived three later | |
| non-improving validations by 07:00 UTC. Both improving Spark runs continue. | |
| Preserve GX10's stopped run and its fleet candidacy. The planned expanded pilot | |
| uses Qwen3-4B-Instruct-2507, FP32 rank 8 / alpha 16, branch batch 1, exact two-pass | |
| gradients, effective batch 4, seed 433, learning rates 2e-5, schedule 3,500 steps. | |
| The original 16 GiB allocation cap and absolute 16:00 UTC training / 18:16:10 UTC | |
| final deadlines remain in force. Startup, progress and process controls are | |
| recorded below after the real pilot and resume checks complete. | |
| ## Verification and execution | |
| Source revision `24b8ccf60d388f9cbb184e03a6ae260a1f5a8b86` is frozen in the clean | |
| detached worktree `/home/andy/ai/opensysone/source/expanded-24b8ccf`. Training uses | |
| the existing isolated Python at `/home/andy/ai/envs/opensysone/bin/python`. | |
| All 86 source tests passed. The independent postbuild audit reconstructed every | |
| added row against its pinned raw source, including the correct answer after | |
| permutation; all 13 downloaded source artifacts matched their hashes. The final | |
| training mix is 50.66% original replay and 49.34% additions. | |
| GX10's original campaign `20260916T193741Z-24h` stopped gracefully at step 4,380, | |
| with trainer and supervisor exit 0. The resumable checkpoint retains all 506 Adam | |
| states, and its selected step 2,500 is unchanged. Its smoke lock was released and | |
| the old trainer disappeared from the GPU process list before loading the pilot. | |
| The playground on port 7466 remains running. | |
| Expanded pilot: `/home/andy/ai/opensysone/runs/20260917T070758Z-train/artifacts`. | |
| It started at 07:07:59 UTC with approximately 100 GiB available memory and OOM | |
| adjustment 0. A frozen copy of the actual step-0 checkpoint verifies all 506 | |
| trainable tensors exactly equal the parent and the Adam state is empty; the new | |
| Python RNG matches seed 433. No parent weights or data signature were rewritten. | |
| Initial FP32 correctness passed, worst probability difference 2.65e-7 against the | |
| 1e-4 tolerance. The first 32 sampled training decisions cover all seven families. | |
| Small evidence is in `results/20260917-expanded-data/`; runtime verification | |
| proofs are also under | |
| `/home/andy/ai/opensysone/fleet-20260917/expanded-startup-verification/`. | |
| Captured initialization weights live outside Git under | |
| `/home/andy/ai/opensysone/runs/20260917T070758Z-expanded-startup-snapshots/`. | |
| The reusable helpers and their tests are also archived as | |
| `scripts/verify_expanded_startup.py` and `scripts/prepare_expansion_backup.py`. | |
| The pilot completed all eight updates and both correctness gates, **exit 0**. | |
| All 512 initial logits/probabilities/log-probabilities exactly match the frozen | |
| parent. Every Adam counter is eight; every loss, gradient and saved tensor is | |
| finite. Median update time is 8.90 seconds. Peak allocation including final | |
| checks is 15.426 GiB, under the 16 GiB cap; final correctness worst difference is | |
| 3.58e-7. Step-8 validation crossfit NLL is 0.170108 and accuracy 94.7266%; its | |
| improvement is below the fixed 0.001 threshold, so branch step 0 remains selected | |
| at 0.170150. Eight updates establish startup, not new-task generalization. | |
| ## Active campaign and controls | |
| Expanded campaign: `/home/andy/ai/opensysone/runs/20260917T072142Z-24h`, launched | |
| 07:21:42 UTC. Supervisor **1630617**, trainer **1630638**, source **24b8ccf**. | |
| It resumes the complete pilot step-8 weights, Adam and RNG. A frozen copy of the | |
| new campaign's initial checkpoint matches every trainable tensor, every Adam | |
| state and Python/Torch/CUDA RNG exactly; all launch processes have OOM adjustment | |
| 0. Final training exit is pending. Check live state rather than relying on PIDs. | |
| Full startup verification passed at **07:28:29 UTC**: all 512 resumed logits, | |
| probabilities and log-probabilities exactly reproduce pilot step 8; the captured | |
| checkpoint preserves every model/Adam/Python/Torch/CUDA state. Updates 9–12 are | |
| finite, with peak allocation 15.38 GiB. At **07:28:58 UTC**, GX10 had reached | |
| step 15, Spark A step 4,599 and Spark B refinement step 1,876. The five-candidate | |
| coordinator was waiting normally, the GUI on 7466 returned HTTP 200, and final | |
| publication watcher PID 1427060 was waiting for completion. Active job exits | |
| remain pending; these startup checks do not claim bitwise equality of future | |
| training trajectories or performance improvements on new tasks. | |
| ```bash | |
| python3 scripts/fleet_status.py --json | |
| tail -n 5 ~/ai/opensysone/runs/20260917T072142Z-24h/training.log | |
| ``` | |
| To stop only this training candidate gracefully: | |
| ```bash | |
| ~/ai/envs/opensysone/bin/python \ | |
| ~/ai/opensysone/source/expanded-24b8ccf/scripts/campaign_status.py \ | |
| --campaign ~/ai/opensysone/runs/20260917T072142Z-24h --stop | |
| ``` | |
| To recover after a verified stop, inspect the last saved step and free memory, | |
| then create a fresh campaign from its training directory with the same deadlines: | |
| ```bash | |
| ~/ai/envs/opensysone/bin/python \ | |
| ~/ai/opensysone/source/expanded-24b8ccf/scripts/launch_24h.py \ | |
| --pilot ~/ai/opensysone/runs/20260917T072142Z-24h/training \ | |
| --train-only --training-deadline 2026-09-17T16:00:00Z \ | |
| --deadline 2026-09-17T18:16:10Z --inference-max-tokens 1024 \ | |
| --selection-metric crossfit_temperature_nll_v1 | |
| ``` | |
| Before restarting a recovered fleet candidate, stop the waiting coordinator, | |
| update that candidate's campaign/training/evidence paths in its backed-up plan, | |
| retain its explicit dataset override and frozen source project path, and resume | |
| the coordinator. Preserve the original stopped candidate and previous artifacts. | |
| The coordinator is `/home/andy/ai/opensysone/runs/20260916T194403396250Z-fleet`, | |
| resumed as PID **1630841** after registering `gx10-4b-expanded` as candidate five. | |
| Use `scripts/fleet_campaign.py --campaign PATH --stop` and then `--resume` from | |
| the main checkout. It will finalize using the winner's verified dataset. Its | |
| absolute cutoffs remain unchanged; never reset them on recovery. | |
| ## Hugging Face backup | |
| The immutable expanded snapshot is | |
| `/home/andy/ai/opensysone/exports/20260917-expanded-pilot-hf-backup`: 41 files, | |
| 390,353,856 bytes. It includes the frozen parent, pilot selected step 0 and | |
| resumable step 8, complete transformed v2 data, diagnostic exclusions, original | |
| lineage, source notices, fixed validation reference and captured five-candidate | |
| fleet plan. Source archives preserve both training revisions and current source. | |
| The first expanded publication completed successfully, exit 0, at 07:24:36 UTC | |
| to the existing private [andyshu/opensysone](https://huggingface.co/andyshu/opensysone) | |
| repository. Payload commit `1a0a7fb679b88043c84325eadbbfc37c83345a31`, verified | |
| pointer commit `8f7cfb622e5dc03ccf2c480a7f7a03b6cb171021`, source `fe7e48f`. | |
| The repository remains private; no new blanket license is assigned. Later source | |
| refreshes use the same verified snapshot, and `CURRENT_SNAPSHOT.json` identifies | |
| the latest archived source revision. Publication run receipts remain under | |
| `~/ai/opensysone/runs/`; `LAST_EXPANDED_HF_PUBLISH` locates the latest attempt. | |
| The existing final-publication watcher still awaits the fleet's separately | |
| calibrated, evaluated winner and will publish `FINAL_MODEL.json` only after its | |
| completion gates pass. | |