# Hugging Face backup The user requested backups in [andyshu/opensysone](https://huggingface.co/andyshu/opensysone). The existing model repository is private; retain its visibility and `license: unknown` metadata. Authentication uses the existing Hugging Face credential store. Never copy credentials into source, logs, manifests or artifacts. **Publication verified on 17 September at 02:54 UTC.** Source, documentation, provenance and all four selected/resumable checkpoint pairs are saved in the private repository. The publisher exited 0 after checking every remote size and hash, then publishing and checking `CURRENT_SNAPSHOT.json` and the model card. The first verified payload commit is `2d5c26ab7929b8c858aeea88842781fdbf79cd54`, with pointer commit `b213728f9acc5e009bc96704db341913d582be4c` and source `93602e6`. `hf-snapshot-publication.json` records the verification. The live `CURRENT_SNAPSHOT.json` is authoritative when documentation/source is refreshed. The replacement credential authenticates as `andyshu` with write access and is stored in the existing login store; token values are excluded from artifacts. The earlier HTTP 403/read-only attempt remains as historical evidence in `hf-initial-artifacts-publication.json`; `hf-write-auth-verified.json` records that authentication was repaired. The final watcher has no inherited `HF_TOKEN`/`HUGGING_FACE_HUB_TOKEN` override; it reads the updated stored login when publication begins. Commands in a shell that retains the old read-only environment token must unset those overrides. ## Initial snapshot Immutable local staging: `/home/andy/ai/opensysone/exports/20260917T022835Z-hf-backup`. Remote payload: `snapshots/20260917T022835Z-hf-backup/`. The payload contains 1,066,618,246 bytes of checkpoints and provenance, excluding the manifest/checksum index. All four selected/resumable pairs were verified on CPU against stable source hashes, model/data provenance and validation evidence. | Candidate | Selected step | Resumable step | | --- | ---: | ---: | | GX10 4B | 2,500 | 2,598 | | Spark A 4B | 2,500 | 2,604 | | Spark B 2B | 2,000 | 6,000 | | Spark B 4B refinement pilot | branch 0, inherited GX10 2,500 | 8 | These are uncalibrated snapshots. The fourth backup is the completed pilot; the longer refinement campaign runs independently. The remote `CURRENT_SNAPSHOT.json` records the verified payload and committed source snapshot. `source/` contains browsable code, documentation and small evidence. Source archives preserve both the current documentation revision and exact `4a60423` training revision. Base weights and raw datasets are excluded; pinned revisions and dataset reconstruction requirements are in the backup manifest and model card. Upload completion is recorded in `results/20260917-fleet-progress/hf-*-publication.json`. Do not infer success from staging alone. Remote verification checks every file's size and LFS SHA-256 or Git blob hash, plus the small manifest bytes. For a later explicit snapshot refresh, use the existing stored login: ```bash env -u HF_TOKEN -u HUGGING_FACE_HUB_TOKEN \ ~/ai/envs/opensysone/bin/python scripts/publish_hf_snapshot.py \ --artifacts /home/andy/ai/opensysone/exports/20260917T022835Z-hf-backup \ --repo-id andyshu/opensysone \ --output /home/andy/ai/opensysone/runs/NEW-UNUSED-hf-snapshot-publish ``` Use a fresh output directory for each attempt and commit source/docs first. This creates a separate immutable publication stage, archives committed HEAD plus each recorded training revision, and publishes `sources//source.tar.gz` with per-file manifests. It checks all remote files before setting `CURRENT_SNAPSHOT.json` and the model card together. It refuses to replace an existing final release. The attempt has a 30-minute timeout and its own `state.json`/`exit_code`; failures record only the error type, without credentials or signed URLs. No retry has run with the known read-only token. ## Automatic final publication `scripts/publish_hf_final.py` runs independently of training and the fleet controller. It waits for a complete, API-ready fleet with exit 0, checks the frozen selection, completed test/holdout metrics, correctness evidence and matching model/deployment/ probe hashes, then creates an immutable export under `~/ai/opensysone/exports`. It uploads to `final//`, verifies the remote payload, and only then updates `FINAL_MODEL.json` and the model card together. It never selects or evaluates a model itself. Failure of this backup does not change the fleet's deployment state. The watcher stops by **2026-09-17 18:46:10 UTC**, 30 minutes after the original delivery deadline. Uploads have three bounded attempts. Training still stops by 16:00 UTC and the final model's deadline remains 18:16:10 UTC. The current watcher path and process are recorded in `HANDOVER.md`. Inspect its `state.json`, `exit_code` and adjacent log. Before stopping, verify `/proc//cmdline` against `state.json.command`, then send SIGTERM to that exact watcher only. Resume with a fresh output directory; an existing immutable export is hash-checked and reused: ```bash env -u HF_TOKEN -u HUGGING_FACE_HUB_TOKEN \ ~/ai/envs/opensysone/bin/python scripts/publish_hf_final.py \ --campaign /home/andy/ai/opensysone/runs/20260916T194403396250Z-fleet \ --repo-id andyshu/opensysone \ --output /home/andy/ai/opensysone/runs/NEW-UNUSED-hf-final-watch --watch ``` The command stays in the foreground; detach with a recorded process/log if leaving it unattended. Never start a second watcher without checking the recorded one.