Download source/docs/operations/handover.md from andyshu/opensysone: direct link, hf CLI and curl.
- Browser
- Download file 27.2 kB
-
https://huggingface.co/andyshu/opensysone/resolve/main/source/docs/operations/handover.md
- Command line
-
hf download hf://andyshu/opensysone/source/docs/operations/handover.md
-
curl -L -o handover.md https://huggingface.co/andyshu/opensysone/resolve/main/source/docs/operations/handover.md
Completed training and profiling campaign
Current state, 2026-09-17 09:22 UTC: training and evaluation are complete. All trainers, final validation jobs and profiling jobs exited 0. No optimizer updates remain active. Both Spark GPUs are idle. Keep all resumable checkpoints; do not resume training as part of this completed request.
The frozen selected Qwen3-4B-Instruct-2507 decision scorer retains Spark B step 1,500 weights through the identical expanded-branch step 0 artifact. All 506 trainable tensors match exactly. The latest expanded 159, Spark A 4,765 and Spark B 2,000 states were freshly validated; none passed the fixed promotion threshold. Five candidates were eligible; selection froze before reserved evaluation at 08:09:51 UTC. The expanded step-159 result is a post-selection diagnostic and does not replace the winner.
The final report, tables and charts records 92.90% vs 84.48% base-verifier accuracy on 2,042 test decisions and 72.92% vs 70.31% on 768 Social IQA holdout decisions. The 95% paired bootstrap accuracy differences are +8.42 pp [6.85, 9.89] and +2.60 pp [0.13, 5.34]. The separate 320-example comparison scores selected/base-verifier/joint-label at 89.06% / 80.94% / 86.25%. Current selected inference is slower in every measured workload: long-state four-choice medians are 3.710 / 3.177 / 0.818 seconds. These are FP32 local warm measurements on an idle Spark, without prefix caching. See profiling-protocol.md and next-steps.md.
Completed runs and remaining services
All run IDs below are relative to /home/andy/ai/opensysone/runs on the named host.
- GX10 fleet
20260916T194403396250Z-fleet: coordinator, evaluation and local harness checks exit 0, completed 09:20:23 UTC. Execution source is07f10e791061a679b829ed1dc5b33897e001d67d, frozen worktree/home/andy/ai/opensysone/source/profile-07f10e7. The original deadline remained 18:16:10 UTC.plan.before-user-wrapup.jsonpreserves the original selection schedule; the final cutoff was advanced to 08:07:30 UTC at the user's request. - The calibrated artifact is
evaluation/model.ptin that fleet run, SHA256e270e3da905604d97bf5a8f380ea308133403d1c4790a5c012cb1c12e9b6f348, temperature 1.745822072, saved before test predictions. The deploy pointer is/home/andy/ai/opensysone/deploy/current.json. - Local Jev-compatible API remains on loopback 18081, PID 1716630,
source
07f10e7, serving the calibrated model.api_probe.jsonrecords real normalized inference;api.logis its log. Inspect/health. To stop, first verify the PID command against fleetstate.json, then send SIGTERM. Restart the recordedapi_commandusing the isolated environment; no training or evaluation resume is needed. Serving exit status remains pending while alive. Hosted Jev inference still requires credentials and has not been measured. - GUI remains on 7466, default Qwen3 4B · Selected, wrapper 1674634,
server 1674635, runtime
20260917T081925Z-playground-selected, launch source00c80dd. Four-model real-browser checks passed. Historical snapshots remain selectable. See playground.md for inspect/stop/restart controls; the active serving processes intentionally have no final exit status yet. - Spark A
20260917T081236Z-spark-a-profile: all three inference methods, exit 0 at 09:08:39 UTC. Spark B20260917T081236Z-spark-b-profile: expanded checkpoint 159, exit 0 at 08:20:35 UTC. Profiling source07f10e7; fixed protocol20260917T081045Z-inference-profile-protocol. All 2,812 predictions and 360 measured timing samples passed audit. Collector and auditor also exited. - Final model publication watcher
20260917T023940Z-hf-final-watchfinished exit 0 at 09:20:52 UTC. Hugging Faceandyshu/opensysoneremains private; verifiedFINAL_MODEL.jsonpointer commit is2082f71beb86740f36f00b82a6eeab64b9e89b61. Its source archive is6729461. Supplemental profiling/checkpoint/report backup usesPROFILE_RESULTS.json; inspect20260917T092700Z-hf-wrapup-evidence/state.jsonandexit_codefor its verified payload/pointer revisions. Pointers are published only after remote integrity checks. The earlier20260917T092300Z-hf-wrapup-evidenceattempt exited 1 on a local progress-callback TypeError before upload; its receipt is retained. The callback was corrected before retry. No credentials or base-model weights belong in the archives.
Final validation source is 18f2b39. Training revisions remain 4a60423
(Sparks/original 4B) and 24b8ccf (expanded GX10). Original/resumable snapshots,
CPU proofs and archival resume commands are in 20260917T075209Z-training-wrapup
and 20260917T075129Z-spark-wrapup on GX10. The former contains the final evidence
file map, selected lineage proof, completion proof and profile-deployment with
collected Spark outputs. The report run is 20260917T092045Z-final-profile-report;
its generator was committed at 6729461. Small copies are committed under
results/20260917-wrapup. The sections below preserve historical campaign detail;
their active-training instructions are superseded by this completed state.
Latest continuation, 2026-09-17 07:29 UTC: the user requested broader training
data. training-data.md records the 80,765-example seven-family
mix, exact protected-split preservation, completed eight-step pilot and new GX10
campaign 20260917T072142Z-24h. It warm-starts from Spark B's selected step 1,500
with fresh Adam, then resumes the verified pilot. GX10's original run stopped
cleanly at step 4,380, retaining best step 2,500. Both Spark 4B runs continue;
the completed 2B remains available. All five candidates are registered.
The expanded run passed exact 512-prediction replay, full optimizer/RNG restore
and subsequent finite updates; it reached step 15 at 07:28:58 UTC. Its data,
pilot weights and source are uploaded and verified in the existing private
Hugging Face repository. The GUI on 7466 and final-publication watcher remain up.
The earlier overnight assessment is in next-steps.md.
The user assigned GX10 and both Sparks to this task, authorized stopping their
workloads, and requested continued experimentation without permission prompts.
SSH key authentication as andy works on 192.168.8.111 (spark-a / spark-d1b4)
and 192.168.8.204 (spark-b / spark-3e2a). The former Qwen serving pair was
stopped cleanly at 19:14 UTC; its files/cache and exact restoration commands are
preserved in fleet.md. All three hosts run independent trials.
The original absolute final deadline remains 2026-09-17 18:16:10 UTC / 19:16:10 BST. Every training supervisor stops by 16:00 UTC / 17:00 BST, leaving 2 h 16 min for selection, calibration, untouched evaluation and the local Jev-compatible API. Never reset that deadline on recovery. Training early stopping can finish sooner. Final evaluation and hosted Jev inference are still pending.
Read plan.md, results-history.md and fleet.md.
The authoritative working source is /home/andy/projects/opensysone on GX10;
there is no hosted Git remote. Do not overwrite it with an older Mac checkout.
Operational documentation is copied to /home/andy/ai/opensysone/gx10-reference
on each host. Read the relevant docs/host.md, docs/training.md, docs/spark-a.md,
docs/spark-b.md and docs/fleet.md before changing machines.
Active runs and source
All run IDs below are relative to /home/andy/ai/opensysone/runs on that host.
The Spark trainers launched from clean source 4a60423. The expanded GX10
trainer uses clean source 24b8ccf in the detached worktree
/home/andy/ai/opensysone/source/expanded-24b8ccf; keep that worktree for its
supervisor and recovery. The main checkout contains current documentation and
backup/verification tools. Running trainers retain their execution revision
and source hashes in their manifests. Inspect live state before
using recorded PIDs. Exit statuses of active jobs remain pending.
| Host | Trial | Campaign | Supervisor / trainer at launch |
|---|---|---|---|
| GX10 | Original 4B, stopped at 4,380; selected 2,500 | 20260916T193741Z-24h |
exited 0 / 0 |
| GX10 | Expanded 4B, LR 0.00002, seed 433, resumed pilot step 8 | 20260917T072142Z-24h |
1630617 / 1630638 |
| spark-a | 4B, LR 0.00003, fresh optimizer then pilot resume | 20260916T194258Z-24h |
327084 / 327116 |
| spark-b | 2B completed at step 6,000; selected step 2,000 | 20260916T193803Z-24h |
exited 0 / 0 |
| spark-b | 4B refinement, LR 0.00001, seed 432 | 20260917T023137Z-24h |
483974 / 484001 |
Fleet coordinator: 20260916T194403396250Z-fleet on GX10, PID 1630841,
source 24b8ccf, running in waiting_for_selection with OOM adjustment 0.
It was stopped before the fifth candidate and its explicit dataset override were
registered, then restarted. The old stop's exit 1 can remain in exit_code while
the new coordinator runs; current process identity/state determines liveness.
It selects the best durable candidate, then runs finalization and serves it on
GX10. The individual campaigns are train_only=true;
they cannot independently evaluate reserved data or publish competing deployments.
Each campaign's training/checkpoint.pt holds resumable optimizer/RNG state;
training/best.pt holds its validation-selected model. Saves occur every 250
steps or 900 seconds, independently of 512-decision validation every 500 steps.
Patience is eight evaluations. A logged update can be newer than its checkpoint.
Three epochs are an upper bound, not a promised completed data pass.
Latest audit 2026-09-17 02:10–02:15 UTC: GX10 step 2,570 / selected 2,500
(93.55% accuracy, 0.188640 crossfit NLL); Spark A step 2,529 / selected 2,500
(92.58%, 0.218012); Spark B 2B finished at 6,000 / selected 2,000 (89.84%,
0.255294). All logged gradients/losses are finite; peak allocations are
15.624 / 15.624 / 8.183 GiB. The 2B final correctness gate passed, worst 6.56e-7.
Small evidence is in results/20260917-fleet-progress/; historical startup proofs
remain in results/20260916-fleet-setup/. Reserved predictions remain untouched.
At 02:38 UTC, GX10/A had logged steps 2,721/2,704, with selected checkpoints
unchanged. The new Spark B campaign replayed all 512 pilot step-8 predictions
exactly, preserved full Adam/RNG state, and resumed finite updates (step 13 in
the fleet snapshot; startup proof covers 9–12). Its selected branch step 0 is
still the frozen GX10 parent. Startup checks passed; final exits remain pending.
Evidence and selection
The fixed selection policy is crossfit_temperature_nll_v1, four source-group-
disjoint validation folds, seed 431. Each fold's temperature is fitted on the other
three; macro-family NLL is scored only on held-out validation predictions. Final
serving temperature is fitted afresh on reserved calibration after the winner is
frozen. Raw NLL and accuracy remain separately reported. No reserved calibration,
test or Social IQA predictions have selected a candidate.
This is a documented validation-driven revision: 4B step 128 scores 89.0625%
accuracy / 0.318518 crossfit NLL, versus step 40's 87.5% / 0.359522. Raw NLL
favored step 40 because step 128 was more overconfident. The accuracy difference
alone is uncertain. Fresh step-178 validation subsequently reached 90.4297%
accuracy / 0.303825 crossfit NLL / 0.404198 raw NLL and became the durable
best; its state and evidence passed the same fleet eligibility checks. See
results/20260916-fleet-setup/selection-diagnostic.json;
independent test/holdout results remain necessary.
GX10's old 20260916T185910Z-24h stopped with a complete step-128 checkpoint;
its trainer exited -9 during subsequent final checks after the supervisor's
30-second grace. No optimizer progress was lost. The next campaign,
20260916T192239Z-24h, restored all trainable weights, Adam and Python/torch/CUDA
RNG exactly, then stopped gracefully at step 178, training exit 0. Its explicit
skipped_on_stop final-check status is not a new correctness pass.
The immutable 20260916T193721Z-selection-parent keeps that step-178 checkpoint
byte-for-byte and reselects the unchanged step-128 best weights under the new
criterion. It preserves the old raw-NLL best separately. Migration proof is in
results/20260916-fleet-setup/selection_migration.json. Do not restart old campaigns.
Both Spark environments passed 21,368 file hashes and 55 exact distribution
versions against GX10. All 13 files in each pinned model were SHA-256 verified.
Spark A reproduced all 512 original 4B pilot predictions exactly before eight
finite updates; its pilot and fresh GPU/HTTP verification exited 0. Spark B passed
fresh GPU/HTTP verification,
reproduced all 512 original 2B predictions exactly, and resumed finite optimizer
updates. Spark A's long campaign also reproduced all 512 step-8 predictions
exactly, preserved all weights/Adam/RNG state, and resumed finite updates. Small proofs are in results/20260916-fleet-setup/. The revised source
passes 37 CPU tests, plus the updated trained-Adam reselection integration.
These wiring and validation checks do not establish held-out generalization.
The exact A step-2,500 / B step-2,000 ensemble-reference artifacts are preserved
on GX10 in 20260917T022201Z-ensemble-reference, outside fleet selection. The
fixed mixed ensemble's small validation NLL advantage is uncertain; see
next-steps.md. This diagnostic is outside the current individual-model selection protocol.
Adoption would require an explicit protocol revision and verified implementation
before any reserved-data evaluation.
Model, data and machine bounds
Pinned Apache-2.0 models are under /home/andy/ai/models/opensysone:
Qwen3-4B-Instruct-2507-cdbee75f, revisioncdbee75f17c01a7cc42f958dc650907174af0554: FP32, rank 8 / alpha 16, 16.518M trainable parameters, 512-token training, exact two-pass gradients.Qwen3.5-2B-15852e8c, revision15852e8c16360a2fea060d615a32b45270f8a8fc: FP32 text decoder, rank 16 / alpha 32, 16.821M trainable parameters, 768-token training.
Use /home/andy/ai/envs/opensysone/bin/python. GX10's isolated environment reuses
existing torch/Transformers read-only; the Sparks have verified isolated copies.
Shared environments are unchanged. BF16 remains blocked by measured numerical
invariance failures. Keep the 16 GiB CUDA allocation cap, at least 24 GiB
MemAvailable before loading, GPU process inspection and oom_score_adj=0.
GX10's small existing router remains; Spark serving jobs remain stopped.
GX10 has no ConnectX; memory pools are separate. No network, swap, earlyoom,
firewall or clock configuration was changed.
Frozen data: /home/andy/ai/opensysone/data/public-decisions-v1-20260916.
Source-group-disjoint SNLI, BoolQ, ARC and four-choice Banking77; Social IQA is
an untrained task-family holdout. Pins/licences/hashes are in
results/public-decisions-v1-manifest.json. The 4B retains 40,915 train / 512
validation / 510 calibration / 2,042 test / 768 holdout; 2B retains 40,937 / 512 /
512 / 2,047 / 768. Validation IDs are identical. Exact deduplication does not
exclude semantic duplicates or pretraining contamination. No customer data.
Inspect, stop and recover
One read-only command checks every registered candidate concurrently, including completed candidates, with exact process identity and no model loading:
python3 scripts/fleet_status.py
python3 scripts/fleet_status.py --json
Use the exact active host/run from the table, or the fleet controls in fleet.md. From the project directory on the relevant host:
~/ai/envs/opensysone/bin/python scripts/campaign_status.py \
--campaign /home/andy/ai/opensysone/runs/20260916T193741Z-24h
tail -n 5 /home/andy/ai/opensysone/runs/20260916T193741Z-24h/training/training.jsonl
Add --stop for a command-verified TERM to the recorded supervisor, orphan child
or API. Wait for exit and lock release before restarting. Training checkpoints
at a safe boundary. Do not start a second model on an occupied host. Resume a
stopped candidate into a fresh campaign on its host:
~/ai/envs/opensysone/bin/python scripts/launch_24h.py \
--pilot /absolute/old/campaign/training --train-only \
--training-deadline 2026-09-17T16:00:00Z \
--deadline 2026-09-17T18:16:10Z --inference-max-tokens 1024 \
--selection-metric crossfit_temperature_nll_v1
Preserve model/data/seed/rank/alpha/learning rates/batches/token limits/schedule/
epochs/two-pass configuration. The launcher restores them from the checkpoint.
Use only trusted project checkpoints. If a candidate path changes, stop the
waiting fleet coordinator, update that candidate in its own plan.json, and
resume it. Editing a plan while the coordinator is running does not reload it.
After selection.json exists, the winner is frozen; recovery must not reselect
after test access. Stopping the waiting coordinator does not stop the independently supervised
independently bounded training jobs; stop each campaign explicitly when needed.
Finalization and Jev harness
The coordinator reconstructs the selected model, fits a scalar temperature on
reserved calibration, checkpoints evaluation/model.pt, then evaluates untouched
test/holdout against the unchanged pretrained scorer with separately fitted base
temperature and source-group uncertainty. It verifies direct inference and a real
HTTP request before publishing /home/andy/ai/opensysone/deploy/current.json.
Success requires fleet exit_code=0, complete evaluation/metrics.json, and
state.json with api_ready=true. Training completion alone is insufficient.
The resulting API is http://127.0.0.1:18081/v1/systemone, inference limit 1,024;
its PID/command remain recorded after the coordinator exits.
jev-api.md documents local, hosted and comparison modes,
optional bearer authentication and Mac SSH tunneling. TYPESAFE_API_KEY is
not configured, so authenticated hosted Jev inference has not been tested.
Local confidence is normalized entropy, not established correctness calibration.
This produces a general-language decision scorer, not a new general-purpose chat
model. Frozen-head/generation controls, new-model prefix caching and the broader
latency matrix remain open.
Completed runs and history
All paths below are under /home/andy/ai/opensysone/runs/.
| Run | Execution source | Exit / result |
|---|---|---|
20260916T182256Z-train |
4b25eec |
1, chat-template return-type setup error before optimizer training; preserved |
20260916T182352Z-train |
f1c9322 |
0, 2B public-data 40-step pilot |
20260916T183240Z-train |
980d881 |
0, exact 512-prediction restart and step 41 |
20260916T183751Z-verify2b |
script SHA in manifest | 0, restored-optimizer longest-input gradients and real authenticated HTTP |
20260916T183823Z-train |
ccbbe6d |
0, selected 4B 40-step pilot |
20260916T185718Z-verify4b |
script SHA in manifest | 0, 4B reload, optimizer-memory, long-context and 255-choice HTTP checks |
20260916T155124Z |
34a993e |
0, original 0.5B synthetic FP32 60-step smoke |
20260916T155314Z |
91019bc |
0, exact 72-prediction restart and step 61 |
20260916T154714Z |
4d6cb0f |
1, BF16 probability-invariance failure; checkpoint preserved |
20260916T161253Z-precision |
staged hashes later b9dd165 |
0, diagnostic completed; BF16 fails |
20260916T161355Z-precision |
409ade4 |
0, expanded diagnosis; all BF16 variants fail |
Small raw results and checkpoint hashes are retained under results/<run-id>;
weights stay under ~/ai. Original 0.5B smoke and verified prefix caching are
unchanged in smoke_train.py/decision_model.py. FP32 passed the original expanded
precision gate at worst 0.00002138; BF16 remains blocked. Synthetic results prove
wiring, not task generalization. New-model prefix caching, frozen-head/generation
controls and the larger latency matrix remain open. Fleet connectivity details
are in fleet-scout.md, including verified numeric SSH addresses. The current fleet allocation supersedes
its earlier serving-occupancy snapshot.
Hugging Face backup and final-publication controls: huggingface.md.
Hugging Face publication watcher
Backup destination: andyshu/opensysone,
private, existing license metadata retained. The initial read-only credential
failed with HTTP 403; its sanitized report remains in
results/20260917-fleet-progress/hf-initial-artifacts-publication.json.
At 02:52 UTC the user-supplied replacement was verified as account andyshu,
role write, and saved to the existing local Hugging Face login store. Token
values are excluded from source, logs and backups. The initial snapshot upload
completed and was verified at 02:54 UTC, exit 0, including source and all four
checkpoint pairs. See hf-write-auth-verified.json and
hf-snapshot-publication.json. The first verified HF pointer commit is
b213728f9acc5e009bc96704db341913d582be4c; remote CURRENT_SNAPSHOT.json records
the authoritative payload/source revisions, including later documentation
refreshes. Local publication state is in
~/ai/opensysone/runs/20260917T025300Z-hf-snapshot-publish; immutable backup
staging remains under ~/ai/opensysone/exports.
An independent final-publication watcher runs on GX10: PID 1427060, source
35d6d8f, OOM adjustment 0, status waiting_for_completion at launch. Its
status directory is /home/andy/ai/opensysone/runs/20260917T023940Z-hf-final-watch;
the adjacent .log file records process output. Exit status remains pending.
Inspect state.json and exit_code; match the exact state.json.command against
/proc/1427060/cmdline before stopping only that watcher with SIGTERM. Restart
with the command in huggingface.md and a new output directory.
The watcher publishes the frozen final model only after completed evaluation and
verified deployment, then checks the remote payload before updating
FINAL_MODEL.json. Its own deadline is 18:46:10 UTC; this does not extend training
or the original model deadline. See the launch proof for the exact command/hash.
The previous watcher (PID 1426447) was deliberately stopped, exit 1, and replaced
with the process above to remove inherited HF_TOKEN/HUGGING_FACE_HUB_TOKEN
overrides. It will read the updated write-capable stored login when final publication
begins. Changing credentials does not require
changing the training jobs, fleet plan, API or repository visibility.
Interactive model playground — 2026-09-17
The user requested a GUI for text plus candidate answers and probabilities. It is
running on GX10 at http://127.0.0.1:7466, PID 1469393, backend source 2d0ff79,
OOM adjustment 0. From the Mac, run ssh -N -L 7466:127.0.0.1:7466 gx10, then
open http://localhost:7466. See playground.md.
Runtime: /home/andy/ai/opensysone/runs/20260917T034059Z-playground-port7466, also recorded
in LAST_PLAYGROUND. launch.json records the exact process command/source
hashes and server.log receives sanitized diagnostics. Exit status is pending
while serving. Inspect /api/status and match /proc/1469393/cmdline against
launch.json.command before sending SIGTERM to this process only. The documented
CLI restarts it from the fixed catalog after the old listener has stopped.
Three immutable snapshots are available: GX10 4B step 2,500, Spark A 4B step 2,500,
and Spark B 2B step 2,000. Their files/hashes and matching provenance are under
the original snapshot runtime, referenced by this runtime's models.json; these probabilities are explicitly uncalibrated. No
reserved evaluation examples were used for GUI testing. One backend resides at
a time; loads, scoring and unloading are serialized on a dedicated worker thread.
The 16 GiB allocation cap and memory/OOM checks remain active. This extra GUI
process shares GPU compute with training, so requests can slow optimizer steps.
Port 18081 remains reserved for final deployment; no firewall/services changed.
All seven backend tests passed. Chromium passed real inference for all three
models, return switching, clipboard JSON, input-edit staleness, duplicate options,
actual tokenizer overflow and mobile layout. Browser script/CSP errors: none.
Cold/switch example requests were 5.08–9.10 seconds; a warm main-model request
was 1.53 seconds. These are individual wiring timings, not latency percentiles
or quality estimates. Real results/screenshots and concurrency observations are
in results/20260917-playground/; the separate frontend fixture report is labeled
as stubbed UI testing. Training and the fleet/final-publication controllers remain
independent of this GUI.
A 45-second observation after GUI verification recorded the GX10 trainer advancing from step 3,063 to 3,068 with finite losses/gradients, a 15.624 GiB allocation peak and 81.3 GiB host memory available. The GUI stayed ready. This confirms continued training during GUI operation; it does not establish zero slowdown or capture all model-switch transients.
At the user's request, the playground moved from port 18082 to 7466. The
previous process received verified SIGTERM and exited; its wait status could not
be collected by the replacement launcher. stop.json in the old runtime records
that observation. The new process starts without a resident model and loads one
on the next scoring request. The page, scripts, styles, model catalog and status
respond on 7466; the old listener is closed. Evidence: results/20260917-playground/port-7466.json.
The frontend now fits the viewport, with Context/Choices/Results tabs on compact
screens and internally scrolling text/results. A fixed action bar and result-copy
footer stay accessible. Very short portrait layouts compact optional content to
retain readable inputs when a keyboard reduces the viewport. Browser fixture
checks pass 13 sizes, including 320×568, 844×390 and 390×360; they check visible
controls, readable input lines, loading, validation, keyboard tabs, resizing,
long result lists, stale results, clipboard and recovery. Fixture probabilities
are not new model evidence. See results/20260917-playground-layout/.
The backend process and model snapshots continue unchanged. Static files are
served directly from web/ with no-store caching, so refresh the browser to use
the layout. frontend-current.json in the active runtime records the current
frontend commit and served-file hashes independently of the backend launch
revision. Training source and processes were not modified for this relayout.