opensysone / source /docs /operations /handover.md
andyshu's picture
Organize verified OpenSysOne publication payload
294f8ea verified
|
Raw History Blame Contribute Delete
27.2 kB

Completed training and profiling campaign

Current state, 2026-09-17 09:22 UTC: training and evaluation are complete. All trainers, final validation jobs and profiling jobs exited 0. No optimizer updates remain active. Both Spark GPUs are idle. Keep all resumable checkpoints; do not resume training as part of this completed request.

The frozen selected Qwen3-4B-Instruct-2507 decision scorer retains Spark B step 1,500 weights through the identical expanded-branch step 0 artifact. All 506 trainable tensors match exactly. The latest expanded 159, Spark A 4,765 and Spark B 2,000 states were freshly validated; none passed the fixed promotion threshold. Five candidates were eligible; selection froze before reserved evaluation at 08:09:51 UTC. The expanded step-159 result is a post-selection diagnostic and does not replace the winner.

The final report, tables and charts records 92.90% vs 84.48% base-verifier accuracy on 2,042 test decisions and 72.92% vs 70.31% on 768 Social IQA holdout decisions. The 95% paired bootstrap accuracy differences are +8.42 pp [6.85, 9.89] and +2.60 pp [0.13, 5.34]. The separate 320-example comparison scores selected/base-verifier/joint-label at 89.06% / 80.94% / 86.25%. Current selected inference is slower in every measured workload: long-state four-choice medians are 3.710 / 3.177 / 0.818 seconds. These are FP32 local warm measurements on an idle Spark, without prefix caching. See profiling-protocol.md and next-steps.md.

Completed runs and remaining services

All run IDs below are relative to /home/andy/ai/opensysone/runs on the named host.

  • GX10 fleet 20260916T194403396250Z-fleet: coordinator, evaluation and local harness checks exit 0, completed 09:20:23 UTC. Execution source is 07f10e791061a679b829ed1dc5b33897e001d67d, frozen worktree /home/andy/ai/opensysone/source/profile-07f10e7. The original deadline remained 18:16:10 UTC. plan.before-user-wrapup.json preserves the original selection schedule; the final cutoff was advanced to 08:07:30 UTC at the user's request.
  • The calibrated artifact is evaluation/model.pt in that fleet run, SHA256 e270e3da905604d97bf5a8f380ea308133403d1c4790a5c012cb1c12e9b6f348, temperature 1.745822072, saved before test predictions. The deploy pointer is /home/andy/ai/opensysone/deploy/current.json.
  • Local Jev-compatible API remains on loopback 18081, PID 1716630, source 07f10e7, serving the calibrated model. api_probe.json records real normalized inference; api.log is its log. Inspect /health. To stop, first verify the PID command against fleet state.json, then send SIGTERM. Restart the recorded api_command using the isolated environment; no training or evaluation resume is needed. Serving exit status remains pending while alive. Hosted Jev inference still requires credentials and has not been measured.
  • GUI remains on 7466, default Qwen3 4B · Selected, wrapper 1674634, server 1674635, runtime 20260917T081925Z-playground-selected, launch source 00c80dd. Four-model real-browser checks passed. Historical snapshots remain selectable. See playground.md for inspect/stop/restart controls; the active serving processes intentionally have no final exit status yet.
  • Spark A 20260917T081236Z-spark-a-profile: all three inference methods, exit 0 at 09:08:39 UTC. Spark B 20260917T081236Z-spark-b-profile: expanded checkpoint 159, exit 0 at 08:20:35 UTC. Profiling source 07f10e7; fixed protocol 20260917T081045Z-inference-profile-protocol. All 2,812 predictions and 360 measured timing samples passed audit. Collector and auditor also exited.
  • Final model publication watcher 20260917T023940Z-hf-final-watch finished exit 0 at 09:20:52 UTC. Hugging Face andyshu/opensysone remains private; verified FINAL_MODEL.json pointer commit is 2082f71beb86740f36f00b82a6eeab64b9e89b61. Its source archive is 6729461. Supplemental profiling/checkpoint/report backup uses PROFILE_RESULTS.json; inspect 20260917T092700Z-hf-wrapup-evidence/state.json and exit_code for its verified payload/pointer revisions. Pointers are published only after remote integrity checks. The earlier 20260917T092300Z-hf-wrapup-evidence attempt exited 1 on a local progress-callback TypeError before upload; its receipt is retained. The callback was corrected before retry. No credentials or base-model weights belong in the archives.

Final validation source is 18f2b39. Training revisions remain 4a60423 (Sparks/original 4B) and 24b8ccf (expanded GX10). Original/resumable snapshots, CPU proofs and archival resume commands are in 20260917T075209Z-training-wrapup and 20260917T075129Z-spark-wrapup on GX10. The former contains the final evidence file map, selected lineage proof, completion proof and profile-deployment with collected Spark outputs. The report run is 20260917T092045Z-final-profile-report; its generator was committed at 6729461. Small copies are committed under results/20260917-wrapup. The sections below preserve historical campaign detail; their active-training instructions are superseded by this completed state.

Latest continuation, 2026-09-17 07:29 UTC: the user requested broader training data. training-data.md records the 80,765-example seven-family mix, exact protected-split preservation, completed eight-step pilot and new GX10 campaign 20260917T072142Z-24h. It warm-starts from Spark B's selected step 1,500 with fresh Adam, then resumes the verified pilot. GX10's original run stopped cleanly at step 4,380, retaining best step 2,500. Both Spark 4B runs continue; the completed 2B remains available. All five candidates are registered. The expanded run passed exact 512-prediction replay, full optimizer/RNG restore and subsequent finite updates; it reached step 15 at 07:28:58 UTC. Its data, pilot weights and source are uploaded and verified in the existing private Hugging Face repository. The GUI on 7466 and final-publication watcher remain up. The earlier overnight assessment is in next-steps.md.

The user assigned GX10 and both Sparks to this task, authorized stopping their workloads, and requested continued experimentation without permission prompts. SSH key authentication as andy works on 192.168.8.111 (spark-a / spark-d1b4) and 192.168.8.204 (spark-b / spark-3e2a). The former Qwen serving pair was stopped cleanly at 19:14 UTC; its files/cache and exact restoration commands are preserved in fleet.md. All three hosts run independent trials.

The original absolute final deadline remains 2026-09-17 18:16:10 UTC / 19:16:10 BST. Every training supervisor stops by 16:00 UTC / 17:00 BST, leaving 2 h 16 min for selection, calibration, untouched evaluation and the local Jev-compatible API. Never reset that deadline on recovery. Training early stopping can finish sooner. Final evaluation and hosted Jev inference are still pending.

Read plan.md, results-history.md and fleet.md. The authoritative working source is /home/andy/projects/opensysone on GX10; there is no hosted Git remote. Do not overwrite it with an older Mac checkout. Operational documentation is copied to /home/andy/ai/opensysone/gx10-reference on each host. Read the relevant docs/host.md, docs/training.md, docs/spark-a.md, docs/spark-b.md and docs/fleet.md before changing machines.

Active runs and source

All run IDs below are relative to /home/andy/ai/opensysone/runs on that host. The Spark trainers launched from clean source 4a60423. The expanded GX10 trainer uses clean source 24b8ccf in the detached worktree /home/andy/ai/opensysone/source/expanded-24b8ccf; keep that worktree for its supervisor and recovery. The main checkout contains current documentation and backup/verification tools. Running trainers retain their execution revision and source hashes in their manifests. Inspect live state before using recorded PIDs. Exit statuses of active jobs remain pending.

Host Trial Campaign Supervisor / trainer at launch
GX10 Original 4B, stopped at 4,380; selected 2,500 20260916T193741Z-24h exited 0 / 0
GX10 Expanded 4B, LR 0.00002, seed 433, resumed pilot step 8 20260917T072142Z-24h 1630617 / 1630638
spark-a 4B, LR 0.00003, fresh optimizer then pilot resume 20260916T194258Z-24h 327084 / 327116
spark-b 2B completed at step 6,000; selected step 2,000 20260916T193803Z-24h exited 0 / 0
spark-b 4B refinement, LR 0.00001, seed 432 20260917T023137Z-24h 483974 / 484001

Fleet coordinator: 20260916T194403396250Z-fleet on GX10, PID 1630841, source 24b8ccf, running in waiting_for_selection with OOM adjustment 0. It was stopped before the fifth candidate and its explicit dataset override were registered, then restarted. The old stop's exit 1 can remain in exit_code while the new coordinator runs; current process identity/state determines liveness. It selects the best durable candidate, then runs finalization and serves it on GX10. The individual campaigns are train_only=true; they cannot independently evaluate reserved data or publish competing deployments.

Each campaign's training/checkpoint.pt holds resumable optimizer/RNG state; training/best.pt holds its validation-selected model. Saves occur every 250 steps or 900 seconds, independently of 512-decision validation every 500 steps. Patience is eight evaluations. A logged update can be newer than its checkpoint. Three epochs are an upper bound, not a promised completed data pass.

Latest audit 2026-09-17 02:10–02:15 UTC: GX10 step 2,570 / selected 2,500 (93.55% accuracy, 0.188640 crossfit NLL); Spark A step 2,529 / selected 2,500 (92.58%, 0.218012); Spark B 2B finished at 6,000 / selected 2,000 (89.84%, 0.255294). All logged gradients/losses are finite; peak allocations are 15.624 / 15.624 / 8.183 GiB. The 2B final correctness gate passed, worst 6.56e-7. Small evidence is in results/20260917-fleet-progress/; historical startup proofs remain in results/20260916-fleet-setup/. Reserved predictions remain untouched. At 02:38 UTC, GX10/A had logged steps 2,721/2,704, with selected checkpoints unchanged. The new Spark B campaign replayed all 512 pilot step-8 predictions exactly, preserved full Adam/RNG state, and resumed finite updates (step 13 in the fleet snapshot; startup proof covers 9–12). Its selected branch step 0 is still the frozen GX10 parent. Startup checks passed; final exits remain pending.

Evidence and selection

The fixed selection policy is crossfit_temperature_nll_v1, four source-group- disjoint validation folds, seed 431. Each fold's temperature is fitted on the other three; macro-family NLL is scored only on held-out validation predictions. Final serving temperature is fitted afresh on reserved calibration after the winner is frozen. Raw NLL and accuracy remain separately reported. No reserved calibration, test or Social IQA predictions have selected a candidate.

This is a documented validation-driven revision: 4B step 128 scores 89.0625% accuracy / 0.318518 crossfit NLL, versus step 40's 87.5% / 0.359522. Raw NLL favored step 40 because step 128 was more overconfident. The accuracy difference alone is uncertain. Fresh step-178 validation subsequently reached 90.4297% accuracy / 0.303825 crossfit NLL / 0.404198 raw NLL and became the durable best; its state and evidence passed the same fleet eligibility checks. See results/20260916-fleet-setup/selection-diagnostic.json; independent test/holdout results remain necessary.

GX10's old 20260916T185910Z-24h stopped with a complete step-128 checkpoint; its trainer exited -9 during subsequent final checks after the supervisor's 30-second grace. No optimizer progress was lost. The next campaign, 20260916T192239Z-24h, restored all trainable weights, Adam and Python/torch/CUDA RNG exactly, then stopped gracefully at step 178, training exit 0. Its explicit skipped_on_stop final-check status is not a new correctness pass. The immutable 20260916T193721Z-selection-parent keeps that step-178 checkpoint byte-for-byte and reselects the unchanged step-128 best weights under the new criterion. It preserves the old raw-NLL best separately. Migration proof is in results/20260916-fleet-setup/selection_migration.json. Do not restart old campaigns.

Both Spark environments passed 21,368 file hashes and 55 exact distribution versions against GX10. All 13 files in each pinned model were SHA-256 verified. Spark A reproduced all 512 original 4B pilot predictions exactly before eight finite updates; its pilot and fresh GPU/HTTP verification exited 0. Spark B passed fresh GPU/HTTP verification, reproduced all 512 original 2B predictions exactly, and resumed finite optimizer updates. Spark A's long campaign also reproduced all 512 step-8 predictions exactly, preserved all weights/Adam/RNG state, and resumed finite updates. Small proofs are in results/20260916-fleet-setup/. The revised source passes 37 CPU tests, plus the updated trained-Adam reselection integration. These wiring and validation checks do not establish held-out generalization.

The exact A step-2,500 / B step-2,000 ensemble-reference artifacts are preserved on GX10 in 20260917T022201Z-ensemble-reference, outside fleet selection. The fixed mixed ensemble's small validation NLL advantage is uncertain; see next-steps.md. This diagnostic is outside the current individual-model selection protocol. Adoption would require an explicit protocol revision and verified implementation before any reserved-data evaluation.

Model, data and machine bounds

Pinned Apache-2.0 models are under /home/andy/ai/models/opensysone:

  • Qwen3-4B-Instruct-2507-cdbee75f, revision cdbee75f17c01a7cc42f958dc650907174af0554: FP32, rank 8 / alpha 16, 16.518M trainable parameters, 512-token training, exact two-pass gradients.
  • Qwen3.5-2B-15852e8c, revision 15852e8c16360a2fea060d615a32b45270f8a8fc: FP32 text decoder, rank 16 / alpha 32, 16.821M trainable parameters, 768-token training.

Use /home/andy/ai/envs/opensysone/bin/python. GX10's isolated environment reuses existing torch/Transformers read-only; the Sparks have verified isolated copies. Shared environments are unchanged. BF16 remains blocked by measured numerical invariance failures. Keep the 16 GiB CUDA allocation cap, at least 24 GiB MemAvailable before loading, GPU process inspection and oom_score_adj=0. GX10's small existing router remains; Spark serving jobs remain stopped. GX10 has no ConnectX; memory pools are separate. No network, swap, earlyoom, firewall or clock configuration was changed.

Frozen data: /home/andy/ai/opensysone/data/public-decisions-v1-20260916. Source-group-disjoint SNLI, BoolQ, ARC and four-choice Banking77; Social IQA is an untrained task-family holdout. Pins/licences/hashes are in results/public-decisions-v1-manifest.json. The 4B retains 40,915 train / 512 validation / 510 calibration / 2,042 test / 768 holdout; 2B retains 40,937 / 512 / 512 / 2,047 / 768. Validation IDs are identical. Exact deduplication does not exclude semantic duplicates or pretraining contamination. No customer data.

Inspect, stop and recover

One read-only command checks every registered candidate concurrently, including completed candidates, with exact process identity and no model loading:

python3 scripts/fleet_status.py
python3 scripts/fleet_status.py --json

Use the exact active host/run from the table, or the fleet controls in fleet.md. From the project directory on the relevant host:

~/ai/envs/opensysone/bin/python scripts/campaign_status.py \
  --campaign /home/andy/ai/opensysone/runs/20260916T193741Z-24h
tail -n 5 /home/andy/ai/opensysone/runs/20260916T193741Z-24h/training/training.jsonl

Add --stop for a command-verified TERM to the recorded supervisor, orphan child or API. Wait for exit and lock release before restarting. Training checkpoints at a safe boundary. Do not start a second model on an occupied host. Resume a stopped candidate into a fresh campaign on its host:

~/ai/envs/opensysone/bin/python scripts/launch_24h.py \
  --pilot /absolute/old/campaign/training --train-only \
  --training-deadline 2026-09-17T16:00:00Z \
  --deadline 2026-09-17T18:16:10Z --inference-max-tokens 1024 \
  --selection-metric crossfit_temperature_nll_v1

Preserve model/data/seed/rank/alpha/learning rates/batches/token limits/schedule/ epochs/two-pass configuration. The launcher restores them from the checkpoint. Use only trusted project checkpoints. If a candidate path changes, stop the waiting fleet coordinator, update that candidate in its own plan.json, and resume it. Editing a plan while the coordinator is running does not reload it. After selection.json exists, the winner is frozen; recovery must not reselect after test access. Stopping the waiting coordinator does not stop the independently supervised independently bounded training jobs; stop each campaign explicitly when needed.

Finalization and Jev harness

The coordinator reconstructs the selected model, fits a scalar temperature on reserved calibration, checkpoints evaluation/model.pt, then evaluates untouched test/holdout against the unchanged pretrained scorer with separately fitted base temperature and source-group uncertainty. It verifies direct inference and a real HTTP request before publishing /home/andy/ai/opensysone/deploy/current.json. Success requires fleet exit_code=0, complete evaluation/metrics.json, and state.json with api_ready=true. Training completion alone is insufficient. The resulting API is http://127.0.0.1:18081/v1/systemone, inference limit 1,024; its PID/command remain recorded after the coordinator exits.

jev-api.md documents local, hosted and comparison modes, optional bearer authentication and Mac SSH tunneling. TYPESAFE_API_KEY is not configured, so authenticated hosted Jev inference has not been tested. Local confidence is normalized entropy, not established correctness calibration. This produces a general-language decision scorer, not a new general-purpose chat model. Frozen-head/generation controls, new-model prefix caching and the broader latency matrix remain open.

Completed runs and history

All paths below are under /home/andy/ai/opensysone/runs/.

Run Execution source Exit / result
20260916T182256Z-train 4b25eec 1, chat-template return-type setup error before optimizer training; preserved
20260916T182352Z-train f1c9322 0, 2B public-data 40-step pilot
20260916T183240Z-train 980d881 0, exact 512-prediction restart and step 41
20260916T183751Z-verify2b script SHA in manifest 0, restored-optimizer longest-input gradients and real authenticated HTTP
20260916T183823Z-train ccbbe6d 0, selected 4B 40-step pilot
20260916T185718Z-verify4b script SHA in manifest 0, 4B reload, optimizer-memory, long-context and 255-choice HTTP checks
20260916T155124Z 34a993e 0, original 0.5B synthetic FP32 60-step smoke
20260916T155314Z 91019bc 0, exact 72-prediction restart and step 61
20260916T154714Z 4d6cb0f 1, BF16 probability-invariance failure; checkpoint preserved
20260916T161253Z-precision staged hashes later b9dd165 0, diagnostic completed; BF16 fails
20260916T161355Z-precision 409ade4 0, expanded diagnosis; all BF16 variants fail

Small raw results and checkpoint hashes are retained under results/<run-id>; weights stay under ~/ai. Original 0.5B smoke and verified prefix caching are unchanged in smoke_train.py/decision_model.py. FP32 passed the original expanded precision gate at worst 0.00002138; BF16 remains blocked. Synthetic results prove wiring, not task generalization. New-model prefix caching, frozen-head/generation controls and the larger latency matrix remain open. Fleet connectivity details are in fleet-scout.md, including verified numeric SSH addresses. The current fleet allocation supersedes its earlier serving-occupancy snapshot.

Hugging Face backup and final-publication controls: huggingface.md.

Hugging Face publication watcher

Backup destination: andyshu/opensysone, private, existing license metadata retained. The initial read-only credential failed with HTTP 403; its sanitized report remains in results/20260917-fleet-progress/hf-initial-artifacts-publication.json. At 02:52 UTC the user-supplied replacement was verified as account andyshu, role write, and saved to the existing local Hugging Face login store. Token values are excluded from source, logs and backups. The initial snapshot upload completed and was verified at 02:54 UTC, exit 0, including source and all four checkpoint pairs. See hf-write-auth-verified.json and hf-snapshot-publication.json. The first verified HF pointer commit is b213728f9acc5e009bc96704db341913d582be4c; remote CURRENT_SNAPSHOT.json records the authoritative payload/source revisions, including later documentation refreshes. Local publication state is in ~/ai/opensysone/runs/20260917T025300Z-hf-snapshot-publish; immutable backup staging remains under ~/ai/opensysone/exports.

An independent final-publication watcher runs on GX10: PID 1427060, source 35d6d8f, OOM adjustment 0, status waiting_for_completion at launch. Its status directory is /home/andy/ai/opensysone/runs/20260917T023940Z-hf-final-watch; the adjacent .log file records process output. Exit status remains pending. Inspect state.json and exit_code; match the exact state.json.command against /proc/1427060/cmdline before stopping only that watcher with SIGTERM. Restart with the command in huggingface.md and a new output directory. The watcher publishes the frozen final model only after completed evaluation and verified deployment, then checks the remote payload before updating FINAL_MODEL.json. Its own deadline is 18:46:10 UTC; this does not extend training or the original model deadline. See the launch proof for the exact command/hash.

The previous watcher (PID 1426447) was deliberately stopped, exit 1, and replaced with the process above to remove inherited HF_TOKEN/HUGGING_FACE_HUB_TOKEN overrides. It will read the updated write-capable stored login when final publication begins. Changing credentials does not require changing the training jobs, fleet plan, API or repository visibility.

Interactive model playground — 2026-09-17

The user requested a GUI for text plus candidate answers and probabilities. It is running on GX10 at http://127.0.0.1:7466, PID 1469393, backend source 2d0ff79, OOM adjustment 0. From the Mac, run ssh -N -L 7466:127.0.0.1:7466 gx10, then open http://localhost:7466. See playground.md.

Runtime: /home/andy/ai/opensysone/runs/20260917T034059Z-playground-port7466, also recorded in LAST_PLAYGROUND. launch.json records the exact process command/source hashes and server.log receives sanitized diagnostics. Exit status is pending while serving. Inspect /api/status and match /proc/1469393/cmdline against launch.json.command before sending SIGTERM to this process only. The documented CLI restarts it from the fixed catalog after the old listener has stopped.

Three immutable snapshots are available: GX10 4B step 2,500, Spark A 4B step 2,500, and Spark B 2B step 2,000. Their files/hashes and matching provenance are under the original snapshot runtime, referenced by this runtime's models.json; these probabilities are explicitly uncalibrated. No reserved evaluation examples were used for GUI testing. One backend resides at a time; loads, scoring and unloading are serialized on a dedicated worker thread. The 16 GiB allocation cap and memory/OOM checks remain active. This extra GUI process shares GPU compute with training, so requests can slow optimizer steps. Port 18081 remains reserved for final deployment; no firewall/services changed.

All seven backend tests passed. Chromium passed real inference for all three models, return switching, clipboard JSON, input-edit staleness, duplicate options, actual tokenizer overflow and mobile layout. Browser script/CSP errors: none. Cold/switch example requests were 5.08–9.10 seconds; a warm main-model request was 1.53 seconds. These are individual wiring timings, not latency percentiles or quality estimates. Real results/screenshots and concurrency observations are in results/20260917-playground/; the separate frontend fixture report is labeled as stubbed UI testing. Training and the fleet/final-publication controllers remain independent of this GUI.

A 45-second observation after GUI verification recorded the GX10 trainer advancing from step 3,063 to 3,068 with finite losses/gradients, a 15.624 GiB allocation peak and 81.3 GiB host memory available. The GUI stayed ready. This confirms continued training during GUI operation; it does not establish zero slowdown or capture all model-switch transients.

At the user's request, the playground moved from port 18082 to 7466. The previous process received verified SIGTERM and exited; its wait status could not be collected by the replacement launcher. stop.json in the old runtime records that observation. The new process starts without a resident model and loads one on the next scoring request. The page, scripts, styles, model catalog and status respond on 7466; the old listener is closed. Evidence: results/20260917-playground/port-7466.json.

The frontend now fits the viewport, with Context/Choices/Results tabs on compact screens and internally scrolling text/results. A fixed action bar and result-copy footer stay accessible. Very short portrait layouts compact optional content to retain readable inputs when a keyboard reduces the viewport. Browser fixture checks pass 13 sizes, including 320×568, 844×390 and 390×360; they check visible controls, readable input lines, loading, validation, keyboard tabs, resizing, long result lists, stale results, clipboard and recovery. Fixture probabilities are not new model evidence. See results/20260917-playground-layout/.

The backend process and model snapshots continue unchanged. Static files are served directly from web/ with no-store caching, so refresh the browser to use the layout. frontend-current.json in the active runtime records the current frontend commit and served-file hashes independently of the backend launch revision. Training source and processes were not modified for this relayout.