opensysone / source /FLEET_RUN.md
andyshu's picture
Back up verified OpenSysOne training snapshot and pinned source
1a0a7fb verified
|
Raw History Blame
14.6 kB

Three-machine campaign

Latest, 2026-09-17 07:22 UTC: GX10 now runs expanded-data 4B candidate 20260917T072142Z-24h, from frozen source worktree /home/andy/ai/opensysone/source/expanded-24b8ccf. Supervisor/trainer PIDs at launch are 1630617 / 1630638, OOM adjustment 0. The original GX10 campaign stopped cleanly at step 4,380, preserving best 2,500. Both Spark 4B runs continue, and the 2B remains completed. The existing fleet now has five candidates, including the explicit v2 dataset override; coordinator PID 1630841 was restarted after the plan update. See EXPANDED_DATA.md for verified evidence, dataset provenance and exact stop/resume commands. Historical startup rows below describe the original fleet and are superseded by this update.

2026-09-17 continuation: NEXT_STEPS.md records overnight findings and the two-Spark assessment. The original 2B campaign finished with exit 0 at step 6,000; its best is step 2,000. Spark B is now running a 4B refinement campaign from GX10 best step 2,500. The four-candidate fleet plan retains the completed 2B and includes the verified refinement campaign.

On 2026-09-16, the user assigned GX10 and both Sparks to OpenSysOne and explicitly authorized stopping existing processes. Key authentication from GX10 works as andy on both 192.168.8.111 (spark-a / spark-d1b4) and 192.168.8.204 (spark-b / spark-3e2a). No additional credentials are needed.

The delivery deadline stays 2026-09-17 18:16:10 UTC / 19:16:10 BST. Training candidates will stop at 16:00 UTC for validation-only model selection followed by calibration, untouched test/holdout evaluation and the local Jev-compatible API. Each training process keeps the 16 GiB CUDA cap. Independent experiments use all three GPUs without requiring an unverified distributed training stack or a ConnectX connection to GX10.

Machine Experiment Current state
GX10 Existing 4B, learning rate 0.0001, exact checkpoint continuation Training resumed at step 178; new validation selected step 178
spark-a 4B initialized from verified pilot weights, learning rate 0.00003, 7,500-step cosine schedule Verification passed; long training campaign running
spark-b 2B completed; new 4B refinement at LR 0.00001 / seed 432 Verified campaign 20260917T023137Z-24h running

All candidates use the frozen public dataset and the same 512 validation decisions. Selection now uses four-fold temperature-crossfit macro-family NLL from durable checkpoints (crossfit_temperature_nll_v1, fixed seed 431). Source groups stay in one fold; each fold's temperature is fitted on the other three. This avoids discarding a better classifier because of recoverable overconfidence. Raw NLL and accuracy are reported separately. Reserved calibration/test/Social IQA predictions do not choose the winner, and the serving temperature is fitted afresh on reserved calibration after selection. Warm initialization is a new experiment with a fresh optimizer, distinguished from an exact resume in its provenance. Model paths remain under ~/ai/models/opensysone; data, checkpoints and setup artifacts under ~/ai/opensysone. Source and small evidence alone belong in this repository.

Reclaimed serving pair

At 19:14 UTC, the exact verified Qwen serving process on spark-a (PID 154121) received SIGTERM and exited. After its memory was released, the exact verified RPC worker on spark-b (PID 114512) received SIGTERM and exited. No force kill was needed. Both GPU process lists were empty. Available memory afterward was 126,665,596,928 bytes on spark-a and 126,821,765,120 bytes on spark-b. The existing model files and RPC cache were preserved.

Full command/cwd/process identity, post-stop memory and GPU inspection are saved on GX10 in ~/ai/opensysone/fleet-20260916/spark-a-service-stop.json and spark-b-service-stop.json. Shared Python, firewall, network interfaces, swap, earlyoom and clocks were not changed. Spark-a still has swap enabled and no earlyoom; the project uses bounded allocations, host-memory checks and oom_score_adj=0 without changing that system configuration.

To restore the old serving pair after all training/evaluation/API GPU jobs on the Sparks have stopped and memory is verified free, start the RPC worker first on spark-b, from /home/andy/ai/apps/llama.cpp/build/bin:

LD_LIBRARY_PATH="$PWD" setsid nohup ./ggml-rpc-server \
  -H 192.168.100.11 -p 50052 -c >/tmp/rpc-server.log 2>&1 </dev/null &

Then on spark-a, from the same binary directory, restore the recorded command:

MODEL_DIR=/home/andy/ai/models/gguf/Qwen3.8-Flash-Next-Uncensored
LD_LIBRARY_PATH="$PWD" setsid nohup ./llama-server \
  -m "$MODEL_DIR/Qwen3.8-Flash-Next-Uncensored-Q8_0-00001-of-00005.gguf" \
  --mmproj "$MODEL_DIR/mmproj-Qwen3.8-Flash-Next-Uncensored-F16.gguf" \
  --rpc 192.168.100.11:50052 -ngl 99 -fa on -c 65536 -ts 70,30 \
  --reasoning off -lv 4 --host 100.114.103.103 --port 18090 \
  --temp 0.7 --top-p 0.8 --top-k 20 --min-p 0 --presence-penalty 1.5 \
  --alias Qwen3.8-Flash-Next-Uncensored-Q8_0-fast \
  >>/home/andy/ai/logs/qfn-server-fast.log 2>&1 </dev/null &

Do not restore it while OpenSysOne is using the Sparks. Exact snapshot commands above retain the prior bindings and tensor split; no network changes are needed.

Setup and live runs

Setup logs and transfer process records are under /home/andy/ai/opensysone/fleet-20260916 on GX10. Live campaign paths and supervisor controls are recorded below.

The isolated Python environments on both Sparks are ready at /home/andy/ai/envs/opensysone/bin/python. Each passed 21,368 file hashes and 55 exact package versions, CPU autograd and both Qwen model-class imports. The stack matches GX10: Python 3.12.3, torch 2.11.0+cu130, Transformers 5.15.0, NumPy 2.5.2 and PyArrow 25.0.1. Full package/file provenance and the source archive are in ~/ai/opensysone/bootstrap on each host. The existing environments were not modified. The archive crossed the existing Spark ConnectX link at 544.6 MB/s; that is a file-transfer observation, not a distributed-training measurement.

The first 32 CPU tests passed on source 6e080e2, including warm initialization, optimizer/RNG resume, graceful stopping, deadline handling, candidate selection and API deployment. Spark source was cloned cleanly at e047343; that commit adds the candidate preparation script and setup evidence. Small proofs, including all model-file SHA-256 hashes, are in results/20260916-fleet-setup/.

The selection revision passes 37 tests, plus the updated real-Adam CPU integration for changing the selection criterion without changing weights or optimizer steps. On the same 512 validation decisions, step 128 has 89.0625% accuracy versus step 40's 87.5%. Raw NLL worsens (0.442683 versus 0.395661), while four-fold crossfit NLL improves (0.318518 versus 0.359522). A source-group bootstrap with temperature refitting supports the crossfit improvement; the accuracy gain alone is uncertain. This is a validation-driven policy revision. The diagnostic, fixed policy and uncertainty are recorded in selection-diagnostic.json in the setup evidence directory. Original raw-selected artifacts are preserved. Fresh step-178 validation then improved to 90.4297% accuracy / 0.303825 crossfit NLL (raw NLL 0.404198), and the durable best was promoted. Independent held-out testing remains pending.

GX10's original campaign saved a complete step-128 checkpoint before its supervisor stopped it. The old trainer was killed after the 30-second termination grace while performing final correctness checks (training exit -9, supervisor status interrupted); no optimizer progress was lost and no test evaluation ran. Checkpoint SHA-256 is 62de5222d2b5f3541d0fa1890eacfa5d647dbb0ab67e43d1013cc3b73156e0b3. The new source skips final checks explicitly on a requested stop, preserving the previously completed correctness proof without reporting a false final pass.

The intermediate main run 20260916T192239Z-24h (source 6e080e2) restored step 128 exactly, then stopped gracefully at step 178 with training exit 0. The selection-parent artifact preserves that checkpoint byte-for-byte and changes only inherited-best selection metadata. selection_migration.json records the hashes and unchanged weights. The old campaigns are stopped; do not restart them.

Active training and coordinator

Active trainer source: 4a60423. Every plan is train_only=true, with selection policy crossfit_temperature_nll_v1, cutoff 2026-09-17 16:00 UTC and final deadline 18:16:10 UTC. All active jobs have OOM adjustment 0. Run IDs are relative to /home/andy/ai/opensysone/runs on the listed machine.

Machine Campaign Supervisor / trainer Startup proof
GX10 20260916T193741Z-24h 1096483 / 1096505 Step-178 state exact; fresh validation selected step 178
spark-a 20260916T194258Z-24h 327084 / 327116 Pilot and resume both reproduce all 512 predictions; new finite updates
spark-b 20260916T193803Z-24h exited 0 / 0 6,000 steps; best 2,000; final correctness passed
spark-b 20260917T023137Z-24h 483974 / 484001 Eight-step pilot and GPU/HTTP verifier exited 0; exact optimizer/RNG resume

Spark B's preparation 20260916T193722Z-spark-b-prepare exited 0. Real verification passed restored-optimizer gradients at the longest 700-token input (8.123 GiB peak), 1,024-token HTTP inference, authentication/length errors and 255 choices (17.73 seconds for the measured request). Spark A's preparation 20260916T192906Z-spark-a-prepare exited 0. Its pilot completed eight finite updates. Fresh verification reproduced 16 decisions exactly, restored Adam for finite longest-input gradients (15.624 GiB peak), matched direct/HTTP at 1,023 tokens, returned expected 401/422 errors, and passed 255 choices (43.31 seconds). These are wiring timings, not latency percentiles. The long campaign then reproduced all 512 step-8 predictions exactly, with weights/Adam/Python/torch/CUDA RNG unchanged. The fixed crossfit criterion chose step 8 (0.358235 NLL); subsequent optimizer updates are finite.

The fleet coordinator is 20260916T194403396250Z-fleet on GX10, PID 1425784, source fed722e. It is running in waiting_for_selection, has OOM adjustment 0, owns .fleet.lock, and waits until the cutoff before stopping the exact recorded candidate jobs. Its own plan.json is the authoritative plan; ~/ai/opensysone/runs/LAST_FLEET_CAMPAIGN points to this run. Its exit status is pending. The shutdown race with naturally finishing training jobs is fixed and covered by a regression test; all nine fleet tests pass. Stable checkpoint hashes, the identical 512 decision IDs, source/model/data compatibility, numerical evidence and recomputed selection scores must pass before a candidate is eligible. An unreachable remote is excluded; an unverified GX10 stop blocks finalization. Selection is written once to selection.json before any reserved-data predictions. Finalization acquires .smoke.lock and runs only the selected model on GX10.

Historical startup snapshot 2026-09-16T19:50:01.290313+00:00: GX10 step 217, spark-a step 17, spark-b step 200. All logged losses/gradients are finite, peak CUDA allocation stays below 16 GiB, and all six supervisor/trainer OOM adjustments were 0. All three jobs and the coordinator were running then; the 2B campaign has since completed, as recorded above. See results/20260916-fleet-setup/fleet-current-status.json.

Inspect and stop

For a compact status table across all registered hosts, run python3 scripts/fleet_status.py on GX10; add --json for saved evidence. It handles unavailable hosts and trailing partial training records independently, loads no models and changes no processes. A new setup pilot appears in this table only after its long campaign is registered in the fleet plan.

From the project directory on the appropriate host:

~/ai/envs/opensysone/bin/python scripts/campaign_status.py \
  --campaign /home/andy/ai/opensysone/runs/20260916T193741Z-24h

Substitute the host's exact run path. Add --stop to request a command-verified stop; confirm supervisor and child have exited before loading another model. Read training.log, training/training.jsonl, state.json and plan.json for progress, heartbeat, exact commands and deadlines. Exit codes are pending until completion. The last logged update may be newer than a durable checkpoint.

For the fleet coordinator, use its exact path with:

~/ai/envs/opensysone/bin/python scripts/fleet_campaign.py \
  --campaign /home/andy/ai/opensysone/runs/20260916T194403396250Z-fleet --status

Use --stop instead of --status to stop the coordinator or its resulting API. Stopping the waiting coordinator does not stop the independent training jobs. Stop each candidate explicitly if the whole fleet must stop. The user has already authorized those operations; no new permission is needed.

Recovery

After a candidate stops and its GPU/lock are free, resume on that host:

~/ai/envs/opensysone/bin/python scripts/launch_24h.py \
  --pilot /absolute/old/campaign/training \
  --train-only --training-deadline 2026-09-17T16:00:00Z \
  --deadline 2026-09-17T18:16:10Z --inference-max-tokens 1024 \
  --selection-metric crossfit_temperature_nll_v1

The new campaign restores optimizer/RNG and preserves the selected best. Keep its parent directory intact for evidence. Before selection is frozen, a replacement campaign must also replace that candidate's campaign and training paths in the coordinator's own plan.json. Stop the waiting coordinator first, edit, then run scripts/fleet_campaign.py --campaign /absolute/fleet/run --resume. The in-memory plan does not reload while it is running. Preserve all absolute deadlines. After selection.json exists, recover that same frozen winner; never reselect using partial test results. Completed evaluation may be reused only if checkpoint/data/source hashes still match.

Success requires fleet exit_code=0, complete evaluation/metrics.json, matching evaluation/model.pt hash, a successful api_probe.json, and api_ready=true. The deployment pointer is /home/andy/ai/opensysone/deploy/current.json and the resulting API is http://127.0.0.1:18081/v1/systemone. See JEV_HARNESS.md; hosted Jev still needs TYPESAFE_API_KEY.