Download source/FLEET_RUN.md from andyshu/opensysone: direct link, hf CLI and curl.
- Browser
- Download file 14.6 kB
-
https://huggingface.co/andyshu/opensysone/resolve/58f289696f58962a8ec98293d7b1abf9fd0c6b8b/source/FLEET_RUN.md
- Command line
-
hf download hf://andyshu/opensysone@58f289696f58962a8ec98293d7b1abf9fd0c6b8b/source/FLEET_RUN.md
-
curl -L -o FLEET_RUN.md https://huggingface.co/andyshu/opensysone/resolve/58f289696f58962a8ec98293d7b1abf9fd0c6b8b/source/FLEET_RUN.md
Three-machine campaign
Latest, 2026-09-17 07:22 UTC: GX10 now runs expanded-data 4B candidate
20260917T072142Z-24h, from frozen source worktree
/home/andy/ai/opensysone/source/expanded-24b8ccf. Supervisor/trainer PIDs at
launch are 1630617 / 1630638, OOM adjustment 0. The original GX10 campaign stopped
cleanly at step 4,380, preserving best 2,500. Both Spark 4B runs continue, and the
2B remains completed. The existing fleet now has five candidates, including
the explicit v2 dataset override; coordinator PID 1630841 was restarted after
the plan update. See EXPANDED_DATA.md for verified evidence,
dataset provenance and exact stop/resume commands. Historical startup rows below
describe the original fleet and are superseded by this update.
2026-09-17 continuation: NEXT_STEPS.md records overnight findings and the two-Spark assessment. The original 2B campaign finished with exit 0 at step 6,000; its best is step 2,000. Spark B is now running a 4B refinement campaign from GX10 best step 2,500. The four-candidate fleet plan retains the completed 2B and includes the verified refinement campaign.
On 2026-09-16, the user assigned GX10 and both Sparks to OpenSysOne and
explicitly authorized stopping existing processes. Key authentication from GX10
works as andy on both 192.168.8.111 (spark-a / spark-d1b4) and
192.168.8.204 (spark-b / spark-3e2a). No additional credentials are needed.
The delivery deadline stays 2026-09-17 18:16:10 UTC / 19:16:10 BST. Training candidates will stop at 16:00 UTC for validation-only model selection followed by calibration, untouched test/holdout evaluation and the local Jev-compatible API. Each training process keeps the 16 GiB CUDA cap. Independent experiments use all three GPUs without requiring an unverified distributed training stack or a ConnectX connection to GX10.
| Machine | Experiment | Current state |
|---|---|---|
| GX10 | Existing 4B, learning rate 0.0001, exact checkpoint continuation | Training resumed at step 178; new validation selected step 178 |
| spark-a | 4B initialized from verified pilot weights, learning rate 0.00003, 7,500-step cosine schedule | Verification passed; long training campaign running |
| spark-b | 2B completed; new 4B refinement at LR 0.00001 / seed 432 | Verified campaign 20260917T023137Z-24h running |
All candidates use the frozen public dataset and the same 512 validation
decisions. Selection now uses four-fold temperature-crossfit macro-family NLL
from durable checkpoints (crossfit_temperature_nll_v1, fixed seed 431). Source
groups stay in one fold; each fold's temperature is fitted on the other three.
This avoids discarding a better classifier because of recoverable overconfidence.
Raw NLL and accuracy are reported separately. Reserved calibration/test/Social
IQA predictions do not choose the winner, and the serving temperature is fitted
afresh on reserved calibration after selection.
Warm initialization is a new experiment with a fresh optimizer, distinguished
from an exact resume in its provenance. Model paths remain under
~/ai/models/opensysone; data, checkpoints and setup artifacts under
~/ai/opensysone. Source and small evidence alone belong in this repository.
Reclaimed serving pair
At 19:14 UTC, the exact verified Qwen serving process on spark-a (PID 154121) received SIGTERM and exited. After its memory was released, the exact verified RPC worker on spark-b (PID 114512) received SIGTERM and exited. No force kill was needed. Both GPU process lists were empty. Available memory afterward was 126,665,596,928 bytes on spark-a and 126,821,765,120 bytes on spark-b. The existing model files and RPC cache were preserved.
Full command/cwd/process identity, post-stop memory and GPU inspection are saved
on GX10 in ~/ai/opensysone/fleet-20260916/spark-a-service-stop.json and
spark-b-service-stop.json. Shared Python, firewall, network interfaces, swap,
earlyoom and clocks were not changed. Spark-a still has swap enabled and no
earlyoom; the project uses bounded allocations, host-memory checks and
oom_score_adj=0 without changing that system configuration.
To restore the old serving pair after all training/evaluation/API GPU jobs on
the Sparks have stopped and memory is verified free, start the RPC worker first
on spark-b, from /home/andy/ai/apps/llama.cpp/build/bin:
LD_LIBRARY_PATH="$PWD" setsid nohup ./ggml-rpc-server \
-H 192.168.100.11 -p 50052 -c >/tmp/rpc-server.log 2>&1 </dev/null &
Then on spark-a, from the same binary directory, restore the recorded command:
MODEL_DIR=/home/andy/ai/models/gguf/Qwen3.8-Flash-Next-Uncensored
LD_LIBRARY_PATH="$PWD" setsid nohup ./llama-server \
-m "$MODEL_DIR/Qwen3.8-Flash-Next-Uncensored-Q8_0-00001-of-00005.gguf" \
--mmproj "$MODEL_DIR/mmproj-Qwen3.8-Flash-Next-Uncensored-F16.gguf" \
--rpc 192.168.100.11:50052 -ngl 99 -fa on -c 65536 -ts 70,30 \
--reasoning off -lv 4 --host 100.114.103.103 --port 18090 \
--temp 0.7 --top-p 0.8 --top-k 20 --min-p 0 --presence-penalty 1.5 \
--alias Qwen3.8-Flash-Next-Uncensored-Q8_0-fast \
>>/home/andy/ai/logs/qfn-server-fast.log 2>&1 </dev/null &
Do not restore it while OpenSysOne is using the Sparks. Exact snapshot commands above retain the prior bindings and tensor split; no network changes are needed.
Setup and live runs
Setup logs and transfer process records are under
/home/andy/ai/opensysone/fleet-20260916 on GX10. Live campaign paths and supervisor controls are recorded below.
The isolated Python environments on both Sparks are ready at
/home/andy/ai/envs/opensysone/bin/python. Each passed 21,368 file hashes and
55 exact package versions, CPU autograd and both Qwen model-class imports.
The stack matches GX10: Python 3.12.3, torch 2.11.0+cu130, Transformers 5.15.0,
NumPy 2.5.2 and PyArrow 25.0.1. Full package/file provenance and the source archive
are in ~/ai/opensysone/bootstrap on each host. The existing environments were
not modified. The archive crossed the existing Spark ConnectX link at 544.6 MB/s;
that is a file-transfer observation, not a distributed-training measurement.
The first 32 CPU tests passed on source 6e080e2, including warm initialization,
optimizer/RNG resume, graceful stopping, deadline handling, candidate selection
and API deployment. Spark source was cloned cleanly at e047343; that commit adds
the candidate preparation script and setup evidence. Small proofs, including
all model-file SHA-256 hashes, are in results/20260916-fleet-setup/.
The selection revision passes 37 tests, plus the updated real-Adam CPU
integration for changing the selection criterion without changing weights or
optimizer steps. On the same 512 validation decisions, step 128 has 89.0625%
accuracy versus step 40's 87.5%. Raw NLL worsens (0.442683 versus 0.395661), while
four-fold crossfit NLL improves (0.318518 versus 0.359522). A source-group bootstrap
with temperature refitting supports the crossfit improvement; the accuracy gain
alone is uncertain. This is a validation-driven policy revision. The diagnostic,
fixed policy and uncertainty are recorded in selection-diagnostic.json in the
setup evidence directory. Original raw-selected artifacts are preserved. Fresh step-178 validation then
improved to 90.4297% accuracy / 0.303825 crossfit NLL (raw NLL 0.404198),
and the durable best was promoted. Independent held-out testing remains pending.
GX10's original campaign saved a complete step-128 checkpoint before its
supervisor stopped it. The old trainer was killed after the 30-second termination
grace while performing final correctness checks (training exit -9, supervisor
status interrupted); no optimizer progress was lost and no test evaluation ran.
Checkpoint SHA-256 is
62de5222d2b5f3541d0fa1890eacfa5d647dbb0ab67e43d1013cc3b73156e0b3.
The new source skips final checks explicitly on a requested stop, preserving the
previously completed correctness proof without reporting a false final pass.
The intermediate main run 20260916T192239Z-24h (source 6e080e2) restored
step 128 exactly, then stopped gracefully at step 178 with training exit 0.
The selection-parent artifact preserves that checkpoint byte-for-byte and changes
only inherited-best selection metadata. selection_migration.json records the
hashes and unchanged weights. The old campaigns are stopped; do not restart them.
Active training and coordinator
Active trainer source: 4a60423. Every plan is train_only=true, with
selection policy crossfit_temperature_nll_v1, cutoff 2026-09-17 16:00 UTC
and final deadline 18:16:10 UTC. All active jobs have OOM adjustment 0.
Run IDs are relative to /home/andy/ai/opensysone/runs on the listed machine.
| Machine | Campaign | Supervisor / trainer | Startup proof |
|---|---|---|---|
| GX10 | 20260916T193741Z-24h |
1096483 / 1096505 | Step-178 state exact; fresh validation selected step 178 |
| spark-a | 20260916T194258Z-24h |
327084 / 327116 | Pilot and resume both reproduce all 512 predictions; new finite updates |
| spark-b | 20260916T193803Z-24h |
exited 0 / 0 | 6,000 steps; best 2,000; final correctness passed |
| spark-b | 20260917T023137Z-24h |
483974 / 484001 | Eight-step pilot and GPU/HTTP verifier exited 0; exact optimizer/RNG resume |
Spark B's preparation 20260916T193722Z-spark-b-prepare exited 0. Real
verification passed restored-optimizer gradients at the longest 700-token input
(8.123 GiB peak), 1,024-token HTTP inference, authentication/length errors and
255 choices (17.73 seconds for the measured request). Spark A's preparation
20260916T192906Z-spark-a-prepare exited 0. Its pilot
completed eight finite updates. Fresh verification reproduced 16 decisions
exactly, restored Adam for finite longest-input gradients (15.624 GiB peak),
matched direct/HTTP at 1,023 tokens, returned expected 401/422 errors, and passed
255 choices (43.31 seconds). These are wiring timings, not latency percentiles.
The long campaign then reproduced all 512 step-8 predictions exactly, with
weights/Adam/Python/torch/CUDA RNG unchanged. The fixed crossfit criterion chose
step 8 (0.358235 NLL); subsequent optimizer updates are finite.
The fleet coordinator is 20260916T194403396250Z-fleet on GX10, PID
1425784, source fed722e. It is running in waiting_for_selection, has
OOM adjustment 0, owns .fleet.lock, and waits until the cutoff before stopping
the exact recorded candidate jobs. Its own plan.json is the authoritative plan;
~/ai/opensysone/runs/LAST_FLEET_CAMPAIGN points to this run. Its exit status is
pending. The shutdown race with naturally finishing training jobs is fixed and
covered by a regression test; all nine fleet tests pass. Stable checkpoint
hashes, the identical 512 decision IDs,
source/model/data compatibility, numerical evidence and recomputed selection
scores must pass before a candidate is eligible. An unreachable remote is
excluded; an unverified GX10 stop blocks finalization. Selection is written once
to selection.json before any reserved-data predictions. Finalization acquires
.smoke.lock and runs only the selected model on GX10.
Historical startup snapshot 2026-09-16T19:50:01.290313+00:00: GX10 step 217, spark-a step 17, spark-b step 200. All logged losses/gradients are finite,
peak CUDA allocation stays below 16 GiB, and all six supervisor/trainer OOM
adjustments were 0. All three jobs and the coordinator were running then; the
2B campaign has since completed, as recorded above. See results/20260916-fleet-setup/fleet-current-status.json.
Inspect and stop
For a compact status table across all registered hosts, run
python3 scripts/fleet_status.py on GX10; add --json for saved evidence.
It handles unavailable hosts and trailing partial training records independently,
loads no models and changes no processes. A new setup pilot appears in this table
only after its long campaign is registered in the fleet plan.
From the project directory on the appropriate host:
~/ai/envs/opensysone/bin/python scripts/campaign_status.py \
--campaign /home/andy/ai/opensysone/runs/20260916T193741Z-24h
Substitute the host's exact run path. Add --stop to request a command-verified
stop; confirm supervisor and child have exited before loading another model.
Read training.log, training/training.jsonl, state.json and plan.json for
progress, heartbeat, exact commands and deadlines. Exit codes are pending until
completion. The last logged update may be newer than a durable checkpoint.
For the fleet coordinator, use its exact path with:
~/ai/envs/opensysone/bin/python scripts/fleet_campaign.py \
--campaign /home/andy/ai/opensysone/runs/20260916T194403396250Z-fleet --status
Use --stop instead of --status to stop the coordinator or its resulting API.
Stopping the waiting coordinator does not stop the independent training jobs.
Stop each candidate explicitly if the whole fleet must stop. The user has already
authorized those operations; no new permission is needed.
Recovery
After a candidate stops and its GPU/lock are free, resume on that host:
~/ai/envs/opensysone/bin/python scripts/launch_24h.py \
--pilot /absolute/old/campaign/training \
--train-only --training-deadline 2026-09-17T16:00:00Z \
--deadline 2026-09-17T18:16:10Z --inference-max-tokens 1024 \
--selection-metric crossfit_temperature_nll_v1
The new campaign restores optimizer/RNG and preserves the selected best. Keep its
parent directory intact for evidence. Before selection is frozen, a replacement
campaign must also replace that candidate's campaign and training paths in
the coordinator's own plan.json. Stop the waiting coordinator first, edit,
then run scripts/fleet_campaign.py --campaign /absolute/fleet/run --resume.
The in-memory plan does not reload while it is running. Preserve all absolute
deadlines. After selection.json exists, recover that same frozen winner; never
reselect using partial test results. Completed evaluation may be reused only if
checkpoint/data/source hashes still match.
Success requires fleet exit_code=0, complete evaluation/metrics.json, matching
evaluation/model.pt hash, a successful api_probe.json, and api_ready=true.
The deployment pointer is /home/andy/ai/opensysone/deploy/current.json and the
resulting API is http://127.0.0.1:18081/v1/systemone. See
JEV_HARNESS.md; hosted Jev still needs TYPESAFE_API_KEY.