opensysone / source /FLEET_RUN.md
andyshu's picture
Back up verified OpenSysOne training snapshot and pinned source
2d5c26a verified
|
Raw History Blame
13.9 kB
# Three-machine campaign
**2026-09-17 continuation:** [NEXT_STEPS.md](NEXT_STEPS.md) records overnight
findings and the two-Spark assessment. The original 2B campaign finished with exit
0 at step 6,000; its best is step 2,000. Spark B is now running a 4B refinement
campaign from GX10 best step 2,500. The four-candidate fleet plan retains the
completed 2B and includes the verified refinement campaign.
On **2026-09-16**, the user assigned GX10 and both Sparks to OpenSysOne and
explicitly authorized stopping existing processes. Key authentication from GX10
works as `andy` on both `192.168.8.111` (spark-a / spark-d1b4) and
`192.168.8.204` (spark-b / spark-3e2a). No additional credentials are needed.
The delivery deadline stays **2026-09-17 18:16:10 UTC / 19:16:10 BST**. Training
candidates will stop at **16:00 UTC** for validation-only model selection followed
by calibration, untouched test/holdout evaluation and the local Jev-compatible API.
Each training process keeps the **16 GiB CUDA cap**. Independent experiments use
all three GPUs without requiring an unverified distributed training stack or a
ConnectX connection to GX10.
| Machine | Experiment | Current state |
| --- | --- | --- |
| GX10 | Existing 4B, learning rate 0.0001, exact checkpoint continuation | Training resumed at step 178; new validation selected step 178 |
| spark-a | 4B initialized from verified pilot weights, learning rate 0.00003, 7,500-step cosine schedule | Verification passed; long training campaign running |
| spark-b | 2B completed; new 4B refinement at LR 0.00001 / seed 432 | Verified campaign `20260917T023137Z-24h` running |
All candidates use the frozen public dataset and the same 512 validation
decisions. Selection now uses **four-fold temperature-crossfit macro-family NLL**
from durable checkpoints (`crossfit_temperature_nll_v1`, fixed seed 431). Source
groups stay in one fold; each fold's temperature is fitted on the other three.
This avoids discarding a better classifier because of recoverable overconfidence.
Raw NLL and accuracy are reported separately. Reserved calibration/test/Social
IQA predictions do not choose the winner, and the serving temperature is fitted
afresh on reserved calibration after selection.
Warm initialization is a new experiment with a fresh optimizer, distinguished
from an exact resume in its provenance. Model paths remain under
`~/ai/models/opensysone`; data, checkpoints and setup artifacts under
`~/ai/opensysone`. Source and small evidence alone belong in this repository.
## Reclaimed serving pair
At **19:14 UTC**, the exact verified Qwen serving process on spark-a (PID 154121)
received SIGTERM and exited. After its memory was released, the exact verified
RPC worker on spark-b (PID 114512) received SIGTERM and exited. No force kill was
needed. Both GPU process lists were empty. Available memory afterward was
126,665,596,928 bytes on spark-a and 126,821,765,120 bytes on spark-b. The existing
model files and RPC cache were preserved.
Full command/cwd/process identity, post-stop memory and GPU inspection are saved
on GX10 in `~/ai/opensysone/fleet-20260916/spark-a-service-stop.json` and
`spark-b-service-stop.json`. Shared Python, firewall, network interfaces, swap,
earlyoom and clocks were not changed. Spark-a still has swap enabled and no
earlyoom; the project uses bounded allocations, host-memory checks and
`oom_score_adj=0` without changing that system configuration.
To restore the old serving pair **after all training/evaluation/API GPU jobs on
the Sparks have stopped and memory is verified free**, start the RPC worker first
on spark-b, from `/home/andy/ai/apps/llama.cpp/build/bin`:
```bash
LD_LIBRARY_PATH="$PWD" setsid nohup ./ggml-rpc-server \
-H 192.168.100.11 -p 50052 -c >/tmp/rpc-server.log 2>&1 </dev/null &
```
Then on spark-a, from the same binary directory, restore the recorded command:
```bash
MODEL_DIR=/home/andy/ai/models/gguf/Qwen3.8-Flash-Next-Uncensored
LD_LIBRARY_PATH="$PWD" setsid nohup ./llama-server \
-m "$MODEL_DIR/Qwen3.8-Flash-Next-Uncensored-Q8_0-00001-of-00005.gguf" \
--mmproj "$MODEL_DIR/mmproj-Qwen3.8-Flash-Next-Uncensored-F16.gguf" \
--rpc 192.168.100.11:50052 -ngl 99 -fa on -c 65536 -ts 70,30 \
--reasoning off -lv 4 --host 100.114.103.103 --port 18090 \
--temp 0.7 --top-p 0.8 --top-k 20 --min-p 0 --presence-penalty 1.5 \
--alias Qwen3.8-Flash-Next-Uncensored-Q8_0-fast \
>>/home/andy/ai/logs/qfn-server-fast.log 2>&1 </dev/null &
```
Do not restore it while OpenSysOne is using the Sparks. Exact snapshot commands
above retain the prior bindings and tensor split; no network changes are needed.
## Setup and live runs
Setup logs and transfer process records are under
`/home/andy/ai/opensysone/fleet-20260916` on GX10. Live campaign paths and supervisor controls are recorded below.
The isolated Python environments on both Sparks are ready at
`/home/andy/ai/envs/opensysone/bin/python`. Each passed **21,368 file hashes and
55 exact package versions**, CPU autograd and both Qwen model-class imports.
The stack matches GX10: Python 3.12.3, torch 2.11.0+cu130, Transformers 5.15.0,
NumPy 2.5.2 and PyArrow 25.0.1. Full package/file provenance and the source archive
are in `~/ai/opensysone/bootstrap` on each host. The existing environments were
not modified. The archive crossed the existing Spark ConnectX link at 544.6 MB/s;
that is a file-transfer observation, not a distributed-training measurement.
The first **32 CPU tests passed** on source `6e080e2`, including warm initialization,
optimizer/RNG resume, graceful stopping, deadline handling, candidate selection
and API deployment. Spark source was cloned cleanly at `e047343`; that commit adds
the candidate preparation script and setup evidence. Small proofs, including
all model-file SHA-256 hashes, are in `results/20260916-fleet-setup/`.
The selection revision passes **37 tests**, plus the updated real-Adam CPU
integration for changing the selection criterion without changing weights or
optimizer steps. On the same 512 validation decisions, step 128 has 89.0625%
accuracy versus step 40's 87.5%. Raw NLL worsens (0.442683 versus 0.395661), while
four-fold crossfit NLL improves (0.318518 versus 0.359522). A source-group bootstrap
with temperature refitting supports the crossfit improvement; the accuracy gain
alone is uncertain. This is a validation-driven policy revision. The diagnostic,
fixed policy and uncertainty are recorded in `selection-diagnostic.json` in the
setup evidence directory. Original raw-selected artifacts are preserved. Fresh step-178 validation then
improved to **90.4297% accuracy / 0.303825 crossfit NLL** (raw NLL 0.404198),
and the durable best was promoted. Independent held-out testing remains pending.
GX10's original campaign saved a complete **step-128** checkpoint before its
supervisor stopped it. The old trainer was killed after the 30-second termination
grace while performing final correctness checks (**training exit -9**, supervisor
status `interrupted`); no optimizer progress was lost and no test evaluation ran.
Checkpoint SHA-256 is
`62de5222d2b5f3541d0fa1890eacfa5d647dbb0ab67e43d1013cc3b73156e0b3`.
The new source skips final checks explicitly on a requested stop, preserving the
previously completed correctness proof without reporting a false final pass.
The intermediate main run `20260916T192239Z-24h` (source `6e080e2`) restored
step 128 exactly, then stopped gracefully at step 178 with **training exit 0**.
The selection-parent artifact preserves that checkpoint byte-for-byte and changes
only inherited-best selection metadata. `selection_migration.json` records the
hashes and unchanged weights. The old campaigns are stopped; do not restart them.
## Active training and coordinator
Active trainer source: **`4a60423`**. Every plan is `train_only=true`, with
selection policy `crossfit_temperature_nll_v1`, cutoff **2026-09-17 16:00 UTC**
and final deadline **18:16:10 UTC**. All active jobs have OOM adjustment 0.
Run IDs are relative to `/home/andy/ai/opensysone/runs` on the listed machine.
| Machine | Campaign | Supervisor / trainer | Startup proof |
| --- | --- | --- | --- |
| GX10 | `20260916T193741Z-24h` | 1096483 / 1096505 | Step-178 state exact; fresh validation selected step 178 |
| spark-a | `20260916T194258Z-24h` | 327084 / 327116 | Pilot and resume both reproduce all 512 predictions; new finite updates |
| spark-b | `20260916T193803Z-24h` | exited 0 / 0 | 6,000 steps; best 2,000; final correctness passed |
| spark-b | `20260917T023137Z-24h` | 483974 / 484001 | Eight-step pilot and GPU/HTTP verifier exited 0; exact optimizer/RNG resume |
Spark B's preparation `20260916T193722Z-spark-b-prepare` exited **0**. Real
verification passed restored-optimizer gradients at the longest 700-token input
(8.123 GiB peak), 1,024-token HTTP inference, authentication/length errors and
255 choices (17.73 seconds for the measured request). Spark A's preparation
`20260916T192906Z-spark-a-prepare` exited **0**. Its pilot
completed eight finite updates. Fresh verification reproduced 16 decisions
exactly, restored Adam for finite longest-input gradients (15.624 GiB peak),
matched direct/HTTP at 1,023 tokens, returned expected 401/422 errors, and passed
255 choices (43.31 seconds). These are wiring timings, not latency percentiles.
The long campaign then reproduced all 512 step-8 predictions exactly, with
weights/Adam/Python/torch/CUDA RNG unchanged. The fixed crossfit criterion chose
step 8 (0.358235 NLL); subsequent optimizer updates are finite.
The fleet coordinator is **`20260916T194403396250Z-fleet` on GX10**, PID
**1425784**, source **`fed722e`**. It is running in `waiting_for_selection`, has
OOM adjustment 0, owns `.fleet.lock`, and waits until the cutoff before stopping
the exact recorded candidate jobs. Its own `plan.json` is the authoritative plan;
`~/ai/opensysone/runs/LAST_FLEET_CAMPAIGN` points to this run. Its exit status is
pending. The shutdown race with naturally finishing training jobs is fixed and
covered by a regression test; all nine fleet tests pass. Stable checkpoint
hashes, the identical 512 decision IDs,
source/model/data compatibility, numerical evidence and recomputed selection
scores must pass before a candidate is eligible. An unreachable remote is
excluded; an unverified GX10 stop blocks finalization. Selection is written once
to `selection.json` before any reserved-data predictions. Finalization acquires
`.smoke.lock` and runs only the selected model on GX10.
Historical startup snapshot **2026-09-16T19:50:01.290313+00:00**: GX10 step 217, spark-a step 17, spark-b step 200. All logged losses/gradients are finite,
peak CUDA allocation stays below 16 GiB, and all six supervisor/trainer OOM
adjustments were 0. All three jobs and the coordinator were running then; the
2B campaign has since completed, as recorded above. See `results/20260916-fleet-setup/fleet-current-status.json`.
### Inspect and stop
For a compact status table across all registered hosts, run
`python3 scripts/fleet_status.py` on GX10; add `--json` for saved evidence.
It handles unavailable hosts and trailing partial training records independently,
loads no models and changes no processes. A new setup pilot appears in this table
only after its long campaign is registered in the fleet plan.
From the project directory on the appropriate host:
```bash
~/ai/envs/opensysone/bin/python scripts/campaign_status.py \
--campaign /home/andy/ai/opensysone/runs/20260916T193741Z-24h
```
Substitute the host's exact run path. Add `--stop` to request a command-verified
stop; confirm supervisor and child have exited before loading another model.
Read `training.log`, `training/training.jsonl`, `state.json` and `plan.json` for
progress, heartbeat, exact commands and deadlines. Exit codes are pending until
completion. The last logged update may be newer than a durable checkpoint.
For the fleet coordinator, use its exact path with:
```bash
~/ai/envs/opensysone/bin/python scripts/fleet_campaign.py \
--campaign /home/andy/ai/opensysone/runs/20260916T194403396250Z-fleet --status
```
Use `--stop` instead of `--status` to stop the coordinator or its resulting API.
**Stopping the waiting coordinator does not stop the independent training jobs.**
Stop each candidate explicitly if the whole fleet must stop. The user has already
authorized those operations; no new permission is needed.
### Recovery
After a candidate stops and its GPU/lock are free, resume on that host:
```bash
~/ai/envs/opensysone/bin/python scripts/launch_24h.py \
--pilot /absolute/old/campaign/training \
--train-only --training-deadline 2026-09-17T16:00:00Z \
--deadline 2026-09-17T18:16:10Z --inference-max-tokens 1024 \
--selection-metric crossfit_temperature_nll_v1
```
The new campaign restores optimizer/RNG and preserves the selected best. Keep its
parent directory intact for evidence. Before selection is frozen, a replacement
campaign must also replace that candidate's `campaign` and `training` paths in
the **coordinator's own `plan.json`**. Stop the waiting coordinator first, edit,
then run `scripts/fleet_campaign.py --campaign /absolute/fleet/run --resume`.
The in-memory plan does not reload while it is running. Preserve all absolute
deadlines. After `selection.json` exists, recover that same frozen winner; never
reselect using partial test results. Completed evaluation may be reused only if
checkpoint/data/source hashes still match.
Success requires fleet `exit_code=0`, complete `evaluation/metrics.json`, matching
`evaluation/model.pt` hash, a successful `api_probe.json`, and `api_ready=true`.
The deployment pointer is `/home/andy/ai/opensysone/deploy/current.json` and the
resulting API is `http://127.0.0.1:18081/v1/systemone`. See
[JEV_HARNESS.md](JEV_HARNESS.md); hosted Jev still needs `TYPESAFE_API_KEY`.