|
Download source/FLEET_RUN.md from andyshu/opensysone: direct link, hf CLI and curl.
- Browser
- Download file 13.9 kB
-
https://huggingface.co/andyshu/opensysone/resolve/f2aba96e0a2b088ef8bc2206a12b20f79a85cc59/source/FLEET_RUN.md
- Command line
-
hf download hf://andyshu/opensysone@f2aba96e0a2b088ef8bc2206a12b20f79a85cc59/source/FLEET_RUN.md
-
curl -L -o FLEET_RUN.md https://huggingface.co/andyshu/opensysone/resolve/f2aba96e0a2b088ef8bc2206a12b20f79a85cc59/source/FLEET_RUN.md
13.9 kB
| # Three-machine campaign | |
| **2026-09-17 continuation:** [NEXT_STEPS.md](NEXT_STEPS.md) records overnight | |
| findings and the two-Spark assessment. The original 2B campaign finished with exit | |
| 0 at step 6,000; its best is step 2,000. Spark B is now running a 4B refinement | |
| campaign from GX10 best step 2,500. The four-candidate fleet plan retains the | |
| completed 2B and includes the verified refinement campaign. | |
| On **2026-09-16**, the user assigned GX10 and both Sparks to OpenSysOne and | |
| explicitly authorized stopping existing processes. Key authentication from GX10 | |
| works as `andy` on both `192.168.8.111` (spark-a / spark-d1b4) and | |
| `192.168.8.204` (spark-b / spark-3e2a). No additional credentials are needed. | |
| The delivery deadline stays **2026-09-17 18:16:10 UTC / 19:16:10 BST**. Training | |
| candidates will stop at **16:00 UTC** for validation-only model selection followed | |
| by calibration, untouched test/holdout evaluation and the local Jev-compatible API. | |
| Each training process keeps the **16 GiB CUDA cap**. Independent experiments use | |
| all three GPUs without requiring an unverified distributed training stack or a | |
| ConnectX connection to GX10. | |
| | Machine | Experiment | Current state | | |
| | --- | --- | --- | | |
| | GX10 | Existing 4B, learning rate 0.0001, exact checkpoint continuation | Training resumed at step 178; new validation selected step 178 | | |
| | spark-a | 4B initialized from verified pilot weights, learning rate 0.00003, 7,500-step cosine schedule | Verification passed; long training campaign running | | |
| | spark-b | 2B completed; new 4B refinement at LR 0.00001 / seed 432 | Verified campaign `20260917T023137Z-24h` running | | |
| All candidates use the frozen public dataset and the same 512 validation | |
| decisions. Selection now uses **four-fold temperature-crossfit macro-family NLL** | |
| from durable checkpoints (`crossfit_temperature_nll_v1`, fixed seed 431). Source | |
| groups stay in one fold; each fold's temperature is fitted on the other three. | |
| This avoids discarding a better classifier because of recoverable overconfidence. | |
| Raw NLL and accuracy are reported separately. Reserved calibration/test/Social | |
| IQA predictions do not choose the winner, and the serving temperature is fitted | |
| afresh on reserved calibration after selection. | |
| Warm initialization is a new experiment with a fresh optimizer, distinguished | |
| from an exact resume in its provenance. Model paths remain under | |
| `~/ai/models/opensysone`; data, checkpoints and setup artifacts under | |
| `~/ai/opensysone`. Source and small evidence alone belong in this repository. | |
| ## Reclaimed serving pair | |
| At **19:14 UTC**, the exact verified Qwen serving process on spark-a (PID 154121) | |
| received SIGTERM and exited. After its memory was released, the exact verified | |
| RPC worker on spark-b (PID 114512) received SIGTERM and exited. No force kill was | |
| needed. Both GPU process lists were empty. Available memory afterward was | |
| 126,665,596,928 bytes on spark-a and 126,821,765,120 bytes on spark-b. The existing | |
| model files and RPC cache were preserved. | |
| Full command/cwd/process identity, post-stop memory and GPU inspection are saved | |
| on GX10 in `~/ai/opensysone/fleet-20260916/spark-a-service-stop.json` and | |
| `spark-b-service-stop.json`. Shared Python, firewall, network interfaces, swap, | |
| earlyoom and clocks were not changed. Spark-a still has swap enabled and no | |
| earlyoom; the project uses bounded allocations, host-memory checks and | |
| `oom_score_adj=0` without changing that system configuration. | |
| To restore the old serving pair **after all training/evaluation/API GPU jobs on | |
| the Sparks have stopped and memory is verified free**, start the RPC worker first | |
| on spark-b, from `/home/andy/ai/apps/llama.cpp/build/bin`: | |
| ```bash | |
| LD_LIBRARY_PATH="$PWD" setsid nohup ./ggml-rpc-server \ | |
| -H 192.168.100.11 -p 50052 -c >/tmp/rpc-server.log 2>&1 </dev/null & | |
| ``` | |
| Then on spark-a, from the same binary directory, restore the recorded command: | |
| ```bash | |
| MODEL_DIR=/home/andy/ai/models/gguf/Qwen3.8-Flash-Next-Uncensored | |
| LD_LIBRARY_PATH="$PWD" setsid nohup ./llama-server \ | |
| -m "$MODEL_DIR/Qwen3.8-Flash-Next-Uncensored-Q8_0-00001-of-00005.gguf" \ | |
| --mmproj "$MODEL_DIR/mmproj-Qwen3.8-Flash-Next-Uncensored-F16.gguf" \ | |
| --rpc 192.168.100.11:50052 -ngl 99 -fa on -c 65536 -ts 70,30 \ | |
| --reasoning off -lv 4 --host 100.114.103.103 --port 18090 \ | |
| --temp 0.7 --top-p 0.8 --top-k 20 --min-p 0 --presence-penalty 1.5 \ | |
| --alias Qwen3.8-Flash-Next-Uncensored-Q8_0-fast \ | |
| >>/home/andy/ai/logs/qfn-server-fast.log 2>&1 </dev/null & | |
| ``` | |
| Do not restore it while OpenSysOne is using the Sparks. Exact snapshot commands | |
| above retain the prior bindings and tensor split; no network changes are needed. | |
| ## Setup and live runs | |
| Setup logs and transfer process records are under | |
| `/home/andy/ai/opensysone/fleet-20260916` on GX10. Live campaign paths and supervisor controls are recorded below. | |
| The isolated Python environments on both Sparks are ready at | |
| `/home/andy/ai/envs/opensysone/bin/python`. Each passed **21,368 file hashes and | |
| 55 exact package versions**, CPU autograd and both Qwen model-class imports. | |
| The stack matches GX10: Python 3.12.3, torch 2.11.0+cu130, Transformers 5.15.0, | |
| NumPy 2.5.2 and PyArrow 25.0.1. Full package/file provenance and the source archive | |
| are in `~/ai/opensysone/bootstrap` on each host. The existing environments were | |
| not modified. The archive crossed the existing Spark ConnectX link at 544.6 MB/s; | |
| that is a file-transfer observation, not a distributed-training measurement. | |
| The first **32 CPU tests passed** on source `6e080e2`, including warm initialization, | |
| optimizer/RNG resume, graceful stopping, deadline handling, candidate selection | |
| and API deployment. Spark source was cloned cleanly at `e047343`; that commit adds | |
| the candidate preparation script and setup evidence. Small proofs, including | |
| all model-file SHA-256 hashes, are in `results/20260916-fleet-setup/`. | |
| The selection revision passes **37 tests**, plus the updated real-Adam CPU | |
| integration for changing the selection criterion without changing weights or | |
| optimizer steps. On the same 512 validation decisions, step 128 has 89.0625% | |
| accuracy versus step 40's 87.5%. Raw NLL worsens (0.442683 versus 0.395661), while | |
| four-fold crossfit NLL improves (0.318518 versus 0.359522). A source-group bootstrap | |
| with temperature refitting supports the crossfit improvement; the accuracy gain | |
| alone is uncertain. This is a validation-driven policy revision. The diagnostic, | |
| fixed policy and uncertainty are recorded in `selection-diagnostic.json` in the | |
| setup evidence directory. Original raw-selected artifacts are preserved. Fresh step-178 validation then | |
| improved to **90.4297% accuracy / 0.303825 crossfit NLL** (raw NLL 0.404198), | |
| and the durable best was promoted. Independent held-out testing remains pending. | |
| GX10's original campaign saved a complete **step-128** checkpoint before its | |
| supervisor stopped it. The old trainer was killed after the 30-second termination | |
| grace while performing final correctness checks (**training exit -9**, supervisor | |
| status `interrupted`); no optimizer progress was lost and no test evaluation ran. | |
| Checkpoint SHA-256 is | |
| `62de5222d2b5f3541d0fa1890eacfa5d647dbb0ab67e43d1013cc3b73156e0b3`. | |
| The new source skips final checks explicitly on a requested stop, preserving the | |
| previously completed correctness proof without reporting a false final pass. | |
| The intermediate main run `20260916T192239Z-24h` (source `6e080e2`) restored | |
| step 128 exactly, then stopped gracefully at step 178 with **training exit 0**. | |
| The selection-parent artifact preserves that checkpoint byte-for-byte and changes | |
| only inherited-best selection metadata. `selection_migration.json` records the | |
| hashes and unchanged weights. The old campaigns are stopped; do not restart them. | |
| ## Active training and coordinator | |
| Active trainer source: **`4a60423`**. Every plan is `train_only=true`, with | |
| selection policy `crossfit_temperature_nll_v1`, cutoff **2026-09-17 16:00 UTC** | |
| and final deadline **18:16:10 UTC**. All active jobs have OOM adjustment 0. | |
| Run IDs are relative to `/home/andy/ai/opensysone/runs` on the listed machine. | |
| | Machine | Campaign | Supervisor / trainer | Startup proof | | |
| | --- | --- | --- | --- | | |
| | GX10 | `20260916T193741Z-24h` | 1096483 / 1096505 | Step-178 state exact; fresh validation selected step 178 | | |
| | spark-a | `20260916T194258Z-24h` | 327084 / 327116 | Pilot and resume both reproduce all 512 predictions; new finite updates | | |
| | spark-b | `20260916T193803Z-24h` | exited 0 / 0 | 6,000 steps; best 2,000; final correctness passed | | |
| | spark-b | `20260917T023137Z-24h` | 483974 / 484001 | Eight-step pilot and GPU/HTTP verifier exited 0; exact optimizer/RNG resume | | |
| Spark B's preparation `20260916T193722Z-spark-b-prepare` exited **0**. Real | |
| verification passed restored-optimizer gradients at the longest 700-token input | |
| (8.123 GiB peak), 1,024-token HTTP inference, authentication/length errors and | |
| 255 choices (17.73 seconds for the measured request). Spark A's preparation | |
| `20260916T192906Z-spark-a-prepare` exited **0**. Its pilot | |
| completed eight finite updates. Fresh verification reproduced 16 decisions | |
| exactly, restored Adam for finite longest-input gradients (15.624 GiB peak), | |
| matched direct/HTTP at 1,023 tokens, returned expected 401/422 errors, and passed | |
| 255 choices (43.31 seconds). These are wiring timings, not latency percentiles. | |
| The long campaign then reproduced all 512 step-8 predictions exactly, with | |
| weights/Adam/Python/torch/CUDA RNG unchanged. The fixed crossfit criterion chose | |
| step 8 (0.358235 NLL); subsequent optimizer updates are finite. | |
| The fleet coordinator is **`20260916T194403396250Z-fleet` on GX10**, PID | |
| **1425784**, source **`fed722e`**. It is running in `waiting_for_selection`, has | |
| OOM adjustment 0, owns `.fleet.lock`, and waits until the cutoff before stopping | |
| the exact recorded candidate jobs. Its own `plan.json` is the authoritative plan; | |
| `~/ai/opensysone/runs/LAST_FLEET_CAMPAIGN` points to this run. Its exit status is | |
| pending. The shutdown race with naturally finishing training jobs is fixed and | |
| covered by a regression test; all nine fleet tests pass. Stable checkpoint | |
| hashes, the identical 512 decision IDs, | |
| source/model/data compatibility, numerical evidence and recomputed selection | |
| scores must pass before a candidate is eligible. An unreachable remote is | |
| excluded; an unverified GX10 stop blocks finalization. Selection is written once | |
| to `selection.json` before any reserved-data predictions. Finalization acquires | |
| `.smoke.lock` and runs only the selected model on GX10. | |
| Historical startup snapshot **2026-09-16T19:50:01.290313+00:00**: GX10 step 217, spark-a step 17, spark-b step 200. All logged losses/gradients are finite, | |
| peak CUDA allocation stays below 16 GiB, and all six supervisor/trainer OOM | |
| adjustments were 0. All three jobs and the coordinator were running then; the | |
| 2B campaign has since completed, as recorded above. See `results/20260916-fleet-setup/fleet-current-status.json`. | |
| ### Inspect and stop | |
| For a compact status table across all registered hosts, run | |
| `python3 scripts/fleet_status.py` on GX10; add `--json` for saved evidence. | |
| It handles unavailable hosts and trailing partial training records independently, | |
| loads no models and changes no processes. A new setup pilot appears in this table | |
| only after its long campaign is registered in the fleet plan. | |
| From the project directory on the appropriate host: | |
| ```bash | |
| ~/ai/envs/opensysone/bin/python scripts/campaign_status.py \ | |
| --campaign /home/andy/ai/opensysone/runs/20260916T193741Z-24h | |
| ``` | |
| Substitute the host's exact run path. Add `--stop` to request a command-verified | |
| stop; confirm supervisor and child have exited before loading another model. | |
| Read `training.log`, `training/training.jsonl`, `state.json` and `plan.json` for | |
| progress, heartbeat, exact commands and deadlines. Exit codes are pending until | |
| completion. The last logged update may be newer than a durable checkpoint. | |
| For the fleet coordinator, use its exact path with: | |
| ```bash | |
| ~/ai/envs/opensysone/bin/python scripts/fleet_campaign.py \ | |
| --campaign /home/andy/ai/opensysone/runs/20260916T194403396250Z-fleet --status | |
| ``` | |
| Use `--stop` instead of `--status` to stop the coordinator or its resulting API. | |
| **Stopping the waiting coordinator does not stop the independent training jobs.** | |
| Stop each candidate explicitly if the whole fleet must stop. The user has already | |
| authorized those operations; no new permission is needed. | |
| ### Recovery | |
| After a candidate stops and its GPU/lock are free, resume on that host: | |
| ```bash | |
| ~/ai/envs/opensysone/bin/python scripts/launch_24h.py \ | |
| --pilot /absolute/old/campaign/training \ | |
| --train-only --training-deadline 2026-09-17T16:00:00Z \ | |
| --deadline 2026-09-17T18:16:10Z --inference-max-tokens 1024 \ | |
| --selection-metric crossfit_temperature_nll_v1 | |
| ``` | |
| The new campaign restores optimizer/RNG and preserves the selected best. Keep its | |
| parent directory intact for evidence. Before selection is frozen, a replacement | |
| campaign must also replace that candidate's `campaign` and `training` paths in | |
| the **coordinator's own `plan.json`**. Stop the waiting coordinator first, edit, | |
| then run `scripts/fleet_campaign.py --campaign /absolute/fleet/run --resume`. | |
| The in-memory plan does not reload while it is running. Preserve all absolute | |
| deadlines. After `selection.json` exists, recover that same frozen winner; never | |
| reselect using partial test results. Completed evaluation may be reused only if | |
| checkpoint/data/source hashes still match. | |
| Success requires fleet `exit_code=0`, complete `evaluation/metrics.json`, matching | |
| `evaluation/model.pt` hash, a successful `api_probe.json`, and `api_ready=true`. | |
| The deployment pointer is `/home/andy/ai/opensysone/deploy/current.json` and the | |
| resulting API is `http://127.0.0.1:18081/v1/systemone`. See | |
| [JEV_HARNESS.md](JEV_HARNESS.md); hosted Jev still needs `TYPESAFE_API_KEY`. | |