# Three-machine campaign **2026-09-17 continuation:** [NEXT_STEPS.md](NEXT_STEPS.md) records overnight findings and the two-Spark assessment. The original 2B campaign finished with exit 0 at step 6,000; its best is step 2,000. Spark B is now running a 4B refinement campaign from GX10 best step 2,500. The four-candidate fleet plan retains the completed 2B and includes the verified refinement campaign. On **2026-09-16**, the user assigned GX10 and both Sparks to OpenSysOne and explicitly authorized stopping existing processes. Key authentication from GX10 works as `andy` on both `192.168.8.111` (spark-a / spark-d1b4) and `192.168.8.204` (spark-b / spark-3e2a). No additional credentials are needed. The delivery deadline stays **2026-09-17 18:16:10 UTC / 19:16:10 BST**. Training candidates will stop at **16:00 UTC** for validation-only model selection followed by calibration, untouched test/holdout evaluation and the local Jev-compatible API. Each training process keeps the **16 GiB CUDA cap**. Independent experiments use all three GPUs without requiring an unverified distributed training stack or a ConnectX connection to GX10. | Machine | Experiment | Current state | | --- | --- | --- | | GX10 | Existing 4B, learning rate 0.0001, exact checkpoint continuation | Training resumed at step 178; new validation selected step 178 | | spark-a | 4B initialized from verified pilot weights, learning rate 0.00003, 7,500-step cosine schedule | Verification passed; long training campaign running | | spark-b | 2B completed; new 4B refinement at LR 0.00001 / seed 432 | Verified campaign `20260917T023137Z-24h` running | All candidates use the frozen public dataset and the same 512 validation decisions. Selection now uses **four-fold temperature-crossfit macro-family NLL** from durable checkpoints (`crossfit_temperature_nll_v1`, fixed seed 431). Source groups stay in one fold; each fold's temperature is fitted on the other three. This avoids discarding a better classifier because of recoverable overconfidence. Raw NLL and accuracy are reported separately. Reserved calibration/test/Social IQA predictions do not choose the winner, and the serving temperature is fitted afresh on reserved calibration after selection. Warm initialization is a new experiment with a fresh optimizer, distinguished from an exact resume in its provenance. Model paths remain under `~/ai/models/opensysone`; data, checkpoints and setup artifacts under `~/ai/opensysone`. Source and small evidence alone belong in this repository. ## Reclaimed serving pair At **19:14 UTC**, the exact verified Qwen serving process on spark-a (PID 154121) received SIGTERM and exited. After its memory was released, the exact verified RPC worker on spark-b (PID 114512) received SIGTERM and exited. No force kill was needed. Both GPU process lists were empty. Available memory afterward was 126,665,596,928 bytes on spark-a and 126,821,765,120 bytes on spark-b. The existing model files and RPC cache were preserved. Full command/cwd/process identity, post-stop memory and GPU inspection are saved on GX10 in `~/ai/opensysone/fleet-20260916/spark-a-service-stop.json` and `spark-b-service-stop.json`. Shared Python, firewall, network interfaces, swap, earlyoom and clocks were not changed. Spark-a still has swap enabled and no earlyoom; the project uses bounded allocations, host-memory checks and `oom_score_adj=0` without changing that system configuration. To restore the old serving pair **after all training/evaluation/API GPU jobs on the Sparks have stopped and memory is verified free**, start the RPC worker first on spark-b, from `/home/andy/ai/apps/llama.cpp/build/bin`: ```bash LD_LIBRARY_PATH="$PWD" setsid nohup ./ggml-rpc-server \ -H 192.168.100.11 -p 50052 -c >/tmp/rpc-server.log 2>&1 >/home/andy/ai/logs/qfn-server-fast.log 2>&1