File size: 13,851 Bytes
2d5c26a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
# Three-machine campaign

**2026-09-17 continuation:** [NEXT_STEPS.md](NEXT_STEPS.md) records overnight
findings and the two-Spark assessment. The original 2B campaign finished with exit
0 at step 6,000; its best is step 2,000. Spark B is now running a 4B refinement
campaign from GX10 best step 2,500. The four-candidate fleet plan retains the
completed 2B and includes the verified refinement campaign.

On **2026-09-16**, the user assigned GX10 and both Sparks to OpenSysOne and
explicitly authorized stopping existing processes. Key authentication from GX10
works as `andy` on both `192.168.8.111` (spark-a / spark-d1b4) and
`192.168.8.204` (spark-b / spark-3e2a). No additional credentials are needed.

The delivery deadline stays **2026-09-17 18:16:10 UTC / 19:16:10 BST**. Training
candidates will stop at **16:00 UTC** for validation-only model selection followed
by calibration, untouched test/holdout evaluation and the local Jev-compatible API.
Each training process keeps the **16 GiB CUDA cap**. Independent experiments use
all three GPUs without requiring an unverified distributed training stack or a
ConnectX connection to GX10.

| Machine | Experiment | Current state |
| --- | --- | --- |
| GX10 | Existing 4B, learning rate 0.0001, exact checkpoint continuation | Training resumed at step 178; new validation selected step 178 |
| spark-a | 4B initialized from verified pilot weights, learning rate 0.00003, 7,500-step cosine schedule | Verification passed; long training campaign running |
| spark-b | 2B completed; new 4B refinement at LR 0.00001 / seed 432 | Verified campaign `20260917T023137Z-24h` running |

All candidates use the frozen public dataset and the same 512 validation
decisions. Selection now uses **four-fold temperature-crossfit macro-family NLL**
from durable checkpoints (`crossfit_temperature_nll_v1`, fixed seed 431). Source
groups stay in one fold; each fold's temperature is fitted on the other three.
This avoids discarding a better classifier because of recoverable overconfidence.
Raw NLL and accuracy are reported separately. Reserved calibration/test/Social
IQA predictions do not choose the winner, and the serving temperature is fitted
afresh on reserved calibration after selection.
Warm initialization is a new experiment with a fresh optimizer, distinguished
from an exact resume in its provenance. Model paths remain under
`~/ai/models/opensysone`; data, checkpoints and setup artifacts under
`~/ai/opensysone`. Source and small evidence alone belong in this repository.

## Reclaimed serving pair

At **19:14 UTC**, the exact verified Qwen serving process on spark-a (PID 154121)
received SIGTERM and exited. After its memory was released, the exact verified
RPC worker on spark-b (PID 114512) received SIGTERM and exited. No force kill was
needed. Both GPU process lists were empty. Available memory afterward was
126,665,596,928 bytes on spark-a and 126,821,765,120 bytes on spark-b. The existing
model files and RPC cache were preserved.

Full command/cwd/process identity, post-stop memory and GPU inspection are saved
on GX10 in `~/ai/opensysone/fleet-20260916/spark-a-service-stop.json` and
`spark-b-service-stop.json`. Shared Python, firewall, network interfaces, swap,
earlyoom and clocks were not changed. Spark-a still has swap enabled and no
earlyoom; the project uses bounded allocations, host-memory checks and
`oom_score_adj=0` without changing that system configuration.

To restore the old serving pair **after all training/evaluation/API GPU jobs on
the Sparks have stopped and memory is verified free**, start the RPC worker first
on spark-b, from `/home/andy/ai/apps/llama.cpp/build/bin`:

```bash
LD_LIBRARY_PATH="$PWD" setsid nohup ./ggml-rpc-server \
  -H 192.168.100.11 -p 50052 -c >/tmp/rpc-server.log 2>&1 </dev/null &
```

Then on spark-a, from the same binary directory, restore the recorded command:

```bash
MODEL_DIR=/home/andy/ai/models/gguf/Qwen3.8-Flash-Next-Uncensored
LD_LIBRARY_PATH="$PWD" setsid nohup ./llama-server \
  -m "$MODEL_DIR/Qwen3.8-Flash-Next-Uncensored-Q8_0-00001-of-00005.gguf" \
  --mmproj "$MODEL_DIR/mmproj-Qwen3.8-Flash-Next-Uncensored-F16.gguf" \
  --rpc 192.168.100.11:50052 -ngl 99 -fa on -c 65536 -ts 70,30 \
  --reasoning off -lv 4 --host 100.114.103.103 --port 18090 \
  --temp 0.7 --top-p 0.8 --top-k 20 --min-p 0 --presence-penalty 1.5 \
  --alias Qwen3.8-Flash-Next-Uncensored-Q8_0-fast \
  >>/home/andy/ai/logs/qfn-server-fast.log 2>&1 </dev/null &
```

Do not restore it while OpenSysOne is using the Sparks. Exact snapshot commands
above retain the prior bindings and tensor split; no network changes are needed.

## Setup and live runs

Setup logs and transfer process records are under
`/home/andy/ai/opensysone/fleet-20260916` on GX10. Live campaign paths and supervisor controls are recorded below.

The isolated Python environments on both Sparks are ready at
`/home/andy/ai/envs/opensysone/bin/python`. Each passed **21,368 file hashes and
55 exact package versions**, CPU autograd and both Qwen model-class imports.
The stack matches GX10: Python 3.12.3, torch 2.11.0+cu130, Transformers 5.15.0,
NumPy 2.5.2 and PyArrow 25.0.1. Full package/file provenance and the source archive
are in `~/ai/opensysone/bootstrap` on each host. The existing environments were
not modified. The archive crossed the existing Spark ConnectX link at 544.6 MB/s;
that is a file-transfer observation, not a distributed-training measurement.

The first **32 CPU tests passed** on source `6e080e2`, including warm initialization,
optimizer/RNG resume, graceful stopping, deadline handling, candidate selection
and API deployment. Spark source was cloned cleanly at `e047343`; that commit adds
the candidate preparation script and setup evidence. Small proofs, including
all model-file SHA-256 hashes, are in `results/20260916-fleet-setup/`.

The selection revision passes **37 tests**, plus the updated real-Adam CPU
integration for changing the selection criterion without changing weights or
optimizer steps. On the same 512 validation decisions, step 128 has 89.0625%
accuracy versus step 40's 87.5%. Raw NLL worsens (0.442683 versus 0.395661), while
four-fold crossfit NLL improves (0.318518 versus 0.359522). A source-group bootstrap
with temperature refitting supports the crossfit improvement; the accuracy gain
alone is uncertain. This is a validation-driven policy revision. The diagnostic,
fixed policy and uncertainty are recorded in `selection-diagnostic.json` in the
setup evidence directory. Original raw-selected artifacts are preserved. Fresh step-178 validation then
improved to **90.4297% accuracy / 0.303825 crossfit NLL** (raw NLL 0.404198),
and the durable best was promoted. Independent held-out testing remains pending.

GX10's original campaign saved a complete **step-128** checkpoint before its
supervisor stopped it. The old trainer was killed after the 30-second termination
grace while performing final correctness checks (**training exit -9**, supervisor
status `interrupted`); no optimizer progress was lost and no test evaluation ran.
Checkpoint SHA-256 is
`62de5222d2b5f3541d0fa1890eacfa5d647dbb0ab67e43d1013cc3b73156e0b3`.
The new source skips final checks explicitly on a requested stop, preserving the
previously completed correctness proof without reporting a false final pass.

The intermediate main run `20260916T192239Z-24h` (source `6e080e2`) restored
step 128 exactly, then stopped gracefully at step 178 with **training exit 0**.
The selection-parent artifact preserves that checkpoint byte-for-byte and changes
only inherited-best selection metadata. `selection_migration.json` records the
hashes and unchanged weights. The old campaigns are stopped; do not restart them.

## Active training and coordinator

Active trainer source: **`4a60423`**. Every plan is `train_only=true`, with
selection policy `crossfit_temperature_nll_v1`, cutoff **2026-09-17 16:00 UTC**
and final deadline **18:16:10 UTC**. All active jobs have OOM adjustment 0.
Run IDs are relative to `/home/andy/ai/opensysone/runs` on the listed machine.

| Machine | Campaign | Supervisor / trainer | Startup proof |
| --- | --- | --- | --- |
| GX10 | `20260916T193741Z-24h` | 1096483 / 1096505 | Step-178 state exact; fresh validation selected step 178 |
| spark-a | `20260916T194258Z-24h` | 327084 / 327116 | Pilot and resume both reproduce all 512 predictions; new finite updates |
| spark-b | `20260916T193803Z-24h` | exited 0 / 0 | 6,000 steps; best 2,000; final correctness passed |
| spark-b | `20260917T023137Z-24h` | 483974 / 484001 | Eight-step pilot and GPU/HTTP verifier exited 0; exact optimizer/RNG resume |

Spark B's preparation `20260916T193722Z-spark-b-prepare` exited **0**. Real
verification passed restored-optimizer gradients at the longest 700-token input
(8.123 GiB peak), 1,024-token HTTP inference, authentication/length errors and
255 choices (17.73 seconds for the measured request). Spark A's preparation
`20260916T192906Z-spark-a-prepare` exited **0**. Its pilot
completed eight finite updates. Fresh verification reproduced 16 decisions
exactly, restored Adam for finite longest-input gradients (15.624 GiB peak),
matched direct/HTTP at 1,023 tokens, returned expected 401/422 errors, and passed
255 choices (43.31 seconds). These are wiring timings, not latency percentiles.
The long campaign then reproduced all 512 step-8 predictions exactly, with
weights/Adam/Python/torch/CUDA RNG unchanged. The fixed crossfit criterion chose
step 8 (0.358235 NLL); subsequent optimizer updates are finite.

The fleet coordinator is **`20260916T194403396250Z-fleet` on GX10**, PID
**1425784**, source **`fed722e`**. It is running in `waiting_for_selection`, has
OOM adjustment 0, owns `.fleet.lock`, and waits until the cutoff before stopping
the exact recorded candidate jobs. Its own `plan.json` is the authoritative plan;
`~/ai/opensysone/runs/LAST_FLEET_CAMPAIGN` points to this run. Its exit status is
pending. The shutdown race with naturally finishing training jobs is fixed and
covered by a regression test; all nine fleet tests pass. Stable checkpoint
hashes, the identical 512 decision IDs,
source/model/data compatibility, numerical evidence and recomputed selection
scores must pass before a candidate is eligible. An unreachable remote is
excluded; an unverified GX10 stop blocks finalization. Selection is written once
to `selection.json` before any reserved-data predictions. Finalization acquires
`.smoke.lock` and runs only the selected model on GX10.

Historical startup snapshot **2026-09-16T19:50:01.290313+00:00**: GX10 step 217, spark-a step 17, spark-b step 200. All logged losses/gradients are finite,
peak CUDA allocation stays below 16 GiB, and all six supervisor/trainer OOM
adjustments were 0. All three jobs and the coordinator were running then; the
2B campaign has since completed, as recorded above. See `results/20260916-fleet-setup/fleet-current-status.json`.

### Inspect and stop

For a compact status table across all registered hosts, run
`python3 scripts/fleet_status.py` on GX10; add `--json` for saved evidence.
It handles unavailable hosts and trailing partial training records independently,
loads no models and changes no processes. A new setup pilot appears in this table
only after its long campaign is registered in the fleet plan.

From the project directory on the appropriate host:

```bash
~/ai/envs/opensysone/bin/python scripts/campaign_status.py \
  --campaign /home/andy/ai/opensysone/runs/20260916T193741Z-24h
```

Substitute the host's exact run path. Add `--stop` to request a command-verified
stop; confirm supervisor and child have exited before loading another model.
Read `training.log`, `training/training.jsonl`, `state.json` and `plan.json` for
progress, heartbeat, exact commands and deadlines. Exit codes are pending until
completion. The last logged update may be newer than a durable checkpoint.

For the fleet coordinator, use its exact path with:

```bash
~/ai/envs/opensysone/bin/python scripts/fleet_campaign.py \
  --campaign /home/andy/ai/opensysone/runs/20260916T194403396250Z-fleet --status
```

Use `--stop` instead of `--status` to stop the coordinator or its resulting API.
**Stopping the waiting coordinator does not stop the independent training jobs.**
Stop each candidate explicitly if the whole fleet must stop. The user has already
authorized those operations; no new permission is needed.

### Recovery

After a candidate stops and its GPU/lock are free, resume on that host:

```bash
~/ai/envs/opensysone/bin/python scripts/launch_24h.py \
  --pilot /absolute/old/campaign/training \
  --train-only --training-deadline 2026-09-17T16:00:00Z \
  --deadline 2026-09-17T18:16:10Z --inference-max-tokens 1024 \
  --selection-metric crossfit_temperature_nll_v1
```

The new campaign restores optimizer/RNG and preserves the selected best. Keep its
parent directory intact for evidence. Before selection is frozen, a replacement
campaign must also replace that candidate's `campaign` and `training` paths in
the **coordinator's own `plan.json`**. Stop the waiting coordinator first, edit,
then run `scripts/fleet_campaign.py --campaign /absolute/fleet/run --resume`.
The in-memory plan does not reload while it is running. Preserve all absolute
deadlines. After `selection.json` exists, recover that same frozen winner; never
reselect using partial test results. Completed evaluation may be reused only if
checkpoint/data/source hashes still match.

Success requires fleet `exit_code=0`, complete `evaluation/metrics.json`, matching
`evaluation/model.pt` hash, a successful `api_probe.json`, and `api_ready=true`.
The deployment pointer is `/home/andy/ai/opensysone/deploy/current.json` and the
resulting API is `http://127.0.0.1:18081/v1/systemone`. See
[JEV_HARNESS.md](JEV_HARNESS.md); hosted Jev still needs `TYPESAFE_API_KEY`.