# OpenSysOne results ## Completed 4B accuracy and speed profile — 2026-09-17 Training and profiling are finished. The selected model is Qwen3-4B-Instruct-2507 with rank-8 LoRA and a learned scalar decision head: Spark B step-1,500 weights, retained unchanged at expanded branch step 0. Selection used validation only. The [complete report](../../results/20260917-wrapup/profile-report/report.md) includes per-family results, all timing cells, machine/source/checkpoint provenance, CSV/JSON data and standalone charts. | Reserved evaluation | Decisions | Selected accuracy | Pretrained verifier accuracy | Selected calibrated NLL | Base calibrated NLL | | --- | ---: | ---: | ---: | ---: | ---: | | Original four-family test | 2,042 | **92.90%** | 84.48% | 0.2051 | 0.4527 | | Social IQA family holdout | 768 | **72.92%** | 70.31% | 0.6783 | 0.7425 | Paired, source-group-stratified 95% bootstrap intervals (400 resamples) put the accuracy gains at **+8.42 pp [6.85, 9.89]** and **+2.60 pp [0.13, 5.34]**. The holdout improvement is modest; this is one task family. Social IQA was excluded from our fine-tuning, but exposure in the pretrained base model is unknown. The test contains four-choice Banking77, BoolQ, ARC and SNLI; this is not a general-intelligence score or a comparison with the hosted Jev service. Temperature 1.745822 was fitted on 510 separate calibration decisions. Selected ECE changes from 4.27% to 1.08% on the test and 15.55% to 8.30% on Social IQA; Brier changes from 0.1172 to 0.1098 and 0.4187 to 0.3797 respectively. Calibration helps these evaluations but does not establish reliability on arbitrary inputs. | Matched profile method | Accuracy, same 320 decisions | Warm median, 128-state-token / four-choice | Warm median, 768-state-token / four-choice | | --- | ---: | ---: | ---: | | Selected scorer | **89.06%** | 0.902 s | 3.710 s | | Pretrained per-option verifier | 80.94% | 0.806 s | 3.177 s | | Pretrained joint answer-label method | 86.25% | **0.213 s** | **0.818 s** | These warm local measurements use the same otherwise idle Spark in FP32, one question per request, and include tokenization/probability construction. They exclude loading, HTTP, generated explanations and concurrent serving. The joint-label method conditions on all options together. The current scorer is **11–17% slower** than the per-option base and **2.18–15.58 times slower** than the joint-label baseline across all 12 workload cells. Ten repetitions per cell make p95 exploratory. A scalar head alone has not made this implementation faster; shared-prefix caching and merged adapters remain future measured experiments. The expanded step-159 checkpoint scores 312/383 (81.46%) expansion diagnostics versus selected 303/383 (79.11%), while losing one answer on the matched 320. This post-selection comparison is descriptive and does not change the winner. All final/resumable states remain preserved. Training, final validation, full evaluation and profiling exited 0. Final correctness differences were at most 2.65e-7 against a 1e-4 tolerance. Four-model GUI browser checks passed. Source/control details are in [handover.md](handover.md); independent profile proofs are in [profile-audit](../../results/20260917-wrapup/profile-audit). ## Historical smoke results — 2026-09-16 The 0.5B model trains, its artifact reconstructs, and FP32 shared-prefix scoring passes correctness checks. **BF16 failed batch invariance on this checkpoint and stack.** The result supports continuing the experiment; it does not establish a useful zero-shot decision model or calibrated deployment probabilities. ## Reference experiment Completed run: `20260916T155124Z`, source commit `34a993e`, exit **0**. Full artifacts: GX10 `/home/andy/ai/opensysone/runs/20260916T155124Z/`. Small artifacts: [results/20260916T155124Z](../../results/20260916T155124Z). Pinned pretrained `Qwen/Qwen2.5-0.5B` revision `060db6499f32faf8b98477b0a26969ef7d8b9987`: 494,033,665 total parameters including the scalar head, **29,825,665 trainable** (final two layers and head). FP32 weights, AdamW state and inference, SDPA; no LM next-token training loss or generated answers. Backbone learning rate 2e-5, head 1e-3, gradient clipping 1, 60 steps, four decisions per batch. The rest of the pretrained backbone is frozen. Invented inventory facts supply 192 train, 48 calibration and 72 test decisions. Each group shares one state across color, seal and quantity questions; candidate orders are shuffled. Entity IDs are disjoint, but templates and underlying fact combinations overlap. These are simple wiring/optimization examples, **not a semantic holdout or a real task-family generalization benchmark**. | Same FP32 run; 72 test decisions | Accuracy | NLL | Brier, class sum | Top-label ECE, 10 bins | | --- | ---: | ---: | ---: | ---: | | Base yes-minus-no token score | 66.7% | 1.090 | 0.574 | 0.290 | | Initial zero scalar head | 27.8% | 1.059 | 0.639 | 0.083 | | Trained scalar | 68.1% | 0.628 | 0.419 | 0.205 | | Trained + calibration-split temperature | 68.1% | 0.568 | 0.370 | 0.141 | The trained model gets **one more example** correct than the matched token baseline. This is not evidence of an accuracy gain. NLL/Brier improve on this tiny synthetic set; the temperature (1.88365) was selected using only the separate 48-example calibration split. There is no basis for a general calibration claim. The uniform head's low ECE despite poor accuracy illustrates why ECE alone is not the selection criterion. NLL uses stable log-softmax, without probability clipping. The optimization loop including periodic saves took **8.80 seconds**; median step was 105 ms. This is partial tuning on very short inputs and is not a full-model training throughput estimate. Maximum allocated CUDA memory over training, eval and the timing grid was **3.43 GiB**, reserved **3.65 GiB**, against a 16 GiB cap. The query-projection probe changed by max 0.000964; scalar weight norm became 0.2667. Full parameter and artifact provenance is in `manifest.json`. ## Shared-prefix correctness and timings The tiny random FP32 CPU model passes four tests, including mixed lengths, chunk sizes 1/2/4/16, candidate permutation, unrelated-question perturbation, gradient flow and equivalence of selected token logits to full vocabulary logits. On the trained GPU model, probability maximum absolute differences were: | Comparison | Difference | | --- | ---: | | Full forward vs shared prefix | 0.00000304 | | Question batch vs isolated question | 0.00000381 | | Candidate permutation, restored order | 0.00000131 | | Repeated prefix call | 0 | | Reset all trainable tensors, reload checkpoint | 0 | The first successful run recorded the original 0.02 tolerance. Its actual errors are below 0.000004. The continuation harness tightens FP32 tolerance to **0.0001**; BF16 retains the original gate so its known failure remains visible. Illustrative end-to-end warm medians, including tokenization, cache copies and device synchronization. One warm-up plus **three measured repeats** per cell; these are not p95 or production claims. Same trained checkpoint/serialized token IDs in both modes, two candidates/question, maximum eight branches per chunk. The full reference is already batched fairly (four two-choice questions at once). | Actual state-prefix tokens | Questions | Full batched forwards | Shared prefix | Speedup | | --- | ---: | ---: | ---: | ---: | | 143 | 1 | 35.9 ms | 47.2 ms | 0.76× | | 143 | 4 | 117.3 ms | 50.6 ms | 2.32× | | 143 | 16 | 464.5 ms | 129.2 ms | 3.60× | | 1,031 | 1 | 312.0 ms | 184.6 ms | 1.69× | | 1,031 | 4 | 1,262.6 ms | 201.6 ms | 6.26× | | 1,031 | 16 | 5,033.1 ms | 336.1 ms | 14.97× | Caching loses on the shortest one-question case. At 1,031 tokens/16 questions, the shared run spends about 156 ms in prefill and 178 ms in branches; single-prefix KV occupies 24.2 MiB before the per-chunk copies. That longer-context point demonstrates amortization in this implementation. The repeated short question is a workload timing probe, not a semantic multi-question benchmark. No generation, constrained decoding, service throughput or 1.5B/3B latency comparison has run. ## Failed BF16 experiment and diagnosis Run `20260916T154714Z`, source `4d6cb0f`, completed its 60 training steps but exited **1** at the correctness gate. The checkpoint and all earlier predictions remain available; no performance conclusion was taken from that failed run. The same trained BF16 weights were evaluated with different precision/backends: | Comparison | BF16 probability difference | Same weights cast to FP32 | | --- | ---: | ---: | | Full vs shared | 0.08544 | 0.00000727 | | Shared vs isolated | 0.09897 | 0.00000519 | | Shared candidate permutation | 0.07889 | 0.00000137 | | Batched full vs separate full calls | 0.05262 | 0.00001433 | SDPA MATH retains the BF16 failure and passes in FP32. This demonstrates precision and batch-shape sensitivity beyond cache handling; it does **not** isolate the root cause to a specific kernel or prove every GB10/model fails in BF16. The BF16 token baseline had different metrics from FP32 and must not be mixed into the matched FP32 comparison above. BF16 AdamW also lacks FP32 master weights in this simple implementation, making small updates prone to rounding. Raw evidence: [parity_diagnosis.json](../../results/20260916T154714Z/parity_diagnosis.json). Reproducer: `scripts/diagnose_parity.py --run `. Keep the FP32 reference; investigate BF16 explicitly before scaling. ## Expanded precision investigation — 2026-09-16 Completed read-only runs `20260916T161253Z-precision` and `20260916T161355Z-precision`, both exit **0**. The second run used clean source commit **`409ade4`**; the first manifest records `94a24e8` with staged additions, whose script hashes correspond to `b9dd165`. Full artifacts are under GX10 `/home/andy/ai/opensysone/runs//artifacts/`; small copies are in [results/20260916T161355Z-precision](../../results/20260916T161355Z-precision). All ablations reconstruct the preserved BF16-trained checkpoint from `20260916T154714Z`; its SHA-256 remained `106efdfb0794e6ca870b7add11c71f06c58281ef46b348305866a85f1e6f6bc8`. The base, data and checkpoint are unchanged. This comparison concerns arithmetic on the same weights, rather than FP32 versus BF16 training quality. It does not evaluate a new task or supply generalization evidence. The expanded test covers **all 24 groups / 72 decisions**, comparing batched full calls with separate question calls, full with cached, cache chunks of 4/16, cached with isolated questions, and restored candidate permutations. The table shows the worst absolute probability difference across these comparisons. Strict reduction sets `allow_bf16_reduced_precision_reduction=False`. FP32 linear casts each decoder linear's inputs and weights to FP32, then casts its output back to BF16; it is an inference diagnostic, not a validated training method. | Arithmetic configuration | Worst probability difference | Groups above BF16's original 0.02 gate | | --- | ---: | ---: | | BF16 default SDPA | 0.238608 | 24/24 | | BF16, strict reduction | 0.213011 | 24/24 | | BF16, math SDPA + strict reduction | 0.168008 | 24/24 | | BF16, FP32 linear + strict reduction | 0.147468 | 24/24 | | BF16, math SDPA + FP32 linear + strict reduction | 0.183657 | 24/24 | | Same weights cast to FP32, default SDPA | 0.00002138 | 0/24 | FP32 also passes the stricter **0.0001** gate. Repeated full and repeated cached calls have exactly zero probability difference in every group/configuration. Full candidate permutations also match exactly; cached permutations can change which branches share a chunk and still fail in BF16. Exit 0 means the diagnostic completed, not that BF16 passed. Final-candidate-token traces for the first serialized group narrow the issue: embeddings and first input normalization match exactly, but default BF16's first query/key projections differ by up to **0.5** between batched/separate calls. Strict reduction removes those initial projection differences in this trace; later differences remain. Combined math attention and FP32 linears reduce the first decoder-layer difference from 0.02344 to 0.00003052, yet the final normalized hidden representation still differs by up to 2.0. This supports shape-dependent numerical differences that propagate through the decoder. It does not isolate every contributing operation or establish a particular kernel defect. The trace samples final candidate tokens, not every token's intermediate representation. Peak CUDA allocation was **1.90 GiB**, reserved **1.94 GiB**, against the 16 GiB cap. The shared environment was unchanged and OOM score adjustment was 0. All four CPU correctness tests passed before execution. Both diagnostic PIDs exited; at 16:15 UTC GX10 again had about 118 GiB available and only the original router GPU process. Continue useful model/data work in FP32; none of these BF16 interventions justifies reopening its correctness gate. ## Public-data 2B adapter pilot — 2026-09-16 Run `/home/andy/ai/opensysone/runs/20260916T182352Z-train/artifacts`, execution source **`f1c9322`**, exited **0** after **40 optimizer steps** (160 decisions), not three completed epochs. The model is pinned `Qwen/Qwen3.5-2B` at `15852e8c16360a2fea060d615a32b45270f8a8fc`. Only its text decoder is retained; the unused vision encoder is discarded before CUDA loading. Rank-16 additive linear adapters and a pretrained yes-minus-no initialized head train **16,821,249 of 1,898,646,337 parameters** in FP32. The frozen data has 40,941 source-group-disjoint train decisions, 512 validation, 512 calibration, 2,048 source test and 768 completely held-out Social IQA decisions. This model's 768-token complete-chat limit excludes four BoolQ train rows and one test row, leaving 40,937/512/512/2,047/768. Banking77 is a four-choice target-plus- three-negative transformation, not a full 77-way benchmark. Source-group splitting does not rule out pretraining contamination or semantic duplicates. | Family | Initial validation accuracy | Step 40 accuracy | | --- | ---: | ---: | | ARC | 74.22% | 81.25% | | Banking77 four-choice | 83.59% | 88.28% | | BoolQ | 64.84% | 79.69% | | SNLI | 64.06% | 82.03% | | All 512 decisions | **71.68%** | **82.81%** | Validation macro-family NLL fell from **0.700136 to 0.498153**. This is validation selection evidence, not untouched test improvement. No calibration, test or Social IQA predictions have been evaluated in this pilot. Median four-decision step was **3.869 s**; the loop including final validation took 298.1 s. Peak CUDA allocation/reservation was **7.746/7.855 GiB**, below the 16 GiB cap. All final permutation/chunk/isolation checks passed the 0.0001 probability gate, with worst difference **0.00000614**. The old repeat label also changed chunk shape; the current source restores the original chunk size before repeat testing. A fresh process in `20260916T183240Z-train`, source **`980d881`**, reconstructed step 40 and reproduced **all 512 raw logits and probabilities exactly**, restored optimizer/RNG, then completed step 41 with finite gradient norm 3.676. It exited **0** and all final parity gates passed, worst difference 0.00000316. Step 41 validation macro NLL was 0.494478. The retained setup failure `20260916T182256Z-train` exited 1 before any optimizer step because Transformers' new chat-template return default was a BatchEncoding; explicit `return_dict=False` fixed it without changing the shared environment. Small raw pilot evidence is in [results/20260916T182352Z-train](../../results/20260916T182352Z-train). Checkpoint SHA-256 is `af5790ae2f2b56477ebbdf6ab9c418d895e48d2bf5a6416e11b6e9863ad1db55`; validation-selected best SHA-256 is `82b4261feb98d3ed56291e4c03304a65da20ce0194a6ad117d113b30d152282e`. New dependencies are isolated in `~/ai/envs/opensysone` (pyarrow 25.0.1), with read-only reuse of the existing torch/Transformers packages. The Jev-compatible stdlib harness and 12 CPU tests pass; real-checkpoint HTTP and longest-input stress are the next gate before the larger campaign. ## Public-data 4B pilot selected for the 24-hour run Run `/home/andy/ai/opensysone/runs/20260916T183823Z-train/artifacts`, clean execution source **`ccbbe6d`**, exited **0** after 40 steps / 160 decisions. The base is `Qwen/Qwen3-4B-Instruct-2507`, pinned to `cdbee75f17c01a7cc42f958dc650907174af0554`, Apache-2.0. Rank-8 adapters (alpha 16) and the pretrained initialized head train **16,517,633 of 4,038,985,729 parameters** in FP32. Exact two-pass categorical gradients keep one candidate graph live; CPU gradients match ordinary CE within 0.000001. Gradient checkpointing is enabled. No quantization or new kernels. | Family | Initial validation accuracy | Step 40 accuracy | Step 40 NLL | | --- | ---: | ---: | ---: | | ARC | 90.63% | 90.63% | 0.374400 | | Banking77 four-choice | 90.63% | 91.41% | 0.229941 | | BoolQ | 82.03% | 84.38% | 0.583093 | | SNLI | 82.81% | 83.59% | 0.395211 | | All 512 validation decisions | **86.52%** | **87.50%** | **0.395661** | Raw validation macro NLL improves from **1.436162 to 0.395661**; the initial readout was severely overconfident. A separately recorded diagnostic fits and scores a temperature on the same validation rows (NLL 0.407737, T 6.9183): it is optimistic validation analysis, not independent calibration. Reserved calibration, test and Social IQA predictions remain unevaluated. The trained 4B validation accuracy and NLL beat the 2B pilot in every family, supporting the larger candidate despite its lower throughput. This does not prove task generalization. The 512-token complete-chat limit retains **40,915 train / 512 validation / 510 calibration / 2,042 test / 768 Social IQA** decisions; it drops 26 train, two calibration and six test BoolQ rows, with no silent truncation. Median four-decision step is **8.956 s**; 55,268 actual branch tokens were processed with no padding overhead. The loop including final validation takes 724.0 s. Initial validation alone takes 327.85 s. Peak CUDA allocated/reserved is **15.510/15.604 GiB** against the 16 GiB cap. OOM adjustment is 0 and about 99 GiB unified RAM remains available with the model loaded. Final correctness passes all 0.0001 gates, worst probability difference **0.00000167**, with exact repeated, isolated and restored-permutation predictions. Checkpoint SHA-256: `e26f75b2396de88311873fac4eb91e1e40d0ec940778ec99f282bcfd96a2e258`. Best SHA-256: `64977ee0b1a6147c6faf59283edea9adf564dd36d53f4580bc20940b94c6764f`. Small raw evidence is in [results/20260916T183823Z-train](../../results/20260916T183823Z-train). Fresh reload, longest-input gradients with restored optimizer state, 1,024-token HTTP inference, and 255-choice HTTP stress **all passed** (verification exit 0). Reload matches all 16 checked validation predictions exactly. Longest training input is 509 tokens and peaks at 15.624 GiB with optimizer state; inference peaks at 15.465 GiB. The long HTTP request has 1,023 tokens in each of two candidate branches and matches direct inference exactly. Invalid-key/oversized-input requests return 401/422. One warm three-question request takes 1.571 s, and one 255-choice request takes 47.042 s; these are wiring stress timings, not latency percentiles or intelligence benchmarks. The checkpoint SHA-256 is unchanged. Evidence: [results/20260916T185718Z-verify4b](../../results/20260916T185718Z-verify4b). All **15 CPU tests pass**, including unequal-source-group bootstrap weighting and the measured evaluation-reserve calculation. The live Jev HTTPS endpoint returns 405 to an unauthenticated GET; no credentials or state were sent and no authenticated hosted inference has been tested. ## Detached 24-hour campaign now running Launched **2026-09-16 18:59:10 UTC** from clean source **`0109eb6`** into `/home/andy/ai/opensysone/runs/20260916T185910Z-24h`. Supervisor PID is **1085496**, current trainer **1085517**; both OOM score adjustments are 0. Training resumes the 4B step-40 checkpoint with optimizer/RNG restored, preserves validation-selected best and all model/data/config signatures, and has passed the initial FP32 correctness gates. Exit is **pending**; the API has not started yet. Fresh restart reproduces **all 512 raw logits and probabilities exactly**; the reference and fresh prediction JSON SHA-256 are both `e671e1508185765552b0f933ba03f356be62143c531d8ef534457d34b1645c9b`. The next four updates, **41–44**, have finite losses/gradients and remain under the cap. Step 41 takes 8.724 s, loss 0.115940, gradient norm 3.81358. This proves reconstruction plus subsequent optimizer updates, not a bitwise interrupted-versus-uninterrupted trajectory comparison. Raw verification is in the launch evidence directory below. The durable checkpoint remains step 40 until the regular save cadence, independently of those logged newer updates. Training ends by **2026-09-17 16:16:10 UTC**, reserving two hours until the final **18:16:10 UTC / 19:16:10 BST** deadline. The reserve estimates 6,640 base/tuned prediction rows at 4,251.8 seconds from measured pilot validation speed, adds 30% plus ten minutes for setup, and keeps a two-hour minimum. Checkpoints save every 250 steps or 900 seconds regardless of evaluation; validation is every 500 steps with patience eight. The three-epoch target is an upper bound. After successful training, the runner loads the best artifact fresh, calibrates only on the 510 reserved known-family decisions, saves a deployable checkpoint before untouched evaluation, records raw/calibrated test and Social IQA metrics against the unchanged pretrained scorer, and starts the loopback API only after complete evaluation and a real-model inference check. Deployment is planned at `http://127.0.0.1:18081`, with 1,024-token inputs. No hosted Jev call runs automatically. Small launch evidence lives in [results/20260916T185910Z-24h-launch](../../results/20260916T185910Z-24h-launch), separate from the completion-results directory reserved by the runner. Current inspection, stop and same-deadline recovery commands are in [handover.md](handover.md). A running job is not a finalized model or successful test result. The frozen-family controls and independent calibration remain the quality gates for final reporting. The complete 15-test suite passed; the new orphan-child stop safeguard also passes an integration test that refuses to terminate a PID when its command line differs from the recorded command. ## Three-machine expansion — 2026-09-16 evening The user assigned GX10 and both Sparks to this task and authorized terminating their workloads. The Spark serving head and RPC worker were stopped in order with verified SIGTERM; both released their GPU allocations and each had about 118 GiB available afterward. Their weights/cache and exact restoration commands are retained. No network or system configuration changed. Both Sparks now have isolated copies of the exact GX10 training dependencies: 21,368 installed file hashes and 55 package versions match. CPU autograd and both Qwen-family imports pass. This initial check verified the environments. Subsequently all pinned model files and real GPU training/HTTP checks passed; see the launch results below. The original GX10 campaign saved step 128 before a requested stop. Its trainer exceeded the 30-second grace while performing final correctness checks and exited -9; the complete step-128 checkpoint and optimizer/RNG are verified intact. The new source records skipped final checks explicitly on a requested stop. It also retains step-specific prediction evidence before publishing each new best artifact and selects a restored checkpoint if its fresh validation improves the best. The intermediate GX10 campaign `20260916T192239Z-24h`, source `6e080e2`, restored step 128 and later stopped gracefully at step 178 with training exit 0. The Spark alternatives are a 4B weights-only warm initialization with fresh Adam, learning rate 0.00003 and 7,500-step cosine horizon, and a longer 2B continuation. The planned fleet cutoff is 2026-09-17 16:00 UTC, leaving 2 h 16 min until the original final deadline. All training remains under 16 GiB per process. All **32 initial fleet CPU tests passed**, including weights-only initialization, optimizer/RNG resume, requested-stop evidence, deadline handling, exact validation-set matching, checkpoint/metric mismatch rejection and API deployment lifecycle. The coordinator recomputes its criterion from all 512 saved validation predictions and freezes selection before calibration/test/holdout. Read-only compatibility checks of the real 4B and 2B pilot artifacts pass, reproducing NLL 0.395661 and 0.498153. Small setup proofs are in [results/20260916-fleet-setup](../../results/20260916-fleet-setup). Live paths, statuses and recovery instructions are in [fleet.md](fleet.md). The subsequent selection revision uses the frozen four-fold source-group-disjoint temperature-crossfit policy `crossfit_temperature_nll_v1` (seed 431, 101 positive temperatures, family-balanced fitting and scoring). Step 128's validation accuracy is **89.0625%**, versus step 40's 87.5%; raw NLL is 0.442683 versus 0.395661. Crossfit NLL reverses that ranking: **0.318518 versus 0.359522**, improving in all four families. A 5,000-replicate paired source-group bootstrap, refitting the temperatures, gives difference -0.041005 with 95% interval [-0.079822, -0.001152]. The accuracy gain alone is uncertain (29 gains, 21 losses; McNemar p=0.322). This supports accounting for recoverable overconfidence during checkpoint selection. It is a validation-driven criterion revision, not independent test evidence. No reserved predictions were read. Original raw-selected checkpoints remain preserved, and final calibration still uses the separate reserved split. All **37 tests pass** after adding policy/selection checks; the updated CPU integration also proves reselection leaves trained weights and Adam steps intact. Raw diagnostic: [selection-diagnostic.json](../../results/20260916-fleet-setup/selection-diagnostic.json). ## Active fleet launch — 2026-09-16 19:44 UTC Three training-only campaigns are active on source **`4a60423`**: GX10 `20260916T193741Z-24h` (4B, LR 0.0001), spark-a `20260916T194258Z-24h` (4B, LR 0.00003, 7,500-step cosine horizon), and spark-b `20260916T193803Z-24h` (2B, LR 0.0001). Each uses the fixed crossfit criterion, 16 GiB allocation cap and 2026-09-17 16:00 UTC cutoff. Training exit statuses remain pending. The fleet coordinator `20260916T194403396250Z-fleet`, source **`6a7b0ed`**, is detached on GX10 and waiting for selection; no reserved-data predictions or final calibration have run. The cutoff shutdown race is covered by a regression test, and all nine fleet tests pass after that fix. GX10 restored step 178's weights, Adam and Python/torch/CUDA RNG exactly. Its fresh 512-decision validation reached **90.4297% accuracy, 0.303825 crossfit NLL, 0.404198 raw NLL**, promoting the durable best beyond step 128. Fresh FP32 correctness passes (worst probability difference 4.77e-7), and resumed updates are finite. These are validation results, not independent test evidence. Spark A reproduced all 512 original 4B pilot predictions exactly before eight finite lower-rate updates (median 8.086 seconds, peak 15.505 GiB). That pilot exited 0; its step-8 accuracy 86.914% / raw NLL 0.405397 did not improve the starting checkpoint. The long-run crossfit selector re-evaluates both inherited best and current checkpoint. Real fresh-artifact verification exited 0: exact 16-decision reload, finite restored-Adam gradients on the longest 509-token input, 15.624 GiB peak, 1,023-token HTTP/direct match, expected 401/422 errors, and 255 choices in 43.31 seconds. The long campaign reproduced all 512 step-8 raw predictions exactly, with identical weights/Adam/RNG. Its fixed crossfit criterion selected step 8 at 0.358235 NLL, and new updates are finite. No independent generalization improvement is claimed for the short pilot. Spark B's preparation exited 0. Fresh verification passed exact reload, restored-Adam gradients at 700 tokens (8.123 GiB peak), 1,024-token inference, authentication/length errors and 255 choices in 17.73 seconds. The long campaign reproduced all 512 original validation predictions exactly, scoring 82.8125% accuracy / 0.476072 crossfit NLL / 0.498153 raw NLL before resumed training. Subsequent finite updates reached step 98 by 19:43:59 UTC. Timing observations are individual wiring checks, not p50/p95 latency measurements. Full small evidence, source revisions, frozen plan and startup state snapshots are under [results/20260916-fleet-setup](../../results/20260916-fleet-setup). Live state, inspection/stop/resume and serving-pair restoration are in [fleet.md](fleet.md). Final calibrated test/holdout metrics and selected-model API deployment are pending; authenticated hosted Jev inference still requires `TYPESAFE_API_KEY`. ## Overnight progress and next experiment — 2026-09-17 The 02:10–02:15 UTC audit found both 4B jobs healthy and improving, while the 2B campaign completed cleanly at **01:59:29 UTC**, training and supervisor exit **0**. All recorded losses/gradients were finite. Peak allocation was 15.624 GiB on each 4B job and 8.183 GiB on the 2B job; the 16 GiB cap remains unchanged. | Candidate | Last audited step | Selected step | Crossfit validation NLL | Selected accuracy | | --- | ---: | ---: | ---: | ---: | | GX10 4B, LR 1e-4 | 2,570 | 2,500 | **0.188640** | **93.55%** | | Spark A 4B, LR 3e-5 | 2,529 | 2,500 | 0.218012 | 92.58% | | Spark B 2B, LR 1e-4 | 6,000 | 2,000 | 0.255294 | 89.84% | These are the same 512 validation decisions, selected with the unchanged fixed crossfit policy. No reserved calibration, test or Social IQA predictions have been read. A higher maximum accuracy at a different step does not override the selection criterion. Spark A improved at all five scheduled validations. Spark B stopped after eight evaluations without a new best; its final step-6,000 score was 0.351593 / 86.91%. Final numerical correctness passed at worst 6.56e-7. The selected step-2,000 and resumable step-6,000 artifacts are preserved. The freed Spark B is training a **fourth candidate**, initialized from a frozen copy of GX10's step-2,500 selected weights (SHA-256 `8956eb6c0cfbb02124aeefd99c3b418c55f55fdb9a64260350622d98dbba1aec`). Fresh Adam, seed 432, LR/head LR 1e-5 and a 5,000-step cosine horizon define a new trajectory. Other model/batch/token/correctness settings and both absolute deadlines stay unchanged. The eight-step pilot started at **02:16:05 UTC**; source `4a60423`. The pinned 4B model copied from Spark A over the existing link passed all 13 file hashes. Warm initialization preserves all 506 trainable tensors exactly and deliberately starts with an empty optimizer. All 512 initial raw predictions match the parent exactly. The eight-step pilot and fresh verifier exited 0: exact 16-decision reload, finite longest-input gradients, 15.623 GiB peak, direct/HTTP agreement at 1,023 tokens, expected 401/422 and 255 choices in 45.44 seconds. These timings are individual wiring checks, not percentiles. Campaign `20260917T023137Z-24h` launched at 02:31:37 UTC, restoring the complete step-8 optimizer/RNG state exactly, and was added as the fourth fleet candidate. Its inherited selected branch step 0 retains the parent score: step 8 scored 0.188576, a change below the fixed 0.001 improvement threshold. The short pilot does not establish a quality gain. [Small audit evidence](../../results/20260917-fleet-progress) records the 22 scheduled validation points, live processes, source revisions, selected-checkpoint hashes and frozen refinement parent. [next-steps.md](../research/next-steps.md) records the decisions and the two-Spark alternatives: independent candidates now, bounded distributed adapter-gradient training or parallel scoring next. Active ConnectX/RoCE and installed NCCL do not establish collective correctness or useful speedup. The current two-pass trainer needs explicit synchronization changes, and its measured peak leaves only about 385 MiB for additional GPU allocations under the cap. ## Fixed validation ensemble diagnostic — 2026-09-17 Saved, identity-matched validation logits were combined with fixed equal weights and the unchanged crossfit-temperature policy; no weights were tuned and no reserved predictions were accessed. GX10 4B + Spark A 4B scores **93.16% / 0.193983 NLL**, worse than GX10 alone (**93.55% / 0.188640**). Spark A 4B + the completed Spark B 2B scores **93.55% / 0.180560**. This more diverse pair shares 19 errors versus 29 for the two-4B pair, but gains eight/losses eight versus GX10. The mixed pair's NLL difference versus GX10 is -0.008080; a 1,000-replicate paired source-group bootstrap with fold-temperature refitting yields 95% interval **[-0.039064, +0.019451]**. No gain over the best single model is established. These are exploratory validation results from already selected checkpoints, not independent generalization evidence. The deployed-candidate protocol remains individual models; ensemble inference and latency have not been implemented or measured. The exact A step-2,500 and B step-2,000 artifacts are frozen on GX10 in `20260917T022201Z-ensemble-reference`, with 134.3 MB copied, stable source hashes and CPU reconstruction/provenance checks. No weights are in Git. [Analysis and provenance](../../results/20260917-fleet-progress/fixed-ensemble-validation.json). ## Remaining gates The active continuation state and checkpoint-resume verification are recorded in [handover.md](handover.md). Public multi-family training and the frozen unseen-family holdout are now implemented; independent calibration/test/holdout metrics await the 24-hour campaign's finalization. Frozen-head and generation controls, new-model prefix caching and the larger latency matrix remain open. The Sparks now host independent candidate experiments; GX10 does not need a ConnectX cable for this selection strategy. Architecture B and distributed training still await quality and profiling evidence in [plan.md](plan.md). ## Expanded public training data — 2026-09-17 Version 2 retains all 40,915 original 4B-compatible training decisions and adds 16,000 HellaSwag, 14,360 PIQA and 9,490 CommonsenseQA decisions: **80,765 total**. All four reserved source files and tokenized sequences match version 1 exactly. The 383 retained new-source diagnostics stay outside training and checkpoint selection. An independent reconstruction audit checked every added source label and shuffled answer position, all downloaded hashes and diagnostic exclusions. See [training-data.md](../research/training-data.md) and its linked small evidence. Clean source `24b8ccf`, pilot `20260917T070758Z-train`: eight finite updates, **exit 0**, all initial/final FP32 gates passed. Frozen Spark B parent step 1,500 reproduces every initial validation logit and probability exactly. Captured step 0 has empty Adam; every final Adam counter is eight. Median update 8.90 seconds, peak allocation including checks 15.426 GiB, worst final probability discrepancy 3.58e-7. The 32 sampled decisions cover all seven task families. Step 8 scores 0.170108 validation crossfit NLL versus parent 0.170150, both 94.7266% accuracy. The difference is below the fixed 0.001 selection threshold; the selected branch remains step 0. This is startup evidence, not a claim of improvement on the added tasks. The new campaign `20260917T072142Z-24h` restores all step-8 model/Adam/Python/Torch/CUDA states exactly and retains the original 16:00 / 18:16:10 UTC deadlines. GX10's former run stopped at step 4,380 with both trainer and supervisor exit 0, preserving its selected step 2,500. The fleet retains all previous candidates and explicitly registers the new dataset. All 86 source tests passed, including rejection of reserved-data changes and unregistered candidate datasets. The actual expanded candidate passed the full fleet eligibility path. Expanded checkpoints and transformed data were uploaded and verified in the existing private Hugging Face repository at 07:24:36 UTC, with exact source revisions, upstream notices and checksums; publication receipts are recorded separately. At 07:28:29 UTC the resumed campaign passed its full startup audit: every one of 512 pilot-step-8 predictions reproduced exactly, full optimizer/RNG state matched, and updates 9–12 were finite under the cap. GX10 reached step 15 by 07:28:58 UTC; both Spark trials, the coordinator, GUI and final-publication watcher remained running. Final campaign evaluation is still pending.