opensysone / source /RESULTS.md
andyshu's picture
Back up verified OpenSysOne training snapshot and pinned source
58f2896 verified
|
Raw History Blame
33.5 kB
# GX10 smoke results — 2026-09-16
The 0.5B model trains, its artifact reconstructs, and FP32 shared-prefix scoring
passes correctness checks. **BF16 failed batch invariance on this checkpoint and
stack.** The result supports continuing the experiment; it does not establish a
useful zero-shot decision model or calibrated deployment probabilities.
## Reference experiment
Completed run: `20260916T155124Z`, source commit `34a993e`, exit **0**.
Full artifacts: GX10 `/home/andy/ai/opensysone/runs/20260916T155124Z/`.
Small artifacts: [results/20260916T155124Z](results/20260916T155124Z/).
Pinned pretrained `Qwen/Qwen2.5-0.5B` revision
`060db6499f32faf8b98477b0a26969ef7d8b9987`: 494,033,665 total parameters including
the scalar head, **29,825,665 trainable** (final two layers and head). FP32 weights,
AdamW state and inference, SDPA; no LM next-token training loss or generated answers.
Backbone learning rate 2e-5, head 1e-3, gradient clipping 1, 60 steps, four decisions
per batch. The rest of the pretrained backbone is frozen.
Invented inventory facts supply 192 train, 48 calibration and 72 test decisions.
Each group shares one state across color, seal and quantity questions; candidate
orders are shuffled. Entity IDs are disjoint, but templates and underlying fact
combinations overlap. These are simple wiring/optimization examples, **not a
semantic holdout or a real task-family generalization benchmark**.
| Same FP32 run; 72 test decisions | Accuracy | NLL | Brier, class sum | Top-label ECE, 10 bins |
| --- | ---: | ---: | ---: | ---: |
| Base yes-minus-no token score | 66.7% | 1.090 | 0.574 | 0.290 |
| Initial zero scalar head | 27.8% | 1.059 | 0.639 | 0.083 |
| Trained scalar | 68.1% | 0.628 | 0.419 | 0.205 |
| Trained + calibration-split temperature | 68.1% | 0.568 | 0.370 | 0.141 |
The trained model gets **one more example** correct than the matched token
baseline. This is not evidence of an accuracy gain. NLL/Brier improve on this
tiny synthetic set; the temperature (1.88365) was selected using only the separate
48-example calibration split. There is no basis for a general calibration claim.
The uniform head's low ECE despite poor accuracy illustrates why ECE alone is
not the selection criterion. NLL uses stable log-softmax, without probability clipping.
The optimization loop including periodic saves took **8.80 seconds**; median step
was 105 ms. This is partial tuning on very short inputs and is not a full-model
training throughput estimate. Maximum allocated CUDA memory over training, eval
and the timing grid was **3.43 GiB**, reserved **3.65 GiB**, against a 16 GiB cap.
The query-projection probe changed by max 0.000964; scalar weight norm became
0.2667. Full parameter and artifact provenance is in `manifest.json`.
## Shared-prefix correctness and timings
The tiny random FP32 CPU model passes four tests, including mixed lengths,
chunk sizes 1/2/4/16, candidate permutation, unrelated-question perturbation,
gradient flow and equivalence of selected token logits to full vocabulary logits.
On the trained GPU model, probability maximum absolute differences were:
| Comparison | Difference |
| --- | ---: |
| Full forward vs shared prefix | 0.00000304 |
| Question batch vs isolated question | 0.00000381 |
| Candidate permutation, restored order | 0.00000131 |
| Repeated prefix call | 0 |
| Reset all trainable tensors, reload checkpoint | 0 |
The first successful run recorded the original 0.02 tolerance. Its actual errors
are below 0.000004. The continuation harness tightens FP32 tolerance to **0.0001**;
BF16 retains the original gate so its known failure remains visible.
Illustrative end-to-end warm medians, including tokenization, cache copies and
device synchronization. One warm-up plus **three measured repeats** per cell;
these are not p95 or production claims. Same trained checkpoint/serialized token
IDs in both modes, two candidates/question, maximum eight branches per chunk.
The full reference is already batched fairly (four two-choice questions at once).
| Actual state-prefix tokens | Questions | Full batched forwards | Shared prefix | Speedup |
| --- | ---: | ---: | ---: | ---: |
| 143 | 1 | 35.9 ms | 47.2 ms | 0.76× |
| 143 | 4 | 117.3 ms | 50.6 ms | 2.32× |
| 143 | 16 | 464.5 ms | 129.2 ms | 3.60× |
| 1,031 | 1 | 312.0 ms | 184.6 ms | 1.69× |
| 1,031 | 4 | 1,262.6 ms | 201.6 ms | 6.26× |
| 1,031 | 16 | 5,033.1 ms | 336.1 ms | 14.97× |
Caching loses on the shortest one-question case. At 1,031 tokens/16 questions,
the shared run spends about 156 ms in prefill and 178 ms in branches; single-prefix
KV occupies 24.2 MiB before the per-chunk copies. That longer-context point
demonstrates amortization in this implementation. The repeated short question is
a workload timing probe, not a semantic multi-question benchmark. No generation,
constrained decoding, service throughput or 1.5B/3B latency comparison has run.
## Failed BF16 experiment and diagnosis
Run `20260916T154714Z`, source `4d6cb0f`, completed its 60 training steps but
exited **1** at the correctness gate. The checkpoint and all earlier predictions
remain available; no performance conclusion was taken from that failed run.
The same trained BF16 weights were evaluated with different precision/backends:
| Comparison | BF16 probability difference | Same weights cast to FP32 |
| --- | ---: | ---: |
| Full vs shared | 0.08544 | 0.00000727 |
| Shared vs isolated | 0.09897 | 0.00000519 |
| Shared candidate permutation | 0.07889 | 0.00000137 |
| Batched full vs separate full calls | 0.05262 | 0.00001433 |
SDPA MATH retains the BF16 failure and passes in FP32. This demonstrates precision
and batch-shape sensitivity beyond cache handling; it does **not** isolate the
root cause to a specific kernel or prove every GB10/model fails in BF16. The
BF16 token baseline had different metrics from FP32 and must not be mixed into
the matched FP32 comparison above. BF16 AdamW also lacks FP32 master weights in
this simple implementation, making small updates prone to rounding.
Raw evidence: [parity_diagnosis.json](results/20260916T154714Z/parity_diagnosis.json).
Reproducer: `scripts/diagnose_parity.py --run <failed-run-directory>`.
Keep the FP32 reference; investigate BF16 explicitly before scaling.
## Expanded precision investigation — 2026-09-16
Completed read-only runs `20260916T161253Z-precision` and
`20260916T161355Z-precision`, both exit **0**. The second run used clean source
commit **`409ade4`**; the first manifest records `94a24e8` with staged additions,
whose script hashes correspond to `b9dd165`. Full artifacts are under GX10
`/home/andy/ai/opensysone/runs/<run-id>/artifacts/`; small copies are in
[results/20260916T161355Z-precision](results/20260916T161355Z-precision/).
All ablations reconstruct the preserved BF16-trained checkpoint from
`20260916T154714Z`; its SHA-256 remained
`106efdfb0794e6ca870b7add11c71f06c58281ef46b348305866a85f1e6f6bc8`.
The base, data and checkpoint are unchanged. This comparison concerns arithmetic
on the same weights, rather than FP32 versus BF16 training quality. It does not
evaluate a new task or supply generalization evidence.
The expanded test covers **all 24 groups / 72 decisions**, comparing batched
full calls with separate question calls, full with cached, cache chunks of 4/16,
cached with isolated questions, and restored candidate permutations. The table
shows the worst absolute probability difference across these comparisons.
Strict reduction sets `allow_bf16_reduced_precision_reduction=False`. FP32 linear
casts each decoder linear's inputs and weights to FP32, then casts its output
back to BF16; it is an inference diagnostic, not a validated training method.
| Arithmetic configuration | Worst probability difference | Groups above BF16's original 0.02 gate |
| --- | ---: | ---: |
| BF16 default SDPA | 0.238608 | 24/24 |
| BF16, strict reduction | 0.213011 | 24/24 |
| BF16, math SDPA + strict reduction | 0.168008 | 24/24 |
| BF16, FP32 linear + strict reduction | 0.147468 | 24/24 |
| BF16, math SDPA + FP32 linear + strict reduction | 0.183657 | 24/24 |
| Same weights cast to FP32, default SDPA | 0.00002138 | 0/24 |
FP32 also passes the stricter **0.0001** gate. Repeated full and repeated cached
calls have exactly zero probability difference in every group/configuration.
Full candidate permutations also match exactly; cached permutations can change
which branches share a chunk and still fail in BF16. Exit 0 means the diagnostic
completed, not that BF16 passed.
Final-candidate-token traces for the first serialized group narrow the issue:
embeddings and first input normalization match exactly, but default BF16's first
query/key projections differ by up to **0.5** between batched/separate calls.
Strict reduction removes those initial projection differences in this trace;
later differences remain. Combined math attention and FP32 linears reduce the
first decoder-layer difference from 0.02344 to 0.00003052, yet the final normalized
hidden representation still differs by up to 2.0. This supports shape-dependent
numerical differences that propagate through the decoder. It does not isolate
every contributing operation or establish a particular kernel defect. The trace
samples final candidate tokens, not every token's intermediate representation.
Peak CUDA allocation was **1.90 GiB**, reserved **1.94 GiB**, against the 16 GiB
cap. The shared environment was unchanged and OOM score adjustment was 0.
All four CPU correctness tests passed before execution. Both diagnostic PIDs
exited; at 16:15 UTC GX10 again had about 118 GiB available and only the original
router GPU process. Continue useful model/data work in FP32; none of these BF16
interventions justifies reopening its correctness gate.
## Public-data 2B adapter pilot — 2026-09-16
Run `/home/andy/ai/opensysone/runs/20260916T182352Z-train/artifacts`,
execution source **`f1c9322`**, exited **0** after **40 optimizer steps**
(160 decisions), not three completed epochs. The model is pinned
`Qwen/Qwen3.5-2B` at `15852e8c16360a2fea060d615a32b45270f8a8fc`.
Only its text decoder is retained; the unused vision encoder is discarded before
CUDA loading. Rank-16 additive linear adapters and a pretrained yes-minus-no
initialized head train **16,821,249 of 1,898,646,337 parameters** in FP32.
The frozen data has 40,941 source-group-disjoint train decisions, 512 validation,
512 calibration, 2,048 source test and 768 completely held-out Social IQA decisions.
This model's 768-token complete-chat limit excludes four BoolQ train rows and one
test row, leaving 40,937/512/512/2,047/768. Banking77 is a four-choice target-plus-
three-negative transformation, not a full 77-way benchmark. Source-group splitting
does not rule out pretraining contamination or semantic duplicates.
| Family | Initial validation accuracy | Step 40 accuracy |
| --- | ---: | ---: |
| ARC | 74.22% | 81.25% |
| Banking77 four-choice | 83.59% | 88.28% |
| BoolQ | 64.84% | 79.69% |
| SNLI | 64.06% | 82.03% |
| All 512 decisions | **71.68%** | **82.81%** |
Validation macro-family NLL fell from **0.700136 to 0.498153**. This is validation
selection evidence, not untouched test improvement. No calibration, test or
Social IQA predictions have been evaluated in this pilot. Median four-decision
step was **3.869 s**; the loop including final validation took 298.1 s.
Peak CUDA allocation/reservation was **7.746/7.855 GiB**, below the 16 GiB cap.
All final permutation/chunk/isolation checks passed the 0.0001 probability gate,
with worst difference **0.00000614**. The old repeat label also changed chunk
shape; the current source restores the original chunk size before repeat testing.
A fresh process in `20260916T183240Z-train`, source **`980d881`**, reconstructed
step 40 and reproduced **all 512 raw logits and probabilities exactly**, restored
optimizer/RNG, then completed step 41 with finite gradient norm 3.676.
It exited **0** and all final parity gates passed, worst difference 0.00000316.
Step 41 validation macro NLL was 0.494478. The retained setup failure
`20260916T182256Z-train` exited 1 before any optimizer step because Transformers'
new chat-template return default was a BatchEncoding; explicit `return_dict=False`
fixed it without changing the shared environment.
Small raw pilot evidence is in [results/20260916T182352Z-train](results/20260916T182352Z-train/).
Checkpoint SHA-256 is
`af5790ae2f2b56477ebbdf6ab9c418d895e48d2bf5a6416e11b6e9863ad1db55`;
validation-selected best SHA-256 is
`82b4261feb98d3ed56291e4c03304a65da20ce0194a6ad117d113b30d152282e`.
New dependencies are isolated in `~/ai/envs/opensysone` (pyarrow 25.0.1), with
read-only reuse of the existing torch/Transformers packages. The Jev-compatible
stdlib harness and 12 CPU tests pass; real-checkpoint HTTP and longest-input
stress are the next gate before the larger campaign.
## Public-data 4B pilot selected for the 24-hour run
Run `/home/andy/ai/opensysone/runs/20260916T183823Z-train/artifacts`, clean execution
source **`ccbbe6d`**, exited **0** after 40 steps / 160 decisions. The base is
`Qwen/Qwen3-4B-Instruct-2507`, pinned to
`cdbee75f17c01a7cc42f958dc650907174af0554`, Apache-2.0.
Rank-8 adapters (alpha 16) and the pretrained initialized head train
**16,517,633 of 4,038,985,729 parameters** in FP32. Exact two-pass categorical
gradients keep one candidate graph live; CPU gradients match ordinary CE within
0.000001. Gradient checkpointing is enabled. No quantization or new kernels.
| Family | Initial validation accuracy | Step 40 accuracy | Step 40 NLL |
| --- | ---: | ---: | ---: |
| ARC | 90.63% | 90.63% | 0.374400 |
| Banking77 four-choice | 90.63% | 91.41% | 0.229941 |
| BoolQ | 82.03% | 84.38% | 0.583093 |
| SNLI | 82.81% | 83.59% | 0.395211 |
| All 512 validation decisions | **86.52%** | **87.50%** | **0.395661** |
Raw validation macro NLL improves from **1.436162 to 0.395661**; the initial
readout was severely overconfident. A separately recorded diagnostic fits and
scores a temperature on the same validation rows (NLL 0.407737, T 6.9183): it is
optimistic validation analysis, not independent calibration. Reserved calibration,
test and Social IQA predictions remain unevaluated. The trained 4B validation
accuracy and NLL beat the 2B pilot in every family, supporting the larger candidate
despite its lower throughput. This does not prove task generalization.
The 512-token complete-chat limit retains **40,915 train / 512 validation /
510 calibration / 2,042 test / 768 Social IQA** decisions; it drops 26 train,
two calibration and six test BoolQ rows, with no silent truncation.
Median four-decision step is **8.956 s**; 55,268 actual branch tokens were
processed with no padding overhead. The loop including final validation takes
724.0 s. Initial validation alone takes 327.85 s. Peak CUDA allocated/reserved
is **15.510/15.604 GiB** against the 16 GiB cap. OOM adjustment is 0 and about
99 GiB unified RAM remains available with the model loaded.
Final correctness passes all 0.0001 gates, worst probability difference
**0.00000167**, with exact repeated, isolated and restored-permutation predictions.
Checkpoint SHA-256:
`e26f75b2396de88311873fac4eb91e1e40d0ec940778ec99f282bcfd96a2e258`.
Best SHA-256:
`64977ee0b1a6147c6faf59283edea9adf564dd36d53f4580bc20940b94c6764f`.
Small raw evidence is in [results/20260916T183823Z-train](results/20260916T183823Z-train/).
Fresh reload, longest-input gradients with restored optimizer state, 1,024-token
HTTP inference, and 255-choice HTTP stress **all passed** (verification exit 0).
Reload matches all 16 checked validation predictions exactly. Longest training
input is 509 tokens and peaks at 15.624 GiB with optimizer state; inference peaks
at 15.465 GiB. The long HTTP request has 1,023 tokens in each of two candidate
branches and matches direct inference exactly. Invalid-key/oversized-input
requests return 401/422. One warm three-question request takes 1.571 s, and one
255-choice request takes 47.042 s; these are wiring stress timings, not latency
percentiles or intelligence benchmarks. The checkpoint SHA-256 is unchanged.
Evidence: [results/20260916T185718Z-verify4b](results/20260916T185718Z-verify4b/).
All **15 CPU tests pass**, including unequal-source-group bootstrap weighting
and the measured evaluation-reserve calculation. The live Jev HTTPS endpoint
returns 405 to an unauthenticated GET; no credentials or state were sent and no
authenticated hosted inference has been tested.
## Detached 24-hour campaign now running
Launched **2026-09-16 18:59:10 UTC** from clean source **`0109eb6`** into
`/home/andy/ai/opensysone/runs/20260916T185910Z-24h`. Supervisor PID is **1085496**,
current trainer **1085517**; both OOM score adjustments are 0. Training resumes the
4B step-40 checkpoint with optimizer/RNG restored, preserves validation-selected
best and all model/data/config signatures, and has passed the initial FP32
correctness gates. Exit is **pending**; the API has not started yet.
Fresh restart reproduces **all 512 raw logits and probabilities exactly**;
the reference and fresh prediction JSON SHA-256 are both
`e671e1508185765552b0f933ba03f356be62143c531d8ef534457d34b1645c9b`.
The next four updates, **41–44**, have finite losses/gradients and remain under
the cap. Step 41 takes 8.724 s, loss 0.115940, gradient norm 3.81358.
This proves reconstruction plus subsequent optimizer updates, not a bitwise
interrupted-versus-uninterrupted trajectory comparison. Raw verification is in
the launch evidence directory below. The durable checkpoint remains step 40
until the regular save cadence, independently of those logged newer updates.
Training ends by **2026-09-17 16:16:10 UTC**, reserving two hours until the final
**18:16:10 UTC / 19:16:10 BST** deadline. The reserve estimates 6,640 base/tuned
prediction rows at 4,251.8 seconds from measured pilot validation speed, adds
30% plus ten minutes for setup, and keeps a two-hour minimum. Checkpoints save
every 250 steps or 900 seconds regardless of evaluation; validation is every
500 steps with patience eight. The three-epoch target is an upper bound.
After successful training, the runner loads the best artifact fresh, calibrates
only on the 510 reserved known-family decisions, saves a deployable checkpoint
before untouched evaluation, records raw/calibrated test and Social IQA metrics
against the unchanged pretrained scorer, and starts the loopback API only after
complete evaluation and a real-model inference check. Deployment is planned at
`http://127.0.0.1:18081`, with 1,024-token inputs. No hosted Jev call runs
automatically. Small launch evidence lives in
[results/20260916T185910Z-24h-launch](results/20260916T185910Z-24h-launch/), separate
from the completion-results directory reserved by the runner.
Current inspection, stop and same-deadline recovery commands are in
[HANDOVER.md](HANDOVER.md). A running job is not a finalized model or successful
test result. The frozen-family controls and independent calibration remain the
quality gates for final reporting. The complete 15-test suite passed; the new
orphan-child stop safeguard also passes an integration test that refuses to
terminate a PID when its command line differs from the recorded command.
## Three-machine expansion — 2026-09-16 evening
The user assigned GX10 and both Sparks to this task and authorized terminating
their workloads. The Spark serving head and RPC worker were stopped in order
with verified SIGTERM; both released their GPU allocations and each had about
118 GiB available afterward. Their weights/cache and exact restoration commands
are retained. No network or system configuration changed.
Both Sparks now have isolated copies of the exact GX10 training dependencies:
21,368 installed file hashes and 55 package versions match. CPU autograd and both
Qwen-family imports pass. This initial check verified the environments. Subsequently all pinned model
files and real GPU training/HTTP checks passed; see the launch results below.
The original GX10 campaign saved step 128 before a requested stop. Its trainer
exceeded the 30-second grace while performing final correctness checks and exited
-9; the complete step-128 checkpoint and optimizer/RNG are verified intact. The
new source records skipped final checks explicitly on a requested stop. It also
retains step-specific prediction evidence before publishing each new best artifact
and selects a restored checkpoint if its fresh validation improves the best.
The intermediate GX10 campaign `20260916T192239Z-24h`, source `6e080e2`, restored
step 128 and later stopped gracefully at step 178 with training exit 0.
The Spark alternatives are a 4B weights-only warm initialization with fresh Adam,
learning rate 0.00003 and 7,500-step cosine horizon, and a longer 2B continuation.
The planned fleet cutoff is 2026-09-17 16:00 UTC, leaving 2 h 16 min until the
original final deadline. All training remains under 16 GiB per process.
All **32 initial fleet CPU tests passed**, including weights-only initialization, optimizer/RNG
resume, requested-stop evidence, deadline handling, exact validation-set matching,
checkpoint/metric mismatch rejection and API deployment lifecycle. The coordinator
recomputes its criterion from all 512 saved validation predictions and freezes
selection before calibration/test/holdout. Read-only compatibility checks of the
real 4B and 2B pilot artifacts pass, reproducing NLL 0.395661 and 0.498153.
Small setup proofs are in [results/20260916-fleet-setup](results/20260916-fleet-setup/).
Live paths, statuses and recovery instructions are in [FLEET_RUN.md](FLEET_RUN.md).
The subsequent selection revision uses the frozen four-fold source-group-disjoint
temperature-crossfit policy `crossfit_temperature_nll_v1` (seed 431, 101 positive
temperatures, family-balanced fitting and scoring). Step 128's validation accuracy
is **89.0625%**, versus step 40's 87.5%; raw NLL is 0.442683 versus 0.395661.
Crossfit NLL reverses that ranking: **0.318518 versus 0.359522**, improving in all
four families. A 5,000-replicate paired source-group bootstrap, refitting the
temperatures, gives difference -0.041005 with 95% interval [-0.079822, -0.001152].
The accuracy gain alone is uncertain (29 gains, 21 losses; McNemar p=0.322).
This supports accounting for recoverable overconfidence during checkpoint
selection. It is a validation-driven criterion revision, not independent test
evidence. No reserved predictions were read. Original raw-selected checkpoints
remain preserved, and final calibration still uses the separate reserved split.
All **37 tests pass** after adding policy/selection checks; the updated CPU
integration also proves reselection leaves trained weights and Adam steps intact.
Raw diagnostic: [selection-diagnostic.json](results/20260916-fleet-setup/selection-diagnostic.json).
## Active fleet launch — 2026-09-16 19:44 UTC
Three training-only campaigns are active on source **`4a60423`**:
GX10 `20260916T193741Z-24h` (4B, LR 0.0001), spark-a
`20260916T194258Z-24h` (4B, LR 0.00003, 7,500-step cosine horizon), and spark-b
`20260916T193803Z-24h` (2B, LR 0.0001). Each uses the fixed crossfit criterion,
16 GiB allocation cap and 2026-09-17 16:00 UTC cutoff. Training exit statuses
remain pending. The fleet coordinator `20260916T194403396250Z-fleet`, source
**`6a7b0ed`**, is detached on GX10 and waiting for selection; no reserved-data
predictions or final calibration have run. The cutoff shutdown race is covered
by a regression test, and all nine fleet tests pass after that fix.
GX10 restored step 178's weights, Adam and Python/torch/CUDA RNG exactly. Its
fresh 512-decision validation reached **90.4297% accuracy, 0.303825 crossfit NLL,
0.404198 raw NLL**, promoting the durable best beyond step 128. Fresh FP32
correctness passes (worst probability difference 4.77e-7), and resumed updates
are finite. These are validation results, not independent test evidence.
Spark A reproduced all 512 original 4B pilot predictions exactly before eight
finite lower-rate updates (median 8.086 seconds, peak 15.505 GiB). That pilot
exited 0; its step-8 accuracy 86.914% / raw NLL 0.405397 did not improve the
starting checkpoint. The long-run crossfit selector re-evaluates both inherited
best and current checkpoint. Real fresh-artifact verification exited 0: exact
16-decision reload, finite restored-Adam gradients on the longest 509-token
input, 15.624 GiB peak, 1,023-token HTTP/direct match, expected 401/422 errors,
and 255 choices in 43.31 seconds. The long campaign reproduced all 512 step-8
raw predictions exactly, with identical weights/Adam/RNG. Its fixed crossfit
criterion selected step 8 at 0.358235 NLL, and new updates are finite. No
independent generalization improvement is claimed for the short pilot.
Spark B's preparation exited 0. Fresh verification passed exact reload,
restored-Adam gradients at 700 tokens (8.123 GiB peak), 1,024-token inference,
authentication/length errors and 255 choices in 17.73 seconds. The long campaign
reproduced all 512 original validation predictions exactly, scoring 82.8125%
accuracy / 0.476072 crossfit NLL / 0.498153 raw NLL before resumed training.
Subsequent finite updates reached step 98 by 19:43:59 UTC. Timing observations
are individual wiring checks, not p50/p95 latency measurements.
Full small evidence, source revisions, frozen plan and startup state snapshots
are under [results/20260916-fleet-setup](results/20260916-fleet-setup/). Live
state, inspection/stop/resume and serving-pair restoration are in
[FLEET_RUN.md](FLEET_RUN.md). Final calibrated test/holdout metrics and selected-model
API deployment are pending; authenticated hosted Jev inference still requires
`TYPESAFE_API_KEY`.
## Overnight progress and next experiment — 2026-09-17
The 02:10–02:15 UTC audit found both 4B jobs healthy and improving, while the 2B
campaign completed cleanly at **01:59:29 UTC**, training and supervisor exit **0**.
All recorded losses/gradients were finite. Peak allocation was 15.624 GiB on each
4B job and 8.183 GiB on the 2B job; the 16 GiB cap remains unchanged.
| Candidate | Last audited step | Selected step | Crossfit validation NLL | Selected accuracy |
| --- | ---: | ---: | ---: | ---: |
| GX10 4B, LR 1e-4 | 2,570 | 2,500 | **0.188640** | **93.55%** |
| Spark A 4B, LR 3e-5 | 2,529 | 2,500 | 0.218012 | 92.58% |
| Spark B 2B, LR 1e-4 | 6,000 | 2,000 | 0.255294 | 89.84% |
These are the same 512 validation decisions, selected with the unchanged fixed
crossfit policy. No reserved calibration, test or Social IQA predictions have
been read. A higher maximum accuracy at a different step does not override the
selection criterion. Spark A improved at all five scheduled validations. Spark B
stopped after eight evaluations without a new best; its final step-6,000 score
was 0.351593 / 86.91%. Final numerical correctness passed at worst 6.56e-7.
The selected step-2,000 and resumable step-6,000 artifacts are preserved.
The freed Spark B is training a **fourth candidate**, initialized from a frozen
copy of GX10's step-2,500 selected weights (SHA-256
`8956eb6c0cfbb02124aeefd99c3b418c55f55fdb9a64260350622d98dbba1aec`).
Fresh Adam, seed 432, LR/head LR 1e-5 and a 5,000-step cosine horizon define a new
trajectory. Other model/batch/token/correctness settings and both absolute
deadlines stay unchanged. The eight-step pilot started at **02:16:05 UTC**;
source `4a60423`. The pinned 4B model copied from Spark A over the existing link
passed all 13 file hashes. Warm initialization preserves all 506 trainable tensors
exactly and deliberately starts with an empty optimizer. All 512 initial raw
predictions match the parent exactly. The eight-step pilot and fresh verifier
exited 0: exact 16-decision reload, finite longest-input gradients, 15.623 GiB
peak, direct/HTTP agreement at 1,023 tokens, expected 401/422 and 255 choices
in 45.44 seconds. These timings are individual wiring checks, not percentiles.
Campaign `20260917T023137Z-24h` launched at 02:31:37 UTC, restoring the complete
step-8 optimizer/RNG state exactly, and was added as the fourth fleet candidate.
Its inherited selected branch step 0 retains the parent score: step 8 scored
0.188576, a change below the fixed 0.001 improvement threshold. The short pilot
does not establish a quality gain.
[Small audit evidence](results/20260917-fleet-progress/) records the 22 scheduled
validation points, live processes, source revisions, selected-checkpoint hashes
and frozen refinement parent. [NEXT_STEPS.md](NEXT_STEPS.md) records the decisions
and the two-Spark alternatives: independent candidates now, bounded distributed
adapter-gradient training or parallel scoring next. Active ConnectX/RoCE and
installed NCCL do not establish collective correctness or useful speedup. The
current two-pass trainer needs explicit synchronization changes, and its measured
peak leaves only about 385 MiB for additional GPU allocations under the cap.
## Fixed validation ensemble diagnostic — 2026-09-17
Saved, identity-matched validation logits were combined with fixed equal weights
and the unchanged crossfit-temperature policy; no weights were tuned and no
reserved predictions were accessed. GX10 4B + Spark A 4B scores **93.16% /
0.193983 NLL**, worse than GX10 alone (**93.55% / 0.188640**). Spark A 4B + the
completed Spark B 2B scores **93.55% / 0.180560**. This more diverse pair shares
19 errors versus 29 for the two-4B pair, but gains eight/losses eight versus GX10.
The mixed pair's NLL difference versus GX10 is -0.008080; a 1,000-replicate paired
source-group bootstrap with fold-temperature refitting yields 95% interval
**[-0.039064, +0.019451]**. No gain over the best single model is established.
These are exploratory validation results from already selected checkpoints,
not independent generalization evidence. The deployed-candidate protocol remains
individual models; ensemble inference and latency have not been implemented or
measured. The exact A step-2,500 and B step-2,000 artifacts are frozen on GX10 in
`20260917T022201Z-ensemble-reference`, with 134.3 MB copied, stable source hashes
and CPU reconstruction/provenance checks. No weights are in Git.
[Analysis and provenance](results/20260917-fleet-progress/fixed-ensemble-validation.json).
## Remaining gates
The active continuation state and checkpoint-resume verification are recorded in
[HANDOVER.md](HANDOVER.md). Public multi-family training and the frozen unseen-family
holdout are now implemented; independent calibration/test/holdout metrics await
the 24-hour campaign's finalization. Frozen-head and generation controls, new-model
prefix caching and the larger latency matrix remain open. The Sparks now host
independent candidate experiments; GX10 does not need a ConnectX cable for this
selection strategy. Architecture B and
distributed training still await quality and profiling evidence in [PLAN.md](PLAN.md).
## Expanded public training data — 2026-09-17
Version 2 retains all 40,915 original 4B-compatible training decisions and adds
16,000 HellaSwag, 14,360 PIQA and 9,490 CommonsenseQA decisions: **80,765 total**.
All four reserved source files and tokenized sequences match version 1 exactly.
The 383 retained new-source diagnostics stay outside training and checkpoint
selection. An independent reconstruction audit checked every added source label
and shuffled answer position, all downloaded hashes and diagnostic exclusions.
See [EXPANDED_DATA.md](EXPANDED_DATA.md) and its linked small evidence.
Clean source `24b8ccf`, pilot `20260917T070758Z-train`: eight finite updates,
**exit 0**, all initial/final FP32 gates passed. Frozen Spark B parent step 1,500
reproduces every initial validation logit and probability exactly. Captured step
0 has empty Adam; every final Adam counter is eight. Median update 8.90 seconds,
peak allocation including checks 15.426 GiB, worst final probability discrepancy
3.58e-7. The 32 sampled decisions cover all seven task families.
Step 8 scores 0.170108 validation crossfit NLL versus parent 0.170150, both
94.7266% accuracy. The difference is below the fixed 0.001 selection threshold;
the selected branch remains step 0. This is startup evidence, not a claim of
improvement on the added tasks. The new campaign `20260917T072142Z-24h` restores
all step-8 model/Adam/Python/Torch/CUDA states exactly and retains the original
16:00 / 18:16:10 UTC deadlines. GX10's former run stopped at step 4,380 with both
trainer and supervisor exit 0, preserving its selected step 2,500. The fleet
retains all previous candidates and explicitly registers the new dataset.
All 86 source tests passed, including rejection of reserved-data changes and
unregistered candidate datasets. The actual expanded candidate passed the full
fleet eligibility path. Expanded checkpoints and transformed data were uploaded
and verified in the existing private Hugging Face repository at 07:24:36 UTC,
with exact source revisions, upstream notices and checksums; publication receipts
are recorded separately.
At 07:28:29 UTC the resumed campaign passed its full startup audit: every one of
512 pilot-step-8 predictions reproduced exactly, full optimizer/RNG state matched,
and updates 9–12 were finite under the cap. GX10 reached step 15 by 07:28:58 UTC;
both Spark trials, the coordinator, GUI and final-publication watcher remained
running. Final campaign evaluation is still pending.