Download source/docs/operations/results-history.md from andyshu/opensysone: direct link, hf CLI and curl.
- Browser
- Download file 36.8 kB
-
https://huggingface.co/andyshu/opensysone/resolve/main/source/docs/operations/results-history.md
- Command line
-
hf download hf://andyshu/opensysone/source/docs/operations/results-history.md
-
curl -L -o results-history.md https://huggingface.co/andyshu/opensysone/resolve/main/source/docs/operations/results-history.md
OpenSysOne results
Completed 4B accuracy and speed profile β 2026-09-17
Training and profiling are finished. The selected model is Qwen3-4B-Instruct-2507 with rank-8 LoRA and a learned scalar decision head: Spark B step-1,500 weights, retained unchanged at expanded branch step 0. Selection used validation only. The complete report includes per-family results, all timing cells, machine/source/checkpoint provenance, CSV/JSON data and standalone charts.
| Reserved evaluation | Decisions | Selected accuracy | Pretrained verifier accuracy | Selected calibrated NLL | Base calibrated NLL |
|---|---|---|---|---|---|
| Original four-family test | 2,042 | 92.90% | 84.48% | 0.2051 | 0.4527 |
| Social IQA family holdout | 768 | 72.92% | 70.31% | 0.6783 | 0.7425 |
Paired, source-group-stratified 95% bootstrap intervals (400 resamples) put the accuracy gains at +8.42 pp [6.85, 9.89] and +2.60 pp [0.13, 5.34]. The holdout improvement is modest; this is one task family. Social IQA was excluded from our fine-tuning, but exposure in the pretrained base model is unknown. The test contains four-choice Banking77, BoolQ, ARC and SNLI; this is not a general-intelligence score or a comparison with the hosted Jev service.
Temperature 1.745822 was fitted on 510 separate calibration decisions. Selected ECE changes from 4.27% to 1.08% on the test and 15.55% to 8.30% on Social IQA; Brier changes from 0.1172 to 0.1098 and 0.4187 to 0.3797 respectively. Calibration helps these evaluations but does not establish reliability on arbitrary inputs.
| Matched profile method | Accuracy, same 320 decisions | Warm median, 128-state-token / four-choice | Warm median, 768-state-token / four-choice |
|---|---|---|---|
| Selected scorer | 89.06% | 0.902 s | 3.710 s |
| Pretrained per-option verifier | 80.94% | 0.806 s | 3.177 s |
| Pretrained joint answer-label method | 86.25% | 0.213 s | 0.818 s |
These warm local measurements use the same otherwise idle Spark in FP32, one question per request, and include tokenization/probability construction. They exclude loading, HTTP, generated explanations and concurrent serving. The joint-label method conditions on all options together. The current scorer is 11β17% slower than the per-option base and 2.18β15.58 times slower than the joint-label baseline across all 12 workload cells. Ten repetitions per cell make p95 exploratory. A scalar head alone has not made this implementation faster; shared-prefix caching and merged adapters remain future measured experiments.
The expanded step-159 checkpoint scores 312/383 (81.46%) expansion diagnostics versus selected 303/383 (79.11%), while losing one answer on the matched 320. This post-selection comparison is descriptive and does not change the winner. All final/resumable states remain preserved. Training, final validation, full evaluation and profiling exited 0. Final correctness differences were at most 2.65e-7 against a 1e-4 tolerance. Four-model GUI browser checks passed. Source/control details are in handover.md; independent profile proofs are in profile-audit.
Historical smoke results β 2026-09-16
The 0.5B model trains, its artifact reconstructs, and FP32 shared-prefix scoring passes correctness checks. BF16 failed batch invariance on this checkpoint and stack. The result supports continuing the experiment; it does not establish a useful zero-shot decision model or calibrated deployment probabilities.
Reference experiment
Completed run: 20260916T155124Z, source commit 34a993e, exit 0.
Full artifacts: GX10 /home/andy/ai/opensysone/runs/20260916T155124Z/.
Small artifacts: results/20260916T155124Z.
Pinned pretrained Qwen/Qwen2.5-0.5B revision
060db6499f32faf8b98477b0a26969ef7d8b9987: 494,033,665 total parameters including
the scalar head, 29,825,665 trainable (final two layers and head). FP32 weights,
AdamW state and inference, SDPA; no LM next-token training loss or generated answers.
Backbone learning rate 2e-5, head 1e-3, gradient clipping 1, 60 steps, four decisions
per batch. The rest of the pretrained backbone is frozen.
Invented inventory facts supply 192 train, 48 calibration and 72 test decisions. Each group shares one state across color, seal and quantity questions; candidate orders are shuffled. Entity IDs are disjoint, but templates and underlying fact combinations overlap. These are simple wiring/optimization examples, not a semantic holdout or a real task-family generalization benchmark.
| Same FP32 run; 72 test decisions | Accuracy | NLL | Brier, class sum | Top-label ECE, 10 bins |
|---|---|---|---|---|
| Base yes-minus-no token score | 66.7% | 1.090 | 0.574 | 0.290 |
| Initial zero scalar head | 27.8% | 1.059 | 0.639 | 0.083 |
| Trained scalar | 68.1% | 0.628 | 0.419 | 0.205 |
| Trained + calibration-split temperature | 68.1% | 0.568 | 0.370 | 0.141 |
The trained model gets one more example correct than the matched token baseline. This is not evidence of an accuracy gain. NLL/Brier improve on this tiny synthetic set; the temperature (1.88365) was selected using only the separate 48-example calibration split. There is no basis for a general calibration claim. The uniform head's low ECE despite poor accuracy illustrates why ECE alone is not the selection criterion. NLL uses stable log-softmax, without probability clipping.
The optimization loop including periodic saves took 8.80 seconds; median step
was 105 ms. This is partial tuning on very short inputs and is not a full-model
training throughput estimate. Maximum allocated CUDA memory over training, eval
and the timing grid was 3.43 GiB, reserved 3.65 GiB, against a 16 GiB cap.
The query-projection probe changed by max 0.000964; scalar weight norm became
0.2667. Full parameter and artifact provenance is in manifest.json.
Shared-prefix correctness and timings
The tiny random FP32 CPU model passes four tests, including mixed lengths, chunk sizes 1/2/4/16, candidate permutation, unrelated-question perturbation, gradient flow and equivalence of selected token logits to full vocabulary logits.
On the trained GPU model, probability maximum absolute differences were:
| Comparison | Difference |
|---|---|
| Full forward vs shared prefix | 0.00000304 |
| Question batch vs isolated question | 0.00000381 |
| Candidate permutation, restored order | 0.00000131 |
| Repeated prefix call | 0 |
| Reset all trainable tensors, reload checkpoint | 0 |
The first successful run recorded the original 0.02 tolerance. Its actual errors are below 0.000004. The continuation harness tightens FP32 tolerance to 0.0001; BF16 retains the original gate so its known failure remains visible.
Illustrative end-to-end warm medians, including tokenization, cache copies and device synchronization. One warm-up plus three measured repeats per cell; these are not p95 or production claims. Same trained checkpoint/serialized token IDs in both modes, two candidates/question, maximum eight branches per chunk. The full reference is already batched fairly (four two-choice questions at once).
| Actual state-prefix tokens | Questions | Full batched forwards | Shared prefix | Speedup |
|---|---|---|---|---|
| 143 | 1 | 35.9 ms | 47.2 ms | 0.76Γ |
| 143 | 4 | 117.3 ms | 50.6 ms | 2.32Γ |
| 143 | 16 | 464.5 ms | 129.2 ms | 3.60Γ |
| 1,031 | 1 | 312.0 ms | 184.6 ms | 1.69Γ |
| 1,031 | 4 | 1,262.6 ms | 201.6 ms | 6.26Γ |
| 1,031 | 16 | 5,033.1 ms | 336.1 ms | 14.97Γ |
Caching loses on the shortest one-question case. At 1,031 tokens/16 questions, the shared run spends about 156 ms in prefill and 178 ms in branches; single-prefix KV occupies 24.2 MiB before the per-chunk copies. That longer-context point demonstrates amortization in this implementation. The repeated short question is a workload timing probe, not a semantic multi-question benchmark. No generation, constrained decoding, service throughput or 1.5B/3B latency comparison has run.
Failed BF16 experiment and diagnosis
Run 20260916T154714Z, source 4d6cb0f, completed its 60 training steps but
exited 1 at the correctness gate. The checkpoint and all earlier predictions
remain available; no performance conclusion was taken from that failed run.
The same trained BF16 weights were evaluated with different precision/backends:
| Comparison | BF16 probability difference | Same weights cast to FP32 |
|---|---|---|
| Full vs shared | 0.08544 | 0.00000727 |
| Shared vs isolated | 0.09897 | 0.00000519 |
| Shared candidate permutation | 0.07889 | 0.00000137 |
| Batched full vs separate full calls | 0.05262 | 0.00001433 |
SDPA MATH retains the BF16 failure and passes in FP32. This demonstrates precision and batch-shape sensitivity beyond cache handling; it does not isolate the root cause to a specific kernel or prove every GB10/model fails in BF16. The BF16 token baseline had different metrics from FP32 and must not be mixed into the matched FP32 comparison above. BF16 AdamW also lacks FP32 master weights in this simple implementation, making small updates prone to rounding.
Raw evidence: parity_diagnosis.json.
Reproducer: scripts/diagnose_parity.py --run <failed-run-directory>.
Keep the FP32 reference; investigate BF16 explicitly before scaling.
Expanded precision investigation β 2026-09-16
Completed read-only runs 20260916T161253Z-precision and
20260916T161355Z-precision, both exit 0. The second run used clean source
commit 409ade4; the first manifest records 94a24e8 with staged additions,
whose script hashes correspond to b9dd165. Full artifacts are under GX10
/home/andy/ai/opensysone/runs/<run-id>/artifacts/; small copies are in
results/20260916T161355Z-precision.
All ablations reconstruct the preserved BF16-trained checkpoint from
20260916T154714Z; its SHA-256 remained
106efdfb0794e6ca870b7add11c71f06c58281ef46b348305866a85f1e6f6bc8.
The base, data and checkpoint are unchanged. This comparison concerns arithmetic
on the same weights, rather than FP32 versus BF16 training quality. It does not
evaluate a new task or supply generalization evidence.
The expanded test covers all 24 groups / 72 decisions, comparing batched
full calls with separate question calls, full with cached, cache chunks of 4/16,
cached with isolated questions, and restored candidate permutations. The table
shows the worst absolute probability difference across these comparisons.
Strict reduction sets allow_bf16_reduced_precision_reduction=False. FP32 linear
casts each decoder linear's inputs and weights to FP32, then casts its output
back to BF16; it is an inference diagnostic, not a validated training method.
| Arithmetic configuration | Worst probability difference | Groups above BF16's original 0.02 gate |
|---|---|---|
| BF16 default SDPA | 0.238608 | 24/24 |
| BF16, strict reduction | 0.213011 | 24/24 |
| BF16, math SDPA + strict reduction | 0.168008 | 24/24 |
| BF16, FP32 linear + strict reduction | 0.147468 | 24/24 |
| BF16, math SDPA + FP32 linear + strict reduction | 0.183657 | 24/24 |
| Same weights cast to FP32, default SDPA | 0.00002138 | 0/24 |
FP32 also passes the stricter 0.0001 gate. Repeated full and repeated cached calls have exactly zero probability difference in every group/configuration. Full candidate permutations also match exactly; cached permutations can change which branches share a chunk and still fail in BF16. Exit 0 means the diagnostic completed, not that BF16 passed.
Final-candidate-token traces for the first serialized group narrow the issue: embeddings and first input normalization match exactly, but default BF16's first query/key projections differ by up to 0.5 between batched/separate calls. Strict reduction removes those initial projection differences in this trace; later differences remain. Combined math attention and FP32 linears reduce the first decoder-layer difference from 0.02344 to 0.00003052, yet the final normalized hidden representation still differs by up to 2.0. This supports shape-dependent numerical differences that propagate through the decoder. It does not isolate every contributing operation or establish a particular kernel defect. The trace samples final candidate tokens, not every token's intermediate representation.
Peak CUDA allocation was 1.90 GiB, reserved 1.94 GiB, against the 16 GiB cap. The shared environment was unchanged and OOM score adjustment was 0. All four CPU correctness tests passed before execution. Both diagnostic PIDs exited; at 16:15 UTC GX10 again had about 118 GiB available and only the original router GPU process. Continue useful model/data work in FP32; none of these BF16 interventions justifies reopening its correctness gate.
Public-data 2B adapter pilot β 2026-09-16
Run /home/andy/ai/opensysone/runs/20260916T182352Z-train/artifacts,
execution source f1c9322, exited 0 after 40 optimizer steps
(160 decisions), not three completed epochs. The model is pinned
Qwen/Qwen3.5-2B at 15852e8c16360a2fea060d615a32b45270f8a8fc.
Only its text decoder is retained; the unused vision encoder is discarded before
CUDA loading. Rank-16 additive linear adapters and a pretrained yes-minus-no
initialized head train 16,821,249 of 1,898,646,337 parameters in FP32.
The frozen data has 40,941 source-group-disjoint train decisions, 512 validation, 512 calibration, 2,048 source test and 768 completely held-out Social IQA decisions. This model's 768-token complete-chat limit excludes four BoolQ train rows and one test row, leaving 40,937/512/512/2,047/768. Banking77 is a four-choice target-plus- three-negative transformation, not a full 77-way benchmark. Source-group splitting does not rule out pretraining contamination or semantic duplicates.
| Family | Initial validation accuracy | Step 40 accuracy |
|---|---|---|
| ARC | 74.22% | 81.25% |
| Banking77 four-choice | 83.59% | 88.28% |
| BoolQ | 64.84% | 79.69% |
| SNLI | 64.06% | 82.03% |
| All 512 decisions | 71.68% | 82.81% |
Validation macro-family NLL fell from 0.700136 to 0.498153. This is validation selection evidence, not untouched test improvement. No calibration, test or Social IQA predictions have been evaluated in this pilot. Median four-decision step was 3.869 s; the loop including final validation took 298.1 s. Peak CUDA allocation/reservation was 7.746/7.855 GiB, below the 16 GiB cap. All final permutation/chunk/isolation checks passed the 0.0001 probability gate, with worst difference 0.00000614. The old repeat label also changed chunk shape; the current source restores the original chunk size before repeat testing.
A fresh process in 20260916T183240Z-train, source 980d881, reconstructed
step 40 and reproduced all 512 raw logits and probabilities exactly, restored
optimizer/RNG, then completed step 41 with finite gradient norm 3.676.
It exited 0 and all final parity gates passed, worst difference 0.00000316.
Step 41 validation macro NLL was 0.494478. The retained setup failure
20260916T182256Z-train exited 1 before any optimizer step because Transformers'
new chat-template return default was a BatchEncoding; explicit return_dict=False
fixed it without changing the shared environment.
Small raw pilot evidence is in results/20260916T182352Z-train.
Checkpoint SHA-256 is
af5790ae2f2b56477ebbdf6ab9c418d895e48d2bf5a6416e11b6e9863ad1db55;
validation-selected best SHA-256 is
82b4261feb98d3ed56291e4c03304a65da20ce0194a6ad117d113b30d152282e.
New dependencies are isolated in ~/ai/envs/opensysone (pyarrow 25.0.1), with
read-only reuse of the existing torch/Transformers packages. The Jev-compatible
stdlib harness and 12 CPU tests pass; real-checkpoint HTTP and longest-input
stress are the next gate before the larger campaign.
Public-data 4B pilot selected for the 24-hour run
Run /home/andy/ai/opensysone/runs/20260916T183823Z-train/artifacts, clean execution
source ccbbe6d, exited 0 after 40 steps / 160 decisions. The base is
Qwen/Qwen3-4B-Instruct-2507, pinned to
cdbee75f17c01a7cc42f958dc650907174af0554, Apache-2.0.
Rank-8 adapters (alpha 16) and the pretrained initialized head train
16,517,633 of 4,038,985,729 parameters in FP32. Exact two-pass categorical
gradients keep one candidate graph live; CPU gradients match ordinary CE within
0.000001. Gradient checkpointing is enabled. No quantization or new kernels.
| Family | Initial validation accuracy | Step 40 accuracy | Step 40 NLL |
|---|---|---|---|
| ARC | 90.63% | 90.63% | 0.374400 |
| Banking77 four-choice | 90.63% | 91.41% | 0.229941 |
| BoolQ | 82.03% | 84.38% | 0.583093 |
| SNLI | 82.81% | 83.59% | 0.395211 |
| All 512 validation decisions | 86.52% | 87.50% | 0.395661 |
Raw validation macro NLL improves from 1.436162 to 0.395661; the initial readout was severely overconfident. A separately recorded diagnostic fits and scores a temperature on the same validation rows (NLL 0.407737, T 6.9183): it is optimistic validation analysis, not independent calibration. Reserved calibration, test and Social IQA predictions remain unevaluated. The trained 4B validation accuracy and NLL beat the 2B pilot in every family, supporting the larger candidate despite its lower throughput. This does not prove task generalization.
The 512-token complete-chat limit retains 40,915 train / 512 validation / 510 calibration / 2,042 test / 768 Social IQA decisions; it drops 26 train, two calibration and six test BoolQ rows, with no silent truncation. Median four-decision step is 8.956 s; 55,268 actual branch tokens were processed with no padding overhead. The loop including final validation takes 724.0 s. Initial validation alone takes 327.85 s. Peak CUDA allocated/reserved is 15.510/15.604 GiB against the 16 GiB cap. OOM adjustment is 0 and about 99 GiB unified RAM remains available with the model loaded. Final correctness passes all 0.0001 gates, worst probability difference 0.00000167, with exact repeated, isolated and restored-permutation predictions.
Checkpoint SHA-256:
e26f75b2396de88311873fac4eb91e1e40d0ec940778ec99f282bcfd96a2e258.
Best SHA-256:
64977ee0b1a6147c6faf59283edea9adf564dd36d53f4580bc20940b94c6764f.
Small raw evidence is in results/20260916T183823Z-train.
Fresh reload, longest-input gradients with restored optimizer state, 1,024-token
HTTP inference, and 255-choice HTTP stress all passed (verification exit 0).
Reload matches all 16 checked validation predictions exactly. Longest training
input is 509 tokens and peaks at 15.624 GiB with optimizer state; inference peaks
at 15.465 GiB. The long HTTP request has 1,023 tokens in each of two candidate
branches and matches direct inference exactly. Invalid-key/oversized-input
requests return 401/422. One warm three-question request takes 1.571 s, and one
255-choice request takes 47.042 s; these are wiring stress timings, not latency
percentiles or intelligence benchmarks. The checkpoint SHA-256 is unchanged.
Evidence: results/20260916T185718Z-verify4b.
All 15 CPU tests pass, including unequal-source-group bootstrap weighting
and the measured evaluation-reserve calculation. The live Jev HTTPS endpoint
returns 405 to an unauthenticated GET; no credentials or state were sent and no
authenticated hosted inference has been tested.
Detached 24-hour campaign now running
Launched 2026-09-16 18:59:10 UTC from clean source 0109eb6 into
/home/andy/ai/opensysone/runs/20260916T185910Z-24h. Supervisor PID is 1085496,
current trainer 1085517; both OOM score adjustments are 0. Training resumes the
4B step-40 checkpoint with optimizer/RNG restored, preserves validation-selected
best and all model/data/config signatures, and has passed the initial FP32
correctness gates. Exit is pending; the API has not started yet.
Fresh restart reproduces all 512 raw logits and probabilities exactly;
the reference and fresh prediction JSON SHA-256 are both
e671e1508185765552b0f933ba03f356be62143c531d8ef534457d34b1645c9b.
The next four updates, 41β44, have finite losses/gradients and remain under
the cap. Step 41 takes 8.724 s, loss 0.115940, gradient norm 3.81358.
This proves reconstruction plus subsequent optimizer updates, not a bitwise
interrupted-versus-uninterrupted trajectory comparison. Raw verification is in
the launch evidence directory below. The durable checkpoint remains step 40
until the regular save cadence, independently of those logged newer updates.
Training ends by 2026-09-17 16:16:10 UTC, reserving two hours until the final 18:16:10 UTC / 19:16:10 BST deadline. The reserve estimates 6,640 base/tuned prediction rows at 4,251.8 seconds from measured pilot validation speed, adds 30% plus ten minutes for setup, and keeps a two-hour minimum. Checkpoints save every 250 steps or 900 seconds regardless of evaluation; validation is every 500 steps with patience eight. The three-epoch target is an upper bound.
After successful training, the runner loads the best artifact fresh, calibrates
only on the 510 reserved known-family decisions, saves a deployable checkpoint
before untouched evaluation, records raw/calibrated test and Social IQA metrics
against the unchanged pretrained scorer, and starts the loopback API only after
complete evaluation and a real-model inference check. Deployment is planned at
http://127.0.0.1:18081, with 1,024-token inputs. No hosted Jev call runs
automatically. Small launch evidence lives in
results/20260916T185910Z-24h-launch, separate
from the completion-results directory reserved by the runner.
Current inspection, stop and same-deadline recovery commands are in handover.md. A running job is not a finalized model or successful test result. The frozen-family controls and independent calibration remain the quality gates for final reporting. The complete 15-test suite passed; the new orphan-child stop safeguard also passes an integration test that refuses to terminate a PID when its command line differs from the recorded command.
Three-machine expansion β 2026-09-16 evening
The user assigned GX10 and both Sparks to this task and authorized terminating their workloads. The Spark serving head and RPC worker were stopped in order with verified SIGTERM; both released their GPU allocations and each had about 118 GiB available afterward. Their weights/cache and exact restoration commands are retained. No network or system configuration changed.
Both Sparks now have isolated copies of the exact GX10 training dependencies: 21,368 installed file hashes and 55 package versions match. CPU autograd and both Qwen-family imports pass. This initial check verified the environments. Subsequently all pinned model files and real GPU training/HTTP checks passed; see the launch results below.
The original GX10 campaign saved step 128 before a requested stop. Its trainer exceeded the 30-second grace while performing final correctness checks and exited -9; the complete step-128 checkpoint and optimizer/RNG are verified intact. The new source records skipped final checks explicitly on a requested stop. It also retains step-specific prediction evidence before publishing each new best artifact and selects a restored checkpoint if its fresh validation improves the best.
The intermediate GX10 campaign 20260916T192239Z-24h, source 6e080e2, restored
step 128 and later stopped gracefully at step 178 with training exit 0.
The Spark alternatives are a 4B weights-only warm initialization with fresh Adam,
learning rate 0.00003 and 7,500-step cosine horizon, and a longer 2B continuation.
The planned fleet cutoff is 2026-09-17 16:00 UTC, leaving 2 h 16 min until the
original final deadline. All training remains under 16 GiB per process.
All 32 initial fleet CPU tests passed, including weights-only initialization, optimizer/RNG resume, requested-stop evidence, deadline handling, exact validation-set matching, checkpoint/metric mismatch rejection and API deployment lifecycle. The coordinator recomputes its criterion from all 512 saved validation predictions and freezes selection before calibration/test/holdout. Read-only compatibility checks of the real 4B and 2B pilot artifacts pass, reproducing NLL 0.395661 and 0.498153. Small setup proofs are in results/20260916-fleet-setup. Live paths, statuses and recovery instructions are in fleet.md.
The subsequent selection revision uses the frozen four-fold source-group-disjoint
temperature-crossfit policy crossfit_temperature_nll_v1 (seed 431, 101 positive
temperatures, family-balanced fitting and scoring). Step 128's validation accuracy
is 89.0625%, versus step 40's 87.5%; raw NLL is 0.442683 versus 0.395661.
Crossfit NLL reverses that ranking: 0.318518 versus 0.359522, improving in all
four families. A 5,000-replicate paired source-group bootstrap, refitting the
temperatures, gives difference -0.041005 with 95% interval [-0.079822, -0.001152].
The accuracy gain alone is uncertain (29 gains, 21 losses; McNemar p=0.322).
This supports accounting for recoverable overconfidence during checkpoint
selection. It is a validation-driven criterion revision, not independent test
evidence. No reserved predictions were read. Original raw-selected checkpoints
remain preserved, and final calibration still uses the separate reserved split.
All 37 tests pass after adding policy/selection checks; the updated CPU
integration also proves reselection leaves trained weights and Adam steps intact.
Raw diagnostic: selection-diagnostic.json.
Active fleet launch β 2026-09-16 19:44 UTC
Three training-only campaigns are active on source 4a60423:
GX10 20260916T193741Z-24h (4B, LR 0.0001), spark-a
20260916T194258Z-24h (4B, LR 0.00003, 7,500-step cosine horizon), and spark-b
20260916T193803Z-24h (2B, LR 0.0001). Each uses the fixed crossfit criterion,
16 GiB allocation cap and 2026-09-17 16:00 UTC cutoff. Training exit statuses
remain pending. The fleet coordinator 20260916T194403396250Z-fleet, source
6a7b0ed, is detached on GX10 and waiting for selection; no reserved-data
predictions or final calibration have run. The cutoff shutdown race is covered
by a regression test, and all nine fleet tests pass after that fix.
GX10 restored step 178's weights, Adam and Python/torch/CUDA RNG exactly. Its fresh 512-decision validation reached 90.4297% accuracy, 0.303825 crossfit NLL, 0.404198 raw NLL, promoting the durable best beyond step 128. Fresh FP32 correctness passes (worst probability difference 4.77e-7), and resumed updates are finite. These are validation results, not independent test evidence.
Spark A reproduced all 512 original 4B pilot predictions exactly before eight finite lower-rate updates (median 8.086 seconds, peak 15.505 GiB). That pilot exited 0; its step-8 accuracy 86.914% / raw NLL 0.405397 did not improve the starting checkpoint. The long-run crossfit selector re-evaluates both inherited best and current checkpoint. Real fresh-artifact verification exited 0: exact 16-decision reload, finite restored-Adam gradients on the longest 509-token input, 15.624 GiB peak, 1,023-token HTTP/direct match, expected 401/422 errors, and 255 choices in 43.31 seconds. The long campaign reproduced all 512 step-8 raw predictions exactly, with identical weights/Adam/RNG. Its fixed crossfit criterion selected step 8 at 0.358235 NLL, and new updates are finite. No independent generalization improvement is claimed for the short pilot.
Spark B's preparation exited 0. Fresh verification passed exact reload, restored-Adam gradients at 700 tokens (8.123 GiB peak), 1,024-token inference, authentication/length errors and 255 choices in 17.73 seconds. The long campaign reproduced all 512 original validation predictions exactly, scoring 82.8125% accuracy / 0.476072 crossfit NLL / 0.498153 raw NLL before resumed training. Subsequent finite updates reached step 98 by 19:43:59 UTC. Timing observations are individual wiring checks, not p50/p95 latency measurements.
Full small evidence, source revisions, frozen plan and startup state snapshots
are under results/20260916-fleet-setup. Live
state, inspection/stop/resume and serving-pair restoration are in
fleet.md. Final calibrated test/holdout metrics and selected-model
API deployment are pending; authenticated hosted Jev inference still requires
TYPESAFE_API_KEY.
Overnight progress and next experiment β 2026-09-17
The 02:10β02:15 UTC audit found both 4B jobs healthy and improving, while the 2B campaign completed cleanly at 01:59:29 UTC, training and supervisor exit 0. All recorded losses/gradients were finite. Peak allocation was 15.624 GiB on each 4B job and 8.183 GiB on the 2B job; the 16 GiB cap remains unchanged.
| Candidate | Last audited step | Selected step | Crossfit validation NLL | Selected accuracy |
|---|---|---|---|---|
| GX10 4B, LR 1e-4 | 2,570 | 2,500 | 0.188640 | 93.55% |
| Spark A 4B, LR 3e-5 | 2,529 | 2,500 | 0.218012 | 92.58% |
| Spark B 2B, LR 1e-4 | 6,000 | 2,000 | 0.255294 | 89.84% |
These are the same 512 validation decisions, selected with the unchanged fixed crossfit policy. No reserved calibration, test or Social IQA predictions have been read. A higher maximum accuracy at a different step does not override the selection criterion. Spark A improved at all five scheduled validations. Spark B stopped after eight evaluations without a new best; its final step-6,000 score was 0.351593 / 86.91%. Final numerical correctness passed at worst 6.56e-7. The selected step-2,000 and resumable step-6,000 artifacts are preserved.
The freed Spark B is training a fourth candidate, initialized from a frozen
copy of GX10's step-2,500 selected weights (SHA-256
8956eb6c0cfbb02124aeefd99c3b418c55f55fdb9a64260350622d98dbba1aec).
Fresh Adam, seed 432, LR/head LR 1e-5 and a 5,000-step cosine horizon define a new
trajectory. Other model/batch/token/correctness settings and both absolute
deadlines stay unchanged. The eight-step pilot started at 02:16:05 UTC;
source 4a60423. The pinned 4B model copied from Spark A over the existing link
passed all 13 file hashes. Warm initialization preserves all 506 trainable tensors
exactly and deliberately starts with an empty optimizer. All 512 initial raw
predictions match the parent exactly. The eight-step pilot and fresh verifier
exited 0: exact 16-decision reload, finite longest-input gradients, 15.623 GiB
peak, direct/HTTP agreement at 1,023 tokens, expected 401/422 and 255 choices
in 45.44 seconds. These timings are individual wiring checks, not percentiles.
Campaign 20260917T023137Z-24h launched at 02:31:37 UTC, restoring the complete
step-8 optimizer/RNG state exactly, and was added as the fourth fleet candidate.
Its inherited selected branch step 0 retains the parent score: step 8 scored
0.188576, a change below the fixed 0.001 improvement threshold. The short pilot
does not establish a quality gain.
Small audit evidence records the 22 scheduled validation points, live processes, source revisions, selected-checkpoint hashes and frozen refinement parent. next-steps.md records the decisions and the two-Spark alternatives: independent candidates now, bounded distributed adapter-gradient training or parallel scoring next. Active ConnectX/RoCE and installed NCCL do not establish collective correctness or useful speedup. The current two-pass trainer needs explicit synchronization changes, and its measured peak leaves only about 385 MiB for additional GPU allocations under the cap.
Fixed validation ensemble diagnostic β 2026-09-17
Saved, identity-matched validation logits were combined with fixed equal weights
and the unchanged crossfit-temperature policy; no weights were tuned and no
reserved predictions were accessed. GX10 4B + Spark A 4B scores 93.16% /
0.193983 NLL, worse than GX10 alone (93.55% / 0.188640). Spark A 4B + the
completed Spark B 2B scores 93.55% / 0.180560. This more diverse pair shares
19 errors versus 29 for the two-4B pair, but gains eight/losses eight versus GX10.
The mixed pair's NLL difference versus GX10 is -0.008080; a 1,000-replicate paired
source-group bootstrap with fold-temperature refitting yields 95% interval
[-0.039064, +0.019451]. No gain over the best single model is established.
These are exploratory validation results from already selected checkpoints,
not independent generalization evidence. The deployed-candidate protocol remains
individual models; ensemble inference and latency have not been implemented or
measured. The exact A step-2,500 and B step-2,000 artifacts are frozen on GX10 in
20260917T022201Z-ensemble-reference, with 134.3 MB copied, stable source hashes
and CPU reconstruction/provenance checks. No weights are in Git.
Analysis and provenance.
Remaining gates
The active continuation state and checkpoint-resume verification are recorded in handover.md. Public multi-family training and the frozen unseen-family holdout are now implemented; independent calibration/test/holdout metrics await the 24-hour campaign's finalization. Frozen-head and generation controls, new-model prefix caching and the larger latency matrix remain open. The Sparks now host independent candidate experiments; GX10 does not need a ConnectX cable for this selection strategy. Architecture B and distributed training still await quality and profiling evidence in plan.md.
Expanded public training data β 2026-09-17
Version 2 retains all 40,915 original 4B-compatible training decisions and adds 16,000 HellaSwag, 14,360 PIQA and 9,490 CommonsenseQA decisions: 80,765 total. All four reserved source files and tokenized sequences match version 1 exactly. The 383 retained new-source diagnostics stay outside training and checkpoint selection. An independent reconstruction audit checked every added source label and shuffled answer position, all downloaded hashes and diagnostic exclusions. See training-data.md and its linked small evidence.
Clean source 24b8ccf, pilot 20260917T070758Z-train: eight finite updates,
exit 0, all initial/final FP32 gates passed. Frozen Spark B parent step 1,500
reproduces every initial validation logit and probability exactly. Captured step
0 has empty Adam; every final Adam counter is eight. Median update 8.90 seconds,
peak allocation including checks 15.426 GiB, worst final probability discrepancy
3.58e-7. The 32 sampled decisions cover all seven task families.
Step 8 scores 0.170108 validation crossfit NLL versus parent 0.170150, both
94.7266% accuracy. The difference is below the fixed 0.001 selection threshold;
the selected branch remains step 0. This is startup evidence, not a claim of
improvement on the added tasks. The new campaign 20260917T072142Z-24h restores
all step-8 model/Adam/Python/Torch/CUDA states exactly and retains the original
16:00 / 18:16:10 UTC deadlines. GX10's former run stopped at step 4,380 with both
trainer and supervisor exit 0, preserving its selected step 2,500. The fleet
retains all previous candidates and explicitly registers the new dataset.
All 86 source tests passed, including rejection of reserved-data changes and unregistered candidate datasets. The actual expanded candidate passed the full fleet eligibility path. Expanded checkpoints and transformed data were uploaded and verified in the existing private Hugging Face repository at 07:24:36 UTC, with exact source revisions, upstream notices and checksums; publication receipts are recorded separately.
At 07:28:29 UTC the resumed campaign passed its full startup audit: every one of 512 pilot-step-8 predictions reproduced exactly, full optimizer/RNG state matched, and updates 9β12 were finite under the cap. GX10 reached step 15 by 07:28:58 UTC; both Spark trials, the coordinator, GUI and final-publication watcher remained running. Final campaign evaluation is still pending.