|
Download source/RESULTS.md from andyshu/opensysone: direct link, hf CLI and curl.
- Browser
- Download file 33.5 kB
-
https://huggingface.co/andyshu/opensysone/resolve/f2d6f8daa15bd21c316e61249f45ac16cbb79d45/source/RESULTS.md
- Command line
-
hf download hf://andyshu/opensysone@f2d6f8daa15bd21c316e61249f45ac16cbb79d45/source/RESULTS.md
-
curl -L -o RESULTS.md https://huggingface.co/andyshu/opensysone/resolve/f2d6f8daa15bd21c316e61249f45ac16cbb79d45/source/RESULTS.md
33.5 kB
| # GX10 smoke results — 2026-09-16 | |
| The 0.5B model trains, its artifact reconstructs, and FP32 shared-prefix scoring | |
| passes correctness checks. **BF16 failed batch invariance on this checkpoint and | |
| stack.** The result supports continuing the experiment; it does not establish a | |
| useful zero-shot decision model or calibrated deployment probabilities. | |
| ## Reference experiment | |
| Completed run: `20260916T155124Z`, source commit `34a993e`, exit **0**. | |
| Full artifacts: GX10 `/home/andy/ai/opensysone/runs/20260916T155124Z/`. | |
| Small artifacts: [results/20260916T155124Z](results/20260916T155124Z/). | |
| Pinned pretrained `Qwen/Qwen2.5-0.5B` revision | |
| `060db6499f32faf8b98477b0a26969ef7d8b9987`: 494,033,665 total parameters including | |
| the scalar head, **29,825,665 trainable** (final two layers and head). FP32 weights, | |
| AdamW state and inference, SDPA; no LM next-token training loss or generated answers. | |
| Backbone learning rate 2e-5, head 1e-3, gradient clipping 1, 60 steps, four decisions | |
| per batch. The rest of the pretrained backbone is frozen. | |
| Invented inventory facts supply 192 train, 48 calibration and 72 test decisions. | |
| Each group shares one state across color, seal and quantity questions; candidate | |
| orders are shuffled. Entity IDs are disjoint, but templates and underlying fact | |
| combinations overlap. These are simple wiring/optimization examples, **not a | |
| semantic holdout or a real task-family generalization benchmark**. | |
| | Same FP32 run; 72 test decisions | Accuracy | NLL | Brier, class sum | Top-label ECE, 10 bins | | |
| | --- | ---: | ---: | ---: | ---: | | |
| | Base yes-minus-no token score | 66.7% | 1.090 | 0.574 | 0.290 | | |
| | Initial zero scalar head | 27.8% | 1.059 | 0.639 | 0.083 | | |
| | Trained scalar | 68.1% | 0.628 | 0.419 | 0.205 | | |
| | Trained + calibration-split temperature | 68.1% | 0.568 | 0.370 | 0.141 | | |
| The trained model gets **one more example** correct than the matched token | |
| baseline. This is not evidence of an accuracy gain. NLL/Brier improve on this | |
| tiny synthetic set; the temperature (1.88365) was selected using only the separate | |
| 48-example calibration split. There is no basis for a general calibration claim. | |
| The uniform head's low ECE despite poor accuracy illustrates why ECE alone is | |
| not the selection criterion. NLL uses stable log-softmax, without probability clipping. | |
| The optimization loop including periodic saves took **8.80 seconds**; median step | |
| was 105 ms. This is partial tuning on very short inputs and is not a full-model | |
| training throughput estimate. Maximum allocated CUDA memory over training, eval | |
| and the timing grid was **3.43 GiB**, reserved **3.65 GiB**, against a 16 GiB cap. | |
| The query-projection probe changed by max 0.000964; scalar weight norm became | |
| 0.2667. Full parameter and artifact provenance is in `manifest.json`. | |
| ## Shared-prefix correctness and timings | |
| The tiny random FP32 CPU model passes four tests, including mixed lengths, | |
| chunk sizes 1/2/4/16, candidate permutation, unrelated-question perturbation, | |
| gradient flow and equivalence of selected token logits to full vocabulary logits. | |
| On the trained GPU model, probability maximum absolute differences were: | |
| | Comparison | Difference | | |
| | --- | ---: | | |
| | Full forward vs shared prefix | 0.00000304 | | |
| | Question batch vs isolated question | 0.00000381 | | |
| | Candidate permutation, restored order | 0.00000131 | | |
| | Repeated prefix call | 0 | | |
| | Reset all trainable tensors, reload checkpoint | 0 | | |
| The first successful run recorded the original 0.02 tolerance. Its actual errors | |
| are below 0.000004. The continuation harness tightens FP32 tolerance to **0.0001**; | |
| BF16 retains the original gate so its known failure remains visible. | |
| Illustrative end-to-end warm medians, including tokenization, cache copies and | |
| device synchronization. One warm-up plus **three measured repeats** per cell; | |
| these are not p95 or production claims. Same trained checkpoint/serialized token | |
| IDs in both modes, two candidates/question, maximum eight branches per chunk. | |
| The full reference is already batched fairly (four two-choice questions at once). | |
| | Actual state-prefix tokens | Questions | Full batched forwards | Shared prefix | Speedup | | |
| | --- | ---: | ---: | ---: | ---: | | |
| | 143 | 1 | 35.9 ms | 47.2 ms | 0.76× | | |
| | 143 | 4 | 117.3 ms | 50.6 ms | 2.32× | | |
| | 143 | 16 | 464.5 ms | 129.2 ms | 3.60× | | |
| | 1,031 | 1 | 312.0 ms | 184.6 ms | 1.69× | | |
| | 1,031 | 4 | 1,262.6 ms | 201.6 ms | 6.26× | | |
| | 1,031 | 16 | 5,033.1 ms | 336.1 ms | 14.97× | | |
| Caching loses on the shortest one-question case. At 1,031 tokens/16 questions, | |
| the shared run spends about 156 ms in prefill and 178 ms in branches; single-prefix | |
| KV occupies 24.2 MiB before the per-chunk copies. That longer-context point | |
| demonstrates amortization in this implementation. The repeated short question is | |
| a workload timing probe, not a semantic multi-question benchmark. No generation, | |
| constrained decoding, service throughput or 1.5B/3B latency comparison has run. | |
| ## Failed BF16 experiment and diagnosis | |
| Run `20260916T154714Z`, source `4d6cb0f`, completed its 60 training steps but | |
| exited **1** at the correctness gate. The checkpoint and all earlier predictions | |
| remain available; no performance conclusion was taken from that failed run. | |
| The same trained BF16 weights were evaluated with different precision/backends: | |
| | Comparison | BF16 probability difference | Same weights cast to FP32 | | |
| | --- | ---: | ---: | | |
| | Full vs shared | 0.08544 | 0.00000727 | | |
| | Shared vs isolated | 0.09897 | 0.00000519 | | |
| | Shared candidate permutation | 0.07889 | 0.00000137 | | |
| | Batched full vs separate full calls | 0.05262 | 0.00001433 | | |
| SDPA MATH retains the BF16 failure and passes in FP32. This demonstrates precision | |
| and batch-shape sensitivity beyond cache handling; it does **not** isolate the | |
| root cause to a specific kernel or prove every GB10/model fails in BF16. The | |
| BF16 token baseline had different metrics from FP32 and must not be mixed into | |
| the matched FP32 comparison above. BF16 AdamW also lacks FP32 master weights in | |
| this simple implementation, making small updates prone to rounding. | |
| Raw evidence: [parity_diagnosis.json](results/20260916T154714Z/parity_diagnosis.json). | |
| Reproducer: `scripts/diagnose_parity.py --run <failed-run-directory>`. | |
| Keep the FP32 reference; investigate BF16 explicitly before scaling. | |
| ## Expanded precision investigation — 2026-09-16 | |
| Completed read-only runs `20260916T161253Z-precision` and | |
| `20260916T161355Z-precision`, both exit **0**. The second run used clean source | |
| commit **`409ade4`**; the first manifest records `94a24e8` with staged additions, | |
| whose script hashes correspond to `b9dd165`. Full artifacts are under GX10 | |
| `/home/andy/ai/opensysone/runs/<run-id>/artifacts/`; small copies are in | |
| [results/20260916T161355Z-precision](results/20260916T161355Z-precision/). | |
| All ablations reconstruct the preserved BF16-trained checkpoint from | |
| `20260916T154714Z`; its SHA-256 remained | |
| `106efdfb0794e6ca870b7add11c71f06c58281ef46b348305866a85f1e6f6bc8`. | |
| The base, data and checkpoint are unchanged. This comparison concerns arithmetic | |
| on the same weights, rather than FP32 versus BF16 training quality. It does not | |
| evaluate a new task or supply generalization evidence. | |
| The expanded test covers **all 24 groups / 72 decisions**, comparing batched | |
| full calls with separate question calls, full with cached, cache chunks of 4/16, | |
| cached with isolated questions, and restored candidate permutations. The table | |
| shows the worst absolute probability difference across these comparisons. | |
| Strict reduction sets `allow_bf16_reduced_precision_reduction=False`. FP32 linear | |
| casts each decoder linear's inputs and weights to FP32, then casts its output | |
| back to BF16; it is an inference diagnostic, not a validated training method. | |
| | Arithmetic configuration | Worst probability difference | Groups above BF16's original 0.02 gate | | |
| | --- | ---: | ---: | | |
| | BF16 default SDPA | 0.238608 | 24/24 | | |
| | BF16, strict reduction | 0.213011 | 24/24 | | |
| | BF16, math SDPA + strict reduction | 0.168008 | 24/24 | | |
| | BF16, FP32 linear + strict reduction | 0.147468 | 24/24 | | |
| | BF16, math SDPA + FP32 linear + strict reduction | 0.183657 | 24/24 | | |
| | Same weights cast to FP32, default SDPA | 0.00002138 | 0/24 | | |
| FP32 also passes the stricter **0.0001** gate. Repeated full and repeated cached | |
| calls have exactly zero probability difference in every group/configuration. | |
| Full candidate permutations also match exactly; cached permutations can change | |
| which branches share a chunk and still fail in BF16. Exit 0 means the diagnostic | |
| completed, not that BF16 passed. | |
| Final-candidate-token traces for the first serialized group narrow the issue: | |
| embeddings and first input normalization match exactly, but default BF16's first | |
| query/key projections differ by up to **0.5** between batched/separate calls. | |
| Strict reduction removes those initial projection differences in this trace; | |
| later differences remain. Combined math attention and FP32 linears reduce the | |
| first decoder-layer difference from 0.02344 to 0.00003052, yet the final normalized | |
| hidden representation still differs by up to 2.0. This supports shape-dependent | |
| numerical differences that propagate through the decoder. It does not isolate | |
| every contributing operation or establish a particular kernel defect. The trace | |
| samples final candidate tokens, not every token's intermediate representation. | |
| Peak CUDA allocation was **1.90 GiB**, reserved **1.94 GiB**, against the 16 GiB | |
| cap. The shared environment was unchanged and OOM score adjustment was 0. | |
| All four CPU correctness tests passed before execution. Both diagnostic PIDs | |
| exited; at 16:15 UTC GX10 again had about 118 GiB available and only the original | |
| router GPU process. Continue useful model/data work in FP32; none of these BF16 | |
| interventions justifies reopening its correctness gate. | |
| ## Public-data 2B adapter pilot — 2026-09-16 | |
| Run `/home/andy/ai/opensysone/runs/20260916T182352Z-train/artifacts`, | |
| execution source **`f1c9322`**, exited **0** after **40 optimizer steps** | |
| (160 decisions), not three completed epochs. The model is pinned | |
| `Qwen/Qwen3.5-2B` at `15852e8c16360a2fea060d615a32b45270f8a8fc`. | |
| Only its text decoder is retained; the unused vision encoder is discarded before | |
| CUDA loading. Rank-16 additive linear adapters and a pretrained yes-minus-no | |
| initialized head train **16,821,249 of 1,898,646,337 parameters** in FP32. | |
| The frozen data has 40,941 source-group-disjoint train decisions, 512 validation, | |
| 512 calibration, 2,048 source test and 768 completely held-out Social IQA decisions. | |
| This model's 768-token complete-chat limit excludes four BoolQ train rows and one | |
| test row, leaving 40,937/512/512/2,047/768. Banking77 is a four-choice target-plus- | |
| three-negative transformation, not a full 77-way benchmark. Source-group splitting | |
| does not rule out pretraining contamination or semantic duplicates. | |
| | Family | Initial validation accuracy | Step 40 accuracy | | |
| | --- | ---: | ---: | | |
| | ARC | 74.22% | 81.25% | | |
| | Banking77 four-choice | 83.59% | 88.28% | | |
| | BoolQ | 64.84% | 79.69% | | |
| | SNLI | 64.06% | 82.03% | | |
| | All 512 decisions | **71.68%** | **82.81%** | | |
| Validation macro-family NLL fell from **0.700136 to 0.498153**. This is validation | |
| selection evidence, not untouched test improvement. No calibration, test or | |
| Social IQA predictions have been evaluated in this pilot. Median four-decision | |
| step was **3.869 s**; the loop including final validation took 298.1 s. | |
| Peak CUDA allocation/reservation was **7.746/7.855 GiB**, below the 16 GiB cap. | |
| All final permutation/chunk/isolation checks passed the 0.0001 probability gate, | |
| with worst difference **0.00000614**. The old repeat label also changed chunk | |
| shape; the current source restores the original chunk size before repeat testing. | |
| A fresh process in `20260916T183240Z-train`, source **`980d881`**, reconstructed | |
| step 40 and reproduced **all 512 raw logits and probabilities exactly**, restored | |
| optimizer/RNG, then completed step 41 with finite gradient norm 3.676. | |
| It exited **0** and all final parity gates passed, worst difference 0.00000316. | |
| Step 41 validation macro NLL was 0.494478. The retained setup failure | |
| `20260916T182256Z-train` exited 1 before any optimizer step because Transformers' | |
| new chat-template return default was a BatchEncoding; explicit `return_dict=False` | |
| fixed it without changing the shared environment. | |
| Small raw pilot evidence is in [results/20260916T182352Z-train](results/20260916T182352Z-train/). | |
| Checkpoint SHA-256 is | |
| `af5790ae2f2b56477ebbdf6ab9c418d895e48d2bf5a6416e11b6e9863ad1db55`; | |
| validation-selected best SHA-256 is | |
| `82b4261feb98d3ed56291e4c03304a65da20ce0194a6ad117d113b30d152282e`. | |
| New dependencies are isolated in `~/ai/envs/opensysone` (pyarrow 25.0.1), with | |
| read-only reuse of the existing torch/Transformers packages. The Jev-compatible | |
| stdlib harness and 12 CPU tests pass; real-checkpoint HTTP and longest-input | |
| stress are the next gate before the larger campaign. | |
| ## Public-data 4B pilot selected for the 24-hour run | |
| Run `/home/andy/ai/opensysone/runs/20260916T183823Z-train/artifacts`, clean execution | |
| source **`ccbbe6d`**, exited **0** after 40 steps / 160 decisions. The base is | |
| `Qwen/Qwen3-4B-Instruct-2507`, pinned to | |
| `cdbee75f17c01a7cc42f958dc650907174af0554`, Apache-2.0. | |
| Rank-8 adapters (alpha 16) and the pretrained initialized head train | |
| **16,517,633 of 4,038,985,729 parameters** in FP32. Exact two-pass categorical | |
| gradients keep one candidate graph live; CPU gradients match ordinary CE within | |
| 0.000001. Gradient checkpointing is enabled. No quantization or new kernels. | |
| | Family | Initial validation accuracy | Step 40 accuracy | Step 40 NLL | | |
| | --- | ---: | ---: | ---: | | |
| | ARC | 90.63% | 90.63% | 0.374400 | | |
| | Banking77 four-choice | 90.63% | 91.41% | 0.229941 | | |
| | BoolQ | 82.03% | 84.38% | 0.583093 | | |
| | SNLI | 82.81% | 83.59% | 0.395211 | | |
| | All 512 validation decisions | **86.52%** | **87.50%** | **0.395661** | | |
| Raw validation macro NLL improves from **1.436162 to 0.395661**; the initial | |
| readout was severely overconfident. A separately recorded diagnostic fits and | |
| scores a temperature on the same validation rows (NLL 0.407737, T 6.9183): it is | |
| optimistic validation analysis, not independent calibration. Reserved calibration, | |
| test and Social IQA predictions remain unevaluated. The trained 4B validation | |
| accuracy and NLL beat the 2B pilot in every family, supporting the larger candidate | |
| despite its lower throughput. This does not prove task generalization. | |
| The 512-token complete-chat limit retains **40,915 train / 512 validation / | |
| 510 calibration / 2,042 test / 768 Social IQA** decisions; it drops 26 train, | |
| two calibration and six test BoolQ rows, with no silent truncation. | |
| Median four-decision step is **8.956 s**; 55,268 actual branch tokens were | |
| processed with no padding overhead. The loop including final validation takes | |
| 724.0 s. Initial validation alone takes 327.85 s. Peak CUDA allocated/reserved | |
| is **15.510/15.604 GiB** against the 16 GiB cap. OOM adjustment is 0 and about | |
| 99 GiB unified RAM remains available with the model loaded. | |
| Final correctness passes all 0.0001 gates, worst probability difference | |
| **0.00000167**, with exact repeated, isolated and restored-permutation predictions. | |
| Checkpoint SHA-256: | |
| `e26f75b2396de88311873fac4eb91e1e40d0ec940778ec99f282bcfd96a2e258`. | |
| Best SHA-256: | |
| `64977ee0b1a6147c6faf59283edea9adf564dd36d53f4580bc20940b94c6764f`. | |
| Small raw evidence is in [results/20260916T183823Z-train](results/20260916T183823Z-train/). | |
| Fresh reload, longest-input gradients with restored optimizer state, 1,024-token | |
| HTTP inference, and 255-choice HTTP stress **all passed** (verification exit 0). | |
| Reload matches all 16 checked validation predictions exactly. Longest training | |
| input is 509 tokens and peaks at 15.624 GiB with optimizer state; inference peaks | |
| at 15.465 GiB. The long HTTP request has 1,023 tokens in each of two candidate | |
| branches and matches direct inference exactly. Invalid-key/oversized-input | |
| requests return 401/422. One warm three-question request takes 1.571 s, and one | |
| 255-choice request takes 47.042 s; these are wiring stress timings, not latency | |
| percentiles or intelligence benchmarks. The checkpoint SHA-256 is unchanged. | |
| Evidence: [results/20260916T185718Z-verify4b](results/20260916T185718Z-verify4b/). | |
| All **15 CPU tests pass**, including unequal-source-group bootstrap weighting | |
| and the measured evaluation-reserve calculation. The live Jev HTTPS endpoint | |
| returns 405 to an unauthenticated GET; no credentials or state were sent and no | |
| authenticated hosted inference has been tested. | |
| ## Detached 24-hour campaign now running | |
| Launched **2026-09-16 18:59:10 UTC** from clean source **`0109eb6`** into | |
| `/home/andy/ai/opensysone/runs/20260916T185910Z-24h`. Supervisor PID is **1085496**, | |
| current trainer **1085517**; both OOM score adjustments are 0. Training resumes the | |
| 4B step-40 checkpoint with optimizer/RNG restored, preserves validation-selected | |
| best and all model/data/config signatures, and has passed the initial FP32 | |
| correctness gates. Exit is **pending**; the API has not started yet. | |
| Fresh restart reproduces **all 512 raw logits and probabilities exactly**; | |
| the reference and fresh prediction JSON SHA-256 are both | |
| `e671e1508185765552b0f933ba03f356be62143c531d8ef534457d34b1645c9b`. | |
| The next four updates, **41–44**, have finite losses/gradients and remain under | |
| the cap. Step 41 takes 8.724 s, loss 0.115940, gradient norm 3.81358. | |
| This proves reconstruction plus subsequent optimizer updates, not a bitwise | |
| interrupted-versus-uninterrupted trajectory comparison. Raw verification is in | |
| the launch evidence directory below. The durable checkpoint remains step 40 | |
| until the regular save cadence, independently of those logged newer updates. | |
| Training ends by **2026-09-17 16:16:10 UTC**, reserving two hours until the final | |
| **18:16:10 UTC / 19:16:10 BST** deadline. The reserve estimates 6,640 base/tuned | |
| prediction rows at 4,251.8 seconds from measured pilot validation speed, adds | |
| 30% plus ten minutes for setup, and keeps a two-hour minimum. Checkpoints save | |
| every 250 steps or 900 seconds regardless of evaluation; validation is every | |
| 500 steps with patience eight. The three-epoch target is an upper bound. | |
| After successful training, the runner loads the best artifact fresh, calibrates | |
| only on the 510 reserved known-family decisions, saves a deployable checkpoint | |
| before untouched evaluation, records raw/calibrated test and Social IQA metrics | |
| against the unchanged pretrained scorer, and starts the loopback API only after | |
| complete evaluation and a real-model inference check. Deployment is planned at | |
| `http://127.0.0.1:18081`, with 1,024-token inputs. No hosted Jev call runs | |
| automatically. Small launch evidence lives in | |
| [results/20260916T185910Z-24h-launch](results/20260916T185910Z-24h-launch/), separate | |
| from the completion-results directory reserved by the runner. | |
| Current inspection, stop and same-deadline recovery commands are in | |
| [HANDOVER.md](HANDOVER.md). A running job is not a finalized model or successful | |
| test result. The frozen-family controls and independent calibration remain the | |
| quality gates for final reporting. The complete 15-test suite passed; the new | |
| orphan-child stop safeguard also passes an integration test that refuses to | |
| terminate a PID when its command line differs from the recorded command. | |
| ## Three-machine expansion — 2026-09-16 evening | |
| The user assigned GX10 and both Sparks to this task and authorized terminating | |
| their workloads. The Spark serving head and RPC worker were stopped in order | |
| with verified SIGTERM; both released their GPU allocations and each had about | |
| 118 GiB available afterward. Their weights/cache and exact restoration commands | |
| are retained. No network or system configuration changed. | |
| Both Sparks now have isolated copies of the exact GX10 training dependencies: | |
| 21,368 installed file hashes and 55 package versions match. CPU autograd and both | |
| Qwen-family imports pass. This initial check verified the environments. Subsequently all pinned model | |
| files and real GPU training/HTTP checks passed; see the launch results below. | |
| The original GX10 campaign saved step 128 before a requested stop. Its trainer | |
| exceeded the 30-second grace while performing final correctness checks and exited | |
| -9; the complete step-128 checkpoint and optimizer/RNG are verified intact. The | |
| new source records skipped final checks explicitly on a requested stop. It also | |
| retains step-specific prediction evidence before publishing each new best artifact | |
| and selects a restored checkpoint if its fresh validation improves the best. | |
| The intermediate GX10 campaign `20260916T192239Z-24h`, source `6e080e2`, restored | |
| step 128 and later stopped gracefully at step 178 with training exit 0. | |
| The Spark alternatives are a 4B weights-only warm initialization with fresh Adam, | |
| learning rate 0.00003 and 7,500-step cosine horizon, and a longer 2B continuation. | |
| The planned fleet cutoff is 2026-09-17 16:00 UTC, leaving 2 h 16 min until the | |
| original final deadline. All training remains under 16 GiB per process. | |
| All **32 initial fleet CPU tests passed**, including weights-only initialization, optimizer/RNG | |
| resume, requested-stop evidence, deadline handling, exact validation-set matching, | |
| checkpoint/metric mismatch rejection and API deployment lifecycle. The coordinator | |
| recomputes its criterion from all 512 saved validation predictions and freezes | |
| selection before calibration/test/holdout. Read-only compatibility checks of the | |
| real 4B and 2B pilot artifacts pass, reproducing NLL 0.395661 and 0.498153. | |
| Small setup proofs are in [results/20260916-fleet-setup](results/20260916-fleet-setup/). | |
| Live paths, statuses and recovery instructions are in [FLEET_RUN.md](FLEET_RUN.md). | |
| The subsequent selection revision uses the frozen four-fold source-group-disjoint | |
| temperature-crossfit policy `crossfit_temperature_nll_v1` (seed 431, 101 positive | |
| temperatures, family-balanced fitting and scoring). Step 128's validation accuracy | |
| is **89.0625%**, versus step 40's 87.5%; raw NLL is 0.442683 versus 0.395661. | |
| Crossfit NLL reverses that ranking: **0.318518 versus 0.359522**, improving in all | |
| four families. A 5,000-replicate paired source-group bootstrap, refitting the | |
| temperatures, gives difference -0.041005 with 95% interval [-0.079822, -0.001152]. | |
| The accuracy gain alone is uncertain (29 gains, 21 losses; McNemar p=0.322). | |
| This supports accounting for recoverable overconfidence during checkpoint | |
| selection. It is a validation-driven criterion revision, not independent test | |
| evidence. No reserved predictions were read. Original raw-selected checkpoints | |
| remain preserved, and final calibration still uses the separate reserved split. | |
| All **37 tests pass** after adding policy/selection checks; the updated CPU | |
| integration also proves reselection leaves trained weights and Adam steps intact. | |
| Raw diagnostic: [selection-diagnostic.json](results/20260916-fleet-setup/selection-diagnostic.json). | |
| ## Active fleet launch — 2026-09-16 19:44 UTC | |
| Three training-only campaigns are active on source **`4a60423`**: | |
| GX10 `20260916T193741Z-24h` (4B, LR 0.0001), spark-a | |
| `20260916T194258Z-24h` (4B, LR 0.00003, 7,500-step cosine horizon), and spark-b | |
| `20260916T193803Z-24h` (2B, LR 0.0001). Each uses the fixed crossfit criterion, | |
| 16 GiB allocation cap and 2026-09-17 16:00 UTC cutoff. Training exit statuses | |
| remain pending. The fleet coordinator `20260916T194403396250Z-fleet`, source | |
| **`6a7b0ed`**, is detached on GX10 and waiting for selection; no reserved-data | |
| predictions or final calibration have run. The cutoff shutdown race is covered | |
| by a regression test, and all nine fleet tests pass after that fix. | |
| GX10 restored step 178's weights, Adam and Python/torch/CUDA RNG exactly. Its | |
| fresh 512-decision validation reached **90.4297% accuracy, 0.303825 crossfit NLL, | |
| 0.404198 raw NLL**, promoting the durable best beyond step 128. Fresh FP32 | |
| correctness passes (worst probability difference 4.77e-7), and resumed updates | |
| are finite. These are validation results, not independent test evidence. | |
| Spark A reproduced all 512 original 4B pilot predictions exactly before eight | |
| finite lower-rate updates (median 8.086 seconds, peak 15.505 GiB). That pilot | |
| exited 0; its step-8 accuracy 86.914% / raw NLL 0.405397 did not improve the | |
| starting checkpoint. The long-run crossfit selector re-evaluates both inherited | |
| best and current checkpoint. Real fresh-artifact verification exited 0: exact | |
| 16-decision reload, finite restored-Adam gradients on the longest 509-token | |
| input, 15.624 GiB peak, 1,023-token HTTP/direct match, expected 401/422 errors, | |
| and 255 choices in 43.31 seconds. The long campaign reproduced all 512 step-8 | |
| raw predictions exactly, with identical weights/Adam/RNG. Its fixed crossfit | |
| criterion selected step 8 at 0.358235 NLL, and new updates are finite. No | |
| independent generalization improvement is claimed for the short pilot. | |
| Spark B's preparation exited 0. Fresh verification passed exact reload, | |
| restored-Adam gradients at 700 tokens (8.123 GiB peak), 1,024-token inference, | |
| authentication/length errors and 255 choices in 17.73 seconds. The long campaign | |
| reproduced all 512 original validation predictions exactly, scoring 82.8125% | |
| accuracy / 0.476072 crossfit NLL / 0.498153 raw NLL before resumed training. | |
| Subsequent finite updates reached step 98 by 19:43:59 UTC. Timing observations | |
| are individual wiring checks, not p50/p95 latency measurements. | |
| Full small evidence, source revisions, frozen plan and startup state snapshots | |
| are under [results/20260916-fleet-setup](results/20260916-fleet-setup/). Live | |
| state, inspection/stop/resume and serving-pair restoration are in | |
| [FLEET_RUN.md](FLEET_RUN.md). Final calibrated test/holdout metrics and selected-model | |
| API deployment are pending; authenticated hosted Jev inference still requires | |
| `TYPESAFE_API_KEY`. | |
| ## Overnight progress and next experiment — 2026-09-17 | |
| The 02:10–02:15 UTC audit found both 4B jobs healthy and improving, while the 2B | |
| campaign completed cleanly at **01:59:29 UTC**, training and supervisor exit **0**. | |
| All recorded losses/gradients were finite. Peak allocation was 15.624 GiB on each | |
| 4B job and 8.183 GiB on the 2B job; the 16 GiB cap remains unchanged. | |
| | Candidate | Last audited step | Selected step | Crossfit validation NLL | Selected accuracy | | |
| | --- | ---: | ---: | ---: | ---: | | |
| | GX10 4B, LR 1e-4 | 2,570 | 2,500 | **0.188640** | **93.55%** | | |
| | Spark A 4B, LR 3e-5 | 2,529 | 2,500 | 0.218012 | 92.58% | | |
| | Spark B 2B, LR 1e-4 | 6,000 | 2,000 | 0.255294 | 89.84% | | |
| These are the same 512 validation decisions, selected with the unchanged fixed | |
| crossfit policy. No reserved calibration, test or Social IQA predictions have | |
| been read. A higher maximum accuracy at a different step does not override the | |
| selection criterion. Spark A improved at all five scheduled validations. Spark B | |
| stopped after eight evaluations without a new best; its final step-6,000 score | |
| was 0.351593 / 86.91%. Final numerical correctness passed at worst 6.56e-7. | |
| The selected step-2,000 and resumable step-6,000 artifacts are preserved. | |
| The freed Spark B is training a **fourth candidate**, initialized from a frozen | |
| copy of GX10's step-2,500 selected weights (SHA-256 | |
| `8956eb6c0cfbb02124aeefd99c3b418c55f55fdb9a64260350622d98dbba1aec`). | |
| Fresh Adam, seed 432, LR/head LR 1e-5 and a 5,000-step cosine horizon define a new | |
| trajectory. Other model/batch/token/correctness settings and both absolute | |
| deadlines stay unchanged. The eight-step pilot started at **02:16:05 UTC**; | |
| source `4a60423`. The pinned 4B model copied from Spark A over the existing link | |
| passed all 13 file hashes. Warm initialization preserves all 506 trainable tensors | |
| exactly and deliberately starts with an empty optimizer. All 512 initial raw | |
| predictions match the parent exactly. The eight-step pilot and fresh verifier | |
| exited 0: exact 16-decision reload, finite longest-input gradients, 15.623 GiB | |
| peak, direct/HTTP agreement at 1,023 tokens, expected 401/422 and 255 choices | |
| in 45.44 seconds. These timings are individual wiring checks, not percentiles. | |
| Campaign `20260917T023137Z-24h` launched at 02:31:37 UTC, restoring the complete | |
| step-8 optimizer/RNG state exactly, and was added as the fourth fleet candidate. | |
| Its inherited selected branch step 0 retains the parent score: step 8 scored | |
| 0.188576, a change below the fixed 0.001 improvement threshold. The short pilot | |
| does not establish a quality gain. | |
| [Small audit evidence](results/20260917-fleet-progress/) records the 22 scheduled | |
| validation points, live processes, source revisions, selected-checkpoint hashes | |
| and frozen refinement parent. [NEXT_STEPS.md](NEXT_STEPS.md) records the decisions | |
| and the two-Spark alternatives: independent candidates now, bounded distributed | |
| adapter-gradient training or parallel scoring next. Active ConnectX/RoCE and | |
| installed NCCL do not establish collective correctness or useful speedup. The | |
| current two-pass trainer needs explicit synchronization changes, and its measured | |
| peak leaves only about 385 MiB for additional GPU allocations under the cap. | |
| ## Fixed validation ensemble diagnostic — 2026-09-17 | |
| Saved, identity-matched validation logits were combined with fixed equal weights | |
| and the unchanged crossfit-temperature policy; no weights were tuned and no | |
| reserved predictions were accessed. GX10 4B + Spark A 4B scores **93.16% / | |
| 0.193983 NLL**, worse than GX10 alone (**93.55% / 0.188640**). Spark A 4B + the | |
| completed Spark B 2B scores **93.55% / 0.180560**. This more diverse pair shares | |
| 19 errors versus 29 for the two-4B pair, but gains eight/losses eight versus GX10. | |
| The mixed pair's NLL difference versus GX10 is -0.008080; a 1,000-replicate paired | |
| source-group bootstrap with fold-temperature refitting yields 95% interval | |
| **[-0.039064, +0.019451]**. No gain over the best single model is established. | |
| These are exploratory validation results from already selected checkpoints, | |
| not independent generalization evidence. The deployed-candidate protocol remains | |
| individual models; ensemble inference and latency have not been implemented or | |
| measured. The exact A step-2,500 and B step-2,000 artifacts are frozen on GX10 in | |
| `20260917T022201Z-ensemble-reference`, with 134.3 MB copied, stable source hashes | |
| and CPU reconstruction/provenance checks. No weights are in Git. | |
| [Analysis and provenance](results/20260917-fleet-progress/fixed-ensemble-validation.json). | |
| ## Remaining gates | |
| The active continuation state and checkpoint-resume verification are recorded in | |
| [HANDOVER.md](HANDOVER.md). Public multi-family training and the frozen unseen-family | |
| holdout are now implemented; independent calibration/test/holdout metrics await | |
| the 24-hour campaign's finalization. Frozen-head and generation controls, new-model | |
| prefix caching and the larger latency matrix remain open. The Sparks now host | |
| independent candidate experiments; GX10 does not need a ConnectX cable for this | |
| selection strategy. Architecture B and | |
| distributed training still await quality and profiling evidence in [PLAN.md](PLAN.md). | |
| ## Expanded public training data — 2026-09-17 | |
| Version 2 retains all 40,915 original 4B-compatible training decisions and adds | |
| 16,000 HellaSwag, 14,360 PIQA and 9,490 CommonsenseQA decisions: **80,765 total**. | |
| All four reserved source files and tokenized sequences match version 1 exactly. | |
| The 383 retained new-source diagnostics stay outside training and checkpoint | |
| selection. An independent reconstruction audit checked every added source label | |
| and shuffled answer position, all downloaded hashes and diagnostic exclusions. | |
| See [EXPANDED_DATA.md](EXPANDED_DATA.md) and its linked small evidence. | |
| Clean source `24b8ccf`, pilot `20260917T070758Z-train`: eight finite updates, | |
| **exit 0**, all initial/final FP32 gates passed. Frozen Spark B parent step 1,500 | |
| reproduces every initial validation logit and probability exactly. Captured step | |
| 0 has empty Adam; every final Adam counter is eight. Median update 8.90 seconds, | |
| peak allocation including checks 15.426 GiB, worst final probability discrepancy | |
| 3.58e-7. The 32 sampled decisions cover all seven task families. | |
| Step 8 scores 0.170108 validation crossfit NLL versus parent 0.170150, both | |
| 94.7266% accuracy. The difference is below the fixed 0.001 selection threshold; | |
| the selected branch remains step 0. This is startup evidence, not a claim of | |
| improvement on the added tasks. The new campaign `20260917T072142Z-24h` restores | |
| all step-8 model/Adam/Python/Torch/CUDA states exactly and retains the original | |
| 16:00 / 18:16:10 UTC deadlines. GX10's former run stopped at step 4,380 with both | |
| trainer and supervisor exit 0, preserving its selected step 2,500. The fleet | |
| retains all previous candidates and explicitly registers the new dataset. | |
| All 86 source tests passed, including rejection of reserved-data changes and | |
| unregistered candidate datasets. The actual expanded candidate passed the full | |
| fleet eligibility path. Expanded checkpoints and transformed data were uploaded | |
| and verified in the existing private Hugging Face repository at 07:24:36 UTC, | |
| with exact source revisions, upstream notices and checksums; publication receipts | |
| are recorded separately. | |
| At 07:28:29 UTC the resumed campaign passed its full startup audit: every one of | |
| 512 pilot-step-8 predictions reproduced exactly, full optimizer/RNG state matched, | |
| and updates 9–12 were finite under the cap. GX10 reached step 15 by 07:28:58 UTC; | |
| both Spark trials, the coordinator, GUI and final-publication watcher remained | |
| running. Final campaign evaluation is still pending. | |