File size: 12,181 Bytes
2d5c26a | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 | # Findings and next steps — 17 September 2026
Continue the two improving 4B runs and use Spark B's freed GPU for a conservative
4B refinement. Preserve the completed 2B candidate. Do not replace the working
training/finalization path with unmeasured distributed training before today's
**16:00 UTC training cutoff / 18:16:10 UTC delivery deadline**. The two Sparks
are promising for a later joint experiment, especially gradient synchronization
or parallel scoring; they remain separate memory pools.
## What the overnight runs established
Audited at **02:10–02:15 UTC**. All three use the same 512 validation decisions
and the fixed `crossfit_temperature_nll_v1` criterion. Lower selection NLL is
better. Accuracy below belongs to the selected checkpoint, not the maximum
accuracy observed at any step.
| Candidate | Latest audited step | Selected step | Selection NLL | Accuracy | Status |
| --- | ---: | ---: | ---: | ---: | --- |
| GX10, 4B, LR 1e-4 | 2,570 | 2,500 | **0.188640** | **93.55%** | Training; latest validation improves best |
| Spark A, 4B, LR 3e-5 | 2,529 | 2,500 | 0.218012 | 92.58% | Training; all five scheduled validations improved |
| Spark B, 2B, LR 1e-4 | 6,000 | 2,000 | 0.255294 | 89.84% | Completed, exit 0, validation early stop |
The initial 4B pilot selected at step 40 scored 87.50% / 0.359522 selection NLL.
Overnight training therefore improved this validation set substantially. It does
not yet prove independent test quality, unseen-family transfer, calibrated
confidence, or Jev equivalence. Reserved calibration/test/Social IQA predictions
remain untouched. The four families are SNLI, BoolQ, ARC and **four-choice**
Banking77, not a complete general-intelligence benchmark.
Spark B's last eight evaluations did not beat step 2,000. Its final step 6,000
scored 0.351593 selection NLL / 86.91% accuracy. Resuming the same trajectory is
poorly supported. All 5,960 resumed updates were finite; final correctness passed
at worst probability difference 6.56e-7. Training and supervisor exited 0 at
**01:59:29 UTC**. The resumable step-6,000 checkpoint and selected step-2,000
artifact remain intact. The 2B peak was 8.183 GiB; both 4B jobs peaked at
15.624 GiB. No runtime errors were found in the audits.
The selection score and accuracy answer different questions. For example, GX10
step 2,000 had 93.95% accuracy but worse selection NLL than step 2,500. Keep the
predeclared criterion rather than switch objectives to whichever number looks
best. Final serving temperature will be fitted on the separate calibration split.
Small reproducible evidence and all 22 scheduled validation points are in
[results/20260917-fleet-progress](results/20260917-fleet-progress/).
[HANDOVER.md](HANDOVER.md) and [FLEET_RUN.md](FLEET_RUN.md) own live paths and controls.
## Actions within this deadline
1. **Continue GX10 and Spark A unchanged.** Both have recent selection-score
improvements. Keep checkpoint cadence, numerical gates and 16 GiB cap.
2. **Refine the stronger checkpoint on Spark B.** Freeze GX10 step 2,500 with
matching validation evidence; warm-start its weights with fresh Adam, seed
432, adapter/head LR 1e-5 and a 5,000-step cosine horizon. Retain rank 8,
alpha 16, global batch four, 512-token training and exact two-pass gradients.
This changes both data order and optimizer trajectory; it is a new experiment,
not an exact optimizer resume. The eight-step pilot and fresh GPU/HTTP verifier
exited 0. Campaign `20260917T023137Z-24h` launched at 02:31:37 UTC with
exactly restored pilot weights, Adam and Python/torch/CUDA RNG. All 512 pilot
predictions replayed exactly before new finite updates. A 5,000-step horizon
is a schedule, not a promise that all those steps fit. Keep the absolute cutoff.
3. **Retain four candidates.** The three original candidates remain in
the fleet plan, including the completed 2B. The verified refinement was added
at 02:34 UTC by restarting only the waiting coordinator. The other training
jobs continued uninterrupted; the previous plan/state/exit remain archived.
4. **Freeze selection, then evaluate once.** At 16:00 UTC, select from verified
durable artifacts using validation only. Fit separate calibration temperature,
compare tuned/pretrained models on untouched test and Social IQA, report proper
scores and source-group uncertainty, and verify real API inference before
publishing the deployment pointer. Preserve the tested GX10 finalizer.
5. **Connect hosted Jev when credentials exist.** The local/hosted/comparison
harness is implemented; authenticated hosted calls remain untested because
`TYPESAFE_API_KEY` is not configured. Local confidence is normalized entropy,
not proof of calibrated correctness.
## Would the two Sparks work better together?
| Use of both Sparks | Benefit to test | Current assessment |
| --- | --- | --- |
| Independent candidates with checkpoint exchange | More optimization diversity without training collectives | Best supported use before today's deadline; already transferring models/checkpoints between hosts |
| Data-parallel adapter training | Faster updates for one model by splitting a global batch across two replicas | Promising next measured experiment; requires trainer and recovery changes |
| Parallel replicas for scoring | More requests or independent evaluation rows; possible ensemble quality gain | Easiest later integration; identical frozen artifacts and row integrity are essential |
| Sharded larger model | Fit parameters that exceed one process's cap | Highest integration/memory risk; no validated larger-model fit or speed result yet |
| Teacher/student pipeline | One model produces training-only soft targets while another learns | Useful if compact-model latency becomes a priority; currently diverts compute from the stronger 4B model |
**Hardware is available, but collective performance is unproven.** Read-only
probes found active ConnectX/RoCE paths, identical PyTorch 2.11.0+cu130 and NCCL
2.28.9, and available distributed/NCCL backends on both Sparks. CUDA was not
initialized by those probes. The infrastructure specification records about
109 Gb/s RDMA per PCIe domain and 188 Gb/s combined; those are historical host
transport measurements, not measured NCCL training throughput. One physical
200 Gb/s port exposes two domain paths; this is not a 400 Gb/s link.
NVIDIA documents that Spark does not support GPUDirect RDMA for ordinary CUDA
device allocations and describes a pinned-host-buffer fallback for verbs
applications. Successful host-memory RoCE tests do not establish direct GPU
buffer transport. [NVIDIA Spark CUDA guidance](https://docs.nvidia.com/dgx/dgx-spark-porting-guide/porting/cuda.html).
The current official two-Spark test procedure targets NCCL 2.30.7-1; the locally
installed 2.28.9 needs its own measurement. This is not evidence that it is broken
and does not justify replacing active environments. [NVIDIA NCCL procedure](https://build.nvidia.com/spark/nccl/stacked-sparks).
**Adapter training has a favorable payload.** Our 16,517,633 trainable FP32
parameters produce **66,070,532 bytes / 63.01 MiB** of gradients per update.
Each Spark would keep a full model replica; only accumulated adapter/head
gradients need synchronization. DDP does not shard input data automatically or
combine RAM. [PyTorch DDP](https://docs.pytorch.org/docs/2.11/generated/torch.nn.parallel.DistributedDataParallel.html).
A nominal bandwidth division is not a measured all-reduce latency or speedup.
**The present trainer cannot simply be wrapped.** It calls custom scoring methods
and performs a backward pass per candidate; candidate counts differ across
examples. A naive wrapper can bypass DDP's forward bookkeeping or issue unmatched
collectives. A deliberate implementation should give each rank two decisions,
accumulate each decision's loss divided by global batch four, SUM gradients once,
then clip and update identical Adam states. Averaging those already globally
normalized gradients would incorrectly halve the update. Data order, failure,
validation, save and resume boundaries also need coordinated rank behavior.
**Memory may be the limiting constraint.** The measured 4B peak leaves only
about **385 MiB** below the cap. Communication buffers and NCCL allocations,
including memory not counted by PyTorch's allocator, need measurement. Sharding
is a separate design: FSDP gathers parameter units for computation, so its peak
is not simply model size divided by two. [PyTorch FSDP2](https://docs.pytorch.org/tutorials/intermediate/FSDP_tutorial.html).
An illustrative 8B FP32 backbone alone consumes about 29.8 GiB in total before
adapters, activations and temporary gathers. Two 16 GiB limits do not demonstrate
that it fits or runs efficiently.
## Next combined experiment and acceptance gates
After the current jobs and selected artifacts are preserved:
1. Test bounded two-rank communication around the real 63 MiB payload, with small
buffers up to 256 MiB and hard timeouts. Record correctness, transport chosen,
latency and both processes' memory. Preserve strict SSH identities and existing
system/network configuration; isolate any extra dependencies.
2. Implement one synchronized global-batch-four update and compare its loss,
gradients, clipping and parameter update against the single-node reference.
Prove fresh reload and an interrupted/resumed update on both ranks.
3. Benchmark 25–100 representative updates plus validation/checkpoint overhead.
Keep both processes below 16 GiB. Adopt it only for a material measured gain,
provisionally at least **1.5× useful decisions/second** at the same global batch.
This threshold is a proposed engineering gate, not an achieved speedup.
4. For joint scoring, start with identical replicas after winner/temperature are
frozen. Prove matching outputs on overlap rows, split by stable IDs, and reject
missing/duplicate results. Keep calibration centralized. Split base/tuned
evaluation in a later finalizer only after merged-output parity is verified;
current delivery continues through the already tested single-host path.
## Measured ensemble check: a concrete use of two replicas
A fixed, untuned **equal mean of raw logits**, followed by the same crossfit
scalar-temperature policy, was checked on identical saved validation IDs. No new
GPU inference or reserved data was used. This is not probability averaging; raw
logit scales still affect each model's influence.
| Fixed model combination | Validation accuracy | Selection NLL |
| --- | ---: | ---: |
| GX10 4B alone, step 2,500 | 93.55% | 0.188640 |
| GX10 4B + Spark A 4B, both step 2,500 | 93.16% | 0.193983 |
| Spark A 4B step 2,500 + Spark B 2B step 2,000 | 93.55% | 0.180560 |
The two 4B members agree on 97.07% of decisions and share 29 errors; their mean
is worse than GX10 alone. The mixed 4B/2B members agree on 89.26% and share 19
errors, so architectural diversity is a useful lead. However, versus GX10, the
mixed ensemble gains eight correct decisions and loses eight. Its NLL difference
is **-0.008080**, with a **95% paired source-group bootstrap interval
[-0.039064, +0.019451]** over 1,000 replicates that refit fold temperatures.
The accuracy-difference interval is ±1.5625 percentage points. This does not
establish a gain over the best single model, and it is still selection-validation
analysis, not independent test evidence.
Both exact checkpoints are preserved under
`~/ai/opensysone/runs/20260917T022201Z-ensemble-reference` on GX10, with
source/copy hashes and CPU reconstruction checks in
[ensemble-reference.json](results/20260917-fleet-progress/ensemble-reference.json),
for a later latency/quality experiment. Do not
add an ensemble to today's deployment based on this small uncertain difference.
The current fleet still selects individual checkpoints. A future ensemble needs
its own frozen artifact rule, common input-length policy, separate calibration,
held-out evaluation, measured latency and two-worker recovery checks.
[Raw analysis and input hashes](results/20260917-fleet-progress/fixed-ensemble-validation.json).
|