Download source/docs/research/next-steps.md from andyshu/opensysone: direct link, hf CLI and curl.
- Browser
- Download file 15.5 kB
-
https://huggingface.co/andyshu/opensysone/resolve/main/source/docs/research/next-steps.md
- Command line
-
hf download hf://andyshu/opensysone/source/docs/research/next-steps.md
-
curl -L -o next-steps.md https://huggingface.co/andyshu/opensysone/resolve/main/source/docs/research/next-steps.md
Findings and next steps — 17 September 2026
After training wrap-up
Training is stopped at the user's request. Preserve the frozen selected 4B model and all final resumable states; the overnight actions below are historical. The completed matched profile gives the selected model 285/320 correct (89.06%), the pretrained per-option verifier 259/320 (80.94%), and the pretrained joint-label method 276/320 (86.25%). These small-sample differences are descriptive.
The current implementation does not demonstrate a speed advantage over the same-sized base model. A 768-token state with one four-choice question takes 3.710 seconds for the trained scorer, 3.177 seconds for the original verifier, and 0.818 seconds for the joint-label method on the same idle Spark in FP32. The trained scorer repeats the context for every option; its unmerged adapters also add work. This comparison excludes model loading, HTTP and generated prose. See profiling-protocol.md for the protocol and final report.
Recommended follow-up experiments, after this completed campaign:
- Reuse the context prefix in the selected 4B inference path. Keep the existing full-forward FP32 result as the reference, prove candidate-order, question-batch and cache-reuse invariance, then remeasure the same 12 timing cells. The earlier 0.5B smoke establishes feasibility only; it is not a speed result for this 4B checkpoint.
- Measure merged adapters separately. Merge the frozen LoRA updates into a deployment copy and verify logits/probabilities against the saved model before timing. The current scorer is 11–17% slower than the unadapted verifier; removing adapter operations is a plausible optimization, not a measured gain. Keep precision changes in a separate correctness-controlled experiment.
- Give expanded-data training an appropriate validation plan. The stopped 159-update branch improves expansion diagnostics from 303/383 to 312/383, while the matched original/holdout sample changes from 285/320 to 284/320. HellaSwag supplies most of the gain. Its original four-family validation criterion did not select the new weights. A future run should predeclare a seven-family validation objective and new untouched test data; these observed diagnostics must not become an unacknowledged selection set.
- Use the Sparks together first as independent scoring replicas. Request sharding is a bounded way to test aggregate throughput while each host holds its own model and memory. Measure one- and two-host completed requests per second and latency under identical load. This does not itself reduce the latency of a single request. Data-parallel LoRA training is a later option: verify gradient/update parity and recovery, then measure communication cost before committing a campaign. No distributed training or pooled-memory speed claim follows from this run.
The two Sparks did useful parallel work in this wrap-up: Spark A measured all three inference methods sequentially without competing GPU work, while Spark B independently evaluated the expanded checkpoint on identical frozen examples. No further training or optimization was started as part of reporting these results.
Historical overnight recommendation
Continue the two improving 4B runs and use Spark B's freed GPU for a conservative 4B refinement. Preserve the completed 2B candidate. Do not replace the working training/finalization path with unmeasured distributed training before today's 16:00 UTC training cutoff / 18:16:10 UTC delivery deadline. The two Sparks are promising for a later joint experiment, especially gradient synchronization or parallel scoring; they remain separate memory pools.
What the overnight runs established
Audited at 02:10–02:15 UTC. All three use the same 512 validation decisions
and the fixed crossfit_temperature_nll_v1 criterion. Lower selection NLL is
better. Accuracy below belongs to the selected checkpoint, not the maximum
accuracy observed at any step.
| Candidate | Latest audited step | Selected step | Selection NLL | Accuracy | Status |
|---|---|---|---|---|---|
| GX10, 4B, LR 1e-4 | 2,570 | 2,500 | 0.188640 | 93.55% | Training; latest validation improves best |
| Spark A, 4B, LR 3e-5 | 2,529 | 2,500 | 0.218012 | 92.58% | Training; all five scheduled validations improved |
| Spark B, 2B, LR 1e-4 | 6,000 | 2,000 | 0.255294 | 89.84% | Completed, exit 0, validation early stop |
The initial 4B pilot selected at step 40 scored 87.50% / 0.359522 selection NLL. Overnight training therefore improved this validation set substantially. It does not yet prove independent test quality, unseen-family transfer, calibrated confidence, or Jev equivalence. Reserved calibration/test/Social IQA predictions remain untouched. The four families are SNLI, BoolQ, ARC and four-choice Banking77, not a complete general-intelligence benchmark.
Spark B's last eight evaluations did not beat step 2,000. Its final step 6,000 scored 0.351593 selection NLL / 86.91% accuracy. Resuming the same trajectory is poorly supported. All 5,960 resumed updates were finite; final correctness passed at worst probability difference 6.56e-7. Training and supervisor exited 0 at 01:59:29 UTC. The resumable step-6,000 checkpoint and selected step-2,000 artifact remain intact. The 2B peak was 8.183 GiB; both 4B jobs peaked at 15.624 GiB. No runtime errors were found in the audits.
The selection score and accuracy answer different questions. For example, GX10 step 2,000 had 93.95% accuracy but worse selection NLL than step 2,500. Keep the predeclared criterion rather than switch objectives to whichever number looks best. Final serving temperature will be fitted on the separate calibration split.
Small reproducible evidence and all 22 scheduled validation points are in results/20260917-fleet-progress. handover.md and fleet.md own live paths and controls.
Actions within this deadline
- Continue GX10 and Spark A unchanged. Both have recent selection-score improvements. Keep checkpoint cadence, numerical gates and 16 GiB cap.
- Refine the stronger checkpoint on Spark B. Freeze GX10 step 2,500 with
matching validation evidence; warm-start its weights with fresh Adam, seed
432, adapter/head LR 1e-5 and a 5,000-step cosine horizon. Retain rank 8,
alpha 16, global batch four, 512-token training and exact two-pass gradients.
This changes both data order and optimizer trajectory; it is a new experiment,
not an exact optimizer resume. The eight-step pilot and fresh GPU/HTTP verifier
exited 0. Campaign
20260917T023137Z-24hlaunched at 02:31:37 UTC with exactly restored pilot weights, Adam and Python/torch/CUDA RNG. All 512 pilot predictions replayed exactly before new finite updates. A 5,000-step horizon is a schedule, not a promise that all those steps fit. Keep the absolute cutoff. - Retain four candidates. The three original candidates remain in the fleet plan, including the completed 2B. The verified refinement was added at 02:34 UTC by restarting only the waiting coordinator. The other training jobs continued uninterrupted; the previous plan/state/exit remain archived.
- Freeze selection, then evaluate once. At 16:00 UTC, select from verified durable artifacts using validation only. Fit separate calibration temperature, compare tuned/pretrained models on untouched test and Social IQA, report proper scores and source-group uncertainty, and verify real API inference before publishing the deployment pointer. Preserve the tested GX10 finalizer.
- Connect hosted Jev when credentials exist. The local/hosted/comparison
harness is implemented; authenticated hosted calls remain untested because
TYPESAFE_API_KEYis not configured. Local confidence is normalized entropy, not proof of calibrated correctness.
Would the two Sparks work better together?
| Use of both Sparks | Benefit to test | Current assessment |
|---|---|---|
| Independent candidates with checkpoint exchange | More optimization diversity without training collectives | Best supported use before today's deadline; already transferring models/checkpoints between hosts |
| Data-parallel adapter training | Faster updates for one model by splitting a global batch across two replicas | Promising next measured experiment; requires trainer and recovery changes |
| Parallel replicas for scoring | More requests or independent evaluation rows; possible ensemble quality gain | Easiest later integration; identical frozen artifacts and row integrity are essential |
| Sharded larger model | Fit parameters that exceed one process's cap | Highest integration/memory risk; no validated larger-model fit or speed result yet |
| Teacher/student pipeline | One model produces training-only soft targets while another learns | Useful if compact-model latency becomes a priority; currently diverts compute from the stronger 4B model |
Hardware is available, but collective performance is unproven. Read-only probes found active ConnectX/RoCE paths, identical PyTorch 2.11.0+cu130 and NCCL 2.28.9, and available distributed/NCCL backends on both Sparks. CUDA was not initialized by those probes. The infrastructure specification records about 109 Gb/s RDMA per PCIe domain and 188 Gb/s combined; those are historical host transport measurements, not measured NCCL training throughput. One physical 200 Gb/s port exposes two domain paths; this is not a 400 Gb/s link.
NVIDIA documents that Spark does not support GPUDirect RDMA for ordinary CUDA device allocations and describes a pinned-host-buffer fallback for verbs applications. Successful host-memory RoCE tests do not establish direct GPU buffer transport. NVIDIA Spark CUDA guidance. The current official two-Spark test procedure targets NCCL 2.30.7-1; the locally installed 2.28.9 needs its own measurement. This is not evidence that it is broken and does not justify replacing active environments. NVIDIA NCCL procedure.
Adapter training has a favorable payload. Our 16,517,633 trainable FP32 parameters produce 66,070,532 bytes / 63.01 MiB of gradients per update. Each Spark would keep a full model replica; only accumulated adapter/head gradients need synchronization. DDP does not shard input data automatically or combine RAM. PyTorch DDP. A nominal bandwidth division is not a measured all-reduce latency or speedup.
The present trainer cannot simply be wrapped. It calls custom scoring methods and performs a backward pass per candidate; candidate counts differ across examples. A naive wrapper can bypass DDP's forward bookkeeping or issue unmatched collectives. A deliberate implementation should give each rank two decisions, accumulate each decision's loss divided by global batch four, SUM gradients once, then clip and update identical Adam states. Averaging those already globally normalized gradients would incorrectly halve the update. Data order, failure, validation, save and resume boundaries also need coordinated rank behavior.
Memory may be the limiting constraint. The measured 4B peak leaves only about 385 MiB below the cap. Communication buffers and NCCL allocations, including memory not counted by PyTorch's allocator, need measurement. Sharding is a separate design: FSDP gathers parameter units for computation, so its peak is not simply model size divided by two. PyTorch FSDP2. An illustrative 8B FP32 backbone alone consumes about 29.8 GiB in total before adapters, activations and temporary gathers. Two 16 GiB limits do not demonstrate that it fits or runs efficiently.
Next combined experiment and acceptance gates
After the current jobs and selected artifacts are preserved:
- Test bounded two-rank communication around the real 63 MiB payload, with small buffers up to 256 MiB and hard timeouts. Record correctness, transport chosen, latency and both processes' memory. Preserve strict SSH identities and existing system/network configuration; isolate any extra dependencies.
- Implement one synchronized global-batch-four update and compare its loss, gradients, clipping and parameter update against the single-node reference. Prove fresh reload and an interrupted/resumed update on both ranks.
- Benchmark 25–100 representative updates plus validation/checkpoint overhead. Keep both processes below 16 GiB. Adopt it only for a material measured gain, provisionally at least 1.5× useful decisions/second at the same global batch. This threshold is a proposed engineering gate, not an achieved speedup.
- For joint scoring, start with identical replicas after winner/temperature are frozen. Prove matching outputs on overlap rows, split by stable IDs, and reject missing/duplicate results. Keep calibration centralized. Split base/tuned evaluation in a later finalizer only after merged-output parity is verified; current delivery continues through the already tested single-host path.
Measured ensemble check: a concrete use of two replicas
A fixed, untuned equal mean of raw logits, followed by the same crossfit scalar-temperature policy, was checked on identical saved validation IDs. No new GPU inference or reserved data was used. This is not probability averaging; raw logit scales still affect each model's influence.
| Fixed model combination | Validation accuracy | Selection NLL |
|---|---|---|
| GX10 4B alone, step 2,500 | 93.55% | 0.188640 |
| GX10 4B + Spark A 4B, both step 2,500 | 93.16% | 0.193983 |
| Spark A 4B step 2,500 + Spark B 2B step 2,000 | 93.55% | 0.180560 |
The two 4B members agree on 97.07% of decisions and share 29 errors; their mean is worse than GX10 alone. The mixed 4B/2B members agree on 89.26% and share 19 errors, so architectural diversity is a useful lead. However, versus GX10, the mixed ensemble gains eight correct decisions and loses eight. Its NLL difference is -0.008080, with a 95% paired source-group bootstrap interval [-0.039064, +0.019451] over 1,000 replicates that refit fold temperatures. The accuracy-difference interval is ±1.5625 percentage points. This does not establish a gain over the best single model, and it is still selection-validation analysis, not independent test evidence.
Both exact checkpoints are preserved under
~/ai/opensysone/runs/20260917T022201Z-ensemble-reference on GX10, with
source/copy hashes and CPU reconstruction checks in
ensemble-reference.json,
for a later latency/quality experiment. Do not
add an ensemble to today's deployment based on this small uncertain difference.
The current fleet still selects individual checkpoints. A future ensemble needs
its own frozen artifact rule, common input-length policy, separate calibration,
held-out evaluation, measured latency and two-worker recovery checks.
Raw analysis and input hashes.