# OpenSysOne: single-node start and GX10 handover ## Continuation decision — 2026-09-17 See [NEXT_STEPS.md](NEXT_STEPS.md) for the overnight findings and the assessment of work shared by the two Sparks. GX10 and Spark A continue improving 4B trials. Spark B's 2B run exited 0 at step 6,000 after validation early stopping; preserve its step-2,000 selected artifact. Use the freed GPU for a fourth candidate from GX10's selected step-2,500 4B weights, with fresh Adam, seed 432 and LR 1e-5. Keep the fixed selection policy, 16 GiB cap and original absolute deadlines. The two Sparks have active ConnectX/RoCE and installed NCCL, but distributed training has no measured correctness or throughput result. Do not interrupt the improving trials to replace their trainer during this delivery window. Next joint experiment: bounded communication and synchronized-gradient parity, followed by 25–100 representative updates at equal global batch and measured memory. Parallel scoring replicas are a simpler later use; preserve the tested GX10 finalizer now. ## Active 24-hour campaign — 2026-09-16 The user now authorizes all three machines for this task, including stopping existing workloads. GX10 continues the main run; the Sparks run independent lower-learning-rate 4B and longer-running 2B candidates. See [FLEET_RUN.md](FLEET_RUN.md) for the active fleet plan and process controls. The existing 16 GiB allocation cap applies to each training process. Select the candidate using the same 512 validation decisions before final calibration/test. The frozen selection criterion is now four-fold source-group-disjoint temperature crossfit macro-family NLL, seed 431, policy `crossfit_temperature_nll_v1`. Fit each fold's scalar temperature on the other three; fit serving temperature afresh on reserved calibration after selection. Raw NLL/accuracy remain separately reported. This validation-driven revision preserves the stronger step-128 classifier that raw NLL discarded because of overconfidence; see the diagnostic in RESULTS.md. The absolute delivery deadline is **2026-09-17 18:16:10 UTC (19:16:10 BST)**, 24 hours from this request. Reserve at least the final two hours for fresh reconstruction, calibration, untouched evaluation and loopback API deployment. Earlier sections below preserve the smoke plan; their 1.5B and leave-idle boundary is superseded. The runner increases that reserve from measured pilot validation time when needed, including both pretrained/tuned passes, 30% margin and ten minutes for setup. Start from a pinned posttrained model, retain its pretrained yes-minus-no readout, and train ordinary FP32 low-rank decoder adapters plus scalar head. BF16 remains blocked by the measured correctness gate. Compare Qwen3.5-2B with Qwen3-4B-Instruct-2507, using the fixed validation crossfit criterion, memory and throughput. The 4B pilot uses rank 8 and an exact two-pass categorical gradient to keep one candidate graph live. The reliable 2B fallback uses rank 16. Keep the 16 GiB CUDA cap and launch free-memory/OOM hardening unchanged. **Initial single-node selection:** the 4B rank-8 candidate, with 87.50% validation accuracy and 0.395661 macro NLL versus the 2B pilot's 82.81% / 0.498153. It beats 2B in every measured validation family. Training limit is 512 complete-chat tokens; separately verified inference limit is 1,024. Longest-input training stress peaks at 15.624 GiB, and real HTTP reload/limits/255-choice checks pass. Frozen source-group-disjoint public data covers SNLI, BoolQ, ARC and four-choice Banking77 routing. Social IQA is a completely untrained task-family holdout. Source pins, licences, raw hashes and split audit are in `results/public-decisions-v1-manifest.json`; data is under `~/ai/opensysone/data/`. Model-specific length filtering is reported, with no silent truncation. Validation chooses checkpoints. Calibration fits only one global temperature; untouched test/holdout evaluation happens in the separate finalization process. Compare the trained scorer with its unchanged pretrained readout, both raw and separately temperature-calibrated, with source-group bootstrap uncertainty. Before launch, prove reconstruction, a subsequent optimizer step, longest-input gradients with restored optimizer state and authenticated HTTP inference using the real checkpoint. Then detach `scripts/launch_24h.py`, preserving optimizer, RNG, source/data/model provenance, checkpoint cadence independent of evaluation, individual child PIDs and an absolute deadline. Select the best validation checkpoint rather than assuming more updates improve intelligence. The standard-library harness supports local inference, Jev HTTP calls and response/timing comparison; see [JEV_HARNESS.md](JEV_HARNESS.md). Hosted calls require `TYPESAFE_API_KEY`; no key is available in the current process environment. Deploy only on loopback and use the existing SSH tunnel for Mac access. This trains a general-language **decision scorer**, not a new general-purpose chat model or a demonstrated substitute for Jev. Generalization and calibration remain evaluation outcomes. Generation baselines, frozen-head controls, shared prefix caching for these new models and the original latency matrix remain open. Consolidated **2026-09-16**. This is the active execution plan. The original proposal is preserved verbatim in [RESEARCH_BRIEF.md](RESEARCH_BRIEF.md). Start a continuation with [HANDOVER.md](HANDOVER.md), then read this file. Build a decision scorer from a pretrained causal Transformer: arbitrary state, question and natural-language candidate go in; one scalar score comes out. Normalize mutually exclusive choices to a distribution. No generated answer or fixed label vocabulary. TypeSafe/Jev architecture claims remain hypotheses; softmax alone does not establish calibration. ## Resources available now | Host | Installed unified RAM | Available at 16:36 BST | Current use | Project role | | --- | ---: | ---: | --- | --- | | GX10 | 121.6 GiB | 118.7 GiB | Idle router; no substantial loaded model | Development, tiny training/evaluation | | spark-a | 121.7 GiB | 28.0 GiB | Qwen3.8-Flash-Next Q8_0 head | Existing serving workload | | spark-b | 121.7 GiB | 22.4 GiB | Same model's RPC worker | Existing serving workload | These are snapshots, not reservations. CPU, GPU, cache and OS share each pool. The table above records the earlier smoke snapshot. The user subsequently assigned all three GB10s to OpenSysOne; at 19:14 UTC the Spark serving pair was stopped and both GPUs were empty, with about 118 GiB available on each host. These remain three separate memory pools. Recheck `free -b` and GPU processes before every run. The Sparks have one physical ConnectX port-0 cable, with two PCIe-domain paths: `192.168.100.10/11` and `192.168.101.10/11`. Both were verified active; earlier fleet tests measured 108.9 Gb/s RDMA per domain, 188 Gb/s aggregate. llama.cpp RPC works; **PyTorch/NCCL training is unverified**. See [FLEET_SCOUT.md](FLEET_SCOUT.md). GX10 uses ordinary Ethernet/Wi-Fi/tailnet, without connected ConnectX. Additional connectivity is expected around **2026-09-18**, per the user; this is an estimate. GX10 can coordinate jobs over SSH today, but should not join the Sparks' collective over a slow network. Even after cabling, verify topology, transport and collective correctness before revising capacity. A two-node DAC does not specify the future three-node topology or create coherent pooled RAM. ## Immediate experiment and handover boundary Finish a bounded smoke on GX10 and leave it free for the next session. - Base: `Qwen/Qwen2.5-0.5B`, revision `060db6499f32faf8b98477b0a26969ef7d8b9987`, Apache-2.0, dense causal decoder. - Method: FP32 backbone, final two layers trainable, FP32 scalar head, categorical cross-entropy. This is partial fine-tuning, not LoRA. - Data: invented inventory facts, three questions per state, shuffled candidate text; 192 train / 48 calibration / 72 test decisions. Disjoint entity groups, same task templates. No customer data. - Bounds: 60 steps, four decisions/batch, short sequences, 16 GiB CUDA cap, 24 GiB available-memory launch gate, 25-minute timeout, checkpoints every ten steps and before evaluation. No long unattended run needed for this phase. - Stack: existing `~/ai/envs/comfy/bin/python`, torch 2.11.0+cu130, Transformers 5.15.0, SDPA; no shared-environment package changes. - Compare base yes-minus-no token logits, initial uniform scalar head, trained scalar and separate-calibration-split global temperature. Uniform output is an optimization sanity baseline, not a competitive classifier. - Time the same trained checkpoint and token IDs: full batched forwards versus cached branching; about 128/1,024 state tokens, 1/4/16 questions, two choices, eight branches/chunk. Save actual lengths and raw warm repetitions, prefill, branch and end-to-end times. This is not yet the complete generation comparison. Completion gates: finite gradients/loss, changed backbone/head weights, checkpoint and optimizer/RNG state, reload/resume verification, strict FP32 tiny-model cache tests and measured BF16 parity/permutation/isolation. Record peak allocated and reserved CUDA memory plus host availability. Save before evaluation can fail. Leave source, model, artifacts, commands, hashes and process state on GX10. Synthetic improvements demonstrate optimization and wiring only. They cannot establish calibration, zero-shot ability, useful judgment or superiority over prompt-and-generate classification. **Precision gate found during the smoke:** BF16 changes probabilities by up to 0.099 when batch composition changes, including uncached forwards. Forcing SDPA MATH does not fix it. Casting the same weights to FP32 reduces discrepancies to about 0.000014 across the tested comparisons. Use FP32 for the reference; BF16 requires an explicit correctness investigation before larger experiments. Preserve the failed run and diagnostic; do not relax tolerances to accept it. **Expanded gate, 2026-09-16:** all 24 synthetic groups fail BF16 even with strict accumulation, math SDPA, FP32 decoder linears, or their combination. Worst probability differences are 0.147–0.239; the same weights cast to FP32 stay below 0.000022. First-layer traces expose shape-dependent projection differences, but correcting those alone does not fix the decoder. See `results/20260916T161355Z-precision/precision.json`. Use FP32 for the next public data/1.5B experiment; further BF16 work should target remaining operations rather than repeat these unsuccessful switches. ## Next working session: first useful 1.5B experiment Budget the next one or two hours for a real data cut and a proven resumed run. 1. Read smoke results and traces. Fix correctness before interpreting speed. Preserve the 0.5B run as a reference and verify fresh checkpoint reconstruction. 2. Pin `Qwen2.5-1.5B` base. Start at 128–1,024 state tokens; increase to 4k after measuring memory. BF16 weights are about 2.9 GiB; training processes every candidate branch. Record exact config rather than assuming context limits. 3. Select public sentiment, entailment and intent/routing sources, checking each licence/version first. Preserve source splits, deduplicate/group before transformations, and reserve an entire further task family plus unseen question/label paraphrases for zero-shot evaluation. Freeze test data early. 4. Compare token scoring, frozen-backbone trained head and tuned scalar on the same data. Start with FP32/SDPA; restore BF16 only after the precision gate. Add LoRA in an isolated pinned PEFT environment if useful; preserve the shared Comfy environment. Defer QLoRA, FP8 and custom kernels. 5. Count all processed branch tokens/padding and measure elapsed step time. Set dataset size and deadline from those observations. Keep evaluation batches small and checkpoint on a cadence independent of evaluation. ## Following 48 hours: prove utility on one node Start with at least 1,000 untouched test decisions across multiple public datasets; increase until proper-score uncertainty is informative. Report per-family counts, accuracy, NLL, multiclass Brier (class sum), declared-bin top-label ECE, reliability and accuracy-versus-coverage. Fit one temperature on a separate calibration set. Evaluate once on test and held-out family. Bootstrap source groups, not augmented rows; use ECE alongside proper scores. Add generation and constrained-output baselines using the same base and inputs. Document prompt, output-token budget, parse/failure policy and timing scope. Distinguish model-load, first-call, warmed and application end-to-end latency. Extend one axis at a time: 1/4/16/64 questions; 2/4/16 choices; 128/1k/4k states. Add 255 choices as one stress point after bounded chunking is proven. Do not run the original full Cartesian product yet. Estimate tail latency with enough repetitions before reporting p95. Account for KV copies, padding and transfers. Recheck permutations, mixed lengths and unrelated-question perturbations. **Scale only after:** repeatable useful accuracy and improved NLL/Brier on at least one untouched task family, no unexplained severe regression elsewhere, and meaningful measured multi-question latency/throughput improvement over a fair baseline. If only familiar label words improve, fix data/objective first. The original <150/<250/<500 ms targets are exploratory, not commitments. ## Later hardware and architecture decisions Schedule a service transition before large Spark training; verify actual memory release. spark-a's active swap and missing earlyoom must be addressed before sustained training. Operational changes belong in the relevant GX10 docs. Before DDP: test CUDA/NCCL all-reduce numerical correctness, transport logs, both directions and realistic message sizes, then a short two-rank optimizer run with checkpoint/resume. Compare useful examples/second with one node and two independent runs. DDP replicates state; it does not combine memory. FSDP is a separate decision, justified by measured memory needs. The 3B class remains the target after the 1.5B gate. Qwen2.5-3B has a separate research licence; select it deliberately or choose another base if deployment requires different terms. A 7B/8B run follows useful scaling evidence. Keep inference local. Primary model/cache links are in [RESEARCH_NOTES.md](RESEARCH_NOTES.md). Stay with architecture A (shared-prefix decoder) until profiling identifies its cost. Every suffix still runs all layers and attends to the state. Batching does not guarantee constant latency. A 1.5B 4k prefix is about 112 MiB KV; 256 physical copies are about 28 GiB before suffixes, weights and workspace. Test architecture B (state encoder plus shallow cross-attention decoder) if branch work/copies dominate and quality passes. Compare at equal data budget. Packed branching needs numerical independence tests; custom kernels need a profiled bottleneck. Soft teacher targets, proper-score losses and quantization calibration ablations follow a reliable baseline. RL is unnecessary initially. ## Evidence and artifacts - [HANDOVER.md](HANDOVER.md): exact continuation commands and run state. - [RESULTS.md](RESULTS.md): measured outcomes and limitations. - [FLEET_SCOUT.md](FLEET_SCOUT.md): live survey and prior bandwidth evidence. - [RESEARCH_NOTES.md](RESEARCH_NOTES.md): primary sources and memory arithmetic. - `results//`: small raw config/data/prediction/correctness/timing files. - GX10 `~/ai/opensysone/runs//`: complete run including checkpoint. - GX10 `~/ai/models/opensysone/`: pinned pretrained weights. Record source commit/hashes, model/data revisions, config, exact software and hardware for every run. No API service is needed for this phase; any future HTTP listener follows the existing loopback/tailnet policy.