# Research notes for the GX10 handover Assessed 2026-09-16. Read alongside `PLAN.md`, the measured run artifacts, and `~/code/gx10/docs/{training,ai-environment}.md`. These are engineering recommendations; they do not claim a reproduction of TypeSafe's architecture or measured model quality. **Measured after the initial research:** the tiny trained Qwen2.5 scorer failed BF16 batch-invariance checks (up to 0.099 probability difference), even with the SDPA MATH backend. The same weights evaluated in FP32 passed at about 1.4e-5 or better across the diagnostic comparisons. The reference smoke therefore uses FP32 parameters/optimizer/inference. BF16 elsewhere in these notes is a proposed future configuration and arithmetic estimate, conditional on fixing this gate. See `results/20260916T154714Z/parity_diagnosis.json` and `RESULTS.md`. ## Scope and available resources The user has confirmed a ConnectX link between `spark-a` and `spark-b`. GX10 is a separate GB10 machine on ordinary networking until approximately 2026-09-18. A link does not pool physical memory. The main agent's live survey additionally found the pair serving a large Qwen model, leaving approximately 28/22 GiB available; GX10 was idle with approximately 119 GiB available. Preserve the distinction between installed capacity and currently usable capacity, and recheck availability before each run. Use GX10 for the present experiment, with a 16 GiB CUDA allocation budget for the smoke. Use the two Sparks later for independent experiments, evaluation, or teacher inference when their existing workload permits. Add distributed training only after the single-node experiment demonstrates a bottleneck that two-node communication could address. GX10 should coordinate the linked pair over SSH rather than join their collective over a slower network. Verify actual RDMA/NCCL transport before assuming the cable is used by a workload. NVIDIA specifies 128 GB coherent system memory, 273 GB/s local bandwidth, and 200 Gbit/s ConnectX-7 networking. The advertised 1 PFLOP is theoretical sparse FP4, so it is not an appropriate BF16 training throughput estimate. Local OS, model serving, CPU allocations, and GPU tensors compete for the unified pool. [NVIDIA specifications](https://www.nvidia.com/en-us/products/workstations/dgx-spark/) ## Model choice | Base model | Parameters | Layers / hidden width / KV heads | BF16 weights, approximately | Role | | --- | ---: | --- | ---: | --- | | Qwen2.5-0.5B | 0.49B | 24 / 896 / 2 | 0.91 GiB | Present train/save/reload/cache smoke | | Qwen2.5-1.5B | 1.54B | 28 / 1536 / 2 | 2.87 GiB | First useful baseline and decision-tuning comparison | | Qwen2.5-3B | 3.09B | 36 / 2048 / 2 | 5.76 GiB | Later research comparison after the 1.5B gate | These are conventional causal decoder models with grouped-query attention, RoPE, SwiGLU, RMSNorm and tied embeddings. Their ordinary KV structure makes them suitable for this first shared-prefix implementation. The 0.5B and 1.5B cards identify Apache-2.0 licensing. [0.5B card](https://huggingface.co/Qwen/Qwen2.5-0.5B), [1.5B card](https://huggingface.co/Qwen/Qwen2.5-1.5B), [3B card](https://huggingface.co/Qwen/Qwen2.5-3B) The 3B variant has the separate Qwen Research licence, whose grant is for non-commercial research/evaluation and requests a separate licence for commercial use. Keep it as a research candidate rather than assuming the 1.5B licence applies to the whole family. [Official 3B licence](https://huggingface.co/Qwen/Qwen2.5-3B/blob/main/LICENSE) The cards state 32,768-token context, but the current 1.5B config at revision `8faed761d45a263340a0528343f099c05c9a4323` specifies 131,072 positions. Pin the downloaded revision and save its config; do not infer long-context quality from that setting. Initial experiments should stay within 128–4,096 state tokens. The configs also expose the large 151,936-token vocabulary. [0.5B config](https://huggingface.co/Qwen/Qwen2.5-0.5B/raw/main/config.json), [1.5B config](https://huggingface.co/Qwen/Qwen2.5-1.5B/resolve/8faed761d45a263340a0528343f099c05c9a4323/config.json), [3B config](https://huggingface.co/Qwen/Qwen2.5-3B/raw/main/config.json) For the current smoke, tuning the last two backbone layers plus a scalar readout is a reasonable bounded test when PEFT is absent. Save the exact trainable names/count and label the method accurately. This establishes that the pipeline optimizes and persists weights; it does not establish that partial tuning is better than LoRA or full fine-tuning. ## Shared-prefix implementation requirements GX10's live environment reported PyTorch `2.11.0+cu130`, Transformers `5.15.0`, Accelerate `1.14.0`, and Hugging Face Hub `1.27.0`. Keep these versioned with the result. Use the installed library's implementation when resolving API details; older cache recipes differ. Transformers 5.15 stores cache tensors in layer objects, and `DynamicLayer.update` changes the cache object by appending keys/values. `batch_repeat_interleave` also mutates it and allocates repeated tensors. Give each branch batch a fresh cache derived from an untouched prefix; cloning only a Python reference is insufficient. A conservative initial approach is a deep copy followed by batch repeat. Treat literal prefix memory sharing as a later optimization: avoiding repeated prefill does not imply avoiding repeated KV storage. [Pinned cache implementation](https://raw.githubusercontent.com/huggingface/transformers/v5.15.0/src/transformers/cache_utils.py) The attention mask must cover cached prefix plus current suffix. In the 5.15 Qwen2 implementation, default positions start at cache length. Supply positions explicitly if padding requires them, and pool the last actual suffix token rather than a right-padding token. For scalar training, call the backbone and apply the readout to that hidden vector; this avoids materializing a vocabulary-sized logit tensor at every sequence position. [5.15 caching documentation](https://huggingface.co/docs/transformers/v5.15.0/cache_explanation), [Pinned Qwen2 implementation](https://raw.githubusercontent.com/huggingface/transformers/v5.15.0/src/transformers/models/qwen2/modeling_qwen2.py) Implementation gates before interpreting benchmarks: 1. Make the uncached reference consume exactly the concatenation of prefix token IDs and suffix token IDs used by the cached path. Text tokenization at a concatenation boundary can otherwise differ. Preserve one canonical serialization in saved metadata. 2. Compare cached batched scores with independent complete forwards. Exercise mixed suffix lengths, a singleton batch, and different chunk sizes. Use a small FP32 reference for strict numerical checks; report actual BF16 discrepancies and declared tolerances. 3. After scoring another question, rescore the first question. Its scores and the saved prefix length/content should be unchanged. Perturb an unrelated branch and repeat. 4. Permute candidates, unpermute outputs, and compare logits/probabilities. Also permute questions and vary batch composition. Candidate order invariance in this independent scorer follows from its design; it is not evidence of semantic generalization. 5. Confirm finite gradients, changed trainable weights, finite loss, and checkpoint reload parity. Test the resume path including optimizer/RNG state before an unattended run. 6. Train complete branches with `use_cache=False` while any backbone layers are trainable. Inference caches reused across optimizer updates would become stale and can remove intended gradient paths. Prefix caching is first an inference optimization. The choices represent mutually exclusive alternatives for softmax. Their outputs are conditional on the supplied set, not probabilities that each description is independently true. A singleton softmax is always 1. For independent multi-label decisions use separately trained sigmoid semantics. With independent scalar scores, `p(A)/p(B)=exp(sA-sB)` cannot change when a third candidate is added; duplicate or overlapping choices are therefore a later semantic stress test. These are algebraic consequences of the proposed interface. ## Conservative memory and time expectations The following are arithmetic estimates, not measurements of the new code. BF16 KV bytes per state token are `2 × layers × KV_heads × head_dim × 2`. For one 4,096-token prefix: | Model | One prefix | 256 physical prefix copies | | --- | ---: | ---: | | 0.5B | 48 MiB | 12 GiB | | 1.5B | 112 MiB | 28 GiB | | 3B | 144 MiB | 36 GiB | These omit suffix KV, temporary attention tensors, allocator reservation and weights. Batch size is questions multiplied by candidates: 64 questions with 255 choices means 16,320 branches. Process branches in bounded chunks and report the chunk size. A full Cartesian sweep of the original plan's largest settings would be premature. Conventional mixed-precision full AdamW fine-tuning at approximately 16 bytes/parameter requires about 7.3/22.9/46.0 GiB for 0.5B/1.5B/3B persistent model and optimizer state, before activations and temporary allocations. Partial tuning uses much less optimizer memory. A scalar scorer also removes the original language-model loss's large vocabulary logit allocation, but its actual peaks must be measured during both train and evaluation. Historical GX10 measurements in `~/code/gx10/docs/training.md` reached roughly 9.5–18 effective TFLOPS for different hybrid models and a different workload. They are useful context, not predictions for this implementation. An illustrative 5–20 TFLOPS range and `6 × parameters × processed_tokens` gives approximately 2.5–10 minutes for 0.5B, 8–31 minutes for 1.5B, or 15–62 minutes for 3B per million processed tokens of full fine-tuning. This is compute-only arithmetic. Count all candidate branches, padding and epochs; one million source-context tokens can become many millions of model-input tokens. Calibrate real elapsed times from the first measured steps before scheduling work. ## Next 1–2 hours and handover gate - Finish the 0.5B smoke on GX10 with small synthetic inputs, short sequences, direct scalar output, and a strict allocation cap. Record untrained and trained results and the trainable parameter subset. Demonstrate save/reload and branch correctness. - Measure a modest grid: state lengths about 128 and 1,024; 1/4/16/64 questions; 2/4 choices. Add one 255-choice stress point only after memory and latency are known. Compare identical input IDs and the same checkpoint on cached and uncached paths. - Save raw timings with warm-up separated, GPU synchronization, repetition count, and prefill/branch/end-to-end components. Report batch/chunk configuration, peak allocated and reserved CUDA memory, and host availability. Include tokenization/copy time in an end-to-end view rather than silently dropping it. - Make GX10 the continuation location: source, frozen model revision, dataset generator and seed, checksums, config, checkpoint, optimizer/RNG state, metrics, logs, exact rerun commands, and process status must be discoverable without the Mac session. Stop the bounded run before handover or document its PID, log, stop and resume commands. Smoke success means the pipeline trains, reloads, and obeys the inference invariants. A lower synthetic loss, higher synthetic accuracy, or attractive synthetic ECE is **not** evidence of calibration, zero-shot generalization, or useful judgment. Preserve these limitations in any summary of the result. ## Following 48 hours Use Qwen2.5-1.5B on one node. Establish a token-logit baseline, a frozen-backbone scalar head baseline, and a tuned scalar model. Use several legally usable public task families with split provenance and source-level deduplication. Hold out entire task families and question/label paraphrases. Fit temperature on a separate calibration split; evaluate once on untouched test splits. Report accuracy, NLL and Brier with uncertainty, and ECE with explicit binning/sample counts rather than using ECE alone as a success gate. The first model-selection gate is a repeatable accuracy/proper-score improvement on held-out tasks plus a latency advantage over independent full forwards at useful branch counts. Compare to generation with a documented prompt, output limit and parse-failure policy. Do not change checkpoints or task formatting between timing baselines. Retain a frozen benchmark sample and test candidate length, descriptions, distractors, permutations, missing information and task shift. Quantization and teacher probabilities remain later ablations: neither should be mixed into the first causally interpretable comparison. Softmax alone never constitutes calibration. ## After connectivity and model-quality gates Re-survey all three machines and verify the actual link topology and collective transport. Choose two-node DDP only if it improves tokens/second per experiment after communication and checkpoint costs. Otherwise use independent runs, teachers and evaluators. Keep inference local for these model sizes. Scale to the 3B-class comparison only after the 1.5B result is worth scaling. The shared KV branch path removes repeated state prefill, but every suffix still traverses all backbone layers and attends to the prefix: sublinear wall-clock growth is an empirical batching/amortization result, not constant total compute. If profiling shows branch compute or replicated KV dominates, compare a shallow cross-attention decision decoder under the same quality tests. Packed branching attention comes after numerical equivalence tests, and custom kernels come after profiles establish a worthwhile bottleneck.