opensysone / source /RESEARCH_NOTES.md
andyshu's picture
Back up verified OpenSysOne training snapshot and pinned source
2d5c26a verified
|
Raw History Blame
13.7 kB
# Research notes for the GX10 handover
Assessed 2026-09-16. Read alongside `PLAN.md`, the measured run artifacts, and
`~/code/gx10/docs/{training,ai-environment}.md`. These are engineering recommendations;
they do not claim a reproduction of TypeSafe's architecture or measured model quality.
**Measured after the initial research:** the tiny trained Qwen2.5 scorer failed
BF16 batch-invariance checks (up to 0.099 probability difference), even with the
SDPA MATH backend. The same weights evaluated in FP32 passed at about 1.4e-5 or
better across the diagnostic comparisons. The reference smoke therefore uses
FP32 parameters/optimizer/inference. BF16 elsewhere in these notes is a proposed
future configuration and arithmetic estimate, conditional on fixing this gate.
See `results/20260916T154714Z/parity_diagnosis.json` and `RESULTS.md`.
## Scope and available resources
The user has confirmed a ConnectX link between `spark-a` and `spark-b`. GX10 is a separate
GB10 machine on ordinary networking until approximately 2026-09-18. A link does not pool
physical memory. The main agent's live survey additionally found the pair serving a large
Qwen model, leaving approximately 28/22 GiB available; GX10 was idle with approximately
119 GiB available. Preserve the distinction between installed capacity and currently usable
capacity, and recheck availability before each run.
Use GX10 for the present experiment, with a 16 GiB CUDA allocation budget for the smoke.
Use the two Sparks later for independent experiments, evaluation, or teacher inference when
their existing workload permits. Add distributed training only after the single-node
experiment demonstrates a bottleneck that two-node communication could address. GX10
should coordinate the linked pair over SSH rather than join their collective over a slower
network. Verify actual RDMA/NCCL transport before assuming the cable is used by a workload.
NVIDIA specifies 128 GB coherent system memory, 273 GB/s local bandwidth, and 200 Gbit/s
ConnectX-7 networking. The advertised 1 PFLOP is theoretical sparse FP4, so it is not an
appropriate BF16 training throughput estimate. Local OS, model serving, CPU allocations,
and GPU tensors compete for the unified pool. [NVIDIA specifications](https://www.nvidia.com/en-us/products/workstations/dgx-spark/)
## Model choice
| Base model | Parameters | Layers / hidden width / KV heads | BF16 weights, approximately | Role |
| --- | ---: | --- | ---: | --- |
| Qwen2.5-0.5B | 0.49B | 24 / 896 / 2 | 0.91 GiB | Present train/save/reload/cache smoke |
| Qwen2.5-1.5B | 1.54B | 28 / 1536 / 2 | 2.87 GiB | First useful baseline and decision-tuning comparison |
| Qwen2.5-3B | 3.09B | 36 / 2048 / 2 | 5.76 GiB | Later research comparison after the 1.5B gate |
These are conventional causal decoder models with grouped-query attention, RoPE, SwiGLU,
RMSNorm and tied embeddings. Their ordinary KV structure makes them suitable for this
first shared-prefix implementation. The 0.5B and 1.5B cards identify Apache-2.0 licensing.
[0.5B card](https://huggingface.co/Qwen/Qwen2.5-0.5B),
[1.5B card](https://huggingface.co/Qwen/Qwen2.5-1.5B),
[3B card](https://huggingface.co/Qwen/Qwen2.5-3B)
The 3B variant has the separate Qwen Research licence, whose grant is for non-commercial
research/evaluation and requests a separate licence for commercial use. Keep it as a
research candidate rather than assuming the 1.5B licence applies to the whole family.
[Official 3B licence](https://huggingface.co/Qwen/Qwen2.5-3B/blob/main/LICENSE)
The cards state 32,768-token context, but the current 1.5B config at revision
`8faed761d45a263340a0528343f099c05c9a4323` specifies 131,072 positions. Pin the downloaded
revision and save its config; do not infer long-context quality from that setting. Initial
experiments should stay within 128–4,096 state tokens. The configs also expose the large
151,936-token vocabulary. [0.5B config](https://huggingface.co/Qwen/Qwen2.5-0.5B/raw/main/config.json),
[1.5B config](https://huggingface.co/Qwen/Qwen2.5-1.5B/resolve/8faed761d45a263340a0528343f099c05c9a4323/config.json),
[3B config](https://huggingface.co/Qwen/Qwen2.5-3B/raw/main/config.json)
For the current smoke, tuning the last two backbone layers plus a scalar readout is a
reasonable bounded test when PEFT is absent. Save the exact trainable names/count and
label the method accurately. This establishes that the pipeline optimizes and persists
weights; it does not establish that partial tuning is better than LoRA or full fine-tuning.
## Shared-prefix implementation requirements
GX10's live environment reported PyTorch `2.11.0+cu130`, Transformers `5.15.0`, Accelerate
`1.14.0`, and Hugging Face Hub `1.27.0`. Keep these versioned with the result. Use the
installed library's implementation when resolving API details; older cache recipes differ.
Transformers 5.15 stores cache tensors in layer objects, and `DynamicLayer.update` changes
the cache object by appending keys/values. `batch_repeat_interleave` also mutates it and
allocates repeated tensors. Give each branch batch a fresh cache derived from an untouched
prefix; cloning only a Python reference is insufficient. A conservative initial approach
is a deep copy followed by batch repeat. Treat literal prefix memory sharing as a later
optimization: avoiding repeated prefill does not imply avoiding repeated KV storage.
[Pinned cache implementation](https://raw.githubusercontent.com/huggingface/transformers/v5.15.0/src/transformers/cache_utils.py)
The attention mask must cover cached prefix plus current suffix. In the 5.15 Qwen2
implementation, default positions start at cache length. Supply positions explicitly if
padding requires them, and pool the last actual suffix token rather than a right-padding
token. For scalar training, call the backbone and apply the readout to that hidden vector;
this avoids materializing a vocabulary-sized logit tensor at every sequence position.
[5.15 caching documentation](https://huggingface.co/docs/transformers/v5.15.0/cache_explanation),
[Pinned Qwen2 implementation](https://raw.githubusercontent.com/huggingface/transformers/v5.15.0/src/transformers/models/qwen2/modeling_qwen2.py)
Implementation gates before interpreting benchmarks:
1. Make the uncached reference consume exactly the concatenation of prefix token IDs and
suffix token IDs used by the cached path. Text tokenization at a concatenation boundary
can otherwise differ. Preserve one canonical serialization in saved metadata.
2. Compare cached batched scores with independent complete forwards. Exercise mixed suffix
lengths, a singleton batch, and different chunk sizes. Use a small FP32 reference for
strict numerical checks; report actual BF16 discrepancies and declared tolerances.
3. After scoring another question, rescore the first question. Its scores and the saved
prefix length/content should be unchanged. Perturb an unrelated branch and repeat.
4. Permute candidates, unpermute outputs, and compare logits/probabilities. Also permute
questions and vary batch composition. Candidate order invariance in this independent
scorer follows from its design; it is not evidence of semantic generalization.
5. Confirm finite gradients, changed trainable weights, finite loss, and checkpoint reload
parity. Test the resume path including optimizer/RNG state before an unattended run.
6. Train complete branches with `use_cache=False` while any backbone layers are trainable.
Inference caches reused across optimizer updates would become stale and can remove
intended gradient paths. Prefix caching is first an inference optimization.
The choices represent mutually exclusive alternatives for softmax. Their outputs are
conditional on the supplied set, not probabilities that each description is independently
true. A singleton softmax is always 1. For independent multi-label decisions use separately
trained sigmoid semantics. With independent scalar scores, `p(A)/p(B)=exp(sA-sB)` cannot
change when a third candidate is added; duplicate or overlapping choices are therefore a
later semantic stress test. These are algebraic consequences of the proposed interface.
## Conservative memory and time expectations
The following are arithmetic estimates, not measurements of the new code. BF16 KV bytes
per state token are `2 × layers × KV_heads × head_dim × 2`. For one 4,096-token prefix:
| Model | One prefix | 256 physical prefix copies |
| --- | ---: | ---: |
| 0.5B | 48 MiB | 12 GiB |
| 1.5B | 112 MiB | 28 GiB |
| 3B | 144 MiB | 36 GiB |
These omit suffix KV, temporary attention tensors, allocator reservation and weights.
Batch size is questions multiplied by candidates: 64 questions with 255 choices means
16,320 branches. Process branches in bounded chunks and report the chunk size. A full
Cartesian sweep of the original plan's largest settings would be premature.
Conventional mixed-precision full AdamW fine-tuning at approximately 16 bytes/parameter
requires about 7.3/22.9/46.0 GiB for 0.5B/1.5B/3B persistent model and optimizer state,
before activations and temporary allocations. Partial tuning uses much less optimizer
memory. A scalar scorer also removes the original language-model loss's large vocabulary
logit allocation, but its actual peaks must be measured during both train and evaluation.
Historical GX10 measurements in `~/code/gx10/docs/training.md` reached roughly 9.5–18
effective TFLOPS for different hybrid models and a different workload. They are useful
context, not predictions for this implementation. An illustrative 5–20 TFLOPS range and
`6 × parameters × processed_tokens` gives approximately 2.5–10 minutes for 0.5B, 8–31
minutes for 1.5B, or 15–62 minutes for 3B per million processed tokens of full fine-tuning.
This is compute-only arithmetic. Count all candidate branches, padding and epochs; one
million source-context tokens can become many millions of model-input tokens. Calibrate
real elapsed times from the first measured steps before scheduling work.
## Next 1–2 hours and handover gate
- Finish the 0.5B smoke on GX10 with small synthetic inputs, short sequences, direct
scalar output, and a strict allocation cap. Record untrained and trained results and
the trainable parameter subset. Demonstrate save/reload and branch correctness.
- Measure a modest grid: state lengths about 128 and 1,024; 1/4/16/64 questions;
2/4 choices. Add one 255-choice stress point only after memory and latency are known.
Compare identical input IDs and the same checkpoint on cached and uncached paths.
- Save raw timings with warm-up separated, GPU synchronization, repetition count, and
prefill/branch/end-to-end components. Report batch/chunk configuration, peak allocated
and reserved CUDA memory, and host availability. Include tokenization/copy time in an
end-to-end view rather than silently dropping it.
- Make GX10 the continuation location: source, frozen model revision, dataset generator
and seed, checksums, config, checkpoint, optimizer/RNG state, metrics, logs, exact rerun
commands, and process status must be discoverable without the Mac session. Stop the
bounded run before handover or document its PID, log, stop and resume commands.
Smoke success means the pipeline trains, reloads, and obeys the inference invariants.
A lower synthetic loss, higher synthetic accuracy, or attractive synthetic ECE is **not**
evidence of calibration, zero-shot generalization, or useful judgment. Preserve these
limitations in any summary of the result.
## Following 48 hours
Use Qwen2.5-1.5B on one node. Establish a token-logit baseline, a frozen-backbone scalar
head baseline, and a tuned scalar model. Use several legally usable public task families
with split provenance and source-level deduplication. Hold out entire task families and
question/label paraphrases. Fit temperature on a separate calibration split; evaluate
once on untouched test splits. Report accuracy, NLL and Brier with uncertainty, and ECE
with explicit binning/sample counts rather than using ECE alone as a success gate.
The first model-selection gate is a repeatable accuracy/proper-score improvement on
held-out tasks plus a latency advantage over independent full forwards at useful branch
counts. Compare to generation with a documented prompt, output limit and parse-failure
policy. Do not change checkpoints or task formatting between timing baselines.
Retain a frozen benchmark sample and test candidate length, descriptions, distractors,
permutations, missing information and task shift. Quantization and teacher probabilities
remain later ablations: neither should be mixed into the first causally interpretable
comparison. Softmax alone never constitutes calibration.
## After connectivity and model-quality gates
Re-survey all three machines and verify the actual link topology and collective transport.
Choose two-node DDP only if it improves tokens/second per experiment after communication
and checkpoint costs. Otherwise use independent runs, teachers and evaluators. Keep
inference local for these model sizes.
Scale to the 3B-class comparison only after the 1.5B result is worth scaling. The shared
KV branch path removes repeated state prefill, but every suffix still traverses all
backbone layers and attends to the prefix: sublinear wall-clock growth is an empirical
batching/amortization result, not constant total compute. If profiling shows branch
compute or replicated KV dominates, compare a shallow cross-attention decision decoder
under the same quality tests. Packed branching attention comes after numerical equivalence
tests, and custom kernels come after profiles establish a worthwhile bottleneck.