|
Download source/RESEARCH_NOTES.md from andyshu/opensysone: direct link, hf CLI and curl.
- Browser
- Download file 13.7 kB
-
https://huggingface.co/andyshu/opensysone/resolve/f2d6f8daa15bd21c316e61249f45ac16cbb79d45/source/RESEARCH_NOTES.md
- Command line
-
hf download hf://andyshu/opensysone@f2d6f8daa15bd21c316e61249f45ac16cbb79d45/source/RESEARCH_NOTES.md
-
curl -L -o RESEARCH_NOTES.md https://huggingface.co/andyshu/opensysone/resolve/f2d6f8daa15bd21c316e61249f45ac16cbb79d45/source/RESEARCH_NOTES.md
13.7 kB
| # Research notes for the GX10 handover | |
| Assessed 2026-09-16. Read alongside `PLAN.md`, the measured run artifacts, and | |
| `~/code/gx10/docs/{training,ai-environment}.md`. These are engineering recommendations; | |
| they do not claim a reproduction of TypeSafe's architecture or measured model quality. | |
| **Measured after the initial research:** the tiny trained Qwen2.5 scorer failed | |
| BF16 batch-invariance checks (up to 0.099 probability difference), even with the | |
| SDPA MATH backend. The same weights evaluated in FP32 passed at about 1.4e-5 or | |
| better across the diagnostic comparisons. The reference smoke therefore uses | |
| FP32 parameters/optimizer/inference. BF16 elsewhere in these notes is a proposed | |
| future configuration and arithmetic estimate, conditional on fixing this gate. | |
| See `results/20260916T154714Z/parity_diagnosis.json` and `RESULTS.md`. | |
| ## Scope and available resources | |
| The user has confirmed a ConnectX link between `spark-a` and `spark-b`. GX10 is a separate | |
| GB10 machine on ordinary networking until approximately 2026-09-18. A link does not pool | |
| physical memory. The main agent's live survey additionally found the pair serving a large | |
| Qwen model, leaving approximately 28/22 GiB available; GX10 was idle with approximately | |
| 119 GiB available. Preserve the distinction between installed capacity and currently usable | |
| capacity, and recheck availability before each run. | |
| Use GX10 for the present experiment, with a 16 GiB CUDA allocation budget for the smoke. | |
| Use the two Sparks later for independent experiments, evaluation, or teacher inference when | |
| their existing workload permits. Add distributed training only after the single-node | |
| experiment demonstrates a bottleneck that two-node communication could address. GX10 | |
| should coordinate the linked pair over SSH rather than join their collective over a slower | |
| network. Verify actual RDMA/NCCL transport before assuming the cable is used by a workload. | |
| NVIDIA specifies 128 GB coherent system memory, 273 GB/s local bandwidth, and 200 Gbit/s | |
| ConnectX-7 networking. The advertised 1 PFLOP is theoretical sparse FP4, so it is not an | |
| appropriate BF16 training throughput estimate. Local OS, model serving, CPU allocations, | |
| and GPU tensors compete for the unified pool. [NVIDIA specifications](https://www.nvidia.com/en-us/products/workstations/dgx-spark/) | |
| ## Model choice | |
| | Base model | Parameters | Layers / hidden width / KV heads | BF16 weights, approximately | Role | | |
| | --- | ---: | --- | ---: | --- | | |
| | Qwen2.5-0.5B | 0.49B | 24 / 896 / 2 | 0.91 GiB | Present train/save/reload/cache smoke | | |
| | Qwen2.5-1.5B | 1.54B | 28 / 1536 / 2 | 2.87 GiB | First useful baseline and decision-tuning comparison | | |
| | Qwen2.5-3B | 3.09B | 36 / 2048 / 2 | 5.76 GiB | Later research comparison after the 1.5B gate | | |
| These are conventional causal decoder models with grouped-query attention, RoPE, SwiGLU, | |
| RMSNorm and tied embeddings. Their ordinary KV structure makes them suitable for this | |
| first shared-prefix implementation. The 0.5B and 1.5B cards identify Apache-2.0 licensing. | |
| [0.5B card](https://huggingface.co/Qwen/Qwen2.5-0.5B), | |
| [1.5B card](https://huggingface.co/Qwen/Qwen2.5-1.5B), | |
| [3B card](https://huggingface.co/Qwen/Qwen2.5-3B) | |
| The 3B variant has the separate Qwen Research licence, whose grant is for non-commercial | |
| research/evaluation and requests a separate licence for commercial use. Keep it as a | |
| research candidate rather than assuming the 1.5B licence applies to the whole family. | |
| [Official 3B licence](https://huggingface.co/Qwen/Qwen2.5-3B/blob/main/LICENSE) | |
| The cards state 32,768-token context, but the current 1.5B config at revision | |
| `8faed761d45a263340a0528343f099c05c9a4323` specifies 131,072 positions. Pin the downloaded | |
| revision and save its config; do not infer long-context quality from that setting. Initial | |
| experiments should stay within 128–4,096 state tokens. The configs also expose the large | |
| 151,936-token vocabulary. [0.5B config](https://huggingface.co/Qwen/Qwen2.5-0.5B/raw/main/config.json), | |
| [1.5B config](https://huggingface.co/Qwen/Qwen2.5-1.5B/resolve/8faed761d45a263340a0528343f099c05c9a4323/config.json), | |
| [3B config](https://huggingface.co/Qwen/Qwen2.5-3B/raw/main/config.json) | |
| For the current smoke, tuning the last two backbone layers plus a scalar readout is a | |
| reasonable bounded test when PEFT is absent. Save the exact trainable names/count and | |
| label the method accurately. This establishes that the pipeline optimizes and persists | |
| weights; it does not establish that partial tuning is better than LoRA or full fine-tuning. | |
| ## Shared-prefix implementation requirements | |
| GX10's live environment reported PyTorch `2.11.0+cu130`, Transformers `5.15.0`, Accelerate | |
| `1.14.0`, and Hugging Face Hub `1.27.0`. Keep these versioned with the result. Use the | |
| installed library's implementation when resolving API details; older cache recipes differ. | |
| Transformers 5.15 stores cache tensors in layer objects, and `DynamicLayer.update` changes | |
| the cache object by appending keys/values. `batch_repeat_interleave` also mutates it and | |
| allocates repeated tensors. Give each branch batch a fresh cache derived from an untouched | |
| prefix; cloning only a Python reference is insufficient. A conservative initial approach | |
| is a deep copy followed by batch repeat. Treat literal prefix memory sharing as a later | |
| optimization: avoiding repeated prefill does not imply avoiding repeated KV storage. | |
| [Pinned cache implementation](https://raw.githubusercontent.com/huggingface/transformers/v5.15.0/src/transformers/cache_utils.py) | |
| The attention mask must cover cached prefix plus current suffix. In the 5.15 Qwen2 | |
| implementation, default positions start at cache length. Supply positions explicitly if | |
| padding requires them, and pool the last actual suffix token rather than a right-padding | |
| token. For scalar training, call the backbone and apply the readout to that hidden vector; | |
| this avoids materializing a vocabulary-sized logit tensor at every sequence position. | |
| [5.15 caching documentation](https://huggingface.co/docs/transformers/v5.15.0/cache_explanation), | |
| [Pinned Qwen2 implementation](https://raw.githubusercontent.com/huggingface/transformers/v5.15.0/src/transformers/models/qwen2/modeling_qwen2.py) | |
| Implementation gates before interpreting benchmarks: | |
| 1. Make the uncached reference consume exactly the concatenation of prefix token IDs and | |
| suffix token IDs used by the cached path. Text tokenization at a concatenation boundary | |
| can otherwise differ. Preserve one canonical serialization in saved metadata. | |
| 2. Compare cached batched scores with independent complete forwards. Exercise mixed suffix | |
| lengths, a singleton batch, and different chunk sizes. Use a small FP32 reference for | |
| strict numerical checks; report actual BF16 discrepancies and declared tolerances. | |
| 3. After scoring another question, rescore the first question. Its scores and the saved | |
| prefix length/content should be unchanged. Perturb an unrelated branch and repeat. | |
| 4. Permute candidates, unpermute outputs, and compare logits/probabilities. Also permute | |
| questions and vary batch composition. Candidate order invariance in this independent | |
| scorer follows from its design; it is not evidence of semantic generalization. | |
| 5. Confirm finite gradients, changed trainable weights, finite loss, and checkpoint reload | |
| parity. Test the resume path including optimizer/RNG state before an unattended run. | |
| 6. Train complete branches with `use_cache=False` while any backbone layers are trainable. | |
| Inference caches reused across optimizer updates would become stale and can remove | |
| intended gradient paths. Prefix caching is first an inference optimization. | |
| The choices represent mutually exclusive alternatives for softmax. Their outputs are | |
| conditional on the supplied set, not probabilities that each description is independently | |
| true. A singleton softmax is always 1. For independent multi-label decisions use separately | |
| trained sigmoid semantics. With independent scalar scores, `p(A)/p(B)=exp(sA-sB)` cannot | |
| change when a third candidate is added; duplicate or overlapping choices are therefore a | |
| later semantic stress test. These are algebraic consequences of the proposed interface. | |
| ## Conservative memory and time expectations | |
| The following are arithmetic estimates, not measurements of the new code. BF16 KV bytes | |
| per state token are `2 × layers × KV_heads × head_dim × 2`. For one 4,096-token prefix: | |
| | Model | One prefix | 256 physical prefix copies | | |
| | --- | ---: | ---: | | |
| | 0.5B | 48 MiB | 12 GiB | | |
| | 1.5B | 112 MiB | 28 GiB | | |
| | 3B | 144 MiB | 36 GiB | | |
| These omit suffix KV, temporary attention tensors, allocator reservation and weights. | |
| Batch size is questions multiplied by candidates: 64 questions with 255 choices means | |
| 16,320 branches. Process branches in bounded chunks and report the chunk size. A full | |
| Cartesian sweep of the original plan's largest settings would be premature. | |
| Conventional mixed-precision full AdamW fine-tuning at approximately 16 bytes/parameter | |
| requires about 7.3/22.9/46.0 GiB for 0.5B/1.5B/3B persistent model and optimizer state, | |
| before activations and temporary allocations. Partial tuning uses much less optimizer | |
| memory. A scalar scorer also removes the original language-model loss's large vocabulary | |
| logit allocation, but its actual peaks must be measured during both train and evaluation. | |
| Historical GX10 measurements in `~/code/gx10/docs/training.md` reached roughly 9.5–18 | |
| effective TFLOPS for different hybrid models and a different workload. They are useful | |
| context, not predictions for this implementation. An illustrative 5–20 TFLOPS range and | |
| `6 × parameters × processed_tokens` gives approximately 2.5–10 minutes for 0.5B, 8–31 | |
| minutes for 1.5B, or 15–62 minutes for 3B per million processed tokens of full fine-tuning. | |
| This is compute-only arithmetic. Count all candidate branches, padding and epochs; one | |
| million source-context tokens can become many millions of model-input tokens. Calibrate | |
| real elapsed times from the first measured steps before scheduling work. | |
| ## Next 1–2 hours and handover gate | |
| - Finish the 0.5B smoke on GX10 with small synthetic inputs, short sequences, direct | |
| scalar output, and a strict allocation cap. Record untrained and trained results and | |
| the trainable parameter subset. Demonstrate save/reload and branch correctness. | |
| - Measure a modest grid: state lengths about 128 and 1,024; 1/4/16/64 questions; | |
| 2/4 choices. Add one 255-choice stress point only after memory and latency are known. | |
| Compare identical input IDs and the same checkpoint on cached and uncached paths. | |
| - Save raw timings with warm-up separated, GPU synchronization, repetition count, and | |
| prefill/branch/end-to-end components. Report batch/chunk configuration, peak allocated | |
| and reserved CUDA memory, and host availability. Include tokenization/copy time in an | |
| end-to-end view rather than silently dropping it. | |
| - Make GX10 the continuation location: source, frozen model revision, dataset generator | |
| and seed, checksums, config, checkpoint, optimizer/RNG state, metrics, logs, exact rerun | |
| commands, and process status must be discoverable without the Mac session. Stop the | |
| bounded run before handover or document its PID, log, stop and resume commands. | |
| Smoke success means the pipeline trains, reloads, and obeys the inference invariants. | |
| A lower synthetic loss, higher synthetic accuracy, or attractive synthetic ECE is **not** | |
| evidence of calibration, zero-shot generalization, or useful judgment. Preserve these | |
| limitations in any summary of the result. | |
| ## Following 48 hours | |
| Use Qwen2.5-1.5B on one node. Establish a token-logit baseline, a frozen-backbone scalar | |
| head baseline, and a tuned scalar model. Use several legally usable public task families | |
| with split provenance and source-level deduplication. Hold out entire task families and | |
| question/label paraphrases. Fit temperature on a separate calibration split; evaluate | |
| once on untouched test splits. Report accuracy, NLL and Brier with uncertainty, and ECE | |
| with explicit binning/sample counts rather than using ECE alone as a success gate. | |
| The first model-selection gate is a repeatable accuracy/proper-score improvement on | |
| held-out tasks plus a latency advantage over independent full forwards at useful branch | |
| counts. Compare to generation with a documented prompt, output limit and parse-failure | |
| policy. Do not change checkpoints or task formatting between timing baselines. | |
| Retain a frozen benchmark sample and test candidate length, descriptions, distractors, | |
| permutations, missing information and task shift. Quantization and teacher probabilities | |
| remain later ablations: neither should be mixed into the first causally interpretable | |
| comparison. Softmax alone never constitutes calibration. | |
| ## After connectivity and model-quality gates | |
| Re-survey all three machines and verify the actual link topology and collective transport. | |
| Choose two-node DDP only if it improves tokens/second per experiment after communication | |
| and checkpoint costs. Otherwise use independent runs, teachers and evaluators. Keep | |
| inference local for these model sizes. | |
| Scale to the 3B-class comparison only after the 1.5B result is worth scaling. The shared | |
| KV branch path removes repeated state prefill, but every suffix still traverses all | |
| backbone layers and attends to the prefix: sublinear wall-clock growth is an empirical | |
| batching/amortization result, not constant total compute. If profiling shows branch | |
| compute or replicated KV dominates, compare a shallow cross-attention decision decoder | |
| under the same quality tests. Packed branching attention comes after numerical equivalence | |
| tests, and custom kernels come after profiles establish a worthwhile bottleneck. | |