# Spark pair: capacity and topology for the OpenSysOne handover Read-only inspection on **2026-09-16, 15:36–15:37 UTC / 16:36–16:37 BST**. Read the GX10 repository's `AGENTS.md`, `README.md`, `docs/fleet.md`, `docs/spark-a.md`, and `docs/spark-b.md` in full before probing. No remote settings, services, files, models, or GPU jobs were changed. No bandwidth benchmark or generation request was run. **Planning consequence:** the Sparks form one connected two-GPU pair, and that pair is currently occupied by a large inference model. GX10 is an independent single-GPU development/training host until its physical connection is added. There is no presently available three-node memory pool. ## Verified live | Fact | spark-a | spark-b | |---|---|---| | Host | `spark-d1b4` | `spark-3e2a` | | Working access from this Mac | `ssh spark-a-ts` | `ssh spark-b-ts` | | Tailnet address | `100.114.103.103` | `100.74.228.31` | | Wi-Fi address | `192.168.8.111` | `192.168.8.204` | | Unified memory, total | 121.69 GiB | 121.69 GiB | | **Unified memory, available** | **28.02 GiB** | **22.44 GiB** | | `free -b` used | 93.66 GiB | 99.24 GiB | | Main resident GPU process | `llama-server`, PID 154121 | `ggml-rpc-server`, PID 114512 | | GPU process allocation (`nvidia-smi`) | 36,888 MiB | 96,083 MiB | | GPU activity at snapshot | 0%, 45 °C, 11.02 W | 0%, 47 °C, 12.78 W | | Root filesystem free | 3.19 TB decimal | 3.59 TB decimal | | earlyoom | **inactive** | active | | Swap | **16 GiB active; 338 MiB used** | none | | SSH child `oom_score_adj` | 0 | 0 | Available RAM is the useful current capacity measure. CPU and GPU share this memory: the GPU allocation column must not be added to host RAM. An idle GPU utilization reading does not mean its resident model has released memory. The RPC worker had about one CPU core busy (`ps` lifetime %CPU 97.1). Direct `ssh spark-a` / `ssh spark-b` from this Mac each timed out after 10 seconds. Both tailnet aliases connected immediately. The remote Wi-Fi IPs and default routes were as documented; the reason the Mac's LAN path timed out was not investigated. Use the tailnet aliases during handover. On **both** boxes, all four Mellanox PCI functions are visible. Both port-0 interfaces report `operstate=up`, `carrier=1`, and `speed=200000`: | Device / interface | spark-a | spark-b | |---|---|---| | `rocep1s0f0` / `enp1s0f0np0` | `192.168.100.10/24` | `192.168.100.11/24` | | `roceP2p1s0f0` / `enP2p1s0f0np0` | `192.168.101.10/24` | `192.168.101.11/24` | `rdma link show` reports both domain-0/domain-2 port-0 devices `ACTIVE`, `LINK_UP`; port 1 is `DOWN`, `DISABLED`. These are two PCIe paths into **one physical 200 Gb/s port**, not two independent 200 Gb/s cables. Private subnets are directly connected without a gateway. Default routing remains over Wi-Fi. The active head command selects `Qwen3.8-Flash-Next-Uncensored-Q8_0-00001-of-00005.gguf`, `--rpc 192.168.100.11:50052`, `-ts 70,30`, `-c 65536`, `--reasoning off`, `--host 100.114.103.103 --port 18090`. `GET /health` returns `{"status":"ok"}`; `GET /v1/models` returns `Qwen3.8-Flash-Next-Uncensored-Q8_0-fast`. No generation was exercised. The head HTTP listener is on its tailnet address; the worker listener is on its ConnectX address. Their current model should remain accounted for when planning any new job. ## Previously measured or reported; not repeated in this scout The authoritative current topology and measurements are in [`~/code/gx10/docs/fleet.md`](/Users/andy/code/gx10/docs/fleet.md). Older summary rows in that repository are stale about spark-b's OTA state, the active model, and some boot/storage facts; the dated detailed sections and live inspection above take precedence. - **User reported today:** the GX10 has no ConnectX connection; connectivity expansion is expected in about two days. This scout did not probe GX10. - **GX10 docs measured 2026-09-16 after a cable-connected reboot:** the pair uses port 0 to port 0 through a 1 m Amphenol passive DAC. Each PCIe domain reaches **108.9 Gb/s RDMA** individually; both together reached **188 Gb/s** (23.5 GB/s decimal). TCP with four streams is approximately 90–97 Gb/s per domain and 171 Gb/s combined in that test. Link speed alone does not prove these rates remain available. - **Cable handling:** hot-plugging has reduced measured throughput to about 13.4 Gb/s RDMA despite a 200000 link-speed reading. A reboot with the cable connected restored throughput. Pulling the cable while a split model was resident wedged its RPC resources even though TCP still looked connected. - **Distributed software evidence:** llama.cpp inference over RPC/RoCE is proven. This does **not** establish working PyTorch/NCCL distributed training, optimizer sharding, or shared coherent memory across the boxes. The documented community NCCL figure is not a measurement on this fleet. - **Serving model:** docs record a 176 GB Q8_0 model with a 52.5 GB host-side embedding table on spark-a. The `70,30` split deliberately sends more layers to RPC0/spark-b; reversing it caused OOMs. MiniMax-M3 was stopped at 14:04 BST to make room; its files remain on spark-a. - **Host policy:** both Sparks' UFW firewalls are documented inactive. That was not rechecked with sudo. Existing tailnet-only HTTP and dedicated ConnectX RPC bindings were verified above. No new listener is needed for the single-host smoke training task. ## What the consolidated plan should assume 1. Use GX10 for the immediate tiny-model training smoke test, with its own live free-memory and running-service preflight. Treat its resources as one GB10 and preserve enough headroom for the handover session. 2. Keep the two Sparks as a separate, currently occupied resource. Starting a distributed training experiment requires a deliberate service transition, memory release verification, spark-a OOM hardening, and a small NCCL/training correctness test before increasing model size. 3. Budget future Spark-pair jobs against two separate ~121.7 GiB unified pools, subtracting OS and safety reserves on each rank. Do not assume a flat 243 GiB allocation, or count GX10 as a third training rank yet. 4. After GX10 connectivity arrives, verify the physical three-node topology, both communication directions, and PyTorch collective correctness before revising the compute budget. The current two-node DAC arrangement does not by itself specify the future topology. To refresh the non-invasive handover checks, run these separately on `spark-a-ts` and `spark-b-ts`: `date -u`, `free -b`, `nvidia-smi --query-compute-apps=pid,process_name,used_memory --format=csv`, `systemctl is-active earlyoom`, `swapon --show --bytes`, `ip -br -4 address`, and `rdma link show`. The inspection did not modify the GX10 documentation repository. ## GX10 control path and source references **Verified 2026-09-16 15:47–15:48 UTC / 16:47–16:48 BST:** GX10 can reach both Sparks directly over their Wi-Fi LAN addresses with its existing SSH keys and existing host trust. Both commands exited 0 and returned the expected hostname plus `2026-09-16T15:48:17Z`: ```bash # Run on GX10. No SSH configuration changes are needed. ssh -o BatchMode=yes -o ConnectTimeout=5 -o StrictHostKeyChecking=yes \ andy@192.168.8.111 'hostname; date -u' # spark-d1b4 ssh -o BatchMode=yes -o ConnectTimeout=5 -o StrictHostKeyChecking=yes \ andy@192.168.8.204 'hostname; date -u' # spark-3e2a ``` GX10 does not currently resolve the short aliases `spark-a` / `spark-b`. The documented tailnet addresses are reachable far enough for SSH host-key checking, but `HostKeyAlias=spark-a` / `spark-b` is not present in GX10's known host trust and strict checks therefore refused those variants. No host keys or SSH configuration were added. The numeric LAN commands above are the verified handover path; the Mac's tailnet aliases are local to the Mac. Neither `/home/andy/projects/gx10` nor `/home/andy/code/gx10` exists on GX10 at the inspection time. A sanitized copy of the infrastructure reference docs under `/home/andy/ai/opensysone/gx10-reference` is therefore useful for the handover; this scout did not create that copy. No other directories were searched, so this establishes the absence of those two standard checkout paths, rather than proving that no reference copy exists anywhere on the host.