opensysone / source /docs /operations /fleet-scout.md
andyshu's picture
Organize verified OpenSysOne publication payload
294f8ea verified
|
Raw History Blame Contribute Delete
8.38 kB

Spark pair: capacity and topology for the OpenSysOne handover

Read-only inspection on 2026-09-16, 15:36–15:37 UTC / 16:36–16:37 BST. Read the GX10 repository's AGENTS.md, README.md, docs/fleet.md, docs/spark-a.md, and docs/spark-b.md in full before probing. No remote settings, services, files, models, or GPU jobs were changed. No bandwidth benchmark or generation request was run.

Planning consequence: the Sparks form one connected two-GPU pair, and that pair is currently occupied by a large inference model. GX10 is an independent single-GPU development/training host until its physical connection is added. There is no presently available three-node memory pool.

Verified live

Fact spark-a spark-b
Host spark-d1b4 spark-3e2a
Working access from this Mac ssh spark-a-ts ssh spark-b-ts
Tailnet address 100.114.103.103 100.74.228.31
Wi-Fi address 192.168.8.111 192.168.8.204
Unified memory, total 121.69 GiB 121.69 GiB
Unified memory, available 28.02 GiB 22.44 GiB
free -b used 93.66 GiB 99.24 GiB
Main resident GPU process llama-server, PID 154121 ggml-rpc-server, PID 114512
GPU process allocation (nvidia-smi) 36,888 MiB 96,083 MiB
GPU activity at snapshot 0%, 45 °C, 11.02 W 0%, 47 °C, 12.78 W
Root filesystem free 3.19 TB decimal 3.59 TB decimal
earlyoom inactive active
Swap 16 GiB active; 338 MiB used none
SSH child oom_score_adj 0 0

Available RAM is the useful current capacity measure. CPU and GPU share this memory: the GPU allocation column must not be added to host RAM. An idle GPU utilization reading does not mean its resident model has released memory. The RPC worker had about one CPU core busy (ps lifetime %CPU 97.1).

Direct ssh spark-a / ssh spark-b from this Mac each timed out after 10 seconds. Both tailnet aliases connected immediately. The remote Wi-Fi IPs and default routes were as documented; the reason the Mac's LAN path timed out was not investigated. Use the tailnet aliases during handover.

On both boxes, all four Mellanox PCI functions are visible. Both port-0 interfaces report operstate=up, carrier=1, and speed=200000:

Device / interface spark-a spark-b
rocep1s0f0 / enp1s0f0np0 192.168.100.10/24 192.168.100.11/24
roceP2p1s0f0 / enP2p1s0f0np0 192.168.101.10/24 192.168.101.11/24

rdma link show reports both domain-0/domain-2 port-0 devices ACTIVE, LINK_UP; port 1 is DOWN, DISABLED. These are two PCIe paths into one physical 200 Gb/s port, not two independent 200 Gb/s cables. Private subnets are directly connected without a gateway. Default routing remains over Wi-Fi.

The active head command selects Qwen3.8-Flash-Next-Uncensored-Q8_0-00001-of-00005.gguf, --rpc 192.168.100.11:50052, -ts 70,30, -c 65536, --reasoning off, --host 100.114.103.103 --port 18090. GET /health returns {"status":"ok"}; GET /v1/models returns Qwen3.8-Flash-Next-Uncensored-Q8_0-fast. No generation was exercised. The head HTTP listener is on its tailnet address; the worker listener is on its ConnectX address. Their current model should remain accounted for when planning any new job.

Previously measured or reported; not repeated in this scout

The authoritative current topology and measurements are in ~/code/gx10/docs/fleet.md. Older summary rows in that repository are stale about spark-b's OTA state, the active model, and some boot/storage facts; the dated detailed sections and live inspection above take precedence.

  • User reported today: the GX10 has no ConnectX connection; connectivity expansion is expected in about two days. This scout did not probe GX10.
  • GX10 docs measured 2026-09-16 after a cable-connected reboot: the pair uses port 0 to port 0 through a 1 m Amphenol passive DAC. Each PCIe domain reaches 108.9 Gb/s RDMA individually; both together reached 188 Gb/s (23.5 GB/s decimal). TCP with four streams is approximately 90–97 Gb/s per domain and 171 Gb/s combined in that test. Link speed alone does not prove these rates remain available.
  • Cable handling: hot-plugging has reduced measured throughput to about 13.4 Gb/s RDMA despite a 200000 link-speed reading. A reboot with the cable connected restored throughput. Pulling the cable while a split model was resident wedged its RPC resources even though TCP still looked connected.
  • Distributed software evidence: llama.cpp inference over RPC/RoCE is proven. This does not establish working PyTorch/NCCL distributed training, optimizer sharding, or shared coherent memory across the boxes. The documented community NCCL figure is not a measurement on this fleet.
  • Serving model: docs record a 176 GB Q8_0 model with a 52.5 GB host-side embedding table on spark-a. The 70,30 split deliberately sends more layers to RPC0/spark-b; reversing it caused OOMs. MiniMax-M3 was stopped at 14:04 BST to make room; its files remain on spark-a.
  • Host policy: both Sparks' UFW firewalls are documented inactive. That was not rechecked with sudo. Existing tailnet-only HTTP and dedicated ConnectX RPC bindings were verified above. No new listener is needed for the single-host smoke training task.

What the consolidated plan should assume

  1. Use GX10 for the immediate tiny-model training smoke test, with its own live free-memory and running-service preflight. Treat its resources as one GB10 and preserve enough headroom for the handover session.
  2. Keep the two Sparks as a separate, currently occupied resource. Starting a distributed training experiment requires a deliberate service transition, memory release verification, spark-a OOM hardening, and a small NCCL/training correctness test before increasing model size.
  3. Budget future Spark-pair jobs against two separate ~121.7 GiB unified pools, subtracting OS and safety reserves on each rank. Do not assume a flat 243 GiB allocation, or count GX10 as a third training rank yet.
  4. After GX10 connectivity arrives, verify the physical three-node topology, both communication directions, and PyTorch collective correctness before revising the compute budget. The current two-node DAC arrangement does not by itself specify the future topology.

To refresh the non-invasive handover checks, run these separately on spark-a-ts and spark-b-ts: date -u, free -b, nvidia-smi --query-compute-apps=pid,process_name,used_memory --format=csv, systemctl is-active earlyoom, swapon --show --bytes, ip -br -4 address, and rdma link show. The inspection did not modify the GX10 documentation repository.

GX10 control path and source references

Verified 2026-09-16 15:47–15:48 UTC / 16:47–16:48 BST: GX10 can reach both Sparks directly over their Wi-Fi LAN addresses with its existing SSH keys and existing host trust. Both commands exited 0 and returned the expected hostname plus 2026-09-16T15:48:17Z:

# Run on GX10. No SSH configuration changes are needed.
ssh -o BatchMode=yes -o ConnectTimeout=5 -o StrictHostKeyChecking=yes \
  andy@192.168.8.111 'hostname; date -u'
# spark-d1b4
ssh -o BatchMode=yes -o ConnectTimeout=5 -o StrictHostKeyChecking=yes \
  andy@192.168.8.204 'hostname; date -u'
# spark-3e2a

GX10 does not currently resolve the short aliases spark-a / spark-b. The documented tailnet addresses are reachable far enough for SSH host-key checking, but HostKeyAlias=spark-a / spark-b is not present in GX10's known host trust and strict checks therefore refused those variants. No host keys or SSH configuration were added. The numeric LAN commands above are the verified handover path; the Mac's tailnet aliases are local to the Mac.

Neither /home/andy/projects/gx10 nor /home/andy/code/gx10 exists on GX10 at the inspection time. A sanitized copy of the infrastructure reference docs under /home/andy/ai/opensysone/gx10-reference is therefore useful for the handover; this scout did not create that copy. No other directories were searched, so this establishes the absence of those two standard checkout paths, rather than proving that no reference copy exists anywhere on the host.