Download source/docs/operations/fleet-scout.md from andyshu/opensysone: direct link, hf CLI and curl.
- Browser
- Download file 8.38 kB
-
https://huggingface.co/andyshu/opensysone/resolve/main/source/docs/operations/fleet-scout.md
- Command line
-
hf download hf://andyshu/opensysone/source/docs/operations/fleet-scout.md
-
curl -L -o fleet-scout.md https://huggingface.co/andyshu/opensysone/resolve/main/source/docs/operations/fleet-scout.md
Spark pair: capacity and topology for the OpenSysOne handover
Read-only inspection on 2026-09-16, 15:36–15:37 UTC / 16:36–16:37 BST.
Read the GX10 repository's AGENTS.md, README.md, docs/fleet.md,
docs/spark-a.md, and docs/spark-b.md in full before probing. No remote settings,
services, files, models, or GPU jobs were changed. No bandwidth benchmark or
generation request was run.
Planning consequence: the Sparks form one connected two-GPU pair, and that pair is currently occupied by a large inference model. GX10 is an independent single-GPU development/training host until its physical connection is added. There is no presently available three-node memory pool.
Verified live
| Fact | spark-a | spark-b |
|---|---|---|
| Host | spark-d1b4 |
spark-3e2a |
| Working access from this Mac | ssh spark-a-ts |
ssh spark-b-ts |
| Tailnet address | 100.114.103.103 |
100.74.228.31 |
| Wi-Fi address | 192.168.8.111 |
192.168.8.204 |
| Unified memory, total | 121.69 GiB | 121.69 GiB |
| Unified memory, available | 28.02 GiB | 22.44 GiB |
free -b used |
93.66 GiB | 99.24 GiB |
| Main resident GPU process | llama-server, PID 154121 |
ggml-rpc-server, PID 114512 |
GPU process allocation (nvidia-smi) |
36,888 MiB | 96,083 MiB |
| GPU activity at snapshot | 0%, 45 °C, 11.02 W | 0%, 47 °C, 12.78 W |
| Root filesystem free | 3.19 TB decimal | 3.59 TB decimal |
| earlyoom | inactive | active |
| Swap | 16 GiB active; 338 MiB used | none |
SSH child oom_score_adj |
0 | 0 |
Available RAM is the useful current capacity measure. CPU and GPU share this
memory: the GPU allocation column must not be added to host RAM. An idle GPU
utilization reading does not mean its resident model has released memory.
The RPC worker had about one CPU core busy (ps lifetime %CPU 97.1).
Direct ssh spark-a / ssh spark-b from this Mac each timed out after 10 seconds.
Both tailnet aliases connected immediately. The remote Wi-Fi IPs and default
routes were as documented; the reason the Mac's LAN path timed out was not
investigated. Use the tailnet aliases during handover.
On both boxes, all four Mellanox PCI functions are visible. Both port-0
interfaces report operstate=up, carrier=1, and speed=200000:
| Device / interface | spark-a | spark-b |
|---|---|---|
rocep1s0f0 / enp1s0f0np0 |
192.168.100.10/24 |
192.168.100.11/24 |
roceP2p1s0f0 / enP2p1s0f0np0 |
192.168.101.10/24 |
192.168.101.11/24 |
rdma link show reports both domain-0/domain-2 port-0 devices ACTIVE,
LINK_UP; port 1 is DOWN, DISABLED. These are two PCIe paths into one
physical 200 Gb/s port, not two independent 200 Gb/s cables. Private subnets
are directly connected without a gateway. Default routing remains over Wi-Fi.
The active head command selects
Qwen3.8-Flash-Next-Uncensored-Q8_0-00001-of-00005.gguf,
--rpc 192.168.100.11:50052, -ts 70,30, -c 65536,
--reasoning off, --host 100.114.103.103 --port 18090.
GET /health returns {"status":"ok"}; GET /v1/models returns
Qwen3.8-Flash-Next-Uncensored-Q8_0-fast. No generation was exercised.
The head HTTP listener is on its tailnet address; the worker listener is on
its ConnectX address. Their current model should remain accounted for when
planning any new job.
Previously measured or reported; not repeated in this scout
The authoritative current topology and measurements are in
~/code/gx10/docs/fleet.md.
Older summary rows in that repository are stale about spark-b's OTA state,
the active model, and some boot/storage facts; the dated detailed sections
and live inspection above take precedence.
- User reported today: the GX10 has no ConnectX connection; connectivity expansion is expected in about two days. This scout did not probe GX10.
- GX10 docs measured 2026-09-16 after a cable-connected reboot: the pair uses port 0 to port 0 through a 1 m Amphenol passive DAC. Each PCIe domain reaches 108.9 Gb/s RDMA individually; both together reached 188 Gb/s (23.5 GB/s decimal). TCP with four streams is approximately 90–97 Gb/s per domain and 171 Gb/s combined in that test. Link speed alone does not prove these rates remain available.
- Cable handling: hot-plugging has reduced measured throughput to about 13.4 Gb/s RDMA despite a 200000 link-speed reading. A reboot with the cable connected restored throughput. Pulling the cable while a split model was resident wedged its RPC resources even though TCP still looked connected.
- Distributed software evidence: llama.cpp inference over RPC/RoCE is proven. This does not establish working PyTorch/NCCL distributed training, optimizer sharding, or shared coherent memory across the boxes. The documented community NCCL figure is not a measurement on this fleet.
- Serving model: docs record a 176 GB Q8_0 model with a 52.5 GB host-side
embedding table on spark-a. The
70,30split deliberately sends more layers to RPC0/spark-b; reversing it caused OOMs. MiniMax-M3 was stopped at 14:04 BST to make room; its files remain on spark-a. - Host policy: both Sparks' UFW firewalls are documented inactive. That was not rechecked with sudo. Existing tailnet-only HTTP and dedicated ConnectX RPC bindings were verified above. No new listener is needed for the single-host smoke training task.
What the consolidated plan should assume
- Use GX10 for the immediate tiny-model training smoke test, with its own live free-memory and running-service preflight. Treat its resources as one GB10 and preserve enough headroom for the handover session.
- Keep the two Sparks as a separate, currently occupied resource. Starting a distributed training experiment requires a deliberate service transition, memory release verification, spark-a OOM hardening, and a small NCCL/training correctness test before increasing model size.
- Budget future Spark-pair jobs against two separate ~121.7 GiB unified pools, subtracting OS and safety reserves on each rank. Do not assume a flat 243 GiB allocation, or count GX10 as a third training rank yet.
- After GX10 connectivity arrives, verify the physical three-node topology, both communication directions, and PyTorch collective correctness before revising the compute budget. The current two-node DAC arrangement does not by itself specify the future topology.
To refresh the non-invasive handover checks, run these separately on
spark-a-ts and spark-b-ts: date -u, free -b,
nvidia-smi --query-compute-apps=pid,process_name,used_memory --format=csv,
systemctl is-active earlyoom, swapon --show --bytes, ip -br -4 address,
and rdma link show. The inspection did not modify the GX10 documentation
repository.
GX10 control path and source references
Verified 2026-09-16 15:47–15:48 UTC / 16:47–16:48 BST: GX10 can reach
both Sparks directly over their Wi-Fi LAN addresses with its existing SSH
keys and existing host trust. Both commands exited 0 and returned the expected
hostname plus 2026-09-16T15:48:17Z:
# Run on GX10. No SSH configuration changes are needed.
ssh -o BatchMode=yes -o ConnectTimeout=5 -o StrictHostKeyChecking=yes \
andy@192.168.8.111 'hostname; date -u'
# spark-d1b4
ssh -o BatchMode=yes -o ConnectTimeout=5 -o StrictHostKeyChecking=yes \
andy@192.168.8.204 'hostname; date -u'
# spark-3e2a
GX10 does not currently resolve the short aliases spark-a / spark-b.
The documented tailnet addresses are reachable far enough for SSH host-key
checking, but HostKeyAlias=spark-a / spark-b is not present in GX10's known
host trust and strict checks therefore refused those variants. No host keys
or SSH configuration were added. The numeric LAN commands above are the
verified handover path; the Mac's tailnet aliases are local to the Mac.
Neither /home/andy/projects/gx10 nor /home/andy/code/gx10 exists on GX10
at the inspection time. A sanitized copy of the infrastructure reference docs
under /home/andy/ai/opensysone/gx10-reference is therefore useful for the
handover; this scout did not create that copy. No other directories were
searched, so this establishes the absence of those two standard checkout
paths, rather than proving that no reference copy exists anywhere on the host.