File size: 1,628 Bytes
d08a55f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
# Scaling MOSAIC correctly

MOSAIC is an **ordered streaming recursion**. The state at time `t+1` depends on the state at `t`. Therefore this benchmark does **not** claim that a single temporal stream can be data-parallelized by splitting adjacent samples across GPUs and averaging states afterward.

The supported scale-out unit is an **independent ordered stream**: independent seeds, hyperparameter trials, sensors, assets, users, trajectories, or dataset shards whose ordering is meaningful within each shard. `benchmarks/run_scaling.py` vectorizes several independent streams per GPU and `torchrun` assigns independent streams to each GPU. Only final scalar metrics are gathered.

For one genuinely huge chronological stream, use `benchmarks/run_memmap.py`; it memory-maps a `.npy` matrix and transfers one sample at a time to a GPU. Future distributed work should first derive and validate a mathematically sound state-merge operator before claiming within-stream data parallelism.

## Scale ladder

1. CPU reference: run the original NumPy tests and smoke benchmarks.
2. DGX Spark: run `./scripts/run_dgx_spark_smoke.sh`.
3. Single GPU sweep: increase `d`, `rank`, `streams-per-gpu`, and `steps`.
4. Multi-GPU node: `NPROC=8 ./scripts/run_multi_gpu.sh --d 8192 --rank 64 --streams-per-gpu 16 --steps 100000`.
5. Multi-node: use the cluster's normal `torchrun` rendezvous parameters and invoke `benchmarks/run_scaling.py` directly.

Report GPU model, PyTorch/CUDA versions, world size, dtype, `d`, `r`, streams/GPU, steps, samples/s, distortion, orthogonality error, minimum SPD eigenvalue, and inverse-consistency error.