# MOSAIC Multi-GPU Scaling Experiments This protocol defines the standardized multi-GPU scaling study for MOSAIC. ## Important interpretation MOSAIC is sequential within an ordered stream. Multi-GPU scaling therefore assigns independent ordered streams to individual GPU processes. It does not partition one chronological stream across GPUs or average independently evolved MOSAIC states. ## Standard GPU-count scaling experiment Use the same workload per GPU so results can be compared across systems: - Dimension: `d = 8192` - Retained rank: `r = 64` - Independent streams per GPU: `8` - Steps per stream: `10000` - GPU counts requested: `2, 4, 8, 16+` This experiment measures aggregate throughput/weak scaling across independent ordered streams. It is not parallel training of a single MOSAIC state. ## Launch command For a single host containing multiple GPUs: ```bash NPROC= ./scripts/run_multi_gpu.sh --d 8192 --rank 64 --streams-per-gpu 8 --steps 10000 --output results/scaling/gpu_d8192_r64.json ``` Replace `` with the number of GPUs being tested, for example `2`, `4`, `8`, or `16`. ## Results to report For each run, please contribute the generated JSON result together with: - GPU model and GPU count - CPU and system RAM - PyTorch version - CUDA version - NCCL version where applicable - Exact launch command - Aggregate samples/second - Numerical invariant diagnostics - Peak GPU memory, if available - Any OOM, numerical instability, or scaling saturation observed Negative results and scaling limits are valuable and should be reported.