Download docs/SCALING.md from kiruluta/MOSAIC-HF-Scaling-Benchmark: direct link, hf CLI and curl.
- Browser
- Download file 1.63 kB
-
https://huggingface.co/kiruluta/MOSAIC-HF-Scaling-Benchmark/resolve/main/docs/SCALING.md
- Command line
-
hf download hf://kiruluta/MOSAIC-HF-Scaling-Benchmark/docs/SCALING.md
-
curl -L -o SCALING.md https://huggingface.co/kiruluta/MOSAIC-HF-Scaling-Benchmark/resolve/main/docs/SCALING.md
Scaling MOSAIC correctly
MOSAIC is an ordered streaming recursion. The state at time t+1 depends on the state at t. Therefore this benchmark does not claim that a single temporal stream can be data-parallelized by splitting adjacent samples across GPUs and averaging states afterward.
The supported scale-out unit is an independent ordered stream: independent seeds, hyperparameter trials, sensors, assets, users, trajectories, or dataset shards whose ordering is meaningful within each shard. benchmarks/run_scaling.py vectorizes several independent streams per GPU and torchrun assigns independent streams to each GPU. Only final scalar metrics are gathered.
For one genuinely huge chronological stream, use benchmarks/run_memmap.py; it memory-maps a .npy matrix and transfers one sample at a time to a GPU. Future distributed work should first derive and validate a mathematically sound state-merge operator before claiming within-stream data parallelism.
Scale ladder
- CPU reference: run the original NumPy tests and smoke benchmarks.
- DGX Spark: run
./scripts/run_dgx_spark_smoke.sh. - Single GPU sweep: increase
d,rank,streams-per-gpu, andsteps. - Multi-GPU node:
NPROC=8 ./scripts/run_multi_gpu.sh --d 8192 --rank 64 --streams-per-gpu 16 --steps 100000. - Multi-node: use the cluster's normal
torchrunrendezvous parameters and invokebenchmarks/run_scaling.pydirectly.
Report GPU model, PyTorch/CUDA versions, world size, dtype, d, r, streams/GPU, steps, samples/s, distortion, orthogonality error, minimum SPD eigenvalue, and inverse-consistency error.