# Collaboration Call: Scale MOSAIC Beyond a Single DGX Spark We are seeking collaborators with access to multi-GPU and multi-node NVIDIA systems to extend the scaling study for: **MOSAIC: Closed-Form Riemannian Streaming Updates on a Low-Rank Covariance Bundle** ## Baseline established on NVIDIA DGX Spark The current repository provides a reproducible single-GPU PyTorch/CUDA benchmark validated on an NVIDIA GB10. Representative measured results include: | Dimension `d` | Rank `r` | Streams/GPU | Throughput | |---:|---:|---:|---:| | 2,048 | 64 | 8 | 12,794 samples/s | | 4,096 | 64 | 8 | 9,014 samples/s | | 8,192 | 64 | 8 | 5,210 samples/s | | 16,384 | 64 | 8 | 2,654 samples/s | At `d=8192`, the measured rank scaling is: | Rank `r` | Throughput | |---:|---:| | 32 | 9,758 samples/s | | 64 | 5,210 samples/s | | 128 | 2,444 samples/s | A sustained `d=8192`, `r=64`, 8-stream run processed **800,000 observations** at approximately **5,199 samples/s** while maintaining the expected numerical invariants. Raw benchmark JSON files are provided under `results/scaling/`. ## Multi-GPU collaboration target We are particularly interested in collaborators who can reproduce and extend these experiments on: - 2 GPUs - 4 GPUs - 8 GPUs - 16+ GPUs - Multi-node GPU clusters - H100/H200, B100/B200, GB200/GB300 and comparable accelerator systems The immediate goal is to characterize throughput, scaling efficiency, numerical stability, memory behavior, and bottlenecks as GPU count, ambient dimension, retained rank, and stream count increase. ## Important scientific constraint MOSAIC is an ordered online algorithm: the state at time `t+1` depends on the state at time `t`. The current distributed benchmark therefore scales over **independent ordered streams**. Each GPU owns one or more complete ordered streams, and only benchmark statistics are aggregated. Do **not** partition one chronological stream across GPUs and average independently evolved MOSAIC states while claiming equivalence to the sequential algorithm. A mathematically justified state-merge or parallel-prefix formulation would itself be an important research contribution. ## High-priority experiments 1. Measure strong and weak scaling from 1 to 2, 4, 8, 16+ GPUs. 2. Determine the saturation point for independent streams per GPU. 3. Explore large `d` and large `r` regimes beyond the DGX Spark baseline. 4. Measure GPU memory usage and communication overhead. 5. Test sustained streams containing millions to billions of observations. 6. Evaluate naturally ordered video, sensor, scientific, financial, and telemetry datasets. 7. Investigate maintained-factorization or iterative alternatives to the dense `O(r^3)` precision solve. 8. Study missingness, heterogeneous noise, and covariance-scale debiasing. 9. Investigate mathematically valid distributed state-merging or parallel-prefix formulations. ## What to contribute Please submit: - raw result JSON files, - exact launch commands, - GPU model and GPU count, - CPU and system-memory information, - PyTorch/CUDA/NCCL versions, - benchmark configuration, - throughput, - numerical-stability diagnostics, - peak GPU memory where available, - plots generated from measured results, - and a short description of any code changes. Positive results, scaling limits, numerical failures, and performance regressions are all useful. The objective is not merely to demonstrate favorable scaling, but to establish where MOSAIC works, where it stops scaling efficiently, and what algorithmic changes are required for genuinely large distributed streaming workloads.