File size: 3,163 Bytes
d08a55f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
4ffc667
 
 
 
 
 
 
 
6ff0c93
4ffc667
d08a55f
 
 
4ffc667
d08a55f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
---
license: mit
tags:
- online-learning
- streaming
- dimensionality-reduction
- riemannian-geometry
- covariance
- benchmark
- pytorch
- multi-gpu
---

# MOSAIC Hugging Face Scaling Benchmark v1

This repository turns the reference MOSAIC implementation into a reproducible benchmark ladder from CPU correctness checks to NVIDIA DGX Spark and multi-GPU throughput experiments.

The manuscript models a fixed-memory state `(mu, [U,S])`, where `U` tracks a retained subspace and `S` is an SPD covariance coordinate in the moving frame. Its key closed form is the square-root-free rank-one affine-invariant covariance update; the empirical gain in the paper is primarily attributed to the noise-regularized precision-preconditioned subspace direction rather than to replacing a first-order covariance retraction by the exact exponential.

## 🚀 Multi-GPU collaborators wanted

The single-GPU baseline has been validated on an **NVIDIA DGX Spark (GB10)**. We are now looking for collaborators with **2, 4, 8, 16+ GPUs and multi-node systems** to measure strong/weak scaling, numerical stability, memory behavior, and performance at larger dimensions and retained ranks.

In particular, access to **H100/H200, B100/B200, GB200/GB300, or comparable multi-GPU systems** would help establish how far MOSAIC can scale beyond the DGX Spark baseline.

The distributed benchmark scales across **independent ordered streams**; it does not incorrectly split one chronological stream and average independently evolved states.

➡️ **See [COLLABORATION.md](COLLABORATION.md) for the collaboration call and [MULTIGPU_EXPERIMENTS.md](MULTIGPU_EXPERIMENTS.md) for the standardized 2/4/8/16+ GPU experiment protocol.**

## DGX Spark quick start

```bash
git clone https://huggingface.co/kiruluta/MOSAIC-HF-Scaling-Benchmark
cd <repo>
python -m venv .venv
source .venv/bin/activate
python -m pip install -U pip
./scripts/run_dgx_spark_smoke.sh
```

## Multi-GPU

```bash
NPROC=8 ./scripts/run_multi_gpu.sh \
  --d 8192 --rank 64 --streams-per-gpu 16 --steps 100000 \
  --output results/scaling/8gpu_d8192_r64.json
```

**Important:** GPUs process independent ordered streams. MOSAIC is sequential within a stream; this repo intentionally does not invent a state-averaging rule.

## Large chronological `.npy` stream

```bash
python benchmarks/run_memmap.py /data/stream.npy --rank 64 --output results/large_stream.json
```

The array must have shape `(samples, features)` and is memory-mapped, so host RAM does not grow with dataset size.

## What collaborators should report

Commit the result JSON plus the exact command and environment metadata. Useful scale studies sweep ambient dimension `d`, retained rank `r`, independent streams/GPU, number of GPUs, dtype, and stream length. See `docs/SCALING.md` and `COLLABORATION.md`.

## Scientific scope

The benchmark preserves the manuscript's claim boundary: the strongest reported result is on controlled drifting anisotropic streams; real-data results are competitive rather than universally dominant. The benchmark is intended to discover the regimes where the method scales and where it fails, not to predeclare a winner.