# High-compute collaboration call SPECTRA-RSI needs independent scaling evidence. The manuscript explicitly treats its numerical gains as hypotheses to test, with dense ground-truth deltas, causal-attribution metrics, optional-stopping stress tests, drift tests, and 20–100-cycle long-horizon studies. ## Highest-value contributions 1. Run larger `n_slices`, `n_experts`, rank, bootstrap, and multi-seed sweeps and upload the resulting JSON files. 2. Port the hot numerical paths to PyTorch/JAX/CuPy and benchmark CPU vs single-GPU vs multi-GPU without changing estimands. 3. Add real open-weight LoRA/MoE-LoRA adapters implementing the `paired_scores` and `probe_delta` interface. 4. Test two model families and two scales, including single-expert gains/regressions, canceling mixtures, interacting experts, non-compressible changes, and out-of-trust-region updates. 5. Run optional-stopping, dictionary-drift, Fisher/second-order dictionary regularization, and 20–100 update-cycle experiments. ## Hardware sought H100/H200, B100/B200, GB200/GB300, MI300X, 2/4/8/16+ GPU servers, and multi-node clusters. Other accelerators are welcome when the environment is fully reported. ## Submission format (no Git required) Download this Hugging Face repository, run the benchmark, then upload result JSON files back to a Hugging Face discussion/community contribution or share them with the maintainer. Include hardware, software versions, command line, wall time, and any code patch as a ZIP if you changed the backend. Do not submit only a screenshot. ## Current reference point and open scaling question The validated reference point is **1 × NVIDIA GB10 (DGX Spark)** at 1,024 capability slices. Sparse single-gain and canceling-mixture cases recover their true expert support exactly in the validated runs, while the genuine `off_dictionary` case triggers immediate dense fallback. The most important current stress-test result is deliberately unresolved: for `broad_noncompressible` at candidate seed 100, the gate returned **accept** while only 6 of 8 true experts were recovered. Contributors should not tune around or discard this case. We specifically want to learn whether increased sensing budgets, bootstrap repetitions, expert count/rank, alternative structured regularization, Fisher/second-order information, or real-model experiments eliminate this failure mode while preserving evaluation efficiency. High-compute contributors are especially encouraged to report both successful and failed runs. Negative results, false accepts, unstable support recovery, poor scaling, memory limits, and dense-fallback behavior are scientifically useful. ## Reproducibility metadata The benchmark records accelerator metadata through PyTorch when available and falls back to `nvidia-smi` on NVIDIA systems. Please also report GPU/accelerator model and count, CPU, RAM, driver/runtime, operating system, and whether the run used bare metal, a container, or a scheduler job. Use `results/validated_results.txt` and `results/validated_leaderboard.csv` as the canonical DGX Spark reference set. Other bundled JSON files may be exploratory or development runs unless listed in the validated manifest.