kiruluta's picture
Upload folder using huggingface_hub
1398681 verified
|
Raw History Blame Contribute Delete
3.21 kB
# High-compute collaboration call
SPECTRA-RSI needs independent scaling evidence. The manuscript explicitly treats its numerical gains as hypotheses to test, with dense ground-truth deltas, causal-attribution metrics, optional-stopping stress tests, drift tests, and 20–100-cycle long-horizon studies.
## Highest-value contributions
1. Run larger `n_slices`, `n_experts`, rank, bootstrap, and multi-seed sweeps and upload the resulting JSON files.
2. Port the hot numerical paths to PyTorch/JAX/CuPy and benchmark CPU vs single-GPU vs multi-GPU without changing estimands.
3. Add real open-weight LoRA/MoE-LoRA adapters implementing the `paired_scores` and `probe_delta` interface.
4. Test two model families and two scales, including single-expert gains/regressions, canceling mixtures, interacting experts, non-compressible changes, and out-of-trust-region updates.
5. Run optional-stopping, dictionary-drift, Fisher/second-order dictionary regularization, and 20–100 update-cycle experiments.
## Hardware sought
H100/H200, B100/B200, GB200/GB300, MI300X, 2/4/8/16+ GPU servers, and multi-node clusters. Other accelerators are welcome when the environment is fully reported.
## Submission format (no Git required)
Download this Hugging Face repository, run the benchmark, then upload result JSON files back to a Hugging Face discussion/community contribution or share them with the maintainer. Include hardware, software versions, command line, wall time, and any code patch as a ZIP if you changed the backend. Do not submit only a screenshot.
## Current reference point and open scaling question
The validated reference point is **1 × NVIDIA GB10 (DGX Spark)** at 1,024 capability slices. Sparse single-gain and canceling-mixture cases recover their true expert support exactly in the validated runs, while the genuine `off_dictionary` case triggers immediate dense fallback.
The most important current stress-test result is deliberately unresolved: for `broad_noncompressible` at candidate seed 100, the gate returned **accept** while only 6 of 8 true experts were recovered. Contributors should not tune around or discard this case. We specifically want to learn whether increased sensing budgets, bootstrap repetitions, expert count/rank, alternative structured regularization, Fisher/second-order information, or real-model experiments eliminate this failure mode while preserving evaluation efficiency.
High-compute contributors are especially encouraged to report both successful and failed runs. Negative results, false accepts, unstable support recovery, poor scaling, memory limits, and dense-fallback behavior are scientifically useful.
## Reproducibility metadata
The benchmark records accelerator metadata through PyTorch when available and falls back to `nvidia-smi` on NVIDIA systems. Please also report GPU/accelerator model and count, CPU, RAM, driver/runtime, operating system, and whether the run used bare metal, a container, or a scheduler job.
Use `results/validated_results.txt` and `results/validated_leaderboard.csv` as the canonical DGX Spark reference set. Other bundled JSON files may be exploratory or development runs unless listed in the validated manifest.