--- license: mit tags: - benchmark - evaluation - recursive-self-improvement - compressed-sensing - mixture-of-experts - lora - scaling --- # SPECTRA-RSI Hugging Face Scaling Benchmark **Counterfactual Spectral Sketching and Anytime-Valid Gating for Modular Recursive Self-Improvement** — benchmark release for independent replication and high-compute scaling. This package converts the reference implementation into a **Hugging Face-downloadable benchmark**. It does **not require Git**. Download the repository files from Hugging Face (or use `hf download`), install locally, run standardized experiments, and return JSON results. The benchmark follows the manuscript's falsifiable protocol: recover base-to-candidate capability deltas, identify perturbed expert groups and effect signs, stress-test anytime-valid acceptance under repeated peeking, measure basis-drift detection, account for total evaluation economics, and ultimately test modular rollback over repeated update cycles. No empirical scaling gain is assumed in advance. ## 🚀 High-compute collaborators wanted SPECTRA-RSI is seeking independent compute collaborators to test the benchmark at substantially larger scales. **Hardware:** H100/H200, B100/B200, GB200/GB300, MI300X, and 2/4/8/16+ GPU or multi-node systems. The goal is explicitly falsification-oriented: determine where structured recovery remains accurate, where it becomes compute-bound, and where its assumptions fail as scale increases. No positive scaling result is assumed in advance. **Priority experiments** - Larger capability-space dimensions and expert counts - More bootstrap repetitions and independent seeds - Known expert gains and regressions - Canceling and interacting expert updates - Broad/non-compressible updates - Out-of-dictionary updates - Dictionary drift - Optional-stopping stress tests - Fisher/second-order regularization - Real open-weight LoRA/MoE-LoRA models - 20–100-cycle long-horizon experiments No Git workflow is required. Download the benchmark, run the requested experiments, and return the raw JSON results with hardware/software metadata. Code changes can be returned as a ZIP patch. See **`COLLABORATION.md`** for the full contribution menu and submission format. ## Quick start — no Git ```bash python -m venv .venv source .venv/bin/activate pip install -e ".[dev]" python -m pytest -q ./scripts/run_smoke.sh ``` Run the standard scaling sweep: ```bash ./scripts/run_scaling_sweep.sh ``` Run a five-seed replication at a larger point: ```bash N_SLICES=2048 N_EXPERTS=64 ./scripts/run_multiseed.sh ``` Or define a custom high-compute point: ```bash python benchmarks/run_scaling.py \ --n-slices 4096 --n-experts 128 --rank 8 \ --m-coarse 800 --m-focused 1200 \ --items-per-row 512 --bootstrap-reps 100 \ --seed 1 --candidate-seed 101 --candidate single_gain \ --output results/my_system_n4096.json ``` ## What to scale The primary axes are capability-space dimension, expert count, expert rank, measurement budget, bootstrap count, seed count, intervention type, and eventually repeated RSI cycles. Report accuracy **and** cost: normalized delta error, expert-support F1, weighted regression recall, false acceptance, evaluated items/tokens, solver/wall time, and accelerator metadata. The manuscript requires dense ground-truth deltas for the research benchmark and explicitly asks for at least two open-weight model families at two scales, LoRA/MoE-LoRA conditions, known expert interventions, optional-stopping stress tests, drift experiments, ablations, and 20–100-cycle long-horizon control experiments. The current package provides the reproducible synthetic core and a stable result format; real-model and accelerator backends are the most valuable next contributions. ## Validated DGX Spark baseline The release candidate has been validated on **1 × NVIDIA GB10 (DGX Spark)** with NVIDIA driver **580.159.03**. The final test suite passes, and the validated benchmark manifest contains **7 runs** at 1,024 capability slices using the finalized runner. | Candidate | Seed | Support F1 | Delta error | Regression recall | Gate | Dense fallback | |---|---:|---:|---:|---:|---|---| | `single_gain` | 99 | 1.000 | 0.3318 | 1.000 | accept | no | | `single_gain` | 100 | 1.000 | 0.3411 | 1.000 | accept | no | | `single_regression` | 99 | 0.667 | 0.4044 | 0.984 | quarantine | no | | `canceling_mixture` | 99 | 1.000 | 0.3344 | 1.000 | quarantine | no | | `broad_noncompressible` | 99 | 0.769 | 0.5666 | 0.923 | quarantine | no | | `broad_noncompressible` | 100 | 0.857 | 0.4642 | 1.000 | **accept** | no | | `off_dictionary` | 99 | — | — | — | — | **yes** | The `off_dictionary` stress case produced a pilot residual of approximately **1.0** and immediately triggered dense fallback before compressed sensing or anchor evaluation, demonstrating the intended out-of-dictionary safeguard. The broad in-dictionary stress test also exposes an important unresolved limitation. At candidate seed 100, the gate returned **accept** even though only 6 of 8 true experts were recovered. This result is intentionally retained rather than tuned away. A central high-compute research question is whether larger sensing budgets, more bootstrap repetitions, alternative structured regularization, Fisher/second-order information, or larger-scale real-model experiments improve this behavior without increasing false rejection. Canonical release results are listed in `results/validated_results.txt` and summarized in `results/validated_leaderboard.csv`. Other JSON files in `results/` should be treated as exploratory/development runs unless they appear in the validated manifest. ## High-compute collaborators wanted We are actively seeking contributors with **H100/H200, B100/B200, GB200/GB300, MI300X, and 2/4/8/16+ GPU or multi-node systems**. The immediate goal is to determine where SPECTRA-RSI's structured recovery remains accurate, where it becomes compute-bound, and where its assumptions fail as dimension, expert count, rank, seeds, and statistical resampling increase. See **`COLLABORATION.md`** for the contribution menu and **`RESULTS_SCHEMA.md`** for the portable JSON contract. Git is optional: contributors can download the benchmark and return result files or a ZIP patch. ## Repository layout - `spectra_rsi/` — reference algorithm implementation - `benchmarks/run_scaling.py` — parameterized benchmark runner - `benchmarks/aggregate_results.py` — portable result aggregation - `scripts/run_smoke.sh` — fast validation - `scripts/run_scaling_sweep.sh` — standard dimension sweep - `scripts/run_multiseed.sh` — replication sweep - `configs/scaling_matrix.md` — requested scaling axes - `results/` — machine-readable contributor results - `COLLABORATION.md` — high-compute call ## Scientific scope This is a research benchmark, not an autonomous deployment system. Compressed discovery does not replace dense catastrophic-risk evaluation or human authorization. Protected anchors and discovery measurements must remain separated, and broad/non-compressible or nonlinear updates should trigger denser evaluation rather than forced sparse attribution.