File size: 7,208 Bytes
1398681 83a7d09 1398681 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 | ---
license: mit
tags:
- benchmark
- evaluation
- recursive-self-improvement
- compressed-sensing
- mixture-of-experts
- lora
- scaling
---
# SPECTRA-RSI Hugging Face Scaling Benchmark
**Counterfactual Spectral Sketching and Anytime-Valid Gating for Modular Recursive Self-Improvement** β benchmark release for independent replication and high-compute scaling.
This package converts the reference implementation into a **Hugging Face-downloadable benchmark**. It does **not require Git**. Download the repository files from Hugging Face (or use `hf download`), install locally, run standardized experiments, and return JSON results.
The benchmark follows the manuscript's falsifiable protocol: recover base-to-candidate capability deltas, identify perturbed expert groups and effect signs, stress-test anytime-valid acceptance under repeated peeking, measure basis-drift detection, account for total evaluation economics, and ultimately test modular rollback over repeated update cycles. No empirical scaling gain is assumed in advance.
## π High-compute collaborators wanted
SPECTRA-RSI is seeking independent compute collaborators to test the benchmark at substantially larger scales.
**Hardware:** H100/H200, B100/B200, GB200/GB300, MI300X, and 2/4/8/16+ GPU or multi-node systems.
The goal is explicitly falsification-oriented: determine where structured recovery remains accurate, where it becomes compute-bound, and where its assumptions fail as scale increases. No positive scaling result is assumed in advance.
**Priority experiments**
- Larger capability-space dimensions and expert counts
- More bootstrap repetitions and independent seeds
- Known expert gains and regressions
- Canceling and interacting expert updates
- Broad/non-compressible updates
- Out-of-dictionary updates
- Dictionary drift
- Optional-stopping stress tests
- Fisher/second-order regularization
- Real open-weight LoRA/MoE-LoRA models
- 20β100-cycle long-horizon experiments
No Git workflow is required. Download the benchmark, run the requested experiments, and return the raw JSON results with hardware/software metadata. Code changes can be returned as a ZIP patch.
See **`COLLABORATION.md`** for the full contribution menu and submission format.
## Quick start β no Git
```bash
python -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
python -m pytest -q
./scripts/run_smoke.sh
```
Run the standard scaling sweep:
```bash
./scripts/run_scaling_sweep.sh
```
Run a five-seed replication at a larger point:
```bash
N_SLICES=2048 N_EXPERTS=64 ./scripts/run_multiseed.sh
```
Or define a custom high-compute point:
```bash
python benchmarks/run_scaling.py \
--n-slices 4096 --n-experts 128 --rank 8 \
--m-coarse 800 --m-focused 1200 \
--items-per-row 512 --bootstrap-reps 100 \
--seed 1 --candidate-seed 101 --candidate single_gain \
--output results/my_system_n4096.json
```
## What to scale
The primary axes are capability-space dimension, expert count, expert rank, measurement budget, bootstrap count, seed count, intervention type, and eventually repeated RSI cycles. Report accuracy **and** cost: normalized delta error, expert-support F1, weighted regression recall, false acceptance, evaluated items/tokens, solver/wall time, and accelerator metadata.
The manuscript requires dense ground-truth deltas for the research benchmark and explicitly asks for at least two open-weight model families at two scales, LoRA/MoE-LoRA conditions, known expert interventions, optional-stopping stress tests, drift experiments, ablations, and 20β100-cycle long-horizon control experiments. The current package provides the reproducible synthetic core and a stable result format; real-model and accelerator backends are the most valuable next contributions.
## Validated DGX Spark baseline
The release candidate has been validated on **1 Γ NVIDIA GB10 (DGX Spark)** with NVIDIA driver **580.159.03**. The final test suite passes, and the validated benchmark manifest contains **7 runs** at 1,024 capability slices using the finalized runner.
| Candidate | Seed | Support F1 | Delta error | Regression recall | Gate | Dense fallback |
|---|---:|---:|---:|---:|---|---|
| `single_gain` | 99 | 1.000 | 0.3318 | 1.000 | accept | no |
| `single_gain` | 100 | 1.000 | 0.3411 | 1.000 | accept | no |
| `single_regression` | 99 | 0.667 | 0.4044 | 0.984 | quarantine | no |
| `canceling_mixture` | 99 | 1.000 | 0.3344 | 1.000 | quarantine | no |
| `broad_noncompressible` | 99 | 0.769 | 0.5666 | 0.923 | quarantine | no |
| `broad_noncompressible` | 100 | 0.857 | 0.4642 | 1.000 | **accept** | no |
| `off_dictionary` | 99 | β | β | β | β | **yes** |
The `off_dictionary` stress case produced a pilot residual of approximately **1.0** and immediately triggered dense fallback before compressed sensing or anchor evaluation, demonstrating the intended out-of-dictionary safeguard.
The broad in-dictionary stress test also exposes an important unresolved limitation. At candidate seed 100, the gate returned **accept** even though only 6 of 8 true experts were recovered. This result is intentionally retained rather than tuned away. A central high-compute research question is whether larger sensing budgets, more bootstrap repetitions, alternative structured regularization, Fisher/second-order information, or larger-scale real-model experiments improve this behavior without increasing false rejection.
Canonical release results are listed in `results/validated_results.txt` and summarized in `results/validated_leaderboard.csv`. Other JSON files in `results/` should be treated as exploratory/development runs unless they appear in the validated manifest.
## High-compute collaborators wanted
We are actively seeking contributors with **H100/H200, B100/B200, GB200/GB300, MI300X, and 2/4/8/16+ GPU or multi-node systems**. The immediate goal is to determine where SPECTRA-RSI's structured recovery remains accurate, where it becomes compute-bound, and where its assumptions fail as dimension, expert count, rank, seeds, and statistical resampling increase.
See **`COLLABORATION.md`** for the contribution menu and **`RESULTS_SCHEMA.md`** for the portable JSON contract. Git is optional: contributors can download the benchmark and return result files or a ZIP patch.
## Repository layout
- `spectra_rsi/` β reference algorithm implementation
- `benchmarks/run_scaling.py` β parameterized benchmark runner
- `benchmarks/aggregate_results.py` β portable result aggregation
- `scripts/run_smoke.sh` β fast validation
- `scripts/run_scaling_sweep.sh` β standard dimension sweep
- `scripts/run_multiseed.sh` β replication sweep
- `configs/scaling_matrix.md` β requested scaling axes
- `results/` β machine-readable contributor results
- `COLLABORATION.md` β high-compute call
## Scientific scope
This is a research benchmark, not an autonomous deployment system. Compressed discovery does not replace dense catastrophic-risk evaluation or human authorization. Protected anchors and discovery measurements must remain separated, and broad/non-compressible or nonlinear updates should trigger denser evaluation rather than forced sparse attribution.
|