Download README.md from kiruluta/SPECTRA-RSI-HF-Scaling-Benchmark: direct link, hf CLI and curl.
- Browser
- Download file 7.21 kB
-
https://huggingface.co/kiruluta/SPECTRA-RSI-HF-Scaling-Benchmark/resolve/main/README.md
- Command line
-
hf download hf://kiruluta/SPECTRA-RSI-HF-Scaling-Benchmark/README.md
-
curl -L -o README.md https://huggingface.co/kiruluta/SPECTRA-RSI-HF-Scaling-Benchmark/resolve/main/README.md
license: mit
tags:
- benchmark
- evaluation
- recursive-self-improvement
- compressed-sensing
- mixture-of-experts
- lora
- scaling
SPECTRA-RSI Hugging Face Scaling Benchmark
Counterfactual Spectral Sketching and Anytime-Valid Gating for Modular Recursive Self-Improvement β benchmark release for independent replication and high-compute scaling.
This package converts the reference implementation into a Hugging Face-downloadable benchmark. It does not require Git. Download the repository files from Hugging Face (or use hf download), install locally, run standardized experiments, and return JSON results.
The benchmark follows the manuscript's falsifiable protocol: recover base-to-candidate capability deltas, identify perturbed expert groups and effect signs, stress-test anytime-valid acceptance under repeated peeking, measure basis-drift detection, account for total evaluation economics, and ultimately test modular rollback over repeated update cycles. No empirical scaling gain is assumed in advance.
π High-compute collaborators wanted
SPECTRA-RSI is seeking independent compute collaborators to test the benchmark at substantially larger scales.
Hardware: H100/H200, B100/B200, GB200/GB300, MI300X, and 2/4/8/16+ GPU or multi-node systems.
The goal is explicitly falsification-oriented: determine where structured recovery remains accurate, where it becomes compute-bound, and where its assumptions fail as scale increases. No positive scaling result is assumed in advance.
Priority experiments
- Larger capability-space dimensions and expert counts
- More bootstrap repetitions and independent seeds
- Known expert gains and regressions
- Canceling and interacting expert updates
- Broad/non-compressible updates
- Out-of-dictionary updates
- Dictionary drift
- Optional-stopping stress tests
- Fisher/second-order regularization
- Real open-weight LoRA/MoE-LoRA models
- 20β100-cycle long-horizon experiments
No Git workflow is required. Download the benchmark, run the requested experiments, and return the raw JSON results with hardware/software metadata. Code changes can be returned as a ZIP patch.
See COLLABORATION.md for the full contribution menu and submission format.
Quick start β no Git
python -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
python -m pytest -q
./scripts/run_smoke.sh
Run the standard scaling sweep:
./scripts/run_scaling_sweep.sh
Run a five-seed replication at a larger point:
N_SLICES=2048 N_EXPERTS=64 ./scripts/run_multiseed.sh
Or define a custom high-compute point:
python benchmarks/run_scaling.py \
--n-slices 4096 --n-experts 128 --rank 8 \
--m-coarse 800 --m-focused 1200 \
--items-per-row 512 --bootstrap-reps 100 \
--seed 1 --candidate-seed 101 --candidate single_gain \
--output results/my_system_n4096.json
What to scale
The primary axes are capability-space dimension, expert count, expert rank, measurement budget, bootstrap count, seed count, intervention type, and eventually repeated RSI cycles. Report accuracy and cost: normalized delta error, expert-support F1, weighted regression recall, false acceptance, evaluated items/tokens, solver/wall time, and accelerator metadata.
The manuscript requires dense ground-truth deltas for the research benchmark and explicitly asks for at least two open-weight model families at two scales, LoRA/MoE-LoRA conditions, known expert interventions, optional-stopping stress tests, drift experiments, ablations, and 20β100-cycle long-horizon control experiments. The current package provides the reproducible synthetic core and a stable result format; real-model and accelerator backends are the most valuable next contributions.
Validated DGX Spark baseline
The release candidate has been validated on 1 Γ NVIDIA GB10 (DGX Spark) with NVIDIA driver 580.159.03. The final test suite passes, and the validated benchmark manifest contains 7 runs at 1,024 capability slices using the finalized runner.
| Candidate | Seed | Support F1 | Delta error | Regression recall | Gate | Dense fallback |
|---|---|---|---|---|---|---|
single_gain |
99 | 1.000 | 0.3318 | 1.000 | accept | no |
single_gain |
100 | 1.000 | 0.3411 | 1.000 | accept | no |
single_regression |
99 | 0.667 | 0.4044 | 0.984 | quarantine | no |
canceling_mixture |
99 | 1.000 | 0.3344 | 1.000 | quarantine | no |
broad_noncompressible |
99 | 0.769 | 0.5666 | 0.923 | quarantine | no |
broad_noncompressible |
100 | 0.857 | 0.4642 | 1.000 | accept | no |
off_dictionary |
99 | β | β | β | β | yes |
The off_dictionary stress case produced a pilot residual of approximately 1.0 and immediately triggered dense fallback before compressed sensing or anchor evaluation, demonstrating the intended out-of-dictionary safeguard.
The broad in-dictionary stress test also exposes an important unresolved limitation. At candidate seed 100, the gate returned accept even though only 6 of 8 true experts were recovered. This result is intentionally retained rather than tuned away. A central high-compute research question is whether larger sensing budgets, more bootstrap repetitions, alternative structured regularization, Fisher/second-order information, or larger-scale real-model experiments improve this behavior without increasing false rejection.
Canonical release results are listed in results/validated_results.txt and summarized in results/validated_leaderboard.csv. Other JSON files in results/ should be treated as exploratory/development runs unless they appear in the validated manifest.
High-compute collaborators wanted
We are actively seeking contributors with H100/H200, B100/B200, GB200/GB300, MI300X, and 2/4/8/16+ GPU or multi-node systems. The immediate goal is to determine where SPECTRA-RSI's structured recovery remains accurate, where it becomes compute-bound, and where its assumptions fail as dimension, expert count, rank, seeds, and statistical resampling increase.
See COLLABORATION.md for the contribution menu and RESULTS_SCHEMA.md for the portable JSON contract. Git is optional: contributors can download the benchmark and return result files or a ZIP patch.
Repository layout
spectra_rsi/β reference algorithm implementationbenchmarks/run_scaling.pyβ parameterized benchmark runnerbenchmarks/aggregate_results.pyβ portable result aggregationscripts/run_smoke.shβ fast validationscripts/run_scaling_sweep.shβ standard dimension sweepscripts/run_multiseed.shβ replication sweepconfigs/scaling_matrix.mdβ requested scaling axesresults/β machine-readable contributor resultsCOLLABORATION.mdβ high-compute call
Scientific scope
This is a research benchmark, not an autonomous deployment system. Compressed discovery does not replace dense catastrophic-risk evaluation or human authorization. Protected anchors and discovery measurements must remain separated, and broad/non-compressible or nonlinear updates should trigger denser evaluation rather than forced sparse attribution.