SPECTRA-RSI Hugging Face Scaling Benchmark

Counterfactual Spectral Sketching and Anytime-Valid Gating for Modular Recursive Self-Improvement β€” benchmark release for independent replication and high-compute scaling.

This package converts the reference implementation into a Hugging Face-downloadable benchmark. It does not require Git. Download the repository files from Hugging Face (or use hf download), install locally, run standardized experiments, and return JSON results.

The benchmark follows the manuscript's falsifiable protocol: recover base-to-candidate capability deltas, identify perturbed expert groups and effect signs, stress-test anytime-valid acceptance under repeated peeking, measure basis-drift detection, account for total evaluation economics, and ultimately test modular rollback over repeated update cycles. No empirical scaling gain is assumed in advance.

πŸš€ High-compute collaborators wanted

SPECTRA-RSI is seeking independent compute collaborators to test the benchmark at substantially larger scales.

Hardware: H100/H200, B100/B200, GB200/GB300, MI300X, and 2/4/8/16+ GPU or multi-node systems.

The goal is explicitly falsification-oriented: determine where structured recovery remains accurate, where it becomes compute-bound, and where its assumptions fail as scale increases. No positive scaling result is assumed in advance.

Priority experiments

  • Larger capability-space dimensions and expert counts
  • More bootstrap repetitions and independent seeds
  • Known expert gains and regressions
  • Canceling and interacting expert updates
  • Broad/non-compressible updates
  • Out-of-dictionary updates
  • Dictionary drift
  • Optional-stopping stress tests
  • Fisher/second-order regularization
  • Real open-weight LoRA/MoE-LoRA models
  • 20–100-cycle long-horizon experiments

No Git workflow is required. Download the benchmark, run the requested experiments, and return the raw JSON results with hardware/software metadata. Code changes can be returned as a ZIP patch.

See COLLABORATION.md for the full contribution menu and submission format.

Quick start β€” no Git

python -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
python -m pytest -q
./scripts/run_smoke.sh

Run the standard scaling sweep:

./scripts/run_scaling_sweep.sh

Run a five-seed replication at a larger point:

N_SLICES=2048 N_EXPERTS=64 ./scripts/run_multiseed.sh

Or define a custom high-compute point:

python benchmarks/run_scaling.py \
  --n-slices 4096 --n-experts 128 --rank 8 \
  --m-coarse 800 --m-focused 1200 \
  --items-per-row 512 --bootstrap-reps 100 \
  --seed 1 --candidate-seed 101 --candidate single_gain \
  --output results/my_system_n4096.json

What to scale

The primary axes are capability-space dimension, expert count, expert rank, measurement budget, bootstrap count, seed count, intervention type, and eventually repeated RSI cycles. Report accuracy and cost: normalized delta error, expert-support F1, weighted regression recall, false acceptance, evaluated items/tokens, solver/wall time, and accelerator metadata.

The manuscript requires dense ground-truth deltas for the research benchmark and explicitly asks for at least two open-weight model families at two scales, LoRA/MoE-LoRA conditions, known expert interventions, optional-stopping stress tests, drift experiments, ablations, and 20–100-cycle long-horizon control experiments. The current package provides the reproducible synthetic core and a stable result format; real-model and accelerator backends are the most valuable next contributions.

Validated DGX Spark baseline

The release candidate has been validated on 1 Γ— NVIDIA GB10 (DGX Spark) with NVIDIA driver 580.159.03. The final test suite passes, and the validated benchmark manifest contains 7 runs at 1,024 capability slices using the finalized runner.

Candidate Seed Support F1 Delta error Regression recall Gate Dense fallback
single_gain 99 1.000 0.3318 1.000 accept no
single_gain 100 1.000 0.3411 1.000 accept no
single_regression 99 0.667 0.4044 0.984 quarantine no
canceling_mixture 99 1.000 0.3344 1.000 quarantine no
broad_noncompressible 99 0.769 0.5666 0.923 quarantine no
broad_noncompressible 100 0.857 0.4642 1.000 accept no
off_dictionary 99 β€” β€” β€” β€” yes

The off_dictionary stress case produced a pilot residual of approximately 1.0 and immediately triggered dense fallback before compressed sensing or anchor evaluation, demonstrating the intended out-of-dictionary safeguard.

The broad in-dictionary stress test also exposes an important unresolved limitation. At candidate seed 100, the gate returned accept even though only 6 of 8 true experts were recovered. This result is intentionally retained rather than tuned away. A central high-compute research question is whether larger sensing budgets, more bootstrap repetitions, alternative structured regularization, Fisher/second-order information, or larger-scale real-model experiments improve this behavior without increasing false rejection.

Canonical release results are listed in results/validated_results.txt and summarized in results/validated_leaderboard.csv. Other JSON files in results/ should be treated as exploratory/development runs unless they appear in the validated manifest.

High-compute collaborators wanted

We are actively seeking contributors with H100/H200, B100/B200, GB200/GB300, MI300X, and 2/4/8/16+ GPU or multi-node systems. The immediate goal is to determine where SPECTRA-RSI's structured recovery remains accurate, where it becomes compute-bound, and where its assumptions fail as dimension, expert count, rank, seeds, and statistical resampling increase.

See COLLABORATION.md for the contribution menu and RESULTS_SCHEMA.md for the portable JSON contract. Git is optional: contributors can download the benchmark and return result files or a ZIP patch.

Repository layout

  • spectra_rsi/ β€” reference algorithm implementation
  • benchmarks/run_scaling.py β€” parameterized benchmark runner
  • benchmarks/aggregate_results.py β€” portable result aggregation
  • scripts/run_smoke.sh β€” fast validation
  • scripts/run_scaling_sweep.sh β€” standard dimension sweep
  • scripts/run_multiseed.sh β€” replication sweep
  • configs/scaling_matrix.md β€” requested scaling axes
  • results/ β€” machine-readable contributor results
  • COLLABORATION.md β€” high-compute call

Scientific scope

This is a research benchmark, not an autonomous deployment system. Compressed discovery does not replace dense catastrophic-risk evaluation or human authorization. Protected anchors and discovery measurements must remain separated, and broad/non-compressible or nonlinear updates should trigger denser evaluation rather than forced sparse attribution.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support