File size: 7,208 Bytes
1398681
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
83a7d09
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1398681
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
---
license: mit
tags:
- benchmark
- evaluation
- recursive-self-improvement
- compressed-sensing
- mixture-of-experts
- lora
- scaling
---
# SPECTRA-RSI Hugging Face Scaling Benchmark

**Counterfactual Spectral Sketching and Anytime-Valid Gating for Modular Recursive Self-Improvement** β€” benchmark release for independent replication and high-compute scaling.

This package converts the reference implementation into a **Hugging Face-downloadable benchmark**. It does **not require Git**. Download the repository files from Hugging Face (or use `hf download`), install locally, run standardized experiments, and return JSON results.

The benchmark follows the manuscript's falsifiable protocol: recover base-to-candidate capability deltas, identify perturbed expert groups and effect signs, stress-test anytime-valid acceptance under repeated peeking, measure basis-drift detection, account for total evaluation economics, and ultimately test modular rollback over repeated update cycles. No empirical scaling gain is assumed in advance.

## πŸš€ High-compute collaborators wanted

SPECTRA-RSI is seeking independent compute collaborators to test the benchmark at substantially larger scales.

**Hardware:** H100/H200, B100/B200, GB200/GB300, MI300X, and 2/4/8/16+ GPU or multi-node systems.

The goal is explicitly falsification-oriented: determine where structured recovery remains accurate, where it becomes compute-bound, and where its assumptions fail as scale increases. No positive scaling result is assumed in advance.

**Priority experiments**
- Larger capability-space dimensions and expert counts
- More bootstrap repetitions and independent seeds
- Known expert gains and regressions
- Canceling and interacting expert updates
- Broad/non-compressible updates
- Out-of-dictionary updates
- Dictionary drift
- Optional-stopping stress tests
- Fisher/second-order regularization
- Real open-weight LoRA/MoE-LoRA models
- 20–100-cycle long-horizon experiments

No Git workflow is required. Download the benchmark, run the requested experiments, and return the raw JSON results with hardware/software metadata. Code changes can be returned as a ZIP patch.

See **`COLLABORATION.md`** for the full contribution menu and submission format.

## Quick start β€” no Git

```bash
python -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
python -m pytest -q
./scripts/run_smoke.sh
```

Run the standard scaling sweep:

```bash
./scripts/run_scaling_sweep.sh
```

Run a five-seed replication at a larger point:

```bash
N_SLICES=2048 N_EXPERTS=64 ./scripts/run_multiseed.sh
```

Or define a custom high-compute point:

```bash
python benchmarks/run_scaling.py \
  --n-slices 4096 --n-experts 128 --rank 8 \
  --m-coarse 800 --m-focused 1200 \
  --items-per-row 512 --bootstrap-reps 100 \
  --seed 1 --candidate-seed 101 --candidate single_gain \
  --output results/my_system_n4096.json
```

## What to scale

The primary axes are capability-space dimension, expert count, expert rank, measurement budget, bootstrap count, seed count, intervention type, and eventually repeated RSI cycles. Report accuracy **and** cost: normalized delta error, expert-support F1, weighted regression recall, false acceptance, evaluated items/tokens, solver/wall time, and accelerator metadata.

The manuscript requires dense ground-truth deltas for the research benchmark and explicitly asks for at least two open-weight model families at two scales, LoRA/MoE-LoRA conditions, known expert interventions, optional-stopping stress tests, drift experiments, ablations, and 20–100-cycle long-horizon control experiments. The current package provides the reproducible synthetic core and a stable result format; real-model and accelerator backends are the most valuable next contributions.

## Validated DGX Spark baseline

The release candidate has been validated on **1 Γ— NVIDIA GB10 (DGX Spark)** with NVIDIA driver **580.159.03**. The final test suite passes, and the validated benchmark manifest contains **7 runs** at 1,024 capability slices using the finalized runner.

| Candidate | Seed | Support F1 | Delta error | Regression recall | Gate | Dense fallback |
|---|---:|---:|---:|---:|---|---|
| `single_gain` | 99 | 1.000 | 0.3318 | 1.000 | accept | no |
| `single_gain` | 100 | 1.000 | 0.3411 | 1.000 | accept | no |
| `single_regression` | 99 | 0.667 | 0.4044 | 0.984 | quarantine | no |
| `canceling_mixture` | 99 | 1.000 | 0.3344 | 1.000 | quarantine | no |
| `broad_noncompressible` | 99 | 0.769 | 0.5666 | 0.923 | quarantine | no |
| `broad_noncompressible` | 100 | 0.857 | 0.4642 | 1.000 | **accept** | no |
| `off_dictionary` | 99 | β€” | β€” | β€” | β€” | **yes** |

The `off_dictionary` stress case produced a pilot residual of approximately **1.0** and immediately triggered dense fallback before compressed sensing or anchor evaluation, demonstrating the intended out-of-dictionary safeguard.

The broad in-dictionary stress test also exposes an important unresolved limitation. At candidate seed 100, the gate returned **accept** even though only 6 of 8 true experts were recovered. This result is intentionally retained rather than tuned away. A central high-compute research question is whether larger sensing budgets, more bootstrap repetitions, alternative structured regularization, Fisher/second-order information, or larger-scale real-model experiments improve this behavior without increasing false rejection.

Canonical release results are listed in `results/validated_results.txt` and summarized in `results/validated_leaderboard.csv`. Other JSON files in `results/` should be treated as exploratory/development runs unless they appear in the validated manifest.

## High-compute collaborators wanted

We are actively seeking contributors with **H100/H200, B100/B200, GB200/GB300, MI300X, and 2/4/8/16+ GPU or multi-node systems**. The immediate goal is to determine where SPECTRA-RSI's structured recovery remains accurate, where it becomes compute-bound, and where its assumptions fail as dimension, expert count, rank, seeds, and statistical resampling increase.

See **`COLLABORATION.md`** for the contribution menu and **`RESULTS_SCHEMA.md`** for the portable JSON contract. Git is optional: contributors can download the benchmark and return result files or a ZIP patch.

## Repository layout

- `spectra_rsi/` β€” reference algorithm implementation
- `benchmarks/run_scaling.py` β€” parameterized benchmark runner
- `benchmarks/aggregate_results.py` β€” portable result aggregation
- `scripts/run_smoke.sh` β€” fast validation
- `scripts/run_scaling_sweep.sh` β€” standard dimension sweep
- `scripts/run_multiseed.sh` β€” replication sweep
- `configs/scaling_matrix.md` β€” requested scaling axes
- `results/` β€” machine-readable contributor results
- `COLLABORATION.md` β€” high-compute call

## Scientific scope

This is a research benchmark, not an autonomous deployment system. Compressed discovery does not replace dense catastrophic-risk evaluation or human authorization. Protected anchors and discovery measurements must remain separated, and broad/non-compressible or nonlinear updates should trigger denser evaluation rather than forced sparse attribution.