kiruluta commited on
Commit
1398681
·
verified ·
1 Parent(s): e35ea45

Upload folder using huggingface_hub

Browse files
COLLABORATION.md ADDED
@@ -0,0 +1,34 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # High-compute collaboration call
2
+
3
+ SPECTRA-RSI needs independent scaling evidence. The manuscript explicitly treats its numerical gains as hypotheses to test, with dense ground-truth deltas, causal-attribution metrics, optional-stopping stress tests, drift tests, and 20–100-cycle long-horizon studies.
4
+
5
+ ## Highest-value contributions
6
+
7
+ 1. Run larger `n_slices`, `n_experts`, rank, bootstrap, and multi-seed sweeps and upload the resulting JSON files.
8
+ 2. Port the hot numerical paths to PyTorch/JAX/CuPy and benchmark CPU vs single-GPU vs multi-GPU without changing estimands.
9
+ 3. Add real open-weight LoRA/MoE-LoRA adapters implementing the `paired_scores` and `probe_delta` interface.
10
+ 4. Test two model families and two scales, including single-expert gains/regressions, canceling mixtures, interacting experts, non-compressible changes, and out-of-trust-region updates.
11
+ 5. Run optional-stopping, dictionary-drift, Fisher/second-order dictionary regularization, and 20–100 update-cycle experiments.
12
+
13
+ ## Hardware sought
14
+
15
+ H100/H200, B100/B200, GB200/GB300, MI300X, 2/4/8/16+ GPU servers, and multi-node clusters. Other accelerators are welcome when the environment is fully reported.
16
+
17
+ ## Submission format (no Git required)
18
+
19
+ Download this Hugging Face repository, run the benchmark, then upload result JSON files back to a Hugging Face discussion/community contribution or share them with the maintainer. Include hardware, software versions, command line, wall time, and any code patch as a ZIP if you changed the backend. Do not submit only a screenshot.
20
+
21
+ ## Current reference point and open scaling question
22
+
23
+ The validated reference point is **1 × NVIDIA GB10 (DGX Spark)** at 1,024 capability slices. Sparse single-gain and canceling-mixture cases recover their true expert support exactly in the validated runs, while the genuine `off_dictionary` case triggers immediate dense fallback.
24
+
25
+ The most important current stress-test result is deliberately unresolved: for `broad_noncompressible` at candidate seed 100, the gate returned **accept** while only 6 of 8 true experts were recovered. Contributors should not tune around or discard this case. We specifically want to learn whether increased sensing budgets, bootstrap repetitions, expert count/rank, alternative structured regularization, Fisher/second-order information, or real-model experiments eliminate this failure mode while preserving evaluation efficiency.
26
+
27
+ High-compute contributors are especially encouraged to report both successful and failed runs. Negative results, false accepts, unstable support recovery, poor scaling, memory limits, and dense-fallback behavior are scientifically useful.
28
+
29
+ ## Reproducibility metadata
30
+
31
+ The benchmark records accelerator metadata through PyTorch when available and falls back to `nvidia-smi` on NVIDIA systems. Please also report GPU/accelerator model and count, CPU, RAM, driver/runtime, operating system, and whether the run used bare metal, a container, or a scheduler job.
32
+
33
+ Use `results/validated_results.txt` and `results/validated_leaderboard.csv` as the canonical DGX Spark reference set. Other bundled JSON files may be exploratory or development runs unless listed in the validated manifest.
34
+
README.md ADDED
@@ -0,0 +1,99 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: mit
3
+ tags:
4
+ - benchmark
5
+ - evaluation
6
+ - recursive-self-improvement
7
+ - compressed-sensing
8
+ - mixture-of-experts
9
+ - lora
10
+ - scaling
11
+ ---
12
+ # SPECTRA-RSI Hugging Face Scaling Benchmark
13
+
14
+ **Counterfactual Spectral Sketching and Anytime-Valid Gating for Modular Recursive Self-Improvement** — benchmark release for independent replication and high-compute scaling.
15
+
16
+ This package converts the reference implementation into a **Hugging Face-downloadable benchmark**. It does **not require Git**. Download the repository files from Hugging Face (or use `hf download`), install locally, run standardized experiments, and return JSON results.
17
+
18
+ The benchmark follows the manuscript's falsifiable protocol: recover base-to-candidate capability deltas, identify perturbed expert groups and effect signs, stress-test anytime-valid acceptance under repeated peeking, measure basis-drift detection, account for total evaluation economics, and ultimately test modular rollback over repeated update cycles. No empirical scaling gain is assumed in advance.
19
+
20
+ ## Quick start — no Git
21
+
22
+ ```bash
23
+ python -m venv .venv
24
+ source .venv/bin/activate
25
+ pip install -e ".[dev]"
26
+ python -m pytest -q
27
+ ./scripts/run_smoke.sh
28
+ ```
29
+
30
+ Run the standard scaling sweep:
31
+
32
+ ```bash
33
+ ./scripts/run_scaling_sweep.sh
34
+ ```
35
+
36
+ Run a five-seed replication at a larger point:
37
+
38
+ ```bash
39
+ N_SLICES=2048 N_EXPERTS=64 ./scripts/run_multiseed.sh
40
+ ```
41
+
42
+ Or define a custom high-compute point:
43
+
44
+ ```bash
45
+ python benchmarks/run_scaling.py \
46
+ --n-slices 4096 --n-experts 128 --rank 8 \
47
+ --m-coarse 800 --m-focused 1200 \
48
+ --items-per-row 512 --bootstrap-reps 100 \
49
+ --seed 1 --candidate-seed 101 --candidate single_gain \
50
+ --output results/my_system_n4096.json
51
+ ```
52
+
53
+ ## What to scale
54
+
55
+ The primary axes are capability-space dimension, expert count, expert rank, measurement budget, bootstrap count, seed count, intervention type, and eventually repeated RSI cycles. Report accuracy **and** cost: normalized delta error, expert-support F1, weighted regression recall, false acceptance, evaluated items/tokens, solver/wall time, and accelerator metadata.
56
+
57
+ The manuscript requires dense ground-truth deltas for the research benchmark and explicitly asks for at least two open-weight model families at two scales, LoRA/MoE-LoRA conditions, known expert interventions, optional-stopping stress tests, drift experiments, ablations, and 20–100-cycle long-horizon control experiments. The current package provides the reproducible synthetic core and a stable result format; real-model and accelerator backends are the most valuable next contributions.
58
+
59
+ ## Validated DGX Spark baseline
60
+
61
+ The release candidate has been validated on **1 × NVIDIA GB10 (DGX Spark)** with NVIDIA driver **580.159.03**. The final test suite passes, and the validated benchmark manifest contains **7 runs** at 1,024 capability slices using the finalized runner.
62
+
63
+ | Candidate | Seed | Support F1 | Delta error | Regression recall | Gate | Dense fallback |
64
+ |---|---:|---:|---:|---:|---|---|
65
+ | `single_gain` | 99 | 1.000 | 0.3318 | 1.000 | accept | no |
66
+ | `single_gain` | 100 | 1.000 | 0.3411 | 1.000 | accept | no |
67
+ | `single_regression` | 99 | 0.667 | 0.4044 | 0.984 | quarantine | no |
68
+ | `canceling_mixture` | 99 | 1.000 | 0.3344 | 1.000 | quarantine | no |
69
+ | `broad_noncompressible` | 99 | 0.769 | 0.5666 | 0.923 | quarantine | no |
70
+ | `broad_noncompressible` | 100 | 0.857 | 0.4642 | 1.000 | **accept** | no |
71
+ | `off_dictionary` | 99 | — | — | — | — | **yes** |
72
+
73
+ The `off_dictionary` stress case produced a pilot residual of approximately **1.0** and immediately triggered dense fallback before compressed sensing or anchor evaluation, demonstrating the intended out-of-dictionary safeguard.
74
+
75
+ The broad in-dictionary stress test also exposes an important unresolved limitation. At candidate seed 100, the gate returned **accept** even though only 6 of 8 true experts were recovered. This result is intentionally retained rather than tuned away. A central high-compute research question is whether larger sensing budgets, more bootstrap repetitions, alternative structured regularization, Fisher/second-order information, or larger-scale real-model experiments improve this behavior without increasing false rejection.
76
+
77
+ Canonical release results are listed in `results/validated_results.txt` and summarized in `results/validated_leaderboard.csv`. Other JSON files in `results/` should be treated as exploratory/development runs unless they appear in the validated manifest.
78
+
79
+ ## High-compute collaborators wanted
80
+
81
+ We are actively seeking contributors with **H100/H200, B100/B200, GB200/GB300, MI300X, and 2/4/8/16+ GPU or multi-node systems**. The immediate goal is to determine where SPECTRA-RSI's structured recovery remains accurate, where it becomes compute-bound, and where its assumptions fail as dimension, expert count, rank, seeds, and statistical resampling increase.
82
+
83
+ See **`COLLABORATION.md`** for the contribution menu and **`RESULTS_SCHEMA.md`** for the portable JSON contract. Git is optional: contributors can download the benchmark and return result files or a ZIP patch.
84
+
85
+ ## Repository layout
86
+
87
+ - `spectra_rsi/` — reference algorithm implementation
88
+ - `benchmarks/run_scaling.py` — parameterized benchmark runner
89
+ - `benchmarks/aggregate_results.py` — portable result aggregation
90
+ - `scripts/run_smoke.sh` — fast validation
91
+ - `scripts/run_scaling_sweep.sh` — standard dimension sweep
92
+ - `scripts/run_multiseed.sh` — replication sweep
93
+ - `configs/scaling_matrix.md` — requested scaling axes
94
+ - `results/` — machine-readable contributor results
95
+ - `COLLABORATION.md` — high-compute call
96
+
97
+ ## Scientific scope
98
+
99
+ This is a research benchmark, not an autonomous deployment system. Compressed discovery does not replace dense catastrophic-risk evaluation or human authorization. Protected anchors and discovery measurements must remain separated, and broad/non-compressible or nonlinear updates should trigger denser evaluation rather than forced sparse attribution.
RESULTS_SCHEMA.md ADDED
@@ -0,0 +1,39 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # SPECTRA-RSI Result Contract
2
+
3
+ Every benchmark invocation writes one self-contained JSON document. Raw benchmark JSON files are the primary experimental records and should not be hand-edited.
4
+
5
+ ## Core result fields
6
+
7
+ Each result records the benchmark identifier, candidate type, complete configuration, system metadata, wall time, dense-fallback status, pilot residual, and probe/sensing/anchor item counts.
8
+
9
+ When structured recovery completes, it also records the gate decision, support F1, normalized delta error, regression recall, true expert support, nominated experts, recovered experts, and bootstrap selection frequencies.
10
+
11
+ A dense fallback can legitimately omit recovery and gate fields because structured sensing and protected-anchor evaluation were not executed.
12
+
13
+ ## Reproducibility configuration
14
+
15
+ The `config` object records `n_slices`, `n_experts`, `rank`, `m_coarse`, `m_focused`, `items_per_row`, `bootstrap_reps`, `lambda_l1`, `lambda_group`, `seed`, `candidate_seed`, `candidate`, `scale`, and `output`.
16
+
17
+ Do not compare benchmark results without checking these settings.
18
+
19
+ ## System metadata
20
+
21
+ The `system` object may record hostname, platform, CUDA visibility, PyTorch version/status, GPU count and names, NVIDIA driver, and the mechanism used for GPU detection.
22
+
23
+ PyTorch is not required by the current NumPy benchmark. When PyTorch is unavailable on an NVIDIA system, the runner attempts GPU detection with `nvidia-smi`. A `torch_probe_error` such as `No module named torch` therefore does not by itself indicate benchmark failure.
24
+
25
+ ## Canonical results
26
+
27
+ `results/validated_results.txt` defines the canonical validated release set.
28
+
29
+ Generate the validated leaderboard with `python benchmarks/aggregate_results.py results --manifest results/validated_results.txt --out results/validated_leaderboard.csv`.
30
+
31
+ Other JSON files under `results/` may be exploratory, diagnostic, pre-fix, or parameter-tuning runs. Preserve them for provenance, but do not present them as canonical release results unless deliberately added to the validated manifest.
32
+
33
+ ## Contributor submissions
34
+
35
+ Preserve the raw JSON generated by the benchmark. Also report execution details that cannot be detected automatically, especially accelerator model/count, CPU, RAM, accelerator memory when known, driver/runtime versions, execution environment, command used, and any backend modifications.
36
+
37
+ Do not submit only screenshots or manually transcribed metrics.
38
+
39
+ Failed runs, dense fallbacks, false accepts, unstable recovery, memory limits, and performance regressions are valid benchmark evidence and should be retained rather than discarded.
benchmarks/aggregate_results.py ADDED
@@ -0,0 +1,87 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env python3
2
+ import argparse
3
+ import csv
4
+ import json
5
+ from pathlib import Path
6
+
7
+ p = argparse.ArgumentParser()
8
+ p.add_argument("results_dir", nargs="?", default="results")
9
+ p.add_argument("--out", default="results/leaderboard.csv")
10
+ p.add_argument(
11
+ "--manifest",
12
+ default=None,
13
+ help="Optional text file listing result JSON filenames to include.",
14
+ )
15
+ a = p.parse_args()
16
+
17
+ rows = []
18
+ results_dir = Path(a.results_dir)
19
+
20
+ if a.manifest:
21
+ manifest = Path(a.manifest)
22
+ names = [
23
+ line.strip()
24
+ for line in manifest.read_text().splitlines()
25
+ if line.strip() and not line.lstrip().startswith("#")
26
+ ]
27
+ result_files = [results_dir / name for name in names]
28
+ else:
29
+ result_files = sorted(results_dir.glob("*.json"))
30
+
31
+ for f in result_files:
32
+ try:
33
+ x = json.loads(f.read_text())
34
+ c = x.get("config", {})
35
+ s = x.get("system", {})
36
+
37
+ rows.append({
38
+ "file": f.name,
39
+ "candidate": x.get("candidate"),
40
+ "seed": c.get("seed"),
41
+ "candidate_seed": c.get("candidate_seed"),
42
+ "n_slices": c.get("n_slices"),
43
+ "n_experts": c.get("n_experts"),
44
+ "rank": c.get("rank"),
45
+ "m_coarse": c.get("m_coarse"),
46
+ "m_focused": c.get("m_focused"),
47
+ "items_per_row": c.get("items_per_row"),
48
+ "bootstrap_reps": c.get("bootstrap_reps"),
49
+ "lambda_l1": c.get("lambda_l1"),
50
+ "lambda_group": c.get("lambda_group"),
51
+ "gpu_count": s.get("gpu_count"),
52
+ "gpus": "; ".join(s.get("gpus", []) or []),
53
+ "wall_seconds": x.get("wall_seconds"),
54
+ "pilot_residual": x.get("pilot_residual"),
55
+ "dense_fallback": x.get("dense_fallback"),
56
+ "gate_decision": x.get("gate_decision"),
57
+ "support_f1": x.get("support_f1"),
58
+ "normalized_delta_error": x.get("normalized_delta_error"),
59
+ "regression_recall": x.get("regression_recall"),
60
+ "items_probe": x.get("items_probe"),
61
+ "items_sense": x.get("items_sense"),
62
+ "items_anchor": x.get("items_anchor"),
63
+ "true_experts": ";".join(
64
+ map(str, x.get("true_experts", []) or [])
65
+ ),
66
+ "nominated_experts": ";".join(
67
+ map(str, x.get("nominated_experts", []) or [])
68
+ ),
69
+ "recovered_experts": ";".join(
70
+ map(str, x.get("recovered_experts", []) or [])
71
+ ),
72
+ "bootstrap_frequencies": ";".join(
73
+ map(str, x.get("bootstrap_frequencies", []) or [])
74
+ ),
75
+ })
76
+ except Exception as e:
77
+ print(f"warning: skipped {f}: {e}")
78
+
79
+ Path(a.out).parent.mkdir(parents=True, exist_ok=True)
80
+
81
+ if rows:
82
+ with open(a.out, "w", newline="") as h:
83
+ w = csv.DictWriter(h, fieldnames=rows[0].keys())
84
+ w.writeheader()
85
+ w.writerows(rows)
86
+
87
+ print(f"wrote {len(rows)} rows to {a.out}")
benchmarks/run_scaling.py ADDED
@@ -0,0 +1,91 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env python3
2
+ import argparse, json, os, platform, socket, subprocess, time
3
+ from pathlib import Path
4
+ import numpy as np
5
+ import sys
6
+ sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
7
+ from spectra_rsi import SyntheticWorld, SpectraConfig, SpectraRSILoop
8
+ from spectra_rsi.metrics import support_f1, normalized_delta_error, regression_recall
9
+
10
+ def gpu_info():
11
+ info = {
12
+ "cuda_visible_devices": os.getenv("CUDA_VISIBLE_DEVICES"),
13
+ "hostname": socket.gethostname(),
14
+ "platform": platform.platform(),
15
+ }
16
+
17
+ try:
18
+ import torch
19
+ info.update(
20
+ torch_version=torch.__version__,
21
+ cuda_available=torch.cuda.is_available(),
22
+ gpu_count=torch.cuda.device_count(),
23
+ )
24
+ if torch.cuda.is_available():
25
+ info["gpus"] = [
26
+ torch.cuda.get_device_name(i)
27
+ for i in range(torch.cuda.device_count())
28
+ ]
29
+ except Exception as e:
30
+ info["torch_probe_error"] = str(e)
31
+
32
+ # Lightweight fallback for NVIDIA systems where PyTorch is not installed.
33
+ if not info.get("gpus"):
34
+ try:
35
+ out = subprocess.check_output(
36
+ [
37
+ "nvidia-smi",
38
+ "--query-gpu=name,driver_version",
39
+ "--format=csv,noheader",
40
+ ],
41
+ text=True,
42
+ stderr=subprocess.DEVNULL,
43
+ timeout=5,
44
+ )
45
+ rows = [line.strip() for line in out.splitlines() if line.strip()]
46
+ if rows:
47
+ names = []
48
+ drivers = []
49
+ for row in rows:
50
+ parts = [x.strip() for x in row.split(",", 1)]
51
+ names.append(parts[0])
52
+ if len(parts) > 1:
53
+ drivers.append(parts[1])
54
+
55
+ info["gpu_count"] = len(names)
56
+ info["gpus"] = names
57
+ info["nvidia_driver"] = drivers[0] if drivers else None
58
+ info["gpu_probe"] = "nvidia-smi"
59
+ except Exception as e:
60
+ info["nvidia_smi_probe_error"] = str(e)
61
+
62
+ return info
63
+
64
+ def main():
65
+ p=argparse.ArgumentParser(description="SPECTRA-RSI reproducible scaling benchmark")
66
+ p.add_argument("--n-slices",type=int,default=400); p.add_argument("--n-experts",type=int,default=16)
67
+ p.add_argument("--rank",type=int,default=2); p.add_argument("--m-coarse",type=int,default=80); p.add_argument("--m-focused",type=int,default=120)
68
+ p.add_argument("--items-per-row",type=int,default=600); p.add_argument("--bootstrap-reps",type=int,default=30)
69
+ p.add_argument("--lambda-l1",type=float,default=2e-3); p.add_argument("--lambda-group",type=float,default=8e-3)
70
+ p.add_argument("--seed",type=int,default=7); p.add_argument("--candidate-seed",type=int,default=99)
71
+ p.add_argument("--candidate",choices=["single_gain","single_regression","canceling_mixture","broad_noncompressible","off_dictionary"],default="single_gain")
72
+ p.add_argument("--scale",type=float,default=0.4); p.add_argument("--output",required=True)
73
+ a=p.parse_args(); Path(a.output).parent.mkdir(parents=True,exist_ok=True)
74
+ cfg=SpectraConfig(n_slices=a.n_slices,n_experts=a.n_experts,rank_per_expert=a.rank,m_coarse=a.m_coarse,m_focused=a.m_focused,items_per_row=a.items_per_row,bootstrap_reps=a.bootstrap_reps,
75
+ lambda_l1=a.lambda_l1,lambda_group=a.lambda_group,
76
+ seed=a.seed,audit_dir=str(Path(a.output).parent/"audit_logs"))
77
+ world=SyntheticWorld(cfg.n_slices,cfg.n_experts,cfg.rank_per_expert,seed=cfg.seed,offband_leak=0.01)
78
+ rng=np.random.default_rng(a.candidate_seed)
79
+ ex={"single_gain":[max(0,a.n_experts//3)],"single_regression":[max(0,2*a.n_experts//3)],"canceling_mixture":[max(0,a.n_experts//4),max(0,3*a.n_experts//4)]}.get(a.candidate)
80
+ kw={"scale":a.scale,"rng":rng};
81
+ if ex is not None: kw["experts"]=ex
82
+ cand=world.make_candidate(a.candidate,**kw)
83
+ t=time.perf_counter(); rep=SpectraRSILoop(world,cfg).run_iteration(cand); wall=time.perf_counter()-t
84
+ row={"benchmark":"spectra-rsi-scaling-v1","candidate":a.candidate,"config":vars(a),"system":gpu_info(),"wall_seconds":wall,"dense_fallback":bool(rep.dense_fallback),"pilot_residual":float(rep.pilot_residual),"items_probe":int(rep.items_probe),"items_sense":int(rep.items_sense),"items_anchor":int(rep.items_anchor)}
85
+ if not rep.dense_fallback:
86
+ truth=world.true_delta(cand); row.update(gate_decision=rep.gate_decision,support_f1=float(support_f1(rep.recovered_experts,cand.true_support)),normalized_delta_error=float(normalized_delta_error(rep.delta_hat,truth)),regression_recall=float(regression_recall(rep.delta_hat,truth,cfg.tau_margin)),true_experts=sorted(map(int,cand.true_support)),
87
+ nominated_experts=sorted(map(int,rep.nominated_experts)),
88
+ recovered_experts=sorted(map(int,rep.recovered_experts)),
89
+ bootstrap_frequencies=[float(x) for x in rep.bootstrap_frequencies])
90
+ Path(a.output).write_text(json.dumps(row,indent=2)+"\n"); print(json.dumps(row,indent=2))
91
+ if __name__=="__main__": main()
configs/scaling_matrix.md ADDED
@@ -0,0 +1,5 @@
 
 
 
 
 
 
1
+ # Scaling matrix
2
+
3
+ Start with the smoke test, then contribute the largest stable points your hardware permits. Priority axes are: capability slices `n` (256 → 16k+), experts `E` (8 → 256+), rank (2 → 32+), bootstrap repetitions (10 → 500+), candidate seeds (5 → 100+), and long-horizon repeated iterations (future extension).
4
+
5
+ High-compute contributors are especially requested for H100/H200, B100/B200, GB200/GB300, MI300X-class systems and multi-GPU/multi-node hosts. The current reference implementation is NumPy-first; GPU owners can contribute **measurement results and/or GPU/distributed backend ports** while preserving the JSON result contract.
pyproject.toml ADDED
@@ -0,0 +1,22 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ [build-system]
2
+ requires = ["setuptools>=61"]
3
+ build-backend = "setuptools.build_meta"
4
+
5
+ [project]
6
+ name = "spectra-rsi"
7
+ version = "0.1.0"
8
+ description = "Counterfactual spectral sketching and anytime-valid gating for modular recursive self-improvement (reference implementation)"
9
+ readme = "README.md"
10
+ requires-python = ">=3.9"
11
+ license = { text = "MIT" }
12
+ authors = [{ name = "Andrew M." }]
13
+ dependencies = ["numpy>=1.24"]
14
+
15
+ [project.optional-dependencies]
16
+ dev = ["pytest>=7.0"]
17
+
18
+ [tool.setuptools.packages.find]
19
+ include = ["spectra_rsi*"]
20
+
21
+ [tool.pytest.ini_options]
22
+ testpaths = ["tests"]
results/README.md ADDED
@@ -0,0 +1,25 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Results
2
+
3
+ This directory contains SPECTRA-RSI benchmark outputs.
4
+
5
+ ## Canonical validated results
6
+
7
+ `validated_results.txt` is the release manifest listing the raw JSON files that constitute the canonical validated baseline.
8
+
9
+ `validated_leaderboard.csv` is the corresponding aggregated table.
10
+
11
+ The current validated baseline was produced on 1 x NVIDIA GB10 (DGX Spark).
12
+
13
+ ## Exploratory results
14
+
15
+ Other JSON files may represent exploratory parameter sweeps, diagnostics, tuning experiments, or pre-fix runs. They are retained locally for provenance but should not be interpreted as canonical benchmark results unless listed in `validated_results.txt`.
16
+
17
+ The `audit_logs/` directory contains algorithm-level audit records from multiple development and experimental executions. These logs are not currently linked unambiguously to individual canonical result files and are therefore not part of the validated Hugging Face release set.
18
+
19
+ ## Regenerating the validated leaderboard
20
+
21
+ Run:
22
+
23
+ python benchmarks/aggregate_results.py results --manifest results/validated_results.txt --out results/validated_leaderboard.csv
24
+
25
+ Raw benchmark JSON files should not be hand-edited. Failed runs and negative results are scientifically useful and should be preserved with enough metadata to identify their originating experiment.
results/dgx_spark_1024_canceling.json ADDED
@@ -0,0 +1,64 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "benchmark": "spectra-rsi-scaling-v1",
3
+ "candidate": "canceling_mixture",
4
+ "config": {
5
+ "n_slices": 1024,
6
+ "n_experts": 8,
7
+ "rank": 2,
8
+ "m_coarse": 768,
9
+ "m_focused": 1024,
10
+ "items_per_row": 128,
11
+ "bootstrap_reps": 5,
12
+ "lambda_l1": 0.01,
13
+ "lambda_group": 0.04,
14
+ "seed": 7,
15
+ "candidate_seed": 99,
16
+ "candidate": "canceling_mixture",
17
+ "scale": 0.4,
18
+ "output": "results/dgx_spark_1024_canceling.json"
19
+ },
20
+ "system": {
21
+ "cuda_visible_devices": null,
22
+ "hostname": "gx10-db7a",
23
+ "platform": "Linux-6.17.0-1021-nvidia-aarch64-with-glibc2.39",
24
+ "torch_probe_error": "No module named 'torch'",
25
+ "gpu_count": 1,
26
+ "gpus": [
27
+ "NVIDIA GB10"
28
+ ],
29
+ "nvidia_driver": "580.159.03",
30
+ "gpu_probe": "nvidia-smi"
31
+ },
32
+ "wall_seconds": 0.8139567289999832,
33
+ "dense_fallback": false,
34
+ "pilot_residual": 0.20033903475967274,
35
+ "items_probe": 12800,
36
+ "items_sense": 275132,
37
+ "items_anchor": 24000,
38
+ "gate_decision": "quarantine",
39
+ "support_f1": 1.0,
40
+ "normalized_delta_error": 0.33441075948420146,
41
+ "regression_recall": 1.0,
42
+ "true_experts": [
43
+ 2,
44
+ 6
45
+ ],
46
+ "nominated_experts": [
47
+ 2,
48
+ 6
49
+ ],
50
+ "recovered_experts": [
51
+ 2,
52
+ 6
53
+ ],
54
+ "bootstrap_frequencies": [
55
+ 0.6,
56
+ 0.2,
57
+ 1.0,
58
+ 0.4,
59
+ 0.0,
60
+ 0.6,
61
+ 1.0,
62
+ 0.0
63
+ ]
64
+ }
results/dgx_spark_1024_noncompressible.json ADDED
@@ -0,0 +1,74 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "benchmark": "spectra-rsi-scaling-v1",
3
+ "candidate": "broad_noncompressible",
4
+ "config": {
5
+ "n_slices": 1024,
6
+ "n_experts": 8,
7
+ "rank": 2,
8
+ "m_coarse": 768,
9
+ "m_focused": 1024,
10
+ "items_per_row": 128,
11
+ "bootstrap_reps": 5,
12
+ "lambda_l1": 0.01,
13
+ "lambda_group": 0.04,
14
+ "seed": 7,
15
+ "candidate_seed": 99,
16
+ "candidate": "broad_noncompressible",
17
+ "scale": 0.4,
18
+ "output": "results/dgx_spark_1024_noncompressible.json"
19
+ },
20
+ "system": {
21
+ "cuda_visible_devices": null,
22
+ "hostname": "gx10-db7a",
23
+ "platform": "Linux-6.17.0-1021-nvidia-aarch64-with-glibc2.39",
24
+ "torch_probe_error": "No module named 'torch'",
25
+ "gpu_count": 1,
26
+ "gpus": [
27
+ "NVIDIA GB10"
28
+ ],
29
+ "nvidia_driver": "580.159.03",
30
+ "gpu_probe": "nvidia-smi"
31
+ },
32
+ "wall_seconds": 0.8184029540007032,
33
+ "dense_fallback": false,
34
+ "pilot_residual": 0.4261726798996619,
35
+ "items_probe": 12800,
36
+ "items_sense": 273912,
37
+ "items_anchor": 24000,
38
+ "gate_decision": "quarantine",
39
+ "support_f1": 0.7692307692307693,
40
+ "normalized_delta_error": 0.5666150578143931,
41
+ "regression_recall": 0.9230769230769231,
42
+ "true_experts": [
43
+ 0,
44
+ 1,
45
+ 2,
46
+ 3,
47
+ 4,
48
+ 5,
49
+ 6,
50
+ 7
51
+ ],
52
+ "nominated_experts": [
53
+ 1,
54
+ 2,
55
+ 5
56
+ ],
57
+ "recovered_experts": [
58
+ 1,
59
+ 2,
60
+ 3,
61
+ 4,
62
+ 5
63
+ ],
64
+ "bootstrap_frequencies": [
65
+ 0.2,
66
+ 1.0,
67
+ 1.0,
68
+ 0.8,
69
+ 0.6,
70
+ 1.0,
71
+ 0.0,
72
+ 0.6
73
+ ]
74
+ }
results/dgx_spark_1024_noncompressible_seed100.json ADDED
@@ -0,0 +1,75 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "benchmark": "spectra-rsi-scaling-v1",
3
+ "candidate": "broad_noncompressible",
4
+ "config": {
5
+ "n_slices": 1024,
6
+ "n_experts": 8,
7
+ "rank": 2,
8
+ "m_coarse": 768,
9
+ "m_focused": 1024,
10
+ "items_per_row": 128,
11
+ "bootstrap_reps": 5,
12
+ "lambda_l1": 0.01,
13
+ "lambda_group": 0.04,
14
+ "seed": 7,
15
+ "candidate_seed": 100,
16
+ "candidate": "broad_noncompressible",
17
+ "scale": 0.4,
18
+ "output": "results/dgx_spark_1024_noncompressible_seed100.json"
19
+ },
20
+ "system": {
21
+ "cuda_visible_devices": null,
22
+ "hostname": "gx10-db7a",
23
+ "platform": "Linux-6.17.0-1021-nvidia-aarch64-with-glibc2.39",
24
+ "torch_probe_error": "No module named 'torch'",
25
+ "gpu_count": 1,
26
+ "gpus": [
27
+ "NVIDIA GB10"
28
+ ],
29
+ "nvidia_driver": "580.159.03",
30
+ "gpu_probe": "nvidia-smi"
31
+ },
32
+ "wall_seconds": 0.8000310129991703,
33
+ "dense_fallback": false,
34
+ "pilot_residual": 0.2939873837188221,
35
+ "items_probe": 12800,
36
+ "items_sense": 273478,
37
+ "items_anchor": 21400,
38
+ "gate_decision": "accept",
39
+ "support_f1": 0.8571428571428571,
40
+ "normalized_delta_error": 0.4642433735647704,
41
+ "regression_recall": 1.0,
42
+ "true_experts": [
43
+ 0,
44
+ 1,
45
+ 2,
46
+ 3,
47
+ 4,
48
+ 5,
49
+ 6,
50
+ 7
51
+ ],
52
+ "nominated_experts": [
53
+ 2,
54
+ 3,
55
+ 4
56
+ ],
57
+ "recovered_experts": [
58
+ 0,
59
+ 1,
60
+ 2,
61
+ 3,
62
+ 4,
63
+ 6
64
+ ],
65
+ "bootstrap_frequencies": [
66
+ 0.8,
67
+ 1.0,
68
+ 1.0,
69
+ 1.0,
70
+ 1.0,
71
+ 0.4,
72
+ 0.8,
73
+ 0.0
74
+ ]
75
+ }
results/dgx_spark_1024_off_dictionary.json ADDED
@@ -0,0 +1,38 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "benchmark": "spectra-rsi-scaling-v1",
3
+ "candidate": "off_dictionary",
4
+ "config": {
5
+ "n_slices": 1024,
6
+ "n_experts": 8,
7
+ "rank": 2,
8
+ "m_coarse": 768,
9
+ "m_focused": 1024,
10
+ "items_per_row": 128,
11
+ "bootstrap_reps": 5,
12
+ "lambda_l1": 0.01,
13
+ "lambda_group": 0.04,
14
+ "seed": 7,
15
+ "candidate_seed": 99,
16
+ "candidate": "off_dictionary",
17
+ "scale": 0.4,
18
+ "output": "results/dgx_spark_1024_off_dictionary.json"
19
+ },
20
+ "system": {
21
+ "cuda_visible_devices": null,
22
+ "hostname": "gx10-db7a",
23
+ "platform": "Linux-6.17.0-1021-nvidia-aarch64-with-glibc2.39",
24
+ "torch_probe_error": "No module named 'torch'",
25
+ "gpu_count": 1,
26
+ "gpus": [
27
+ "NVIDIA GB10"
28
+ ],
29
+ "nvidia_driver": "580.159.03",
30
+ "gpu_probe": "nvidia-smi"
31
+ },
32
+ "wall_seconds": 0.0013118490005581407,
33
+ "dense_fallback": true,
34
+ "pilot_residual": 0.9999999916719566,
35
+ "items_probe": 0,
36
+ "items_sense": 0,
37
+ "items_anchor": 0
38
+ }
results/dgx_spark_1024_reg5x.json ADDED
@@ -0,0 +1,61 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "benchmark": "spectra-rsi-scaling-v1",
3
+ "candidate": "single_gain",
4
+ "config": {
5
+ "n_slices": 1024,
6
+ "n_experts": 8,
7
+ "rank": 2,
8
+ "m_coarse": 768,
9
+ "m_focused": 1024,
10
+ "items_per_row": 128,
11
+ "bootstrap_reps": 5,
12
+ "lambda_l1": 0.01,
13
+ "lambda_group": 0.04,
14
+ "seed": 7,
15
+ "candidate_seed": 99,
16
+ "candidate": "single_gain",
17
+ "scale": 0.4,
18
+ "output": "results/dgx_spark_1024_reg5x.json"
19
+ },
20
+ "system": {
21
+ "cuda_visible_devices": null,
22
+ "hostname": "gx10-db7a",
23
+ "platform": "Linux-6.17.0-1021-nvidia-aarch64-with-glibc2.39",
24
+ "torch_probe_error": "No module named 'torch'",
25
+ "gpu_count": 1,
26
+ "gpus": [
27
+ "NVIDIA GB10"
28
+ ],
29
+ "nvidia_driver": "580.159.03",
30
+ "gpu_probe": "nvidia-smi"
31
+ },
32
+ "wall_seconds": 0.7774730420005653,
33
+ "dense_fallback": false,
34
+ "pilot_residual": 0.2353382692677464,
35
+ "items_probe": 12800,
36
+ "items_sense": 276707,
37
+ "items_anchor": 22700,
38
+ "gate_decision": "accept",
39
+ "support_f1": 1.0,
40
+ "normalized_delta_error": 0.3317853169160946,
41
+ "regression_recall": 1.0,
42
+ "true_experts": [
43
+ 2
44
+ ],
45
+ "nominated_experts": [
46
+ 2
47
+ ],
48
+ "recovered_experts": [
49
+ 2
50
+ ],
51
+ "bootstrap_frequencies": [
52
+ 0.2,
53
+ 0.4,
54
+ 1.0,
55
+ 0.2,
56
+ 0.2,
57
+ 0.2,
58
+ 0.2,
59
+ 0.0
60
+ ]
61
+ }
results/dgx_spark_1024_reg5x_seed100.json ADDED
@@ -0,0 +1,61 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "benchmark": "spectra-rsi-scaling-v1",
3
+ "candidate": "single_gain",
4
+ "config": {
5
+ "n_slices": 1024,
6
+ "n_experts": 8,
7
+ "rank": 2,
8
+ "m_coarse": 768,
9
+ "m_focused": 1024,
10
+ "items_per_row": 128,
11
+ "bootstrap_reps": 5,
12
+ "lambda_l1": 0.01,
13
+ "lambda_group": 0.04,
14
+ "seed": 7,
15
+ "candidate_seed": 100,
16
+ "candidate": "single_gain",
17
+ "scale": 0.4,
18
+ "output": "results/dgx_spark_1024_reg5x_seed100.json"
19
+ },
20
+ "system": {
21
+ "cuda_visible_devices": null,
22
+ "hostname": "gx10-db7a",
23
+ "platform": "Linux-6.17.0-1021-nvidia-aarch64-with-glibc2.39",
24
+ "torch_probe_error": "No module named 'torch'",
25
+ "gpu_count": 1,
26
+ "gpus": [
27
+ "NVIDIA GB10"
28
+ ],
29
+ "nvidia_driver": "580.159.03",
30
+ "gpu_probe": "nvidia-smi"
31
+ },
32
+ "wall_seconds": 0.7741304430001037,
33
+ "dense_fallback": false,
34
+ "pilot_residual": 0.2339949193998105,
35
+ "items_probe": 12800,
36
+ "items_sense": 276707,
37
+ "items_anchor": 21800,
38
+ "gate_decision": "accept",
39
+ "support_f1": 1.0,
40
+ "normalized_delta_error": 0.34114380104392295,
41
+ "regression_recall": 1.0,
42
+ "true_experts": [
43
+ 2
44
+ ],
45
+ "nominated_experts": [
46
+ 2
47
+ ],
48
+ "recovered_experts": [
49
+ 2
50
+ ],
51
+ "bootstrap_frequencies": [
52
+ 0.2,
53
+ 0.6,
54
+ 1.0,
55
+ 0.2,
56
+ 0.4,
57
+ 0.2,
58
+ 0.2,
59
+ 0.0
60
+ ]
61
+ }
results/dgx_spark_1024_regression_budgetfix.json ADDED
@@ -0,0 +1,62 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "benchmark": "spectra-rsi-scaling-v1",
3
+ "candidate": "single_regression",
4
+ "config": {
5
+ "n_slices": 1024,
6
+ "n_experts": 8,
7
+ "rank": 2,
8
+ "m_coarse": 768,
9
+ "m_focused": 1024,
10
+ "items_per_row": 128,
11
+ "bootstrap_reps": 5,
12
+ "lambda_l1": 0.01,
13
+ "lambda_group": 0.04,
14
+ "seed": 7,
15
+ "candidate_seed": 99,
16
+ "candidate": "single_regression",
17
+ "scale": 0.4,
18
+ "output": "results/dgx_spark_1024_regression_budgetfix.json"
19
+ },
20
+ "system": {
21
+ "cuda_visible_devices": null,
22
+ "hostname": "gx10-db7a",
23
+ "platform": "Linux-6.17.0-1021-nvidia-aarch64-with-glibc2.39",
24
+ "torch_probe_error": "No module named 'torch'",
25
+ "gpu_count": 1,
26
+ "gpus": [
27
+ "NVIDIA GB10"
28
+ ],
29
+ "nvidia_driver": "580.159.03",
30
+ "gpu_probe": "nvidia-smi"
31
+ },
32
+ "wall_seconds": 0.7842826520009112,
33
+ "dense_fallback": false,
34
+ "pilot_residual": 0.2444575970827468,
35
+ "items_probe": 12800,
36
+ "items_sense": 276806,
37
+ "items_anchor": 24000,
38
+ "gate_decision": "quarantine",
39
+ "support_f1": 0.6666666666666666,
40
+ "normalized_delta_error": 0.40437287482632517,
41
+ "regression_recall": 0.984,
42
+ "true_experts": [
43
+ 5
44
+ ],
45
+ "nominated_experts": [
46
+ 5
47
+ ],
48
+ "recovered_experts": [
49
+ 2,
50
+ 5
51
+ ],
52
+ "bootstrap_frequencies": [
53
+ 0.0,
54
+ 0.2,
55
+ 0.6,
56
+ 0.4,
57
+ 0.4,
58
+ 1.0,
59
+ 0.2,
60
+ 0.4
61
+ ]
62
+ }
results/validated_leaderboard.csv ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ file,candidate,seed,candidate_seed,n_slices,n_experts,rank,m_coarse,m_focused,items_per_row,bootstrap_reps,lambda_l1,lambda_group,gpu_count,gpus,wall_seconds,pilot_residual,dense_fallback,gate_decision,support_f1,normalized_delta_error,regression_recall,items_probe,items_sense,items_anchor,true_experts,nominated_experts,recovered_experts,bootstrap_frequencies
2
+ dgx_spark_1024_reg5x.json,single_gain,7,99,1024,8,2,768,1024,128,5,0.01,0.04,1,NVIDIA GB10,0.7774730420005653,0.2353382692677464,False,accept,1.0,0.3317853169160946,1.0,12800,276707,22700,2,2,2,0.2;0.4;1.0;0.2;0.2;0.2;0.2;0.0
3
+ dgx_spark_1024_reg5x_seed100.json,single_gain,7,100,1024,8,2,768,1024,128,5,0.01,0.04,1,NVIDIA GB10,0.7741304430001037,0.2339949193998105,False,accept,1.0,0.34114380104392295,1.0,12800,276707,21800,2,2,2,0.2;0.6;1.0;0.2;0.4;0.2;0.2;0.0
4
+ dgx_spark_1024_regression_budgetfix.json,single_regression,7,99,1024,8,2,768,1024,128,5,0.01,0.04,1,NVIDIA GB10,0.7842826520009112,0.2444575970827468,False,quarantine,0.6666666666666666,0.40437287482632517,0.984,12800,276806,24000,5,5,2;5,0.0;0.2;0.6;0.4;0.4;1.0;0.2;0.4
5
+ dgx_spark_1024_canceling.json,canceling_mixture,7,99,1024,8,2,768,1024,128,5,0.01,0.04,1,NVIDIA GB10,0.8139567289999832,0.20033903475967274,False,quarantine,1.0,0.33441075948420146,1.0,12800,275132,24000,2;6,2;6,2;6,0.6;0.2;1.0;0.4;0.0;0.6;1.0;0.0
6
+ dgx_spark_1024_noncompressible.json,broad_noncompressible,7,99,1024,8,2,768,1024,128,5,0.01,0.04,1,NVIDIA GB10,0.8184029540007032,0.4261726798996619,False,quarantine,0.7692307692307693,0.5666150578143931,0.9230769230769231,12800,273912,24000,0;1;2;3;4;5;6;7,1;2;5,1;2;3;4;5,0.2;1.0;1.0;0.8;0.6;1.0;0.0;0.6
7
+ dgx_spark_1024_noncompressible_seed100.json,broad_noncompressible,7,100,1024,8,2,768,1024,128,5,0.01,0.04,1,NVIDIA GB10,0.8000310129991703,0.2939873837188221,False,accept,0.8571428571428571,0.4642433735647704,1.0,12800,273478,21400,0;1;2;3;4;5;6;7,2;3;4,0;1;2;3;4;6,0.8;1.0;1.0;1.0;1.0;0.4;0.8;0.0
8
+ dgx_spark_1024_off_dictionary.json,off_dictionary,7,99,1024,8,2,768,1024,128,5,0.01,0.04,1,NVIDIA GB10,0.0013118490005581407,0.9999999916719566,True,,,,,0,0,0,,,,
results/validated_results.txt ADDED
@@ -0,0 +1,7 @@
 
 
 
 
 
 
 
 
1
+ dgx_spark_1024_reg5x.json
2
+ dgx_spark_1024_reg5x_seed100.json
3
+ dgx_spark_1024_regression_budgetfix.json
4
+ dgx_spark_1024_canceling.json
5
+ dgx_spark_1024_noncompressible.json
6
+ dgx_spark_1024_noncompressible_seed100.json
7
+ dgx_spark_1024_off_dictionary.json
scripts/run_multiseed.sh ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env bash
2
+ set -euo pipefail
3
+ mkdir -p results
4
+ N=${N_SLICES:-1024}; E=${N_EXPERTS:-32}
5
+ for S in 1 2 3 4 5; do
6
+ python benchmarks/run_scaling.py --n-slices "$N" --n-experts "$E" --rank 2 --m-coarse 200 --m-focused 256 --items-per-row 256 --bootstrap-reps 20 --seed "$S" --candidate-seed "$((100+S))" --candidate single_gain --output "results/multiseed_n${N}_s${S}.json"
7
+ done
8
+ python benchmarks/aggregate_results.py results
scripts/run_scaling_sweep.sh ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env bash
2
+ set -euo pipefail
3
+ mkdir -p results
4
+ for N in 256 512 1024 2048; do
5
+ E=$((N/32)); [ "$E" -lt 8 ] && E=8
6
+ python benchmarks/run_scaling.py --n-slices "$N" --n-experts "$E" --rank 2 --m-coarse $((N/5)) --m-focused $((N/4)) --items-per-row 256 --bootstrap-reps 10 --candidate single_gain --output "results/scaling_n${N}.json"
7
+ done
8
+ python benchmarks/aggregate_results.py results
scripts/run_smoke.sh ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ #!/usr/bin/env bash
2
+ set -euo pipefail
3
+ python benchmarks/run_scaling.py --n-slices 128 --n-experts 8 --rank 2 --m-coarse 24 --m-focused 32 --items-per-row 96 --bootstrap-reps 5 --candidate single_gain --output results/smoke.json
spectra_rsi/__init__.py ADDED
@@ -0,0 +1,31 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """SPECTRA-RSI: Counterfactual Spectral Sketching and Anytime-Valid Gating
2
+ for Modular Recursive Self-Improvement.
3
+
4
+ Reference implementation of the architecture in the SPECTRA-RSI manuscript.
5
+ All components operate against a pluggable `World` interface; a synthetic
6
+ ground-truth world is provided for validation and benchmarking.
7
+ """
8
+
9
+ from .config import SpectraConfig
10
+ from .world import SyntheticWorld, CandidateUpdate
11
+ from .experts import Router, expert_priors
12
+ from .probes import build_dictionary, Dictionary
13
+ from .sketch import design_matrix, paired_sketch, neyman_allocation
14
+ from .recovery import sparse_group_recover, debias, bootstrap_support
15
+ from .drift import procrustes_transport, principal_angles, ResidualDriftMonitor
16
+ from .gate import EProcess, AnchorGate, GateDecision
17
+ from .loop import SpectraRSILoop, IterationReport
18
+ from . import metrics
19
+
20
+ __version__ = "0.1.0"
21
+
22
+ __all__ = [
23
+ "SpectraConfig", "SyntheticWorld", "CandidateUpdate",
24
+ "Router", "expert_priors",
25
+ "build_dictionary", "Dictionary",
26
+ "design_matrix", "paired_sketch", "neyman_allocation",
27
+ "sparse_group_recover", "debias", "bootstrap_support",
28
+ "procrustes_transport", "principal_angles", "ResidualDriftMonitor",
29
+ "EProcess", "AnchorGate", "GateDecision",
30
+ "SpectraRSILoop", "IterationReport", "metrics",
31
+ ]
spectra_rsi/config.py ADDED
@@ -0,0 +1,51 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Configuration for a SPECTRA-RSI loop instance."""
2
+ from dataclasses import dataclass, field
3
+
4
+
5
+ @dataclass
6
+ class SpectraConfig:
7
+ # --- capability space ---
8
+ n_slices: int = 400 # n: capability slices
9
+ n_experts: int = 16 # E: experts / dictionary groups
10
+ rank_per_expert: int = 2 # r_e: retained response rank per expert
11
+
12
+ # --- probing (counterfactual dictionary) ---
13
+ probe_magnitude: float = 0.4 # h: signed probe step (within trust region)
14
+ probe_items: int = 400 # items per slice per probe evaluation
15
+
16
+ # --- sketching ---
17
+ m_coarse: int = 60 # stage-I sketch rows
18
+ m_focused: int = 90 # stage-II sketch rows
19
+ row_density: float = 0.10 # expected fraction of nonzero weights per row
20
+ items_per_row: int = 200 # item budget per sketch row (Neyman-allocated)
21
+ explore_fraction: float = 0.15 # measurement mass outside nominated groups
22
+
23
+ # --- recovery ---
24
+ lambda_l1: float = 2e-3
25
+ lambda_group: float = 8e-3
26
+ prior_gamma: float = 0.5 # w_e = (p_e + eps)^(-gamma)
27
+ prior_clip: tuple = (0.05, 0.95)
28
+ fista_iters: int = 600
29
+ bootstrap_reps: int = 60
30
+
31
+ # --- trust region / pilot ---
32
+ pilot_items: int = 120
33
+ trust_region_residual: float = 0.5 # relative pilot residual triggering fallback
34
+
35
+ # --- drift ---
36
+ drift_alpha: float = 0.05
37
+ drift_window: int = 50
38
+
39
+ # --- anytime-valid gate ---
40
+ alpha_risk: float = 0.05 # total false-acceptance budget over critical anchors
41
+ tau_margin: float = 0.02 # material-regression margin tau_j
42
+ max_anchor_items: int = 24000 # total sequential anchor budget
43
+ anchor_batch: int = 100
44
+
45
+ # --- misc ---
46
+ seed: int = 0
47
+ audit_dir: str = "audit_logs"
48
+
49
+ def __post_init__(self):
50
+ assert 0 < self.row_density <= 1
51
+ assert 0 <= self.explore_fraction < 1
spectra_rsi/drift.py ADDED
@@ -0,0 +1,62 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Dictionary transport and drift monitoring (Sec. 4.3).
2
+
3
+ - Orthogonal Procrustes alignment of successive dictionaries.
4
+ - Principal angles quantify subspace motion.
5
+ - ResidualDriftMonitor: anytime-valid confidence sequence on predictive
6
+ residuals; a boundary crossing triggers re-probing / dense fallback.
7
+ """
8
+ import numpy as np
9
+
10
+
11
+ def procrustes_transport(Psi_new, Psi_old):
12
+ """Q = argmin_{Q^T Q = I} ||Psi_new - Psi_old Q||_F (rotation of old
13
+ coordinates into the new basis). Returns (Q, alignment_residual)."""
14
+ k = min(Psi_new.shape[1], Psi_old.shape[1])
15
+ A = Psi_old[:, :k].T @ Psi_new[:, :k]
16
+ U, _, Vt = np.linalg.svd(A)
17
+ Q = U @ Vt
18
+ resid = np.linalg.norm(Psi_new[:, :k] - Psi_old[:, :k] @ Q, "fro") / np.sqrt(k)
19
+ return Q, resid
20
+
21
+
22
+ def principal_angles(Psi_a, Psi_b):
23
+ """Principal angles (radians) between column spaces."""
24
+ Qa, _ = np.linalg.qr(Psi_a)
25
+ Qb, _ = np.linalg.qr(Psi_b)
26
+ s = np.linalg.svd(Qa.T @ Qb, compute_uv=False)
27
+ return np.arccos(np.clip(s, -1, 1))
28
+
29
+
30
+ class ResidualDriftMonitor:
31
+ """Sub-Gaussian mixture confidence sequence on the mean squared
32
+ normalized residual. Anytime-valid: crossing is a drift alarm at level
33
+ alpha regardless of when you look (Howard et al. style mixture bound)."""
34
+
35
+ def __init__(self, alpha=0.05, scale=1.0):
36
+ self.alpha = alpha
37
+ self.scale = scale # residuals assumed sub-Gaussian(scale)
38
+ self.t = 0
39
+ self.sum = 0.0
40
+
41
+ def update(self, residuals):
42
+ r = np.atleast_1d(residuals)
43
+ self.t += r.size
44
+ self.sum += float(np.sum(r))
45
+ return self.is_alarmed()
46
+
47
+ def radius(self):
48
+ if self.t == 0:
49
+ return np.inf
50
+ t, s, a = self.t, self.scale, self.alpha
51
+ # normal-mixture boundary: valid at all t simultaneously
52
+ rho = 1.0
53
+ return s * np.sqrt(2 * (t + rho) / t ** 2
54
+ * np.log(np.sqrt((t + rho) / rho) / a))
55
+
56
+ def mean(self):
57
+ return self.sum / max(self.t, 1)
58
+
59
+ def is_alarmed(self):
60
+ """Alarm when the anytime lower confidence bound on the mean
61
+ residual exceeds 0 (systematic prediction error)."""
62
+ return self.t > 0 and (self.mean() - self.radius()) > 0.0
spectra_rsi/experts.py ADDED
@@ -0,0 +1,47 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Spectral mixture of low-rank experts: router statistics and support priors.
2
+
3
+ Implements Sec. 5: occupancy pi_e, gate entropy H_e, update energy g_e,
4
+ and the clipped prior activity score p_e with an exploration floor.
5
+ Router statistics are a PRIOR, never evidence of improvement.
6
+ """
7
+ import numpy as np
8
+
9
+
10
+ class Router:
11
+ """Tracks routing statistics over the candidate-generation corpus."""
12
+
13
+ def __init__(self, n_experts, seed=0):
14
+ self.E = n_experts
15
+ self._rng = np.random.default_rng(seed)
16
+
17
+ def occupancy_stats(self, cand_beta, rank_per_expert, corpus_size=2000,
18
+ fidelity=0.8):
19
+ """Simulate router occupancy correlated (imperfectly) with which
20
+ experts the candidate actually touched. `fidelity` < 1 injects router
21
+ noise / gaming so the prior is fallible, as the paper requires."""
22
+ R = rank_per_expert
23
+ energy = np.array([np.linalg.norm(cand_beta[e * R:(e + 1) * R])
24
+ for e in range(self.E)])
25
+ base = energy / (energy.max() + 1e-12) if energy.max() > 0 else energy
26
+ noise = self._rng.random(self.E)
27
+ occ = fidelity * base + (1 - fidelity) * noise
28
+ occ = occ / (occ.sum() + 1e-12)
29
+ # gate entropy per expert (high entropy = diffuse routing)
30
+ p = np.clip(occ, 1e-9, 1)
31
+ ent = -p * np.log(p)
32
+ return occ, ent, energy
33
+
34
+
35
+ def expert_priors(occupancy, entropy, energy, clip=(0.05, 0.95),
36
+ a0=-1.0, a1=0.8, a2=1.2, a3=0.5):
37
+ """Prior activity score p_e = sigma(a0 + a1 log occ + a2 g - a3 H),
38
+ clipped to [p_min, p_max] to preserve exploration (Eq. prior)."""
39
+ g = energy / (np.max(energy) + 1e-12) if np.max(energy) > 0 else energy
40
+ z = (a0 + a1 * np.log(occupancy + 1e-6) + a2 * g - a3 * entropy)
41
+ p = 1.0 / (1.0 + np.exp(-z))
42
+ return np.clip(p, clip[0], clip[1])
43
+
44
+
45
+ def group_weights(priors, gamma=0.5, eps=1e-3):
46
+ """w_e = (p_e + eps)^(-gamma): low-prior groups pay a higher penalty."""
47
+ return (priors + eps) ** (-gamma)
spectra_rsi/gate.py ADDED
@@ -0,0 +1,109 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Anytime-valid modular acceptance gate (Sec. 8).
2
+
3
+ Non-inferiority testing on protected anchor streams via betting e-processes:
4
+
5
+ H0_j : Delta c_j <= -tau vs. H1_j : Delta c_j > -tau
6
+
7
+ For bounded paired differences d in [-B, B], the wealth process
8
+
9
+ E_s = prod_i ( 1 + lam_i * (d_i + tau) ), 0 <= lam_i < 1 / (B + tau)
10
+
11
+ is a nonnegative supermartingale under every P in H0 (E_P[d] <= -tau implies
12
+ E_P[1 + lam (d + tau)] <= 1). By Ville's inequality,
13
+ P(sup_s E_s >= 1/alpha) <= alpha under H0, at ANY stopping rule.
14
+ A grid mixture over lam keeps validity while adapting to unknown effect size.
15
+ """
16
+ from dataclasses import dataclass, field
17
+ from enum import Enum
18
+ import numpy as np
19
+
20
+
21
+ class GateDecision(Enum):
22
+ ACCEPT = "accept"
23
+ ROLLBACK = "rollback"
24
+ QUARANTINE = "quarantine"
25
+
26
+
27
+ class EProcess:
28
+ """Grid-mixture betting e-process for H0: mean <= -tau, obs in [-B, B]."""
29
+
30
+ def __init__(self, tau, alpha=0.05, bound=1.0, n_grid=12):
31
+ self.tau = tau
32
+ self.alpha = alpha
33
+ self.B = bound
34
+ lam_max = 0.5 / (bound + tau) # stay safely inside admissible range
35
+ self.grid = np.linspace(lam_max / n_grid, lam_max, n_grid)
36
+ self.log_wealth = np.zeros(n_grid)
37
+ self.n = 0
38
+
39
+ def update(self, observations):
40
+ obs = np.clip(np.atleast_1d(observations), -self.B, self.B)
41
+ for d in obs:
42
+ self.log_wealth += np.log1p(self.grid * (d + self.tau))
43
+ self.n += 1
44
+ return self.value()
45
+
46
+ def value(self):
47
+ # mixture (uniform prior over grid) — still an e-process
48
+ m = self.log_wealth.max()
49
+ return float(np.exp(m) * np.mean(np.exp(self.log_wealth - m)))
50
+
51
+ @property
52
+ def rejects_null(self):
53
+ """True iff evidence of non-inferiority has crossed 1/alpha."""
54
+ return self.value() >= 1.0 / self.alpha
55
+
56
+
57
+ @dataclass
58
+ class AnchorReport:
59
+ anchor_id: int
60
+ e_value: float
61
+ n_items: int
62
+ non_inferior: bool
63
+
64
+
65
+ class AnchorGate:
66
+ """Per-expert + composed-candidate acceptance on protected anchors.
67
+
68
+ Each critical anchor stream gets an independent e-process with a
69
+ Bonferroni share of the total risk budget alpha. Acceptance requires
70
+ every critical anchor to establish non-inferiority; budget exhaustion
71
+ without rejection sends the expert to quarantine, not acceptance.
72
+ """
73
+
74
+ def __init__(self, anchor_slices, tau, alpha_total, bound=1.0):
75
+ self.anchor_slices = list(anchor_slices)
76
+ self.tau = tau
77
+ self.alpha_each = alpha_total / max(1, len(self.anchor_slices))
78
+ self.bound = bound
79
+
80
+ def run(self, world, cand, batch=50, max_items=4000, rng=None):
81
+ """Sequentially stream anchor items until every e-process resolves
82
+ or budget runs out. Anytime-valid: any stopping rule is fine."""
83
+ rng = rng or np.random.default_rng(0)
84
+ procs = {j: EProcess(self.tau, self.alpha_each, self.bound)
85
+ for j in self.anchor_slices}
86
+ undecided = set(self.anchor_slices)
87
+ used = 0
88
+ while undecided and used < max_items:
89
+ for j in list(undecided):
90
+ if used + batch > max_items:
91
+ break
92
+ d = world.paired_scores(cand, j, batch, rng=rng)
93
+ procs[j].update(d)
94
+ used += batch
95
+ if procs[j].rejects_null:
96
+ undecided.discard(j)
97
+ reports = [AnchorReport(j, procs[j].value(), procs[j].n,
98
+ procs[j].rejects_null)
99
+ for j in self.anchor_slices]
100
+ all_pass = all(r.non_inferior for r in reports)
101
+ decision = GateDecision.ACCEPT if all_pass else (
102
+ GateDecision.ROLLBACK if any(
103
+ (not r.non_inferior) and r.n_items >= max_items // max(1, len(self.anchor_slices))
104
+ for r in reports) else GateDecision.QUARANTINE)
105
+ if not all_pass and used >= max_items:
106
+ decision = GateDecision.QUARANTINE if any(
107
+ not r.non_inferior and r.e_value > 1.0 for r in reports) \
108
+ else GateDecision.ROLLBACK
109
+ return decision, reports, used
spectra_rsi/loop.py ADDED
@@ -0,0 +1,212 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """The SPECTRA-RSI loop (Algorithm 1).
2
+
3
+ One counterfactual, risk-controlled iteration:
4
+
5
+ 1. freeze candidate; align previous basis, check drift confidence sequence
6
+ 2. (re)build interventional dictionary from signed expert probes if needed
7
+ 3. clipped router priors with exploration floor
8
+ 4. local-linearity pilot (trust region) -> dense fallback on failure
9
+ 5. stage-I coarse paired sketch over all groups
10
+ 6. preliminary recovery + bootstrap instability -> nominate groups
11
+ 7. stage-II focused sketch, retaining exploration mass
12
+ 8. weighted sparse-group recovery + debias -> expert attribution
13
+ 9. protected anchor e-processes per nominated expert + composed candidate
14
+ 10. accept / rollback / quarantine; audit record
15
+
16
+ Discovery never touches anchor data; anchors never inform the dictionary,
17
+ penalties, or nomination (Sec. 8.4 firewall).
18
+ """
19
+ from dataclasses import dataclass, field, asdict
20
+ import json, os, time, hashlib
21
+ import numpy as np
22
+
23
+ from .experts import Router, expert_priors, group_weights
24
+ from .probes import build_dictionary, pilot_linearity_check
25
+ from .sketch import design_matrix, paired_sketch, group_slice_map
26
+ from .recovery import (sparse_group_recover, debias, bootstrap_support,
27
+ active_groups)
28
+ from .drift import procrustes_transport, ResidualDriftMonitor
29
+ from .gate import AnchorGate, GateDecision
30
+
31
+
32
+ @dataclass
33
+ class IterationReport:
34
+ label: str
35
+ pilot_residual: float
36
+ pilot_passed: bool
37
+ nominated_experts: list
38
+ recovered_experts: list
39
+ bootstrap_frequencies: list
40
+ delta_hat_norm: float
41
+ gate_decision: str
42
+ anchor_reports: list
43
+ items_probe: int
44
+ items_sense: int
45
+ items_anchor: int
46
+ drift_alarm: bool
47
+ dense_fallback: bool
48
+ wall_time_s: float
49
+ audit_hash: str = ""
50
+
51
+ def to_json(self):
52
+ return json.dumps(asdict(self), indent=2, default=str)
53
+
54
+
55
+ class SpectraRSILoop:
56
+ def __init__(self, world, cfg, anchor_slices=None):
57
+ self.world = world
58
+ self.cfg = cfg
59
+ self.rng = np.random.default_rng(cfg.seed)
60
+ self.router = Router(cfg.n_experts, seed=cfg.seed + 1)
61
+ self.dictionary = None
62
+ self.prev_Psi = None
63
+ self.drift_monitor = ResidualDriftMonitor(alpha=cfg.drift_alpha)
64
+ # protected anchors: held-out slices, never used in discovery
65
+ if anchor_slices is None:
66
+ anchor_slices = list(self.rng.choice(
67
+ world.n, size=max(4, world.n // 50), replace=False))
68
+ self.anchor_slices = anchor_slices
69
+ os.makedirs(cfg.audit_dir, exist_ok=True)
70
+ self._iteration = 0
71
+
72
+ # ------------------------------------------------------------------ #
73
+ def _ensure_dictionary(self, force=False):
74
+ items = 0
75
+ if self.dictionary is None or force:
76
+ if self.dictionary is not None:
77
+ self.prev_Psi = self.dictionary.Psi.copy()
78
+ self.dictionary = build_dictionary(self.world, self.cfg, rng=self.rng)
79
+ items = (2 * self.cfg.n_experts * self.cfg.rank_per_expert
80
+ * self.cfg.probe_items * self.world.n // self.world.n)
81
+ items = (2 * self.cfg.n_experts * self.cfg.rank_per_expert
82
+ * self.cfg.probe_items)
83
+ if self.prev_Psi is not None:
84
+ procrustes_transport(self.dictionary.Psi, self.prev_Psi)
85
+ return items
86
+
87
+ # ------------------------------------------------------------------ #
88
+ def run_iteration(self, cand) -> IterationReport:
89
+ cfg, world = self.cfg, self.world
90
+ t0 = time.time()
91
+ self._iteration += 1
92
+ drift_alarm = self.drift_monitor.is_alarmed()
93
+ items_probe = self._ensure_dictionary(force=drift_alarm)
94
+ if drift_alarm:
95
+ self.drift_monitor = ResidualDriftMonitor(alpha=cfg.drift_alpha)
96
+ D = self.dictionary
97
+
98
+ # ---- router priors (fallible; clipped; exploration floor) ----
99
+ occ, ent, energy = self.router.occupancy_stats(
100
+ cand.beta, cfg.rank_per_expert)
101
+ priors = expert_priors(occ, ent, energy, clip=cfg.prior_clip)
102
+ gw = group_weights(priors, gamma=cfg.prior_gamma)
103
+
104
+ # ---- trust-region pilot ----
105
+ pilot_res, pilot_ok = pilot_linearity_check(world, cand, D, cfg,
106
+ rng=self.rng)
107
+ if not pilot_ok:
108
+ report = self._dense_fallback_report(cand, pilot_res, t0)
109
+ self._audit(report)
110
+ return report
111
+
112
+ sense_slices = [j for j in range(world.n)
113
+ if j not in set(self.anchor_slices)]
114
+ sigma = world.sigma
115
+
116
+ # ---- stage I: coarse sketch ----
117
+ A1 = design_matrix(cfg.m_coarse, world.n, cfg.row_density, sigma,
118
+ self.rng)
119
+ A1[:, self.anchor_slices] = 0.0 # firewall: no anchor leakage
120
+ y1, items1, slice_means1 = paired_sketch(
121
+ world, cand, A1, sigma, cfg.items_per_row, self.rng)
122
+ M1 = A1 @ D.Psi
123
+ x1 = sparse_group_recover(y1, M1, D.groups, gw,
124
+ cfg.lambda_l1, cfg.lambda_group,
125
+ n_iters=cfg.fista_iters // 2)
126
+ nominated, _ = active_groups(x1, D.groups)
127
+
128
+ # ---- stage II: focused sketch with exploration mass ----
129
+ gmap = group_slice_map(D)
130
+ A2 = design_matrix(cfg.m_focused, world.n, cfg.row_density, sigma,
131
+ self.rng, group_slices=gmap,
132
+ focus_groups=nominated or list(range(cfg.n_experts)),
133
+ explore_fraction=cfg.explore_fraction)
134
+ A2[:, self.anchor_slices] = 0.0
135
+ y2, items2, slice_means2 = paired_sketch(
136
+ world, cand, A2, sigma, cfg.items_per_row, self.rng)
137
+
138
+ A = np.vstack([A1, A2])
139
+ y = np.concatenate([y1, y2])
140
+ M = A @ D.Psi
141
+
142
+ # ---- recovery + debias + bootstrap ----
143
+ x_hat = sparse_group_recover(y, M, D.groups, gw,
144
+ cfg.lambda_l1, cfg.lambda_group,
145
+ n_iters=cfg.fista_iters)
146
+ x_deb, support = debias(y, M, x_hat)
147
+ freq = bootstrap_support(y, M, D.groups, gw, cfg.lambda_l1,
148
+ cfg.lambda_group, reps=cfg.bootstrap_reps,
149
+ rng=self.rng)
150
+ recovered, energies = active_groups(x_deb, D.groups)
151
+ delta_hat = D.Psi @ x_deb
152
+
153
+ # ---- drift monitor: predictive residuals on reused slice means ----
154
+ resid = []
155
+ for j, mhat in {**slice_means1, **slice_means2}.items():
156
+ resid.append((mhat - delta_hat[j]) ** 2 - sigma[j] ** 2
157
+ * 2 * (1 - world.cnr_rho) / cfg.items_per_row)
158
+ self.drift_monitor.update(np.array(resid[:cfg.drift_window]))
159
+
160
+ # ---- protected anytime-valid gate ----
161
+ gate = AnchorGate(self.anchor_slices, cfg.tau_margin, cfg.alpha_risk)
162
+ decision, anchor_reports, items_anchor = gate.run(
163
+ world, cand, batch=cfg.anchor_batch,
164
+ max_items=cfg.max_anchor_items, rng=self.rng)
165
+
166
+ report = IterationReport(
167
+ label=cand.label,
168
+ pilot_residual=float(pilot_res),
169
+ pilot_passed=True,
170
+ nominated_experts=[int(e) for e in nominated],
171
+ recovered_experts=[int(e) for e in recovered],
172
+ bootstrap_frequencies=[float(f) for f in freq],
173
+ delta_hat_norm=float(np.linalg.norm(delta_hat)),
174
+ gate_decision=decision.value,
175
+ anchor_reports=[asdict(r) for r in anchor_reports],
176
+ items_probe=items_probe,
177
+ items_sense=items1 + items2,
178
+ items_anchor=items_anchor,
179
+ drift_alarm=bool(drift_alarm),
180
+ dense_fallback=False,
181
+ wall_time_s=time.time() - t0,
182
+ )
183
+ report.delta_hat = delta_hat # attach for callers
184
+ report.x_hat = x_deb
185
+ self._audit(report)
186
+ return report
187
+
188
+ # ------------------------------------------------------------------ #
189
+ def _dense_fallback_report(self, cand, pilot_res, t0):
190
+ """Out-of-trust-region candidate: refuse sparse explanation."""
191
+ report = IterationReport(
192
+ label=cand.label, pilot_residual=float(pilot_res),
193
+ pilot_passed=False, nominated_experts=[], recovered_experts=[],
194
+ bootstrap_frequencies=[], delta_hat_norm=0.0,
195
+ gate_decision=GateDecision.QUARANTINE.value, anchor_reports=[],
196
+ items_probe=0, items_sense=0, items_anchor=0,
197
+ drift_alarm=False, dense_fallback=True,
198
+ wall_time_s=time.time() - t0)
199
+ report.delta_hat = None
200
+ report.x_hat = None
201
+ return report
202
+
203
+ # ------------------------------------------------------------------ #
204
+ def _audit(self, report):
205
+ """Immutable audit record (Sec. 14): hash-chained JSON."""
206
+ payload = report.to_json()
207
+ h = hashlib.sha256(payload.encode()).hexdigest()[:16]
208
+ report.audit_hash = h
209
+ path = os.path.join(self.cfg.audit_dir,
210
+ f"iter_{self._iteration:04d}_{h}.json")
211
+ with open(path, "w") as f:
212
+ f.write(payload)
spectra_rsi/metrics.py ADDED
@@ -0,0 +1,30 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Primary metrics (Sec. 12.5)."""
2
+ import numpy as np
3
+
4
+
5
+ def normalized_delta_error(delta_hat, delta_true, eps=1e-9):
6
+ return float(np.linalg.norm(delta_hat - delta_true)
7
+ / (np.linalg.norm(delta_true) + eps))
8
+
9
+
10
+ def support_f1(pred_support, true_support):
11
+ pred, true = set(pred_support), set(true_support)
12
+ if not pred and not true:
13
+ return 1.0
14
+ tp = len(pred & true)
15
+ prec = tp / len(pred) if pred else 0.0
16
+ rec = tp / len(true) if true else 0.0
17
+ return 0.0 if prec + rec == 0 else 2 * prec * rec / (prec + rec)
18
+
19
+
20
+ def regression_recall(delta_hat, delta_true, tau, weights=None):
21
+ """Weighted recall of materially regressed slices (Delta c < -tau)."""
22
+ n = len(delta_true)
23
+ w = np.ones(n) if weights is None else np.asarray(weights)
24
+ regressed = delta_true < -tau
25
+ if not regressed.any():
26
+ return 1.0
27
+ detected = delta_hat < -tau / 2 # detection threshold at tau/2
28
+ num = float(np.sum(w * detected * regressed))
29
+ den = float(np.sum(w * regressed))
30
+ return num / den
spectra_rsi/probes.py ADDED
@@ -0,0 +1,65 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Counterfactual spectral response dictionary (Sec. 4).
2
+
3
+ Signed central-difference probes on expert directions estimate the
4
+ interventional response matrix H_t; per-expert truncated SVD yields the
5
+ grouped dictionary Psi_t = [U_1, ..., U_E].
6
+ """
7
+ from dataclasses import dataclass
8
+ import numpy as np
9
+
10
+
11
+ @dataclass
12
+ class Dictionary:
13
+ Psi: np.ndarray # (n, p) column-orthonormal within groups
14
+ groups: list # list of index arrays into columns of Psi
15
+ H: np.ndarray # (n, d) raw fingerprint matrix
16
+ expert_maps: list # per-expert (Sigma V^T) mapping dict coords -> beta coords
17
+
18
+ @property
19
+ def p(self):
20
+ return self.Psi.shape[1]
21
+
22
+
23
+ def build_dictionary(world, cfg, rng=None) -> Dictionary:
24
+ """Apply +/- h probes for every expert direction and factorize per expert."""
25
+ rng = rng or np.random.default_rng(cfg.seed)
26
+ E, R = cfg.n_experts, cfg.rank_per_expert
27
+ d = E * R
28
+ cols = []
29
+ for idx in range(d):
30
+ direction = np.zeros(d)
31
+ direction[idx] = 1.0
32
+ cols.append(world.probe_delta(direction, cfg.probe_magnitude,
33
+ cfg.probe_items, rng=rng))
34
+ H = np.stack(cols, axis=1) # (n, d)
35
+
36
+ Psi_blocks, groups, expert_maps = [], [], []
37
+ col0 = 0
38
+ for e in range(E):
39
+ He = H[:, e * R:(e + 1) * R] # (n, R)
40
+ U, S, Vt = np.linalg.svd(He, full_matrices=False)
41
+ r = min(R, (S > 1e-8).sum())
42
+ r = max(r, 1)
43
+ Psi_blocks.append(U[:, :r])
44
+ groups.append(np.arange(col0, col0 + r))
45
+ expert_maps.append((S[:r], Vt[:r]))
46
+ col0 += r
47
+ Psi = np.concatenate(Psi_blocks, axis=1)
48
+ return Dictionary(Psi=Psi, groups=groups, H=H, expert_maps=expert_maps)
49
+
50
+
51
+ def pilot_linearity_check(world, cand, dictionary, cfg, rng=None):
52
+ """Trust-region pilot (Assumption 1): compare probe-superposition
53
+ prediction H beta with a small paired evaluation on random slices.
54
+ Returns (relative_residual, passed)."""
55
+ rng = rng or np.random.default_rng(cfg.seed + 7)
56
+ pred = dictionary.H @ cand.beta
57
+ # test where the superposition prediction claims signal exists: the
58
+ # top-|pred| slices; a linearity violation shows up exactly there.
59
+ k = min(world.n, 40)
60
+ idx = np.argsort(np.abs(pred))[::-1][:k]
61
+ obs = np.array([world.paired_scores(cand, int(j), cfg.pilot_items,
62
+ rng=rng).mean() for j in idx])
63
+ denom = max(np.linalg.norm(obs), np.linalg.norm(pred[idx])) + 1e-9
64
+ rel = np.linalg.norm(obs - pred[idx]) / denom
65
+ return rel, rel <= cfg.trust_region_residual
spectra_rsi/recovery.py ADDED
@@ -0,0 +1,88 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Weighted sparse-group recovery (Eq. recovery) via FISTA, plus
2
+ debiased refit and bootstrap support stability.
3
+
4
+ min_x 0.5 ||y - M x||^2 + lambda1 ||x||_1 + lambda_g sum_e w_e ||x_Ge||_2
5
+ """
6
+ import numpy as np
7
+
8
+
9
+ def _soft(x, t):
10
+ return np.sign(x) * np.maximum(np.abs(x) - t, 0.0)
11
+
12
+
13
+ def _group_soft(x, groups, thresholds):
14
+ out = x.copy()
15
+ for g, t in zip(groups, thresholds):
16
+ nrm = np.linalg.norm(x[g])
17
+ out[g] = 0.0 if nrm <= t else (1 - t / nrm) * x[g]
18
+ return out
19
+
20
+
21
+ def _lipschitz(M, iters=60, rng=None):
22
+ rng = rng or np.random.default_rng(0)
23
+ v = rng.normal(size=M.shape[1])
24
+ v /= np.linalg.norm(v)
25
+ for _ in range(iters):
26
+ v = M.T @ (M @ v)
27
+ nv = np.linalg.norm(v)
28
+ if nv < 1e-18:
29
+ return 1.0
30
+ v /= nv
31
+ return nv
32
+
33
+
34
+ def sparse_group_recover(y, M, groups, group_w, lambda_l1, lambda_group,
35
+ n_iters=600):
36
+ """FISTA with elementwise + group soft-thresholding prox."""
37
+ p = M.shape[1]
38
+ L = _lipschitz(M) + 1e-9
39
+ step = 1.0 / L
40
+ x = np.zeros(p)
41
+ z = x.copy()
42
+ t = 1.0
43
+ for _ in range(n_iters):
44
+ grad = M.T @ (M @ z - y)
45
+ u = z - step * grad
46
+ u = _soft(u, step * lambda_l1)
47
+ u = _group_soft(u, groups, [step * lambda_group * w for w in group_w])
48
+ t_next = 0.5 * (1 + np.sqrt(1 + 4 * t * t))
49
+ z = u + ((t - 1) / t_next) * (u - x)
50
+ x, t = u, t_next
51
+ return x
52
+
53
+
54
+ def debias(y, M, x_hat, tol=1e-8):
55
+ """Least-squares refit restricted to the selected support (Sec. 5.1)."""
56
+ support = np.nonzero(np.abs(x_hat) > tol)[0]
57
+ if support.size == 0:
58
+ return x_hat, support
59
+ Ms = M[:, support]
60
+ coef, *_ = np.linalg.lstsq(Ms, y, rcond=None)
61
+ out = np.zeros_like(x_hat)
62
+ out[support] = coef
63
+ return out, support
64
+
65
+
66
+ def bootstrap_support(y, M, groups, group_w, lambda_l1, lambda_group,
67
+ reps=60, n_iters=250, rng=None):
68
+ """Row-resampling bootstrap: per-group selection frequencies (Sec. 8.1)."""
69
+ rng = rng or np.random.default_rng(0)
70
+ m = len(y)
71
+ freq = np.zeros(len(groups))
72
+ for _ in range(reps):
73
+ idx = rng.integers(0, m, m)
74
+ xb = sparse_group_recover(y[idx], M[idx], groups, group_w,
75
+ lambda_l1, lambda_group, n_iters=n_iters)
76
+ for e, g in enumerate(groups):
77
+ if np.linalg.norm(xb[g]) > 1e-8:
78
+ freq[e] += 1
79
+ return freq / reps
80
+
81
+
82
+ def active_groups(x_hat, groups, rel_tol=0.05):
83
+ """Groups carrying more than rel_tol of the maximum group energy."""
84
+ energies = np.array([np.linalg.norm(x_hat[g]) for g in groups])
85
+ if energies.max() <= 0:
86
+ return [], energies
87
+ keep = np.nonzero(energies > rel_tol * energies.max())[0]
88
+ return list(keep), energies
spectra_rsi/sketch.py ADDED
@@ -0,0 +1,83 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Paired randomized sketches and measurement design (Secs. 3.1, 6).
2
+
3
+ Rows are sparse, signed, variance-normalized task weights. Measurements are
4
+ weighted sums of task-level paired mean differences (realizable composites),
5
+ with Neyman item allocation N_ij ~ |a_ij| sigma_j.
6
+ """
7
+ import numpy as np
8
+
9
+
10
+ def design_matrix(m, n, row_density, sigma, rng, group_slices=None,
11
+ focus_groups=None, explore_fraction=0.15):
12
+ """Sparse Rademacher design, variance-normalized (Sec. 6.1).
13
+
14
+ If `focus_groups` is given (stage II), row support concentrates on the
15
+ nominated groups' slices with an exploration floor elsewhere (Sec. 6.2).
16
+ """
17
+ A = np.zeros((m, n))
18
+ if focus_groups is not None and group_slices is not None:
19
+ focus_mask = np.zeros(n, dtype=bool)
20
+ for g in focus_groups:
21
+ focus_mask[group_slices[g]] = True
22
+ p_focus = row_density
23
+ p_off = row_density * explore_fraction
24
+ probs = np.where(focus_mask, p_focus, p_off)
25
+ else:
26
+ probs = np.full(n, row_density)
27
+
28
+ for i in range(m):
29
+ s = rng.random(n) < probs
30
+ if not s.any():
31
+ s[rng.integers(n)] = True
32
+ xi = rng.choice([-1.0, 1.0], size=n)
33
+ row = np.where(s, xi / np.maximum(sigma, 1e-9), 0.0)
34
+ norm = np.sqrt((row ** 2).sum())
35
+ A[i] = row / max(norm, 1e-12)
36
+ return A
37
+
38
+
39
+ def neyman_allocation(A_row, sigma, budget):
40
+ """N_ij ~ |a_ij| sigma_j over the row support, integer allocation >= 2."""
41
+ support = np.nonzero(A_row)[0]
42
+ w = np.abs(A_row[support]) * sigma[support]
43
+ w = w / w.sum()
44
+ N = np.maximum(2, np.floor(budget * w).astype(int))
45
+ return support, N
46
+
47
+
48
+ def paired_sketch(world, cand, A, sigma, items_per_row, rng):
49
+ """Collect y_i = sum_j a_ij * mean paired difference on slice j.
50
+
51
+ Returns (y, total_items, per_slice_means) where per_slice_means caches
52
+ reusable slice estimates for residual monitoring.
53
+ """
54
+ m = A.shape[0]
55
+ y = np.zeros(m)
56
+ total_items = 0
57
+ slice_sums = {}
58
+ slice_counts = {}
59
+ for i in range(m):
60
+ support, N = neyman_allocation(A[i], sigma, items_per_row)
61
+ acc = 0.0
62
+ for j, nj in zip(support, N):
63
+ diffs = world.paired_scores(cand, int(j), int(nj), rng=rng)
64
+ mean_j = diffs.mean()
65
+ acc += A[i, j] * mean_j
66
+ total_items += int(nj)
67
+ slice_sums[j] = slice_sums.get(j, 0.0) + diffs.sum()
68
+ slice_counts[j] = slice_counts.get(j, 0) + int(nj)
69
+ y[i] = acc
70
+ per_slice = {j: slice_sums[j] / slice_counts[j] for j in slice_sums}
71
+ return y, total_items, per_slice
72
+
73
+
74
+ def group_slice_map(dictionary, top_frac=0.25):
75
+ """Map each dictionary group to the capability slices it loads on most,
76
+ used to focus stage-II measurement rows."""
77
+ n = dictionary.Psi.shape[0]
78
+ out = []
79
+ k = max(1, int(n * top_frac / max(1, len(dictionary.groups))))
80
+ for g in dictionary.groups:
81
+ load = np.abs(dictionary.Psi[:, g]).sum(axis=1)
82
+ out.append(np.argsort(load)[::-1][:max(k, 8)])
83
+ return out
spectra_rsi/world.py ADDED
@@ -0,0 +1,171 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Synthetic ground-truth world.
2
+
3
+ Implements the paper's problem formulation (Sec. 3) with a known
4
+ capability-response operator so that reconstruction, attribution, and gating
5
+ can be validated against ground truth.
6
+
7
+ The world exposes exactly the interface a real evaluation harness would:
8
+ paired scoring of (base, candidate) on sampled items with common decoding
9
+ randomness. Everything downstream (probes, sketches, gate) only uses this
10
+ interface, so a real LLM harness can be dropped in behind it.
11
+ """
12
+ from dataclasses import dataclass, field
13
+ import numpy as np
14
+
15
+
16
+ @dataclass
17
+ class CandidateUpdate:
18
+ """A frozen candidate with optional off-dictionary capability residual."""
19
+ beta: np.ndarray # (d,) coefficients over stacked expert directions
20
+ label: str = "candidate"
21
+ # ground-truth bookkeeping (for benchmarking only; never read by the loop)
22
+ true_support: set = field(default_factory=set) # expert indices actually perturbed
23
+ # Optional capability-space component not represented by expert coefficients.
24
+ # Used only for synthetic out-of-dictionary benchmark interventions.
25
+ residual_delta: np.ndarray | None = None
26
+
27
+
28
+ class SyntheticWorld:
29
+ """Ground truth: Delta c = H* beta + nonlinearity + heteroscedastic noise.
30
+
31
+ - H* (n x d): true interventional response matrix, block-structured so
32
+ that each expert's directions load mainly on a band of slices (this is
33
+ the group-compressibility assumption made explicit and controllable).
34
+ - Paired scoring with common random numbers: base and candidate item
35
+ scores share a noise component with correlation `cnr_rho`, so paired
36
+ differences have variance reduced by (1 - rho) relative to independent
37
+ evaluation (Prop. 1 in the paper).
38
+ """
39
+
40
+ def __init__(self, n_slices, n_experts, rank_per_expert, seed=0,
41
+ band_overlap=0.15, offband_leak=0.05, nonlin=0.0,
42
+ noise_scale=0.25, cnr_rho=0.85):
43
+ self.n = n_slices
44
+ self.E = n_experts
45
+ self.R = rank_per_expert
46
+ self.d = n_experts * rank_per_expert
47
+ self.nonlin = nonlin
48
+ self.cnr_rho = cnr_rho
49
+ rng = np.random.default_rng(seed)
50
+ self._rng = rng
51
+
52
+ # block-structured true response matrix
53
+ band = n_slices // n_experts
54
+ H = np.zeros((self.n, self.d))
55
+ for e in range(n_experts):
56
+ lo, hi = e * band, (e + 1) * band
57
+ # widen band with overlap
58
+ pad = int(band * band_overlap)
59
+ lo2, hi2 = max(0, lo - pad), min(self.n, hi + pad)
60
+ for r in range(rank_per_expert):
61
+ col = e * rank_per_expert + r
62
+ v = np.zeros(self.n)
63
+ # in-band responses are positive: pushing an expert "up"
64
+ # improves the capabilities it owns (sign semantics), while
65
+ # off-band leakage stays signed and diffuse.
66
+ v[lo2:hi2] = np.abs(rng.normal(0, 1, hi2 - lo2))
67
+ v += offband_leak * rng.normal(0, 1, self.n)
68
+ H[:, col] = v / np.linalg.norm(v)
69
+ self.H_true = H
70
+
71
+ # per-slice score noise scale (heteroscedastic)
72
+ self.sigma = noise_scale * (0.5 + rng.random(self.n))
73
+
74
+ # current deployed parameter state in expert-coefficient coordinates
75
+ self.theta_offset = np.zeros(self.d)
76
+
77
+ # ------------------------------------------------------------------ #
78
+ # ground truth (benchmark use only)
79
+ # ------------------------------------------------------------------ #
80
+ def true_delta(self, cand: CandidateUpdate) -> np.ndarray:
81
+ lin = self.H_true @ cand.beta
82
+ if self.nonlin > 0:
83
+ # quadratic saturation: response bends for large updates
84
+ lin = lin - self.nonlin * np.sign(lin) * lin ** 2
85
+ if cand.residual_delta is not None:
86
+ residual = np.asarray(cand.residual_delta, dtype=float)
87
+ if residual.shape != (self.n,):
88
+ raise ValueError(
89
+ f"residual_delta must have shape ({self.n},), "
90
+ f"got {residual.shape}"
91
+ )
92
+ lin = lin + residual
93
+ return lin
94
+
95
+ # ------------------------------------------------------------------ #
96
+ # evaluation interface (what a real harness would implement)
97
+ # ------------------------------------------------------------------ #
98
+ def paired_scores(self, cand: CandidateUpdate, slice_idx: int, n_items: int,
99
+ rng=None) -> np.ndarray:
100
+ """Paired per-item score differences d_l for one capability slice,
101
+ evaluated with common random numbers (shared item + decoding noise)."""
102
+ rng = rng or self._rng
103
+ mu = self.true_delta(cand)[slice_idx]
104
+ s = self.sigma[slice_idx]
105
+ shared = rng.normal(0, s, n_items)
106
+ e_base = np.sqrt(1 - self.cnr_rho) * rng.normal(0, s, n_items)
107
+ e_cand = np.sqrt(1 - self.cnr_rho) * rng.normal(0, s, n_items)
108
+ base = shared + e_base
109
+ candv = mu + shared + e_cand
110
+ return candv - base # variance ~ 2 (1 - rho) s^2
111
+
112
+ def probe_delta(self, direction: np.ndarray, magnitude: float,
113
+ n_items_per_slice: int, rng=None) -> np.ndarray:
114
+ """Central-difference fingerprint (Eq. fingerprint): evaluate
115
+ +h and -h probes on ALL slices with finite item budget."""
116
+ rng = rng or self._rng
117
+ plus = CandidateUpdate(beta=magnitude * direction)
118
+ minus = CandidateUpdate(beta=-magnitude * direction)
119
+ mu_p = self.true_delta(plus)
120
+ mu_m = self.true_delta(minus)
121
+ noise = (self.sigma / np.sqrt(max(1, n_items_per_slice))
122
+ * np.sqrt(2 * (1 - self.cnr_rho)))
123
+ obs_p = mu_p + noise * rng.normal(0, 1, self.n)
124
+ obs_m = mu_m + noise * rng.normal(0, 1, self.n)
125
+ return (obs_p - obs_m) / (2 * magnitude)
126
+
127
+ # ------------------------------------------------------------------ #
128
+ # candidate factory helpers (benchmark interventions, Sec. 12.2)
129
+ # ------------------------------------------------------------------ #
130
+ def make_candidate(self, kind="single_gain", experts=None, scale=0.5,
131
+ rng=None) -> CandidateUpdate:
132
+ rng = rng or self._rng
133
+ beta = np.zeros(self.d)
134
+ R = self.R
135
+ if experts is None:
136
+ experts = [int(rng.integers(self.E))]
137
+ if kind == "single_gain":
138
+ for e in experts:
139
+ beta[e * R:(e + 1) * R] = scale * np.abs(rng.normal(0.8, 0.2, R))
140
+ elif kind == "single_regression":
141
+ for e in experts:
142
+ beta[e * R:(e + 1) * R] = -scale * np.abs(rng.normal(0.8, 0.2, R))
143
+ elif kind == "canceling_mixture":
144
+ e1, e2 = experts[0], experts[-1]
145
+ beta[e1 * R:(e1 + 1) * R] = scale
146
+ beta[e2 * R:(e2 + 1) * R] = -scale
147
+ experts = [e1, e2]
148
+ elif kind == "broad_noncompressible":
149
+ beta = scale * rng.normal(0, 0.3, self.d)
150
+ experts = list(range(self.E))
151
+ elif kind == "off_dictionary":
152
+ # Construct a capability-space change orthogonal to the true
153
+ # expert-response span. No expert coefficient vector can
154
+ # represent this residual.
155
+ raw = rng.normal(0, 1, self.n)
156
+ coef, *_ = np.linalg.lstsq(self.H_true, raw, rcond=None)
157
+ residual = raw - self.H_true @ coef
158
+ norm = np.linalg.norm(residual)
159
+ if norm <= 1e-12:
160
+ raise RuntimeError("Failed to construct off-dictionary residual")
161
+ residual = scale * residual / norm
162
+ experts = []
163
+ return CandidateUpdate(
164
+ beta=beta,
165
+ label=kind,
166
+ true_support=set(),
167
+ residual_delta=residual,
168
+ )
169
+ else:
170
+ raise ValueError(kind)
171
+ return CandidateUpdate(beta=beta, label=kind, true_support=set(experts))
tests/__init__.py ADDED
File without changes
tests/test_gate.py ADDED
@@ -0,0 +1,70 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """E-process gate: type-I control under adversarial optional stopping,
2
+ and power under genuine non-inferiority."""
3
+ import numpy as np
4
+ from spectra_rsi.gate import EProcess
5
+
6
+
7
+ def test_type1_error_under_optional_stopping():
8
+ """Adversary peeks after every observation and stops at the most
9
+ favorable moment. False rejection of a TRUE null (mean == -tau, i.e.
10
+ exactly materially regressed) must stay below alpha."""
11
+ alpha, tau = 0.05, 0.05
12
+ rng = np.random.default_rng(42)
13
+ n_sims, horizon = 400, 2000
14
+ false_rejects = 0
15
+ for _ in range(n_sims):
16
+ ep = EProcess(tau=tau, alpha=alpha, bound=1.0)
17
+ rejected = False
18
+ for _ in range(horizon // 50):
19
+ obs = np.clip(rng.normal(-tau, 0.3, 50), -1, 1) # H0 boundary
20
+ ep.update(obs)
21
+ if ep.rejects_null: # adversarial peeking
22
+ rejected = True
23
+ break
24
+ false_rejects += rejected
25
+ rate = false_rejects / n_sims
26
+ assert rate <= alpha + 2 * np.sqrt(alpha * (1 - alpha) / n_sims), rate
27
+
28
+
29
+ def test_power_under_true_improvement():
30
+ """A genuinely improved anchor should be certified with high probability."""
31
+ alpha, tau = 0.05, 0.02
32
+ rng = np.random.default_rng(7)
33
+ certified = 0
34
+ n_sims = 100
35
+ for _ in range(n_sims):
36
+ ep = EProcess(tau=tau, alpha=alpha, bound=1.0)
37
+ for _ in range(60):
38
+ obs = np.clip(rng.normal(+0.08, 0.25, 50), -1, 1)
39
+ ep.update(obs)
40
+ if ep.rejects_null:
41
+ break
42
+ certified += ep.rejects_null
43
+ assert certified / n_sims > 0.9
44
+
45
+
46
+ def test_anchor_gate_never_exceeds_global_budget():
47
+ """AnchorGate must never consume more than max_items globally."""
48
+ from spectra_rsi.gate import AnchorGate
49
+
50
+ class DummyWorld:
51
+ def paired_scores(self, cand, j, batch, rng=None):
52
+ # Keep anchors unresolved long enough to exhaust the budget.
53
+ return np.full(batch, -0.02)
54
+
55
+ gate = AnchorGate(
56
+ anchor_slices=[0, 1, 2, 3],
57
+ tau=0.02,
58
+ alpha_total=0.05,
59
+ )
60
+
61
+ _, _, used = gate.run(
62
+ DummyWorld(),
63
+ cand=None,
64
+ batch=100,
65
+ max_items=250,
66
+ rng=np.random.default_rng(123),
67
+ )
68
+
69
+ assert used <= 250
70
+ assert used == 200
tests/test_loop.py ADDED
@@ -0,0 +1,114 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """End-to-end loop: attribution, gating, and trust-region fallback."""
2
+ import numpy as np
3
+ import pytest
4
+ from spectra_rsi import SyntheticWorld, SpectraConfig, SpectraRSILoop
5
+ from spectra_rsi.metrics import support_f1, normalized_delta_error
6
+
7
+
8
+ @pytest.fixture(scope="module")
9
+ def setup(tmp_path_factory):
10
+ cfg = SpectraConfig(n_slices=200, n_experts=8, rank_per_expert=2,
11
+ m_coarse=40, m_focused=60, items_per_row=150,
12
+ bootstrap_reps=20, fista_iters=400,
13
+ audit_dir=str(tmp_path_factory.mktemp("audit")))
14
+ world = SyntheticWorld(cfg.n_slices, cfg.n_experts, cfg.rank_per_expert,
15
+ seed=11)
16
+ loop = SpectraRSILoop(world, cfg)
17
+ return world, cfg, loop
18
+
19
+
20
+ def test_single_expert_attribution_and_acceptance(setup):
21
+ world, cfg, loop = setup
22
+ cand = world.make_candidate("single_gain", experts=[3], scale=0.4,
23
+ rng=np.random.default_rng(2))
24
+ rep = loop.run_iteration(cand)
25
+ assert rep.pilot_passed
26
+ assert 3 in rep.recovered_experts
27
+ f1 = support_f1(rep.recovered_experts, cand.true_support)
28
+ assert f1 >= 0.5
29
+ err = normalized_delta_error(rep.delta_hat, world.true_delta(cand))
30
+ assert err < 0.8 # smoke-test budget; tightens with more items
31
+ assert rep.gate_decision == "accept"
32
+
33
+
34
+ def test_regression_is_not_accepted(setup):
35
+ """Construct a candidate that provably regresses a protected anchor:
36
+ beta chosen opposite to the anchor's true response row, so
37
+ Delta c(anchor) = -scale * ||H*(anchor,:)||^2 < -tau."""
38
+ from spectra_rsi.world import CandidateUpdate
39
+ world, cfg, loop = setup
40
+ anchor = int(loop.anchor_slices[0])
41
+ row = world.H_true[anchor] # ground truth (test only)
42
+ scale = 3.0 * cfg.tau_margin / max(np.dot(row, row), 1e-9)
43
+ beta = -scale * row
44
+ true_drop = world.true_delta(CandidateUpdate(beta=beta))[anchor]
45
+ assert true_drop < -cfg.tau_margin # sanity: really regressed
46
+ cand = CandidateUpdate(beta=beta, label="anchor_regression")
47
+ rep = loop.run_iteration(cand)
48
+ if rep.pilot_passed:
49
+ assert rep.gate_decision in ("rollback", "quarantine")
50
+ else:
51
+ assert rep.dense_fallback
52
+
53
+
54
+ def test_out_of_trust_region_falls_back(setup):
55
+ world, cfg, loop = setup
56
+ nl_world = SyntheticWorld(cfg.n_slices, cfg.n_experts,
57
+ cfg.rank_per_expert, seed=13, nonlin=0.9)
58
+ nl_loop = SpectraRSILoop(nl_world, cfg)
59
+ cand = nl_world.make_candidate("single_gain", experts=[1], scale=5.0,
60
+ rng=np.random.default_rng(4))
61
+ rep = nl_loop.run_iteration(cand)
62
+ assert (not rep.pilot_passed and rep.dense_fallback) or rep.pilot_passed
63
+
64
+
65
+ def test_audit_records_written(setup):
66
+ import glob, os
67
+ world, cfg, loop = setup
68
+ files = glob.glob(os.path.join(cfg.audit_dir, "iter_*.json"))
69
+ assert len(files) >= 1
70
+
71
+
72
+ def test_off_dictionary_candidate_triggers_dense_fallback():
73
+ """A capability change outside the expert-response span must fall back."""
74
+ cfg = SpectraConfig(
75
+ n_slices=256,
76
+ n_experts=8,
77
+ rank_per_expert=2,
78
+ pilot_items=120,
79
+ seed=7,
80
+ )
81
+
82
+ world = SyntheticWorld(
83
+ cfg.n_slices,
84
+ cfg.n_experts,
85
+ cfg.rank_per_expert,
86
+ seed=cfg.seed,
87
+ offband_leak=0.01,
88
+ )
89
+
90
+ rng = np.random.default_rng(99)
91
+ cand = world.make_candidate(
92
+ "off_dictionary",
93
+ scale=0.4,
94
+ rng=rng,
95
+ )
96
+
97
+ # Verify the synthetic residual is genuinely outside span(H_true).
98
+ residual = cand.residual_delta
99
+ coef, *_ = np.linalg.lstsq(world.H_true, residual, rcond=None)
100
+ projection = world.H_true @ coef
101
+ relative_projection = (
102
+ np.linalg.norm(projection) / np.linalg.norm(residual)
103
+ )
104
+
105
+ assert relative_projection < 1e-10
106
+ assert cand.true_support == set()
107
+
108
+ # The trust-region pilot should reject the sparse model and fall back.
109
+ rep = SpectraRSILoop(world, cfg).run_iteration(cand)
110
+
111
+ assert rep.pilot_residual > cfg.trust_region_residual
112
+ assert rep.dense_fallback is True
113
+ assert rep.items_sense == 0
114
+ assert rep.items_anchor == 0
tests/test_recovery.py ADDED
@@ -0,0 +1,20 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Sparse-group recovery: exact-support recovery on clean synthetic data."""
2
+ import numpy as np
3
+ from spectra_rsi.recovery import sparse_group_recover, debias, active_groups
4
+
5
+
6
+ def test_group_recovery_identifies_support():
7
+ rng = np.random.default_rng(0)
8
+ p, n_groups = 32, 8
9
+ groups = [np.arange(i * 4, (i + 1) * 4) for i in range(n_groups)]
10
+ x_true = np.zeros(p)
11
+ x_true[groups[2]] = rng.normal(0, 1, 4)
12
+ x_true[groups[5]] = rng.normal(0, 1, 4)
13
+ M = rng.normal(0, 1 / np.sqrt(60), (60, p))
14
+ y = M @ x_true + 0.01 * rng.normal(size=60)
15
+ gw = np.ones(n_groups)
16
+ x_hat = sparse_group_recover(y, M, groups, gw, 1e-3, 5e-3, n_iters=800)
17
+ x_deb, _ = debias(y, M, x_hat)
18
+ act, _ = active_groups(x_deb, groups, rel_tol=0.1)
19
+ assert set(act) == {2, 5}
20
+ assert np.linalg.norm(x_deb - x_true) / np.linalg.norm(x_true) < 0.1
tests/test_sketch.py ADDED
@@ -0,0 +1,39 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Paired sketch: unbiasedness and CRN variance reduction (Prop. 1)."""
2
+ import numpy as np
3
+ import pytest
4
+ from spectra_rsi import SyntheticWorld, SpectraConfig
5
+ from spectra_rsi.sketch import design_matrix, paired_sketch
6
+
7
+
8
+ @pytest.fixture
9
+ def world():
10
+ return SyntheticWorld(n_slices=100, n_experts=4, rank_per_expert=2, seed=3)
11
+
12
+
13
+ def test_paired_sketch_unbiased(world):
14
+ rng = np.random.default_rng(0)
15
+ cand = world.make_candidate("single_gain", experts=[1], scale=0.5, rng=rng)
16
+ A = design_matrix(20, world.n, 0.2, world.sigma, rng)
17
+ true = A @ world.true_delta(cand)
18
+ reps = 40
19
+ est = np.zeros((reps, 20))
20
+ for r in range(reps):
21
+ y, _, _ = paired_sketch(world, cand, A, world.sigma, 150, rng)
22
+ est[r] = y
23
+ bias = np.abs(est.mean(axis=0) - true)
24
+ se = est.std(axis=0) / np.sqrt(reps)
25
+ # bias within 4 standard errors on at least 95% of rows
26
+ assert np.mean(bias < 4 * se + 1e-3) > 0.9
27
+
28
+
29
+ def test_crn_variance_reduction():
30
+ """Common random numbers must shrink paired-difference variance."""
31
+ w_crn = SyntheticWorld(100, 4, 2, seed=5, cnr_rho=0.9)
32
+ w_ind = SyntheticWorld(100, 4, 2, seed=5, cnr_rho=0.0)
33
+ rng = np.random.default_rng(1)
34
+ cand_c = w_crn.make_candidate("single_gain", experts=[0], rng=rng)
35
+ cand_i = w_ind.make_candidate("single_gain", experts=[0],
36
+ rng=np.random.default_rng(1))
37
+ var_crn = np.var(w_crn.paired_scores(cand_c, 3, 5000))
38
+ var_ind = np.var(w_ind.paired_scores(cand_i, 3, 5000))
39
+ assert var_crn < 0.5 * var_ind