Upload folder using huggingface_hub
Browse files- COLLABORATION.md +34 -0
- README.md +99 -0
- RESULTS_SCHEMA.md +39 -0
- benchmarks/aggregate_results.py +87 -0
- benchmarks/run_scaling.py +91 -0
- configs/scaling_matrix.md +5 -0
- pyproject.toml +22 -0
- results/README.md +25 -0
- results/dgx_spark_1024_canceling.json +64 -0
- results/dgx_spark_1024_noncompressible.json +74 -0
- results/dgx_spark_1024_noncompressible_seed100.json +75 -0
- results/dgx_spark_1024_off_dictionary.json +38 -0
- results/dgx_spark_1024_reg5x.json +61 -0
- results/dgx_spark_1024_reg5x_seed100.json +61 -0
- results/dgx_spark_1024_regression_budgetfix.json +62 -0
- results/validated_leaderboard.csv +8 -0
- results/validated_results.txt +7 -0
- scripts/run_multiseed.sh +8 -0
- scripts/run_scaling_sweep.sh +8 -0
- scripts/run_smoke.sh +3 -0
- spectra_rsi/__init__.py +31 -0
- spectra_rsi/config.py +51 -0
- spectra_rsi/drift.py +62 -0
- spectra_rsi/experts.py +47 -0
- spectra_rsi/gate.py +109 -0
- spectra_rsi/loop.py +212 -0
- spectra_rsi/metrics.py +30 -0
- spectra_rsi/probes.py +65 -0
- spectra_rsi/recovery.py +88 -0
- spectra_rsi/sketch.py +83 -0
- spectra_rsi/world.py +171 -0
- tests/__init__.py +0 -0
- tests/test_gate.py +70 -0
- tests/test_loop.py +114 -0
- tests/test_recovery.py +20 -0
- tests/test_sketch.py +39 -0
COLLABORATION.md
ADDED
|
@@ -0,0 +1,34 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# High-compute collaboration call
|
| 2 |
+
|
| 3 |
+
SPECTRA-RSI needs independent scaling evidence. The manuscript explicitly treats its numerical gains as hypotheses to test, with dense ground-truth deltas, causal-attribution metrics, optional-stopping stress tests, drift tests, and 20–100-cycle long-horizon studies.
|
| 4 |
+
|
| 5 |
+
## Highest-value contributions
|
| 6 |
+
|
| 7 |
+
1. Run larger `n_slices`, `n_experts`, rank, bootstrap, and multi-seed sweeps and upload the resulting JSON files.
|
| 8 |
+
2. Port the hot numerical paths to PyTorch/JAX/CuPy and benchmark CPU vs single-GPU vs multi-GPU without changing estimands.
|
| 9 |
+
3. Add real open-weight LoRA/MoE-LoRA adapters implementing the `paired_scores` and `probe_delta` interface.
|
| 10 |
+
4. Test two model families and two scales, including single-expert gains/regressions, canceling mixtures, interacting experts, non-compressible changes, and out-of-trust-region updates.
|
| 11 |
+
5. Run optional-stopping, dictionary-drift, Fisher/second-order dictionary regularization, and 20–100 update-cycle experiments.
|
| 12 |
+
|
| 13 |
+
## Hardware sought
|
| 14 |
+
|
| 15 |
+
H100/H200, B100/B200, GB200/GB300, MI300X, 2/4/8/16+ GPU servers, and multi-node clusters. Other accelerators are welcome when the environment is fully reported.
|
| 16 |
+
|
| 17 |
+
## Submission format (no Git required)
|
| 18 |
+
|
| 19 |
+
Download this Hugging Face repository, run the benchmark, then upload result JSON files back to a Hugging Face discussion/community contribution or share them with the maintainer. Include hardware, software versions, command line, wall time, and any code patch as a ZIP if you changed the backend. Do not submit only a screenshot.
|
| 20 |
+
|
| 21 |
+
## Current reference point and open scaling question
|
| 22 |
+
|
| 23 |
+
The validated reference point is **1 × NVIDIA GB10 (DGX Spark)** at 1,024 capability slices. Sparse single-gain and canceling-mixture cases recover their true expert support exactly in the validated runs, while the genuine `off_dictionary` case triggers immediate dense fallback.
|
| 24 |
+
|
| 25 |
+
The most important current stress-test result is deliberately unresolved: for `broad_noncompressible` at candidate seed 100, the gate returned **accept** while only 6 of 8 true experts were recovered. Contributors should not tune around or discard this case. We specifically want to learn whether increased sensing budgets, bootstrap repetitions, expert count/rank, alternative structured regularization, Fisher/second-order information, or real-model experiments eliminate this failure mode while preserving evaluation efficiency.
|
| 26 |
+
|
| 27 |
+
High-compute contributors are especially encouraged to report both successful and failed runs. Negative results, false accepts, unstable support recovery, poor scaling, memory limits, and dense-fallback behavior are scientifically useful.
|
| 28 |
+
|
| 29 |
+
## Reproducibility metadata
|
| 30 |
+
|
| 31 |
+
The benchmark records accelerator metadata through PyTorch when available and falls back to `nvidia-smi` on NVIDIA systems. Please also report GPU/accelerator model and count, CPU, RAM, driver/runtime, operating system, and whether the run used bare metal, a container, or a scheduler job.
|
| 32 |
+
|
| 33 |
+
Use `results/validated_results.txt` and `results/validated_leaderboard.csv` as the canonical DGX Spark reference set. Other bundled JSON files may be exploratory or development runs unless listed in the validated manifest.
|
| 34 |
+
|
README.md
ADDED
|
@@ -0,0 +1,99 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: mit
|
| 3 |
+
tags:
|
| 4 |
+
- benchmark
|
| 5 |
+
- evaluation
|
| 6 |
+
- recursive-self-improvement
|
| 7 |
+
- compressed-sensing
|
| 8 |
+
- mixture-of-experts
|
| 9 |
+
- lora
|
| 10 |
+
- scaling
|
| 11 |
+
---
|
| 12 |
+
# SPECTRA-RSI Hugging Face Scaling Benchmark
|
| 13 |
+
|
| 14 |
+
**Counterfactual Spectral Sketching and Anytime-Valid Gating for Modular Recursive Self-Improvement** — benchmark release for independent replication and high-compute scaling.
|
| 15 |
+
|
| 16 |
+
This package converts the reference implementation into a **Hugging Face-downloadable benchmark**. It does **not require Git**. Download the repository files from Hugging Face (or use `hf download`), install locally, run standardized experiments, and return JSON results.
|
| 17 |
+
|
| 18 |
+
The benchmark follows the manuscript's falsifiable protocol: recover base-to-candidate capability deltas, identify perturbed expert groups and effect signs, stress-test anytime-valid acceptance under repeated peeking, measure basis-drift detection, account for total evaluation economics, and ultimately test modular rollback over repeated update cycles. No empirical scaling gain is assumed in advance.
|
| 19 |
+
|
| 20 |
+
## Quick start — no Git
|
| 21 |
+
|
| 22 |
+
```bash
|
| 23 |
+
python -m venv .venv
|
| 24 |
+
source .venv/bin/activate
|
| 25 |
+
pip install -e ".[dev]"
|
| 26 |
+
python -m pytest -q
|
| 27 |
+
./scripts/run_smoke.sh
|
| 28 |
+
```
|
| 29 |
+
|
| 30 |
+
Run the standard scaling sweep:
|
| 31 |
+
|
| 32 |
+
```bash
|
| 33 |
+
./scripts/run_scaling_sweep.sh
|
| 34 |
+
```
|
| 35 |
+
|
| 36 |
+
Run a five-seed replication at a larger point:
|
| 37 |
+
|
| 38 |
+
```bash
|
| 39 |
+
N_SLICES=2048 N_EXPERTS=64 ./scripts/run_multiseed.sh
|
| 40 |
+
```
|
| 41 |
+
|
| 42 |
+
Or define a custom high-compute point:
|
| 43 |
+
|
| 44 |
+
```bash
|
| 45 |
+
python benchmarks/run_scaling.py \
|
| 46 |
+
--n-slices 4096 --n-experts 128 --rank 8 \
|
| 47 |
+
--m-coarse 800 --m-focused 1200 \
|
| 48 |
+
--items-per-row 512 --bootstrap-reps 100 \
|
| 49 |
+
--seed 1 --candidate-seed 101 --candidate single_gain \
|
| 50 |
+
--output results/my_system_n4096.json
|
| 51 |
+
```
|
| 52 |
+
|
| 53 |
+
## What to scale
|
| 54 |
+
|
| 55 |
+
The primary axes are capability-space dimension, expert count, expert rank, measurement budget, bootstrap count, seed count, intervention type, and eventually repeated RSI cycles. Report accuracy **and** cost: normalized delta error, expert-support F1, weighted regression recall, false acceptance, evaluated items/tokens, solver/wall time, and accelerator metadata.
|
| 56 |
+
|
| 57 |
+
The manuscript requires dense ground-truth deltas for the research benchmark and explicitly asks for at least two open-weight model families at two scales, LoRA/MoE-LoRA conditions, known expert interventions, optional-stopping stress tests, drift experiments, ablations, and 20–100-cycle long-horizon control experiments. The current package provides the reproducible synthetic core and a stable result format; real-model and accelerator backends are the most valuable next contributions.
|
| 58 |
+
|
| 59 |
+
## Validated DGX Spark baseline
|
| 60 |
+
|
| 61 |
+
The release candidate has been validated on **1 × NVIDIA GB10 (DGX Spark)** with NVIDIA driver **580.159.03**. The final test suite passes, and the validated benchmark manifest contains **7 runs** at 1,024 capability slices using the finalized runner.
|
| 62 |
+
|
| 63 |
+
| Candidate | Seed | Support F1 | Delta error | Regression recall | Gate | Dense fallback |
|
| 64 |
+
|---|---:|---:|---:|---:|---|---|
|
| 65 |
+
| `single_gain` | 99 | 1.000 | 0.3318 | 1.000 | accept | no |
|
| 66 |
+
| `single_gain` | 100 | 1.000 | 0.3411 | 1.000 | accept | no |
|
| 67 |
+
| `single_regression` | 99 | 0.667 | 0.4044 | 0.984 | quarantine | no |
|
| 68 |
+
| `canceling_mixture` | 99 | 1.000 | 0.3344 | 1.000 | quarantine | no |
|
| 69 |
+
| `broad_noncompressible` | 99 | 0.769 | 0.5666 | 0.923 | quarantine | no |
|
| 70 |
+
| `broad_noncompressible` | 100 | 0.857 | 0.4642 | 1.000 | **accept** | no |
|
| 71 |
+
| `off_dictionary` | 99 | — | — | — | — | **yes** |
|
| 72 |
+
|
| 73 |
+
The `off_dictionary` stress case produced a pilot residual of approximately **1.0** and immediately triggered dense fallback before compressed sensing or anchor evaluation, demonstrating the intended out-of-dictionary safeguard.
|
| 74 |
+
|
| 75 |
+
The broad in-dictionary stress test also exposes an important unresolved limitation. At candidate seed 100, the gate returned **accept** even though only 6 of 8 true experts were recovered. This result is intentionally retained rather than tuned away. A central high-compute research question is whether larger sensing budgets, more bootstrap repetitions, alternative structured regularization, Fisher/second-order information, or larger-scale real-model experiments improve this behavior without increasing false rejection.
|
| 76 |
+
|
| 77 |
+
Canonical release results are listed in `results/validated_results.txt` and summarized in `results/validated_leaderboard.csv`. Other JSON files in `results/` should be treated as exploratory/development runs unless they appear in the validated manifest.
|
| 78 |
+
|
| 79 |
+
## High-compute collaborators wanted
|
| 80 |
+
|
| 81 |
+
We are actively seeking contributors with **H100/H200, B100/B200, GB200/GB300, MI300X, and 2/4/8/16+ GPU or multi-node systems**. The immediate goal is to determine where SPECTRA-RSI's structured recovery remains accurate, where it becomes compute-bound, and where its assumptions fail as dimension, expert count, rank, seeds, and statistical resampling increase.
|
| 82 |
+
|
| 83 |
+
See **`COLLABORATION.md`** for the contribution menu and **`RESULTS_SCHEMA.md`** for the portable JSON contract. Git is optional: contributors can download the benchmark and return result files or a ZIP patch.
|
| 84 |
+
|
| 85 |
+
## Repository layout
|
| 86 |
+
|
| 87 |
+
- `spectra_rsi/` — reference algorithm implementation
|
| 88 |
+
- `benchmarks/run_scaling.py` — parameterized benchmark runner
|
| 89 |
+
- `benchmarks/aggregate_results.py` — portable result aggregation
|
| 90 |
+
- `scripts/run_smoke.sh` — fast validation
|
| 91 |
+
- `scripts/run_scaling_sweep.sh` — standard dimension sweep
|
| 92 |
+
- `scripts/run_multiseed.sh` — replication sweep
|
| 93 |
+
- `configs/scaling_matrix.md` — requested scaling axes
|
| 94 |
+
- `results/` — machine-readable contributor results
|
| 95 |
+
- `COLLABORATION.md` — high-compute call
|
| 96 |
+
|
| 97 |
+
## Scientific scope
|
| 98 |
+
|
| 99 |
+
This is a research benchmark, not an autonomous deployment system. Compressed discovery does not replace dense catastrophic-risk evaluation or human authorization. Protected anchors and discovery measurements must remain separated, and broad/non-compressible or nonlinear updates should trigger denser evaluation rather than forced sparse attribution.
|
RESULTS_SCHEMA.md
ADDED
|
@@ -0,0 +1,39 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# SPECTRA-RSI Result Contract
|
| 2 |
+
|
| 3 |
+
Every benchmark invocation writes one self-contained JSON document. Raw benchmark JSON files are the primary experimental records and should not be hand-edited.
|
| 4 |
+
|
| 5 |
+
## Core result fields
|
| 6 |
+
|
| 7 |
+
Each result records the benchmark identifier, candidate type, complete configuration, system metadata, wall time, dense-fallback status, pilot residual, and probe/sensing/anchor item counts.
|
| 8 |
+
|
| 9 |
+
When structured recovery completes, it also records the gate decision, support F1, normalized delta error, regression recall, true expert support, nominated experts, recovered experts, and bootstrap selection frequencies.
|
| 10 |
+
|
| 11 |
+
A dense fallback can legitimately omit recovery and gate fields because structured sensing and protected-anchor evaluation were not executed.
|
| 12 |
+
|
| 13 |
+
## Reproducibility configuration
|
| 14 |
+
|
| 15 |
+
The `config` object records `n_slices`, `n_experts`, `rank`, `m_coarse`, `m_focused`, `items_per_row`, `bootstrap_reps`, `lambda_l1`, `lambda_group`, `seed`, `candidate_seed`, `candidate`, `scale`, and `output`.
|
| 16 |
+
|
| 17 |
+
Do not compare benchmark results without checking these settings.
|
| 18 |
+
|
| 19 |
+
## System metadata
|
| 20 |
+
|
| 21 |
+
The `system` object may record hostname, platform, CUDA visibility, PyTorch version/status, GPU count and names, NVIDIA driver, and the mechanism used for GPU detection.
|
| 22 |
+
|
| 23 |
+
PyTorch is not required by the current NumPy benchmark. When PyTorch is unavailable on an NVIDIA system, the runner attempts GPU detection with `nvidia-smi`. A `torch_probe_error` such as `No module named torch` therefore does not by itself indicate benchmark failure.
|
| 24 |
+
|
| 25 |
+
## Canonical results
|
| 26 |
+
|
| 27 |
+
`results/validated_results.txt` defines the canonical validated release set.
|
| 28 |
+
|
| 29 |
+
Generate the validated leaderboard with `python benchmarks/aggregate_results.py results --manifest results/validated_results.txt --out results/validated_leaderboard.csv`.
|
| 30 |
+
|
| 31 |
+
Other JSON files under `results/` may be exploratory, diagnostic, pre-fix, or parameter-tuning runs. Preserve them for provenance, but do not present them as canonical release results unless deliberately added to the validated manifest.
|
| 32 |
+
|
| 33 |
+
## Contributor submissions
|
| 34 |
+
|
| 35 |
+
Preserve the raw JSON generated by the benchmark. Also report execution details that cannot be detected automatically, especially accelerator model/count, CPU, RAM, accelerator memory when known, driver/runtime versions, execution environment, command used, and any backend modifications.
|
| 36 |
+
|
| 37 |
+
Do not submit only screenshots or manually transcribed metrics.
|
| 38 |
+
|
| 39 |
+
Failed runs, dense fallbacks, false accepts, unstable recovery, memory limits, and performance regressions are valid benchmark evidence and should be retained rather than discarded.
|
benchmarks/aggregate_results.py
ADDED
|
@@ -0,0 +1,87 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
#!/usr/bin/env python3
|
| 2 |
+
import argparse
|
| 3 |
+
import csv
|
| 4 |
+
import json
|
| 5 |
+
from pathlib import Path
|
| 6 |
+
|
| 7 |
+
p = argparse.ArgumentParser()
|
| 8 |
+
p.add_argument("results_dir", nargs="?", default="results")
|
| 9 |
+
p.add_argument("--out", default="results/leaderboard.csv")
|
| 10 |
+
p.add_argument(
|
| 11 |
+
"--manifest",
|
| 12 |
+
default=None,
|
| 13 |
+
help="Optional text file listing result JSON filenames to include.",
|
| 14 |
+
)
|
| 15 |
+
a = p.parse_args()
|
| 16 |
+
|
| 17 |
+
rows = []
|
| 18 |
+
results_dir = Path(a.results_dir)
|
| 19 |
+
|
| 20 |
+
if a.manifest:
|
| 21 |
+
manifest = Path(a.manifest)
|
| 22 |
+
names = [
|
| 23 |
+
line.strip()
|
| 24 |
+
for line in manifest.read_text().splitlines()
|
| 25 |
+
if line.strip() and not line.lstrip().startswith("#")
|
| 26 |
+
]
|
| 27 |
+
result_files = [results_dir / name for name in names]
|
| 28 |
+
else:
|
| 29 |
+
result_files = sorted(results_dir.glob("*.json"))
|
| 30 |
+
|
| 31 |
+
for f in result_files:
|
| 32 |
+
try:
|
| 33 |
+
x = json.loads(f.read_text())
|
| 34 |
+
c = x.get("config", {})
|
| 35 |
+
s = x.get("system", {})
|
| 36 |
+
|
| 37 |
+
rows.append({
|
| 38 |
+
"file": f.name,
|
| 39 |
+
"candidate": x.get("candidate"),
|
| 40 |
+
"seed": c.get("seed"),
|
| 41 |
+
"candidate_seed": c.get("candidate_seed"),
|
| 42 |
+
"n_slices": c.get("n_slices"),
|
| 43 |
+
"n_experts": c.get("n_experts"),
|
| 44 |
+
"rank": c.get("rank"),
|
| 45 |
+
"m_coarse": c.get("m_coarse"),
|
| 46 |
+
"m_focused": c.get("m_focused"),
|
| 47 |
+
"items_per_row": c.get("items_per_row"),
|
| 48 |
+
"bootstrap_reps": c.get("bootstrap_reps"),
|
| 49 |
+
"lambda_l1": c.get("lambda_l1"),
|
| 50 |
+
"lambda_group": c.get("lambda_group"),
|
| 51 |
+
"gpu_count": s.get("gpu_count"),
|
| 52 |
+
"gpus": "; ".join(s.get("gpus", []) or []),
|
| 53 |
+
"wall_seconds": x.get("wall_seconds"),
|
| 54 |
+
"pilot_residual": x.get("pilot_residual"),
|
| 55 |
+
"dense_fallback": x.get("dense_fallback"),
|
| 56 |
+
"gate_decision": x.get("gate_decision"),
|
| 57 |
+
"support_f1": x.get("support_f1"),
|
| 58 |
+
"normalized_delta_error": x.get("normalized_delta_error"),
|
| 59 |
+
"regression_recall": x.get("regression_recall"),
|
| 60 |
+
"items_probe": x.get("items_probe"),
|
| 61 |
+
"items_sense": x.get("items_sense"),
|
| 62 |
+
"items_anchor": x.get("items_anchor"),
|
| 63 |
+
"true_experts": ";".join(
|
| 64 |
+
map(str, x.get("true_experts", []) or [])
|
| 65 |
+
),
|
| 66 |
+
"nominated_experts": ";".join(
|
| 67 |
+
map(str, x.get("nominated_experts", []) or [])
|
| 68 |
+
),
|
| 69 |
+
"recovered_experts": ";".join(
|
| 70 |
+
map(str, x.get("recovered_experts", []) or [])
|
| 71 |
+
),
|
| 72 |
+
"bootstrap_frequencies": ";".join(
|
| 73 |
+
map(str, x.get("bootstrap_frequencies", []) or [])
|
| 74 |
+
),
|
| 75 |
+
})
|
| 76 |
+
except Exception as e:
|
| 77 |
+
print(f"warning: skipped {f}: {e}")
|
| 78 |
+
|
| 79 |
+
Path(a.out).parent.mkdir(parents=True, exist_ok=True)
|
| 80 |
+
|
| 81 |
+
if rows:
|
| 82 |
+
with open(a.out, "w", newline="") as h:
|
| 83 |
+
w = csv.DictWriter(h, fieldnames=rows[0].keys())
|
| 84 |
+
w.writeheader()
|
| 85 |
+
w.writerows(rows)
|
| 86 |
+
|
| 87 |
+
print(f"wrote {len(rows)} rows to {a.out}")
|
benchmarks/run_scaling.py
ADDED
|
@@ -0,0 +1,91 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
#!/usr/bin/env python3
|
| 2 |
+
import argparse, json, os, platform, socket, subprocess, time
|
| 3 |
+
from pathlib import Path
|
| 4 |
+
import numpy as np
|
| 5 |
+
import sys
|
| 6 |
+
sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
|
| 7 |
+
from spectra_rsi import SyntheticWorld, SpectraConfig, SpectraRSILoop
|
| 8 |
+
from spectra_rsi.metrics import support_f1, normalized_delta_error, regression_recall
|
| 9 |
+
|
| 10 |
+
def gpu_info():
|
| 11 |
+
info = {
|
| 12 |
+
"cuda_visible_devices": os.getenv("CUDA_VISIBLE_DEVICES"),
|
| 13 |
+
"hostname": socket.gethostname(),
|
| 14 |
+
"platform": platform.platform(),
|
| 15 |
+
}
|
| 16 |
+
|
| 17 |
+
try:
|
| 18 |
+
import torch
|
| 19 |
+
info.update(
|
| 20 |
+
torch_version=torch.__version__,
|
| 21 |
+
cuda_available=torch.cuda.is_available(),
|
| 22 |
+
gpu_count=torch.cuda.device_count(),
|
| 23 |
+
)
|
| 24 |
+
if torch.cuda.is_available():
|
| 25 |
+
info["gpus"] = [
|
| 26 |
+
torch.cuda.get_device_name(i)
|
| 27 |
+
for i in range(torch.cuda.device_count())
|
| 28 |
+
]
|
| 29 |
+
except Exception as e:
|
| 30 |
+
info["torch_probe_error"] = str(e)
|
| 31 |
+
|
| 32 |
+
# Lightweight fallback for NVIDIA systems where PyTorch is not installed.
|
| 33 |
+
if not info.get("gpus"):
|
| 34 |
+
try:
|
| 35 |
+
out = subprocess.check_output(
|
| 36 |
+
[
|
| 37 |
+
"nvidia-smi",
|
| 38 |
+
"--query-gpu=name,driver_version",
|
| 39 |
+
"--format=csv,noheader",
|
| 40 |
+
],
|
| 41 |
+
text=True,
|
| 42 |
+
stderr=subprocess.DEVNULL,
|
| 43 |
+
timeout=5,
|
| 44 |
+
)
|
| 45 |
+
rows = [line.strip() for line in out.splitlines() if line.strip()]
|
| 46 |
+
if rows:
|
| 47 |
+
names = []
|
| 48 |
+
drivers = []
|
| 49 |
+
for row in rows:
|
| 50 |
+
parts = [x.strip() for x in row.split(",", 1)]
|
| 51 |
+
names.append(parts[0])
|
| 52 |
+
if len(parts) > 1:
|
| 53 |
+
drivers.append(parts[1])
|
| 54 |
+
|
| 55 |
+
info["gpu_count"] = len(names)
|
| 56 |
+
info["gpus"] = names
|
| 57 |
+
info["nvidia_driver"] = drivers[0] if drivers else None
|
| 58 |
+
info["gpu_probe"] = "nvidia-smi"
|
| 59 |
+
except Exception as e:
|
| 60 |
+
info["nvidia_smi_probe_error"] = str(e)
|
| 61 |
+
|
| 62 |
+
return info
|
| 63 |
+
|
| 64 |
+
def main():
|
| 65 |
+
p=argparse.ArgumentParser(description="SPECTRA-RSI reproducible scaling benchmark")
|
| 66 |
+
p.add_argument("--n-slices",type=int,default=400); p.add_argument("--n-experts",type=int,default=16)
|
| 67 |
+
p.add_argument("--rank",type=int,default=2); p.add_argument("--m-coarse",type=int,default=80); p.add_argument("--m-focused",type=int,default=120)
|
| 68 |
+
p.add_argument("--items-per-row",type=int,default=600); p.add_argument("--bootstrap-reps",type=int,default=30)
|
| 69 |
+
p.add_argument("--lambda-l1",type=float,default=2e-3); p.add_argument("--lambda-group",type=float,default=8e-3)
|
| 70 |
+
p.add_argument("--seed",type=int,default=7); p.add_argument("--candidate-seed",type=int,default=99)
|
| 71 |
+
p.add_argument("--candidate",choices=["single_gain","single_regression","canceling_mixture","broad_noncompressible","off_dictionary"],default="single_gain")
|
| 72 |
+
p.add_argument("--scale",type=float,default=0.4); p.add_argument("--output",required=True)
|
| 73 |
+
a=p.parse_args(); Path(a.output).parent.mkdir(parents=True,exist_ok=True)
|
| 74 |
+
cfg=SpectraConfig(n_slices=a.n_slices,n_experts=a.n_experts,rank_per_expert=a.rank,m_coarse=a.m_coarse,m_focused=a.m_focused,items_per_row=a.items_per_row,bootstrap_reps=a.bootstrap_reps,
|
| 75 |
+
lambda_l1=a.lambda_l1,lambda_group=a.lambda_group,
|
| 76 |
+
seed=a.seed,audit_dir=str(Path(a.output).parent/"audit_logs"))
|
| 77 |
+
world=SyntheticWorld(cfg.n_slices,cfg.n_experts,cfg.rank_per_expert,seed=cfg.seed,offband_leak=0.01)
|
| 78 |
+
rng=np.random.default_rng(a.candidate_seed)
|
| 79 |
+
ex={"single_gain":[max(0,a.n_experts//3)],"single_regression":[max(0,2*a.n_experts//3)],"canceling_mixture":[max(0,a.n_experts//4),max(0,3*a.n_experts//4)]}.get(a.candidate)
|
| 80 |
+
kw={"scale":a.scale,"rng":rng};
|
| 81 |
+
if ex is not None: kw["experts"]=ex
|
| 82 |
+
cand=world.make_candidate(a.candidate,**kw)
|
| 83 |
+
t=time.perf_counter(); rep=SpectraRSILoop(world,cfg).run_iteration(cand); wall=time.perf_counter()-t
|
| 84 |
+
row={"benchmark":"spectra-rsi-scaling-v1","candidate":a.candidate,"config":vars(a),"system":gpu_info(),"wall_seconds":wall,"dense_fallback":bool(rep.dense_fallback),"pilot_residual":float(rep.pilot_residual),"items_probe":int(rep.items_probe),"items_sense":int(rep.items_sense),"items_anchor":int(rep.items_anchor)}
|
| 85 |
+
if not rep.dense_fallback:
|
| 86 |
+
truth=world.true_delta(cand); row.update(gate_decision=rep.gate_decision,support_f1=float(support_f1(rep.recovered_experts,cand.true_support)),normalized_delta_error=float(normalized_delta_error(rep.delta_hat,truth)),regression_recall=float(regression_recall(rep.delta_hat,truth,cfg.tau_margin)),true_experts=sorted(map(int,cand.true_support)),
|
| 87 |
+
nominated_experts=sorted(map(int,rep.nominated_experts)),
|
| 88 |
+
recovered_experts=sorted(map(int,rep.recovered_experts)),
|
| 89 |
+
bootstrap_frequencies=[float(x) for x in rep.bootstrap_frequencies])
|
| 90 |
+
Path(a.output).write_text(json.dumps(row,indent=2)+"\n"); print(json.dumps(row,indent=2))
|
| 91 |
+
if __name__=="__main__": main()
|
configs/scaling_matrix.md
ADDED
|
@@ -0,0 +1,5 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Scaling matrix
|
| 2 |
+
|
| 3 |
+
Start with the smoke test, then contribute the largest stable points your hardware permits. Priority axes are: capability slices `n` (256 → 16k+), experts `E` (8 → 256+), rank (2 → 32+), bootstrap repetitions (10 → 500+), candidate seeds (5 → 100+), and long-horizon repeated iterations (future extension).
|
| 4 |
+
|
| 5 |
+
High-compute contributors are especially requested for H100/H200, B100/B200, GB200/GB300, MI300X-class systems and multi-GPU/multi-node hosts. The current reference implementation is NumPy-first; GPU owners can contribute **measurement results and/or GPU/distributed backend ports** while preserving the JSON result contract.
|
pyproject.toml
ADDED
|
@@ -0,0 +1,22 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
[build-system]
|
| 2 |
+
requires = ["setuptools>=61"]
|
| 3 |
+
build-backend = "setuptools.build_meta"
|
| 4 |
+
|
| 5 |
+
[project]
|
| 6 |
+
name = "spectra-rsi"
|
| 7 |
+
version = "0.1.0"
|
| 8 |
+
description = "Counterfactual spectral sketching and anytime-valid gating for modular recursive self-improvement (reference implementation)"
|
| 9 |
+
readme = "README.md"
|
| 10 |
+
requires-python = ">=3.9"
|
| 11 |
+
license = { text = "MIT" }
|
| 12 |
+
authors = [{ name = "Andrew M." }]
|
| 13 |
+
dependencies = ["numpy>=1.24"]
|
| 14 |
+
|
| 15 |
+
[project.optional-dependencies]
|
| 16 |
+
dev = ["pytest>=7.0"]
|
| 17 |
+
|
| 18 |
+
[tool.setuptools.packages.find]
|
| 19 |
+
include = ["spectra_rsi*"]
|
| 20 |
+
|
| 21 |
+
[tool.pytest.ini_options]
|
| 22 |
+
testpaths = ["tests"]
|
results/README.md
ADDED
|
@@ -0,0 +1,25 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Results
|
| 2 |
+
|
| 3 |
+
This directory contains SPECTRA-RSI benchmark outputs.
|
| 4 |
+
|
| 5 |
+
## Canonical validated results
|
| 6 |
+
|
| 7 |
+
`validated_results.txt` is the release manifest listing the raw JSON files that constitute the canonical validated baseline.
|
| 8 |
+
|
| 9 |
+
`validated_leaderboard.csv` is the corresponding aggregated table.
|
| 10 |
+
|
| 11 |
+
The current validated baseline was produced on 1 x NVIDIA GB10 (DGX Spark).
|
| 12 |
+
|
| 13 |
+
## Exploratory results
|
| 14 |
+
|
| 15 |
+
Other JSON files may represent exploratory parameter sweeps, diagnostics, tuning experiments, or pre-fix runs. They are retained locally for provenance but should not be interpreted as canonical benchmark results unless listed in `validated_results.txt`.
|
| 16 |
+
|
| 17 |
+
The `audit_logs/` directory contains algorithm-level audit records from multiple development and experimental executions. These logs are not currently linked unambiguously to individual canonical result files and are therefore not part of the validated Hugging Face release set.
|
| 18 |
+
|
| 19 |
+
## Regenerating the validated leaderboard
|
| 20 |
+
|
| 21 |
+
Run:
|
| 22 |
+
|
| 23 |
+
python benchmarks/aggregate_results.py results --manifest results/validated_results.txt --out results/validated_leaderboard.csv
|
| 24 |
+
|
| 25 |
+
Raw benchmark JSON files should not be hand-edited. Failed runs and negative results are scientifically useful and should be preserved with enough metadata to identify their originating experiment.
|
results/dgx_spark_1024_canceling.json
ADDED
|
@@ -0,0 +1,64 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"benchmark": "spectra-rsi-scaling-v1",
|
| 3 |
+
"candidate": "canceling_mixture",
|
| 4 |
+
"config": {
|
| 5 |
+
"n_slices": 1024,
|
| 6 |
+
"n_experts": 8,
|
| 7 |
+
"rank": 2,
|
| 8 |
+
"m_coarse": 768,
|
| 9 |
+
"m_focused": 1024,
|
| 10 |
+
"items_per_row": 128,
|
| 11 |
+
"bootstrap_reps": 5,
|
| 12 |
+
"lambda_l1": 0.01,
|
| 13 |
+
"lambda_group": 0.04,
|
| 14 |
+
"seed": 7,
|
| 15 |
+
"candidate_seed": 99,
|
| 16 |
+
"candidate": "canceling_mixture",
|
| 17 |
+
"scale": 0.4,
|
| 18 |
+
"output": "results/dgx_spark_1024_canceling.json"
|
| 19 |
+
},
|
| 20 |
+
"system": {
|
| 21 |
+
"cuda_visible_devices": null,
|
| 22 |
+
"hostname": "gx10-db7a",
|
| 23 |
+
"platform": "Linux-6.17.0-1021-nvidia-aarch64-with-glibc2.39",
|
| 24 |
+
"torch_probe_error": "No module named 'torch'",
|
| 25 |
+
"gpu_count": 1,
|
| 26 |
+
"gpus": [
|
| 27 |
+
"NVIDIA GB10"
|
| 28 |
+
],
|
| 29 |
+
"nvidia_driver": "580.159.03",
|
| 30 |
+
"gpu_probe": "nvidia-smi"
|
| 31 |
+
},
|
| 32 |
+
"wall_seconds": 0.8139567289999832,
|
| 33 |
+
"dense_fallback": false,
|
| 34 |
+
"pilot_residual": 0.20033903475967274,
|
| 35 |
+
"items_probe": 12800,
|
| 36 |
+
"items_sense": 275132,
|
| 37 |
+
"items_anchor": 24000,
|
| 38 |
+
"gate_decision": "quarantine",
|
| 39 |
+
"support_f1": 1.0,
|
| 40 |
+
"normalized_delta_error": 0.33441075948420146,
|
| 41 |
+
"regression_recall": 1.0,
|
| 42 |
+
"true_experts": [
|
| 43 |
+
2,
|
| 44 |
+
6
|
| 45 |
+
],
|
| 46 |
+
"nominated_experts": [
|
| 47 |
+
2,
|
| 48 |
+
6
|
| 49 |
+
],
|
| 50 |
+
"recovered_experts": [
|
| 51 |
+
2,
|
| 52 |
+
6
|
| 53 |
+
],
|
| 54 |
+
"bootstrap_frequencies": [
|
| 55 |
+
0.6,
|
| 56 |
+
0.2,
|
| 57 |
+
1.0,
|
| 58 |
+
0.4,
|
| 59 |
+
0.0,
|
| 60 |
+
0.6,
|
| 61 |
+
1.0,
|
| 62 |
+
0.0
|
| 63 |
+
]
|
| 64 |
+
}
|
results/dgx_spark_1024_noncompressible.json
ADDED
|
@@ -0,0 +1,74 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"benchmark": "spectra-rsi-scaling-v1",
|
| 3 |
+
"candidate": "broad_noncompressible",
|
| 4 |
+
"config": {
|
| 5 |
+
"n_slices": 1024,
|
| 6 |
+
"n_experts": 8,
|
| 7 |
+
"rank": 2,
|
| 8 |
+
"m_coarse": 768,
|
| 9 |
+
"m_focused": 1024,
|
| 10 |
+
"items_per_row": 128,
|
| 11 |
+
"bootstrap_reps": 5,
|
| 12 |
+
"lambda_l1": 0.01,
|
| 13 |
+
"lambda_group": 0.04,
|
| 14 |
+
"seed": 7,
|
| 15 |
+
"candidate_seed": 99,
|
| 16 |
+
"candidate": "broad_noncompressible",
|
| 17 |
+
"scale": 0.4,
|
| 18 |
+
"output": "results/dgx_spark_1024_noncompressible.json"
|
| 19 |
+
},
|
| 20 |
+
"system": {
|
| 21 |
+
"cuda_visible_devices": null,
|
| 22 |
+
"hostname": "gx10-db7a",
|
| 23 |
+
"platform": "Linux-6.17.0-1021-nvidia-aarch64-with-glibc2.39",
|
| 24 |
+
"torch_probe_error": "No module named 'torch'",
|
| 25 |
+
"gpu_count": 1,
|
| 26 |
+
"gpus": [
|
| 27 |
+
"NVIDIA GB10"
|
| 28 |
+
],
|
| 29 |
+
"nvidia_driver": "580.159.03",
|
| 30 |
+
"gpu_probe": "nvidia-smi"
|
| 31 |
+
},
|
| 32 |
+
"wall_seconds": 0.8184029540007032,
|
| 33 |
+
"dense_fallback": false,
|
| 34 |
+
"pilot_residual": 0.4261726798996619,
|
| 35 |
+
"items_probe": 12800,
|
| 36 |
+
"items_sense": 273912,
|
| 37 |
+
"items_anchor": 24000,
|
| 38 |
+
"gate_decision": "quarantine",
|
| 39 |
+
"support_f1": 0.7692307692307693,
|
| 40 |
+
"normalized_delta_error": 0.5666150578143931,
|
| 41 |
+
"regression_recall": 0.9230769230769231,
|
| 42 |
+
"true_experts": [
|
| 43 |
+
0,
|
| 44 |
+
1,
|
| 45 |
+
2,
|
| 46 |
+
3,
|
| 47 |
+
4,
|
| 48 |
+
5,
|
| 49 |
+
6,
|
| 50 |
+
7
|
| 51 |
+
],
|
| 52 |
+
"nominated_experts": [
|
| 53 |
+
1,
|
| 54 |
+
2,
|
| 55 |
+
5
|
| 56 |
+
],
|
| 57 |
+
"recovered_experts": [
|
| 58 |
+
1,
|
| 59 |
+
2,
|
| 60 |
+
3,
|
| 61 |
+
4,
|
| 62 |
+
5
|
| 63 |
+
],
|
| 64 |
+
"bootstrap_frequencies": [
|
| 65 |
+
0.2,
|
| 66 |
+
1.0,
|
| 67 |
+
1.0,
|
| 68 |
+
0.8,
|
| 69 |
+
0.6,
|
| 70 |
+
1.0,
|
| 71 |
+
0.0,
|
| 72 |
+
0.6
|
| 73 |
+
]
|
| 74 |
+
}
|
results/dgx_spark_1024_noncompressible_seed100.json
ADDED
|
@@ -0,0 +1,75 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"benchmark": "spectra-rsi-scaling-v1",
|
| 3 |
+
"candidate": "broad_noncompressible",
|
| 4 |
+
"config": {
|
| 5 |
+
"n_slices": 1024,
|
| 6 |
+
"n_experts": 8,
|
| 7 |
+
"rank": 2,
|
| 8 |
+
"m_coarse": 768,
|
| 9 |
+
"m_focused": 1024,
|
| 10 |
+
"items_per_row": 128,
|
| 11 |
+
"bootstrap_reps": 5,
|
| 12 |
+
"lambda_l1": 0.01,
|
| 13 |
+
"lambda_group": 0.04,
|
| 14 |
+
"seed": 7,
|
| 15 |
+
"candidate_seed": 100,
|
| 16 |
+
"candidate": "broad_noncompressible",
|
| 17 |
+
"scale": 0.4,
|
| 18 |
+
"output": "results/dgx_spark_1024_noncompressible_seed100.json"
|
| 19 |
+
},
|
| 20 |
+
"system": {
|
| 21 |
+
"cuda_visible_devices": null,
|
| 22 |
+
"hostname": "gx10-db7a",
|
| 23 |
+
"platform": "Linux-6.17.0-1021-nvidia-aarch64-with-glibc2.39",
|
| 24 |
+
"torch_probe_error": "No module named 'torch'",
|
| 25 |
+
"gpu_count": 1,
|
| 26 |
+
"gpus": [
|
| 27 |
+
"NVIDIA GB10"
|
| 28 |
+
],
|
| 29 |
+
"nvidia_driver": "580.159.03",
|
| 30 |
+
"gpu_probe": "nvidia-smi"
|
| 31 |
+
},
|
| 32 |
+
"wall_seconds": 0.8000310129991703,
|
| 33 |
+
"dense_fallback": false,
|
| 34 |
+
"pilot_residual": 0.2939873837188221,
|
| 35 |
+
"items_probe": 12800,
|
| 36 |
+
"items_sense": 273478,
|
| 37 |
+
"items_anchor": 21400,
|
| 38 |
+
"gate_decision": "accept",
|
| 39 |
+
"support_f1": 0.8571428571428571,
|
| 40 |
+
"normalized_delta_error": 0.4642433735647704,
|
| 41 |
+
"regression_recall": 1.0,
|
| 42 |
+
"true_experts": [
|
| 43 |
+
0,
|
| 44 |
+
1,
|
| 45 |
+
2,
|
| 46 |
+
3,
|
| 47 |
+
4,
|
| 48 |
+
5,
|
| 49 |
+
6,
|
| 50 |
+
7
|
| 51 |
+
],
|
| 52 |
+
"nominated_experts": [
|
| 53 |
+
2,
|
| 54 |
+
3,
|
| 55 |
+
4
|
| 56 |
+
],
|
| 57 |
+
"recovered_experts": [
|
| 58 |
+
0,
|
| 59 |
+
1,
|
| 60 |
+
2,
|
| 61 |
+
3,
|
| 62 |
+
4,
|
| 63 |
+
6
|
| 64 |
+
],
|
| 65 |
+
"bootstrap_frequencies": [
|
| 66 |
+
0.8,
|
| 67 |
+
1.0,
|
| 68 |
+
1.0,
|
| 69 |
+
1.0,
|
| 70 |
+
1.0,
|
| 71 |
+
0.4,
|
| 72 |
+
0.8,
|
| 73 |
+
0.0
|
| 74 |
+
]
|
| 75 |
+
}
|
results/dgx_spark_1024_off_dictionary.json
ADDED
|
@@ -0,0 +1,38 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"benchmark": "spectra-rsi-scaling-v1",
|
| 3 |
+
"candidate": "off_dictionary",
|
| 4 |
+
"config": {
|
| 5 |
+
"n_slices": 1024,
|
| 6 |
+
"n_experts": 8,
|
| 7 |
+
"rank": 2,
|
| 8 |
+
"m_coarse": 768,
|
| 9 |
+
"m_focused": 1024,
|
| 10 |
+
"items_per_row": 128,
|
| 11 |
+
"bootstrap_reps": 5,
|
| 12 |
+
"lambda_l1": 0.01,
|
| 13 |
+
"lambda_group": 0.04,
|
| 14 |
+
"seed": 7,
|
| 15 |
+
"candidate_seed": 99,
|
| 16 |
+
"candidate": "off_dictionary",
|
| 17 |
+
"scale": 0.4,
|
| 18 |
+
"output": "results/dgx_spark_1024_off_dictionary.json"
|
| 19 |
+
},
|
| 20 |
+
"system": {
|
| 21 |
+
"cuda_visible_devices": null,
|
| 22 |
+
"hostname": "gx10-db7a",
|
| 23 |
+
"platform": "Linux-6.17.0-1021-nvidia-aarch64-with-glibc2.39",
|
| 24 |
+
"torch_probe_error": "No module named 'torch'",
|
| 25 |
+
"gpu_count": 1,
|
| 26 |
+
"gpus": [
|
| 27 |
+
"NVIDIA GB10"
|
| 28 |
+
],
|
| 29 |
+
"nvidia_driver": "580.159.03",
|
| 30 |
+
"gpu_probe": "nvidia-smi"
|
| 31 |
+
},
|
| 32 |
+
"wall_seconds": 0.0013118490005581407,
|
| 33 |
+
"dense_fallback": true,
|
| 34 |
+
"pilot_residual": 0.9999999916719566,
|
| 35 |
+
"items_probe": 0,
|
| 36 |
+
"items_sense": 0,
|
| 37 |
+
"items_anchor": 0
|
| 38 |
+
}
|
results/dgx_spark_1024_reg5x.json
ADDED
|
@@ -0,0 +1,61 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"benchmark": "spectra-rsi-scaling-v1",
|
| 3 |
+
"candidate": "single_gain",
|
| 4 |
+
"config": {
|
| 5 |
+
"n_slices": 1024,
|
| 6 |
+
"n_experts": 8,
|
| 7 |
+
"rank": 2,
|
| 8 |
+
"m_coarse": 768,
|
| 9 |
+
"m_focused": 1024,
|
| 10 |
+
"items_per_row": 128,
|
| 11 |
+
"bootstrap_reps": 5,
|
| 12 |
+
"lambda_l1": 0.01,
|
| 13 |
+
"lambda_group": 0.04,
|
| 14 |
+
"seed": 7,
|
| 15 |
+
"candidate_seed": 99,
|
| 16 |
+
"candidate": "single_gain",
|
| 17 |
+
"scale": 0.4,
|
| 18 |
+
"output": "results/dgx_spark_1024_reg5x.json"
|
| 19 |
+
},
|
| 20 |
+
"system": {
|
| 21 |
+
"cuda_visible_devices": null,
|
| 22 |
+
"hostname": "gx10-db7a",
|
| 23 |
+
"platform": "Linux-6.17.0-1021-nvidia-aarch64-with-glibc2.39",
|
| 24 |
+
"torch_probe_error": "No module named 'torch'",
|
| 25 |
+
"gpu_count": 1,
|
| 26 |
+
"gpus": [
|
| 27 |
+
"NVIDIA GB10"
|
| 28 |
+
],
|
| 29 |
+
"nvidia_driver": "580.159.03",
|
| 30 |
+
"gpu_probe": "nvidia-smi"
|
| 31 |
+
},
|
| 32 |
+
"wall_seconds": 0.7774730420005653,
|
| 33 |
+
"dense_fallback": false,
|
| 34 |
+
"pilot_residual": 0.2353382692677464,
|
| 35 |
+
"items_probe": 12800,
|
| 36 |
+
"items_sense": 276707,
|
| 37 |
+
"items_anchor": 22700,
|
| 38 |
+
"gate_decision": "accept",
|
| 39 |
+
"support_f1": 1.0,
|
| 40 |
+
"normalized_delta_error": 0.3317853169160946,
|
| 41 |
+
"regression_recall": 1.0,
|
| 42 |
+
"true_experts": [
|
| 43 |
+
2
|
| 44 |
+
],
|
| 45 |
+
"nominated_experts": [
|
| 46 |
+
2
|
| 47 |
+
],
|
| 48 |
+
"recovered_experts": [
|
| 49 |
+
2
|
| 50 |
+
],
|
| 51 |
+
"bootstrap_frequencies": [
|
| 52 |
+
0.2,
|
| 53 |
+
0.4,
|
| 54 |
+
1.0,
|
| 55 |
+
0.2,
|
| 56 |
+
0.2,
|
| 57 |
+
0.2,
|
| 58 |
+
0.2,
|
| 59 |
+
0.0
|
| 60 |
+
]
|
| 61 |
+
}
|
results/dgx_spark_1024_reg5x_seed100.json
ADDED
|
@@ -0,0 +1,61 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"benchmark": "spectra-rsi-scaling-v1",
|
| 3 |
+
"candidate": "single_gain",
|
| 4 |
+
"config": {
|
| 5 |
+
"n_slices": 1024,
|
| 6 |
+
"n_experts": 8,
|
| 7 |
+
"rank": 2,
|
| 8 |
+
"m_coarse": 768,
|
| 9 |
+
"m_focused": 1024,
|
| 10 |
+
"items_per_row": 128,
|
| 11 |
+
"bootstrap_reps": 5,
|
| 12 |
+
"lambda_l1": 0.01,
|
| 13 |
+
"lambda_group": 0.04,
|
| 14 |
+
"seed": 7,
|
| 15 |
+
"candidate_seed": 100,
|
| 16 |
+
"candidate": "single_gain",
|
| 17 |
+
"scale": 0.4,
|
| 18 |
+
"output": "results/dgx_spark_1024_reg5x_seed100.json"
|
| 19 |
+
},
|
| 20 |
+
"system": {
|
| 21 |
+
"cuda_visible_devices": null,
|
| 22 |
+
"hostname": "gx10-db7a",
|
| 23 |
+
"platform": "Linux-6.17.0-1021-nvidia-aarch64-with-glibc2.39",
|
| 24 |
+
"torch_probe_error": "No module named 'torch'",
|
| 25 |
+
"gpu_count": 1,
|
| 26 |
+
"gpus": [
|
| 27 |
+
"NVIDIA GB10"
|
| 28 |
+
],
|
| 29 |
+
"nvidia_driver": "580.159.03",
|
| 30 |
+
"gpu_probe": "nvidia-smi"
|
| 31 |
+
},
|
| 32 |
+
"wall_seconds": 0.7741304430001037,
|
| 33 |
+
"dense_fallback": false,
|
| 34 |
+
"pilot_residual": 0.2339949193998105,
|
| 35 |
+
"items_probe": 12800,
|
| 36 |
+
"items_sense": 276707,
|
| 37 |
+
"items_anchor": 21800,
|
| 38 |
+
"gate_decision": "accept",
|
| 39 |
+
"support_f1": 1.0,
|
| 40 |
+
"normalized_delta_error": 0.34114380104392295,
|
| 41 |
+
"regression_recall": 1.0,
|
| 42 |
+
"true_experts": [
|
| 43 |
+
2
|
| 44 |
+
],
|
| 45 |
+
"nominated_experts": [
|
| 46 |
+
2
|
| 47 |
+
],
|
| 48 |
+
"recovered_experts": [
|
| 49 |
+
2
|
| 50 |
+
],
|
| 51 |
+
"bootstrap_frequencies": [
|
| 52 |
+
0.2,
|
| 53 |
+
0.6,
|
| 54 |
+
1.0,
|
| 55 |
+
0.2,
|
| 56 |
+
0.4,
|
| 57 |
+
0.2,
|
| 58 |
+
0.2,
|
| 59 |
+
0.0
|
| 60 |
+
]
|
| 61 |
+
}
|
results/dgx_spark_1024_regression_budgetfix.json
ADDED
|
@@ -0,0 +1,62 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"benchmark": "spectra-rsi-scaling-v1",
|
| 3 |
+
"candidate": "single_regression",
|
| 4 |
+
"config": {
|
| 5 |
+
"n_slices": 1024,
|
| 6 |
+
"n_experts": 8,
|
| 7 |
+
"rank": 2,
|
| 8 |
+
"m_coarse": 768,
|
| 9 |
+
"m_focused": 1024,
|
| 10 |
+
"items_per_row": 128,
|
| 11 |
+
"bootstrap_reps": 5,
|
| 12 |
+
"lambda_l1": 0.01,
|
| 13 |
+
"lambda_group": 0.04,
|
| 14 |
+
"seed": 7,
|
| 15 |
+
"candidate_seed": 99,
|
| 16 |
+
"candidate": "single_regression",
|
| 17 |
+
"scale": 0.4,
|
| 18 |
+
"output": "results/dgx_spark_1024_regression_budgetfix.json"
|
| 19 |
+
},
|
| 20 |
+
"system": {
|
| 21 |
+
"cuda_visible_devices": null,
|
| 22 |
+
"hostname": "gx10-db7a",
|
| 23 |
+
"platform": "Linux-6.17.0-1021-nvidia-aarch64-with-glibc2.39",
|
| 24 |
+
"torch_probe_error": "No module named 'torch'",
|
| 25 |
+
"gpu_count": 1,
|
| 26 |
+
"gpus": [
|
| 27 |
+
"NVIDIA GB10"
|
| 28 |
+
],
|
| 29 |
+
"nvidia_driver": "580.159.03",
|
| 30 |
+
"gpu_probe": "nvidia-smi"
|
| 31 |
+
},
|
| 32 |
+
"wall_seconds": 0.7842826520009112,
|
| 33 |
+
"dense_fallback": false,
|
| 34 |
+
"pilot_residual": 0.2444575970827468,
|
| 35 |
+
"items_probe": 12800,
|
| 36 |
+
"items_sense": 276806,
|
| 37 |
+
"items_anchor": 24000,
|
| 38 |
+
"gate_decision": "quarantine",
|
| 39 |
+
"support_f1": 0.6666666666666666,
|
| 40 |
+
"normalized_delta_error": 0.40437287482632517,
|
| 41 |
+
"regression_recall": 0.984,
|
| 42 |
+
"true_experts": [
|
| 43 |
+
5
|
| 44 |
+
],
|
| 45 |
+
"nominated_experts": [
|
| 46 |
+
5
|
| 47 |
+
],
|
| 48 |
+
"recovered_experts": [
|
| 49 |
+
2,
|
| 50 |
+
5
|
| 51 |
+
],
|
| 52 |
+
"bootstrap_frequencies": [
|
| 53 |
+
0.0,
|
| 54 |
+
0.2,
|
| 55 |
+
0.6,
|
| 56 |
+
0.4,
|
| 57 |
+
0.4,
|
| 58 |
+
1.0,
|
| 59 |
+
0.2,
|
| 60 |
+
0.4
|
| 61 |
+
]
|
| 62 |
+
}
|
results/validated_leaderboard.csv
ADDED
|
@@ -0,0 +1,8 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
file,candidate,seed,candidate_seed,n_slices,n_experts,rank,m_coarse,m_focused,items_per_row,bootstrap_reps,lambda_l1,lambda_group,gpu_count,gpus,wall_seconds,pilot_residual,dense_fallback,gate_decision,support_f1,normalized_delta_error,regression_recall,items_probe,items_sense,items_anchor,true_experts,nominated_experts,recovered_experts,bootstrap_frequencies
|
| 2 |
+
dgx_spark_1024_reg5x.json,single_gain,7,99,1024,8,2,768,1024,128,5,0.01,0.04,1,NVIDIA GB10,0.7774730420005653,0.2353382692677464,False,accept,1.0,0.3317853169160946,1.0,12800,276707,22700,2,2,2,0.2;0.4;1.0;0.2;0.2;0.2;0.2;0.0
|
| 3 |
+
dgx_spark_1024_reg5x_seed100.json,single_gain,7,100,1024,8,2,768,1024,128,5,0.01,0.04,1,NVIDIA GB10,0.7741304430001037,0.2339949193998105,False,accept,1.0,0.34114380104392295,1.0,12800,276707,21800,2,2,2,0.2;0.6;1.0;0.2;0.4;0.2;0.2;0.0
|
| 4 |
+
dgx_spark_1024_regression_budgetfix.json,single_regression,7,99,1024,8,2,768,1024,128,5,0.01,0.04,1,NVIDIA GB10,0.7842826520009112,0.2444575970827468,False,quarantine,0.6666666666666666,0.40437287482632517,0.984,12800,276806,24000,5,5,2;5,0.0;0.2;0.6;0.4;0.4;1.0;0.2;0.4
|
| 5 |
+
dgx_spark_1024_canceling.json,canceling_mixture,7,99,1024,8,2,768,1024,128,5,0.01,0.04,1,NVIDIA GB10,0.8139567289999832,0.20033903475967274,False,quarantine,1.0,0.33441075948420146,1.0,12800,275132,24000,2;6,2;6,2;6,0.6;0.2;1.0;0.4;0.0;0.6;1.0;0.0
|
| 6 |
+
dgx_spark_1024_noncompressible.json,broad_noncompressible,7,99,1024,8,2,768,1024,128,5,0.01,0.04,1,NVIDIA GB10,0.8184029540007032,0.4261726798996619,False,quarantine,0.7692307692307693,0.5666150578143931,0.9230769230769231,12800,273912,24000,0;1;2;3;4;5;6;7,1;2;5,1;2;3;4;5,0.2;1.0;1.0;0.8;0.6;1.0;0.0;0.6
|
| 7 |
+
dgx_spark_1024_noncompressible_seed100.json,broad_noncompressible,7,100,1024,8,2,768,1024,128,5,0.01,0.04,1,NVIDIA GB10,0.8000310129991703,0.2939873837188221,False,accept,0.8571428571428571,0.4642433735647704,1.0,12800,273478,21400,0;1;2;3;4;5;6;7,2;3;4,0;1;2;3;4;6,0.8;1.0;1.0;1.0;1.0;0.4;0.8;0.0
|
| 8 |
+
dgx_spark_1024_off_dictionary.json,off_dictionary,7,99,1024,8,2,768,1024,128,5,0.01,0.04,1,NVIDIA GB10,0.0013118490005581407,0.9999999916719566,True,,,,,0,0,0,,,,
|
results/validated_results.txt
ADDED
|
@@ -0,0 +1,7 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
dgx_spark_1024_reg5x.json
|
| 2 |
+
dgx_spark_1024_reg5x_seed100.json
|
| 3 |
+
dgx_spark_1024_regression_budgetfix.json
|
| 4 |
+
dgx_spark_1024_canceling.json
|
| 5 |
+
dgx_spark_1024_noncompressible.json
|
| 6 |
+
dgx_spark_1024_noncompressible_seed100.json
|
| 7 |
+
dgx_spark_1024_off_dictionary.json
|
scripts/run_multiseed.sh
ADDED
|
@@ -0,0 +1,8 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
#!/usr/bin/env bash
|
| 2 |
+
set -euo pipefail
|
| 3 |
+
mkdir -p results
|
| 4 |
+
N=${N_SLICES:-1024}; E=${N_EXPERTS:-32}
|
| 5 |
+
for S in 1 2 3 4 5; do
|
| 6 |
+
python benchmarks/run_scaling.py --n-slices "$N" --n-experts "$E" --rank 2 --m-coarse 200 --m-focused 256 --items-per-row 256 --bootstrap-reps 20 --seed "$S" --candidate-seed "$((100+S))" --candidate single_gain --output "results/multiseed_n${N}_s${S}.json"
|
| 7 |
+
done
|
| 8 |
+
python benchmarks/aggregate_results.py results
|
scripts/run_scaling_sweep.sh
ADDED
|
@@ -0,0 +1,8 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
#!/usr/bin/env bash
|
| 2 |
+
set -euo pipefail
|
| 3 |
+
mkdir -p results
|
| 4 |
+
for N in 256 512 1024 2048; do
|
| 5 |
+
E=$((N/32)); [ "$E" -lt 8 ] && E=8
|
| 6 |
+
python benchmarks/run_scaling.py --n-slices "$N" --n-experts "$E" --rank 2 --m-coarse $((N/5)) --m-focused $((N/4)) --items-per-row 256 --bootstrap-reps 10 --candidate single_gain --output "results/scaling_n${N}.json"
|
| 7 |
+
done
|
| 8 |
+
python benchmarks/aggregate_results.py results
|
scripts/run_smoke.sh
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
#!/usr/bin/env bash
|
| 2 |
+
set -euo pipefail
|
| 3 |
+
python benchmarks/run_scaling.py --n-slices 128 --n-experts 8 --rank 2 --m-coarse 24 --m-focused 32 --items-per-row 96 --bootstrap-reps 5 --candidate single_gain --output results/smoke.json
|
spectra_rsi/__init__.py
ADDED
|
@@ -0,0 +1,31 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""SPECTRA-RSI: Counterfactual Spectral Sketching and Anytime-Valid Gating
|
| 2 |
+
for Modular Recursive Self-Improvement.
|
| 3 |
+
|
| 4 |
+
Reference implementation of the architecture in the SPECTRA-RSI manuscript.
|
| 5 |
+
All components operate against a pluggable `World` interface; a synthetic
|
| 6 |
+
ground-truth world is provided for validation and benchmarking.
|
| 7 |
+
"""
|
| 8 |
+
|
| 9 |
+
from .config import SpectraConfig
|
| 10 |
+
from .world import SyntheticWorld, CandidateUpdate
|
| 11 |
+
from .experts import Router, expert_priors
|
| 12 |
+
from .probes import build_dictionary, Dictionary
|
| 13 |
+
from .sketch import design_matrix, paired_sketch, neyman_allocation
|
| 14 |
+
from .recovery import sparse_group_recover, debias, bootstrap_support
|
| 15 |
+
from .drift import procrustes_transport, principal_angles, ResidualDriftMonitor
|
| 16 |
+
from .gate import EProcess, AnchorGate, GateDecision
|
| 17 |
+
from .loop import SpectraRSILoop, IterationReport
|
| 18 |
+
from . import metrics
|
| 19 |
+
|
| 20 |
+
__version__ = "0.1.0"
|
| 21 |
+
|
| 22 |
+
__all__ = [
|
| 23 |
+
"SpectraConfig", "SyntheticWorld", "CandidateUpdate",
|
| 24 |
+
"Router", "expert_priors",
|
| 25 |
+
"build_dictionary", "Dictionary",
|
| 26 |
+
"design_matrix", "paired_sketch", "neyman_allocation",
|
| 27 |
+
"sparse_group_recover", "debias", "bootstrap_support",
|
| 28 |
+
"procrustes_transport", "principal_angles", "ResidualDriftMonitor",
|
| 29 |
+
"EProcess", "AnchorGate", "GateDecision",
|
| 30 |
+
"SpectraRSILoop", "IterationReport", "metrics",
|
| 31 |
+
]
|
spectra_rsi/config.py
ADDED
|
@@ -0,0 +1,51 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Configuration for a SPECTRA-RSI loop instance."""
|
| 2 |
+
from dataclasses import dataclass, field
|
| 3 |
+
|
| 4 |
+
|
| 5 |
+
@dataclass
|
| 6 |
+
class SpectraConfig:
|
| 7 |
+
# --- capability space ---
|
| 8 |
+
n_slices: int = 400 # n: capability slices
|
| 9 |
+
n_experts: int = 16 # E: experts / dictionary groups
|
| 10 |
+
rank_per_expert: int = 2 # r_e: retained response rank per expert
|
| 11 |
+
|
| 12 |
+
# --- probing (counterfactual dictionary) ---
|
| 13 |
+
probe_magnitude: float = 0.4 # h: signed probe step (within trust region)
|
| 14 |
+
probe_items: int = 400 # items per slice per probe evaluation
|
| 15 |
+
|
| 16 |
+
# --- sketching ---
|
| 17 |
+
m_coarse: int = 60 # stage-I sketch rows
|
| 18 |
+
m_focused: int = 90 # stage-II sketch rows
|
| 19 |
+
row_density: float = 0.10 # expected fraction of nonzero weights per row
|
| 20 |
+
items_per_row: int = 200 # item budget per sketch row (Neyman-allocated)
|
| 21 |
+
explore_fraction: float = 0.15 # measurement mass outside nominated groups
|
| 22 |
+
|
| 23 |
+
# --- recovery ---
|
| 24 |
+
lambda_l1: float = 2e-3
|
| 25 |
+
lambda_group: float = 8e-3
|
| 26 |
+
prior_gamma: float = 0.5 # w_e = (p_e + eps)^(-gamma)
|
| 27 |
+
prior_clip: tuple = (0.05, 0.95)
|
| 28 |
+
fista_iters: int = 600
|
| 29 |
+
bootstrap_reps: int = 60
|
| 30 |
+
|
| 31 |
+
# --- trust region / pilot ---
|
| 32 |
+
pilot_items: int = 120
|
| 33 |
+
trust_region_residual: float = 0.5 # relative pilot residual triggering fallback
|
| 34 |
+
|
| 35 |
+
# --- drift ---
|
| 36 |
+
drift_alpha: float = 0.05
|
| 37 |
+
drift_window: int = 50
|
| 38 |
+
|
| 39 |
+
# --- anytime-valid gate ---
|
| 40 |
+
alpha_risk: float = 0.05 # total false-acceptance budget over critical anchors
|
| 41 |
+
tau_margin: float = 0.02 # material-regression margin tau_j
|
| 42 |
+
max_anchor_items: int = 24000 # total sequential anchor budget
|
| 43 |
+
anchor_batch: int = 100
|
| 44 |
+
|
| 45 |
+
# --- misc ---
|
| 46 |
+
seed: int = 0
|
| 47 |
+
audit_dir: str = "audit_logs"
|
| 48 |
+
|
| 49 |
+
def __post_init__(self):
|
| 50 |
+
assert 0 < self.row_density <= 1
|
| 51 |
+
assert 0 <= self.explore_fraction < 1
|
spectra_rsi/drift.py
ADDED
|
@@ -0,0 +1,62 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Dictionary transport and drift monitoring (Sec. 4.3).
|
| 2 |
+
|
| 3 |
+
- Orthogonal Procrustes alignment of successive dictionaries.
|
| 4 |
+
- Principal angles quantify subspace motion.
|
| 5 |
+
- ResidualDriftMonitor: anytime-valid confidence sequence on predictive
|
| 6 |
+
residuals; a boundary crossing triggers re-probing / dense fallback.
|
| 7 |
+
"""
|
| 8 |
+
import numpy as np
|
| 9 |
+
|
| 10 |
+
|
| 11 |
+
def procrustes_transport(Psi_new, Psi_old):
|
| 12 |
+
"""Q = argmin_{Q^T Q = I} ||Psi_new - Psi_old Q||_F (rotation of old
|
| 13 |
+
coordinates into the new basis). Returns (Q, alignment_residual)."""
|
| 14 |
+
k = min(Psi_new.shape[1], Psi_old.shape[1])
|
| 15 |
+
A = Psi_old[:, :k].T @ Psi_new[:, :k]
|
| 16 |
+
U, _, Vt = np.linalg.svd(A)
|
| 17 |
+
Q = U @ Vt
|
| 18 |
+
resid = np.linalg.norm(Psi_new[:, :k] - Psi_old[:, :k] @ Q, "fro") / np.sqrt(k)
|
| 19 |
+
return Q, resid
|
| 20 |
+
|
| 21 |
+
|
| 22 |
+
def principal_angles(Psi_a, Psi_b):
|
| 23 |
+
"""Principal angles (radians) between column spaces."""
|
| 24 |
+
Qa, _ = np.linalg.qr(Psi_a)
|
| 25 |
+
Qb, _ = np.linalg.qr(Psi_b)
|
| 26 |
+
s = np.linalg.svd(Qa.T @ Qb, compute_uv=False)
|
| 27 |
+
return np.arccos(np.clip(s, -1, 1))
|
| 28 |
+
|
| 29 |
+
|
| 30 |
+
class ResidualDriftMonitor:
|
| 31 |
+
"""Sub-Gaussian mixture confidence sequence on the mean squared
|
| 32 |
+
normalized residual. Anytime-valid: crossing is a drift alarm at level
|
| 33 |
+
alpha regardless of when you look (Howard et al. style mixture bound)."""
|
| 34 |
+
|
| 35 |
+
def __init__(self, alpha=0.05, scale=1.0):
|
| 36 |
+
self.alpha = alpha
|
| 37 |
+
self.scale = scale # residuals assumed sub-Gaussian(scale)
|
| 38 |
+
self.t = 0
|
| 39 |
+
self.sum = 0.0
|
| 40 |
+
|
| 41 |
+
def update(self, residuals):
|
| 42 |
+
r = np.atleast_1d(residuals)
|
| 43 |
+
self.t += r.size
|
| 44 |
+
self.sum += float(np.sum(r))
|
| 45 |
+
return self.is_alarmed()
|
| 46 |
+
|
| 47 |
+
def radius(self):
|
| 48 |
+
if self.t == 0:
|
| 49 |
+
return np.inf
|
| 50 |
+
t, s, a = self.t, self.scale, self.alpha
|
| 51 |
+
# normal-mixture boundary: valid at all t simultaneously
|
| 52 |
+
rho = 1.0
|
| 53 |
+
return s * np.sqrt(2 * (t + rho) / t ** 2
|
| 54 |
+
* np.log(np.sqrt((t + rho) / rho) / a))
|
| 55 |
+
|
| 56 |
+
def mean(self):
|
| 57 |
+
return self.sum / max(self.t, 1)
|
| 58 |
+
|
| 59 |
+
def is_alarmed(self):
|
| 60 |
+
"""Alarm when the anytime lower confidence bound on the mean
|
| 61 |
+
residual exceeds 0 (systematic prediction error)."""
|
| 62 |
+
return self.t > 0 and (self.mean() - self.radius()) > 0.0
|
spectra_rsi/experts.py
ADDED
|
@@ -0,0 +1,47 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Spectral mixture of low-rank experts: router statistics and support priors.
|
| 2 |
+
|
| 3 |
+
Implements Sec. 5: occupancy pi_e, gate entropy H_e, update energy g_e,
|
| 4 |
+
and the clipped prior activity score p_e with an exploration floor.
|
| 5 |
+
Router statistics are a PRIOR, never evidence of improvement.
|
| 6 |
+
"""
|
| 7 |
+
import numpy as np
|
| 8 |
+
|
| 9 |
+
|
| 10 |
+
class Router:
|
| 11 |
+
"""Tracks routing statistics over the candidate-generation corpus."""
|
| 12 |
+
|
| 13 |
+
def __init__(self, n_experts, seed=0):
|
| 14 |
+
self.E = n_experts
|
| 15 |
+
self._rng = np.random.default_rng(seed)
|
| 16 |
+
|
| 17 |
+
def occupancy_stats(self, cand_beta, rank_per_expert, corpus_size=2000,
|
| 18 |
+
fidelity=0.8):
|
| 19 |
+
"""Simulate router occupancy correlated (imperfectly) with which
|
| 20 |
+
experts the candidate actually touched. `fidelity` < 1 injects router
|
| 21 |
+
noise / gaming so the prior is fallible, as the paper requires."""
|
| 22 |
+
R = rank_per_expert
|
| 23 |
+
energy = np.array([np.linalg.norm(cand_beta[e * R:(e + 1) * R])
|
| 24 |
+
for e in range(self.E)])
|
| 25 |
+
base = energy / (energy.max() + 1e-12) if energy.max() > 0 else energy
|
| 26 |
+
noise = self._rng.random(self.E)
|
| 27 |
+
occ = fidelity * base + (1 - fidelity) * noise
|
| 28 |
+
occ = occ / (occ.sum() + 1e-12)
|
| 29 |
+
# gate entropy per expert (high entropy = diffuse routing)
|
| 30 |
+
p = np.clip(occ, 1e-9, 1)
|
| 31 |
+
ent = -p * np.log(p)
|
| 32 |
+
return occ, ent, energy
|
| 33 |
+
|
| 34 |
+
|
| 35 |
+
def expert_priors(occupancy, entropy, energy, clip=(0.05, 0.95),
|
| 36 |
+
a0=-1.0, a1=0.8, a2=1.2, a3=0.5):
|
| 37 |
+
"""Prior activity score p_e = sigma(a0 + a1 log occ + a2 g - a3 H),
|
| 38 |
+
clipped to [p_min, p_max] to preserve exploration (Eq. prior)."""
|
| 39 |
+
g = energy / (np.max(energy) + 1e-12) if np.max(energy) > 0 else energy
|
| 40 |
+
z = (a0 + a1 * np.log(occupancy + 1e-6) + a2 * g - a3 * entropy)
|
| 41 |
+
p = 1.0 / (1.0 + np.exp(-z))
|
| 42 |
+
return np.clip(p, clip[0], clip[1])
|
| 43 |
+
|
| 44 |
+
|
| 45 |
+
def group_weights(priors, gamma=0.5, eps=1e-3):
|
| 46 |
+
"""w_e = (p_e + eps)^(-gamma): low-prior groups pay a higher penalty."""
|
| 47 |
+
return (priors + eps) ** (-gamma)
|
spectra_rsi/gate.py
ADDED
|
@@ -0,0 +1,109 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Anytime-valid modular acceptance gate (Sec. 8).
|
| 2 |
+
|
| 3 |
+
Non-inferiority testing on protected anchor streams via betting e-processes:
|
| 4 |
+
|
| 5 |
+
H0_j : Delta c_j <= -tau vs. H1_j : Delta c_j > -tau
|
| 6 |
+
|
| 7 |
+
For bounded paired differences d in [-B, B], the wealth process
|
| 8 |
+
|
| 9 |
+
E_s = prod_i ( 1 + lam_i * (d_i + tau) ), 0 <= lam_i < 1 / (B + tau)
|
| 10 |
+
|
| 11 |
+
is a nonnegative supermartingale under every P in H0 (E_P[d] <= -tau implies
|
| 12 |
+
E_P[1 + lam (d + tau)] <= 1). By Ville's inequality,
|
| 13 |
+
P(sup_s E_s >= 1/alpha) <= alpha under H0, at ANY stopping rule.
|
| 14 |
+
A grid mixture over lam keeps validity while adapting to unknown effect size.
|
| 15 |
+
"""
|
| 16 |
+
from dataclasses import dataclass, field
|
| 17 |
+
from enum import Enum
|
| 18 |
+
import numpy as np
|
| 19 |
+
|
| 20 |
+
|
| 21 |
+
class GateDecision(Enum):
|
| 22 |
+
ACCEPT = "accept"
|
| 23 |
+
ROLLBACK = "rollback"
|
| 24 |
+
QUARANTINE = "quarantine"
|
| 25 |
+
|
| 26 |
+
|
| 27 |
+
class EProcess:
|
| 28 |
+
"""Grid-mixture betting e-process for H0: mean <= -tau, obs in [-B, B]."""
|
| 29 |
+
|
| 30 |
+
def __init__(self, tau, alpha=0.05, bound=1.0, n_grid=12):
|
| 31 |
+
self.tau = tau
|
| 32 |
+
self.alpha = alpha
|
| 33 |
+
self.B = bound
|
| 34 |
+
lam_max = 0.5 / (bound + tau) # stay safely inside admissible range
|
| 35 |
+
self.grid = np.linspace(lam_max / n_grid, lam_max, n_grid)
|
| 36 |
+
self.log_wealth = np.zeros(n_grid)
|
| 37 |
+
self.n = 0
|
| 38 |
+
|
| 39 |
+
def update(self, observations):
|
| 40 |
+
obs = np.clip(np.atleast_1d(observations), -self.B, self.B)
|
| 41 |
+
for d in obs:
|
| 42 |
+
self.log_wealth += np.log1p(self.grid * (d + self.tau))
|
| 43 |
+
self.n += 1
|
| 44 |
+
return self.value()
|
| 45 |
+
|
| 46 |
+
def value(self):
|
| 47 |
+
# mixture (uniform prior over grid) — still an e-process
|
| 48 |
+
m = self.log_wealth.max()
|
| 49 |
+
return float(np.exp(m) * np.mean(np.exp(self.log_wealth - m)))
|
| 50 |
+
|
| 51 |
+
@property
|
| 52 |
+
def rejects_null(self):
|
| 53 |
+
"""True iff evidence of non-inferiority has crossed 1/alpha."""
|
| 54 |
+
return self.value() >= 1.0 / self.alpha
|
| 55 |
+
|
| 56 |
+
|
| 57 |
+
@dataclass
|
| 58 |
+
class AnchorReport:
|
| 59 |
+
anchor_id: int
|
| 60 |
+
e_value: float
|
| 61 |
+
n_items: int
|
| 62 |
+
non_inferior: bool
|
| 63 |
+
|
| 64 |
+
|
| 65 |
+
class AnchorGate:
|
| 66 |
+
"""Per-expert + composed-candidate acceptance on protected anchors.
|
| 67 |
+
|
| 68 |
+
Each critical anchor stream gets an independent e-process with a
|
| 69 |
+
Bonferroni share of the total risk budget alpha. Acceptance requires
|
| 70 |
+
every critical anchor to establish non-inferiority; budget exhaustion
|
| 71 |
+
without rejection sends the expert to quarantine, not acceptance.
|
| 72 |
+
"""
|
| 73 |
+
|
| 74 |
+
def __init__(self, anchor_slices, tau, alpha_total, bound=1.0):
|
| 75 |
+
self.anchor_slices = list(anchor_slices)
|
| 76 |
+
self.tau = tau
|
| 77 |
+
self.alpha_each = alpha_total / max(1, len(self.anchor_slices))
|
| 78 |
+
self.bound = bound
|
| 79 |
+
|
| 80 |
+
def run(self, world, cand, batch=50, max_items=4000, rng=None):
|
| 81 |
+
"""Sequentially stream anchor items until every e-process resolves
|
| 82 |
+
or budget runs out. Anytime-valid: any stopping rule is fine."""
|
| 83 |
+
rng = rng or np.random.default_rng(0)
|
| 84 |
+
procs = {j: EProcess(self.tau, self.alpha_each, self.bound)
|
| 85 |
+
for j in self.anchor_slices}
|
| 86 |
+
undecided = set(self.anchor_slices)
|
| 87 |
+
used = 0
|
| 88 |
+
while undecided and used < max_items:
|
| 89 |
+
for j in list(undecided):
|
| 90 |
+
if used + batch > max_items:
|
| 91 |
+
break
|
| 92 |
+
d = world.paired_scores(cand, j, batch, rng=rng)
|
| 93 |
+
procs[j].update(d)
|
| 94 |
+
used += batch
|
| 95 |
+
if procs[j].rejects_null:
|
| 96 |
+
undecided.discard(j)
|
| 97 |
+
reports = [AnchorReport(j, procs[j].value(), procs[j].n,
|
| 98 |
+
procs[j].rejects_null)
|
| 99 |
+
for j in self.anchor_slices]
|
| 100 |
+
all_pass = all(r.non_inferior for r in reports)
|
| 101 |
+
decision = GateDecision.ACCEPT if all_pass else (
|
| 102 |
+
GateDecision.ROLLBACK if any(
|
| 103 |
+
(not r.non_inferior) and r.n_items >= max_items // max(1, len(self.anchor_slices))
|
| 104 |
+
for r in reports) else GateDecision.QUARANTINE)
|
| 105 |
+
if not all_pass and used >= max_items:
|
| 106 |
+
decision = GateDecision.QUARANTINE if any(
|
| 107 |
+
not r.non_inferior and r.e_value > 1.0 for r in reports) \
|
| 108 |
+
else GateDecision.ROLLBACK
|
| 109 |
+
return decision, reports, used
|
spectra_rsi/loop.py
ADDED
|
@@ -0,0 +1,212 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""The SPECTRA-RSI loop (Algorithm 1).
|
| 2 |
+
|
| 3 |
+
One counterfactual, risk-controlled iteration:
|
| 4 |
+
|
| 5 |
+
1. freeze candidate; align previous basis, check drift confidence sequence
|
| 6 |
+
2. (re)build interventional dictionary from signed expert probes if needed
|
| 7 |
+
3. clipped router priors with exploration floor
|
| 8 |
+
4. local-linearity pilot (trust region) -> dense fallback on failure
|
| 9 |
+
5. stage-I coarse paired sketch over all groups
|
| 10 |
+
6. preliminary recovery + bootstrap instability -> nominate groups
|
| 11 |
+
7. stage-II focused sketch, retaining exploration mass
|
| 12 |
+
8. weighted sparse-group recovery + debias -> expert attribution
|
| 13 |
+
9. protected anchor e-processes per nominated expert + composed candidate
|
| 14 |
+
10. accept / rollback / quarantine; audit record
|
| 15 |
+
|
| 16 |
+
Discovery never touches anchor data; anchors never inform the dictionary,
|
| 17 |
+
penalties, or nomination (Sec. 8.4 firewall).
|
| 18 |
+
"""
|
| 19 |
+
from dataclasses import dataclass, field, asdict
|
| 20 |
+
import json, os, time, hashlib
|
| 21 |
+
import numpy as np
|
| 22 |
+
|
| 23 |
+
from .experts import Router, expert_priors, group_weights
|
| 24 |
+
from .probes import build_dictionary, pilot_linearity_check
|
| 25 |
+
from .sketch import design_matrix, paired_sketch, group_slice_map
|
| 26 |
+
from .recovery import (sparse_group_recover, debias, bootstrap_support,
|
| 27 |
+
active_groups)
|
| 28 |
+
from .drift import procrustes_transport, ResidualDriftMonitor
|
| 29 |
+
from .gate import AnchorGate, GateDecision
|
| 30 |
+
|
| 31 |
+
|
| 32 |
+
@dataclass
|
| 33 |
+
class IterationReport:
|
| 34 |
+
label: str
|
| 35 |
+
pilot_residual: float
|
| 36 |
+
pilot_passed: bool
|
| 37 |
+
nominated_experts: list
|
| 38 |
+
recovered_experts: list
|
| 39 |
+
bootstrap_frequencies: list
|
| 40 |
+
delta_hat_norm: float
|
| 41 |
+
gate_decision: str
|
| 42 |
+
anchor_reports: list
|
| 43 |
+
items_probe: int
|
| 44 |
+
items_sense: int
|
| 45 |
+
items_anchor: int
|
| 46 |
+
drift_alarm: bool
|
| 47 |
+
dense_fallback: bool
|
| 48 |
+
wall_time_s: float
|
| 49 |
+
audit_hash: str = ""
|
| 50 |
+
|
| 51 |
+
def to_json(self):
|
| 52 |
+
return json.dumps(asdict(self), indent=2, default=str)
|
| 53 |
+
|
| 54 |
+
|
| 55 |
+
class SpectraRSILoop:
|
| 56 |
+
def __init__(self, world, cfg, anchor_slices=None):
|
| 57 |
+
self.world = world
|
| 58 |
+
self.cfg = cfg
|
| 59 |
+
self.rng = np.random.default_rng(cfg.seed)
|
| 60 |
+
self.router = Router(cfg.n_experts, seed=cfg.seed + 1)
|
| 61 |
+
self.dictionary = None
|
| 62 |
+
self.prev_Psi = None
|
| 63 |
+
self.drift_monitor = ResidualDriftMonitor(alpha=cfg.drift_alpha)
|
| 64 |
+
# protected anchors: held-out slices, never used in discovery
|
| 65 |
+
if anchor_slices is None:
|
| 66 |
+
anchor_slices = list(self.rng.choice(
|
| 67 |
+
world.n, size=max(4, world.n // 50), replace=False))
|
| 68 |
+
self.anchor_slices = anchor_slices
|
| 69 |
+
os.makedirs(cfg.audit_dir, exist_ok=True)
|
| 70 |
+
self._iteration = 0
|
| 71 |
+
|
| 72 |
+
# ------------------------------------------------------------------ #
|
| 73 |
+
def _ensure_dictionary(self, force=False):
|
| 74 |
+
items = 0
|
| 75 |
+
if self.dictionary is None or force:
|
| 76 |
+
if self.dictionary is not None:
|
| 77 |
+
self.prev_Psi = self.dictionary.Psi.copy()
|
| 78 |
+
self.dictionary = build_dictionary(self.world, self.cfg, rng=self.rng)
|
| 79 |
+
items = (2 * self.cfg.n_experts * self.cfg.rank_per_expert
|
| 80 |
+
* self.cfg.probe_items * self.world.n // self.world.n)
|
| 81 |
+
items = (2 * self.cfg.n_experts * self.cfg.rank_per_expert
|
| 82 |
+
* self.cfg.probe_items)
|
| 83 |
+
if self.prev_Psi is not None:
|
| 84 |
+
procrustes_transport(self.dictionary.Psi, self.prev_Psi)
|
| 85 |
+
return items
|
| 86 |
+
|
| 87 |
+
# ------------------------------------------------------------------ #
|
| 88 |
+
def run_iteration(self, cand) -> IterationReport:
|
| 89 |
+
cfg, world = self.cfg, self.world
|
| 90 |
+
t0 = time.time()
|
| 91 |
+
self._iteration += 1
|
| 92 |
+
drift_alarm = self.drift_monitor.is_alarmed()
|
| 93 |
+
items_probe = self._ensure_dictionary(force=drift_alarm)
|
| 94 |
+
if drift_alarm:
|
| 95 |
+
self.drift_monitor = ResidualDriftMonitor(alpha=cfg.drift_alpha)
|
| 96 |
+
D = self.dictionary
|
| 97 |
+
|
| 98 |
+
# ---- router priors (fallible; clipped; exploration floor) ----
|
| 99 |
+
occ, ent, energy = self.router.occupancy_stats(
|
| 100 |
+
cand.beta, cfg.rank_per_expert)
|
| 101 |
+
priors = expert_priors(occ, ent, energy, clip=cfg.prior_clip)
|
| 102 |
+
gw = group_weights(priors, gamma=cfg.prior_gamma)
|
| 103 |
+
|
| 104 |
+
# ---- trust-region pilot ----
|
| 105 |
+
pilot_res, pilot_ok = pilot_linearity_check(world, cand, D, cfg,
|
| 106 |
+
rng=self.rng)
|
| 107 |
+
if not pilot_ok:
|
| 108 |
+
report = self._dense_fallback_report(cand, pilot_res, t0)
|
| 109 |
+
self._audit(report)
|
| 110 |
+
return report
|
| 111 |
+
|
| 112 |
+
sense_slices = [j for j in range(world.n)
|
| 113 |
+
if j not in set(self.anchor_slices)]
|
| 114 |
+
sigma = world.sigma
|
| 115 |
+
|
| 116 |
+
# ---- stage I: coarse sketch ----
|
| 117 |
+
A1 = design_matrix(cfg.m_coarse, world.n, cfg.row_density, sigma,
|
| 118 |
+
self.rng)
|
| 119 |
+
A1[:, self.anchor_slices] = 0.0 # firewall: no anchor leakage
|
| 120 |
+
y1, items1, slice_means1 = paired_sketch(
|
| 121 |
+
world, cand, A1, sigma, cfg.items_per_row, self.rng)
|
| 122 |
+
M1 = A1 @ D.Psi
|
| 123 |
+
x1 = sparse_group_recover(y1, M1, D.groups, gw,
|
| 124 |
+
cfg.lambda_l1, cfg.lambda_group,
|
| 125 |
+
n_iters=cfg.fista_iters // 2)
|
| 126 |
+
nominated, _ = active_groups(x1, D.groups)
|
| 127 |
+
|
| 128 |
+
# ---- stage II: focused sketch with exploration mass ----
|
| 129 |
+
gmap = group_slice_map(D)
|
| 130 |
+
A2 = design_matrix(cfg.m_focused, world.n, cfg.row_density, sigma,
|
| 131 |
+
self.rng, group_slices=gmap,
|
| 132 |
+
focus_groups=nominated or list(range(cfg.n_experts)),
|
| 133 |
+
explore_fraction=cfg.explore_fraction)
|
| 134 |
+
A2[:, self.anchor_slices] = 0.0
|
| 135 |
+
y2, items2, slice_means2 = paired_sketch(
|
| 136 |
+
world, cand, A2, sigma, cfg.items_per_row, self.rng)
|
| 137 |
+
|
| 138 |
+
A = np.vstack([A1, A2])
|
| 139 |
+
y = np.concatenate([y1, y2])
|
| 140 |
+
M = A @ D.Psi
|
| 141 |
+
|
| 142 |
+
# ---- recovery + debias + bootstrap ----
|
| 143 |
+
x_hat = sparse_group_recover(y, M, D.groups, gw,
|
| 144 |
+
cfg.lambda_l1, cfg.lambda_group,
|
| 145 |
+
n_iters=cfg.fista_iters)
|
| 146 |
+
x_deb, support = debias(y, M, x_hat)
|
| 147 |
+
freq = bootstrap_support(y, M, D.groups, gw, cfg.lambda_l1,
|
| 148 |
+
cfg.lambda_group, reps=cfg.bootstrap_reps,
|
| 149 |
+
rng=self.rng)
|
| 150 |
+
recovered, energies = active_groups(x_deb, D.groups)
|
| 151 |
+
delta_hat = D.Psi @ x_deb
|
| 152 |
+
|
| 153 |
+
# ---- drift monitor: predictive residuals on reused slice means ----
|
| 154 |
+
resid = []
|
| 155 |
+
for j, mhat in {**slice_means1, **slice_means2}.items():
|
| 156 |
+
resid.append((mhat - delta_hat[j]) ** 2 - sigma[j] ** 2
|
| 157 |
+
* 2 * (1 - world.cnr_rho) / cfg.items_per_row)
|
| 158 |
+
self.drift_monitor.update(np.array(resid[:cfg.drift_window]))
|
| 159 |
+
|
| 160 |
+
# ---- protected anytime-valid gate ----
|
| 161 |
+
gate = AnchorGate(self.anchor_slices, cfg.tau_margin, cfg.alpha_risk)
|
| 162 |
+
decision, anchor_reports, items_anchor = gate.run(
|
| 163 |
+
world, cand, batch=cfg.anchor_batch,
|
| 164 |
+
max_items=cfg.max_anchor_items, rng=self.rng)
|
| 165 |
+
|
| 166 |
+
report = IterationReport(
|
| 167 |
+
label=cand.label,
|
| 168 |
+
pilot_residual=float(pilot_res),
|
| 169 |
+
pilot_passed=True,
|
| 170 |
+
nominated_experts=[int(e) for e in nominated],
|
| 171 |
+
recovered_experts=[int(e) for e in recovered],
|
| 172 |
+
bootstrap_frequencies=[float(f) for f in freq],
|
| 173 |
+
delta_hat_norm=float(np.linalg.norm(delta_hat)),
|
| 174 |
+
gate_decision=decision.value,
|
| 175 |
+
anchor_reports=[asdict(r) for r in anchor_reports],
|
| 176 |
+
items_probe=items_probe,
|
| 177 |
+
items_sense=items1 + items2,
|
| 178 |
+
items_anchor=items_anchor,
|
| 179 |
+
drift_alarm=bool(drift_alarm),
|
| 180 |
+
dense_fallback=False,
|
| 181 |
+
wall_time_s=time.time() - t0,
|
| 182 |
+
)
|
| 183 |
+
report.delta_hat = delta_hat # attach for callers
|
| 184 |
+
report.x_hat = x_deb
|
| 185 |
+
self._audit(report)
|
| 186 |
+
return report
|
| 187 |
+
|
| 188 |
+
# ------------------------------------------------------------------ #
|
| 189 |
+
def _dense_fallback_report(self, cand, pilot_res, t0):
|
| 190 |
+
"""Out-of-trust-region candidate: refuse sparse explanation."""
|
| 191 |
+
report = IterationReport(
|
| 192 |
+
label=cand.label, pilot_residual=float(pilot_res),
|
| 193 |
+
pilot_passed=False, nominated_experts=[], recovered_experts=[],
|
| 194 |
+
bootstrap_frequencies=[], delta_hat_norm=0.0,
|
| 195 |
+
gate_decision=GateDecision.QUARANTINE.value, anchor_reports=[],
|
| 196 |
+
items_probe=0, items_sense=0, items_anchor=0,
|
| 197 |
+
drift_alarm=False, dense_fallback=True,
|
| 198 |
+
wall_time_s=time.time() - t0)
|
| 199 |
+
report.delta_hat = None
|
| 200 |
+
report.x_hat = None
|
| 201 |
+
return report
|
| 202 |
+
|
| 203 |
+
# ------------------------------------------------------------------ #
|
| 204 |
+
def _audit(self, report):
|
| 205 |
+
"""Immutable audit record (Sec. 14): hash-chained JSON."""
|
| 206 |
+
payload = report.to_json()
|
| 207 |
+
h = hashlib.sha256(payload.encode()).hexdigest()[:16]
|
| 208 |
+
report.audit_hash = h
|
| 209 |
+
path = os.path.join(self.cfg.audit_dir,
|
| 210 |
+
f"iter_{self._iteration:04d}_{h}.json")
|
| 211 |
+
with open(path, "w") as f:
|
| 212 |
+
f.write(payload)
|
spectra_rsi/metrics.py
ADDED
|
@@ -0,0 +1,30 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Primary metrics (Sec. 12.5)."""
|
| 2 |
+
import numpy as np
|
| 3 |
+
|
| 4 |
+
|
| 5 |
+
def normalized_delta_error(delta_hat, delta_true, eps=1e-9):
|
| 6 |
+
return float(np.linalg.norm(delta_hat - delta_true)
|
| 7 |
+
/ (np.linalg.norm(delta_true) + eps))
|
| 8 |
+
|
| 9 |
+
|
| 10 |
+
def support_f1(pred_support, true_support):
|
| 11 |
+
pred, true = set(pred_support), set(true_support)
|
| 12 |
+
if not pred and not true:
|
| 13 |
+
return 1.0
|
| 14 |
+
tp = len(pred & true)
|
| 15 |
+
prec = tp / len(pred) if pred else 0.0
|
| 16 |
+
rec = tp / len(true) if true else 0.0
|
| 17 |
+
return 0.0 if prec + rec == 0 else 2 * prec * rec / (prec + rec)
|
| 18 |
+
|
| 19 |
+
|
| 20 |
+
def regression_recall(delta_hat, delta_true, tau, weights=None):
|
| 21 |
+
"""Weighted recall of materially regressed slices (Delta c < -tau)."""
|
| 22 |
+
n = len(delta_true)
|
| 23 |
+
w = np.ones(n) if weights is None else np.asarray(weights)
|
| 24 |
+
regressed = delta_true < -tau
|
| 25 |
+
if not regressed.any():
|
| 26 |
+
return 1.0
|
| 27 |
+
detected = delta_hat < -tau / 2 # detection threshold at tau/2
|
| 28 |
+
num = float(np.sum(w * detected * regressed))
|
| 29 |
+
den = float(np.sum(w * regressed))
|
| 30 |
+
return num / den
|
spectra_rsi/probes.py
ADDED
|
@@ -0,0 +1,65 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Counterfactual spectral response dictionary (Sec. 4).
|
| 2 |
+
|
| 3 |
+
Signed central-difference probes on expert directions estimate the
|
| 4 |
+
interventional response matrix H_t; per-expert truncated SVD yields the
|
| 5 |
+
grouped dictionary Psi_t = [U_1, ..., U_E].
|
| 6 |
+
"""
|
| 7 |
+
from dataclasses import dataclass
|
| 8 |
+
import numpy as np
|
| 9 |
+
|
| 10 |
+
|
| 11 |
+
@dataclass
|
| 12 |
+
class Dictionary:
|
| 13 |
+
Psi: np.ndarray # (n, p) column-orthonormal within groups
|
| 14 |
+
groups: list # list of index arrays into columns of Psi
|
| 15 |
+
H: np.ndarray # (n, d) raw fingerprint matrix
|
| 16 |
+
expert_maps: list # per-expert (Sigma V^T) mapping dict coords -> beta coords
|
| 17 |
+
|
| 18 |
+
@property
|
| 19 |
+
def p(self):
|
| 20 |
+
return self.Psi.shape[1]
|
| 21 |
+
|
| 22 |
+
|
| 23 |
+
def build_dictionary(world, cfg, rng=None) -> Dictionary:
|
| 24 |
+
"""Apply +/- h probes for every expert direction and factorize per expert."""
|
| 25 |
+
rng = rng or np.random.default_rng(cfg.seed)
|
| 26 |
+
E, R = cfg.n_experts, cfg.rank_per_expert
|
| 27 |
+
d = E * R
|
| 28 |
+
cols = []
|
| 29 |
+
for idx in range(d):
|
| 30 |
+
direction = np.zeros(d)
|
| 31 |
+
direction[idx] = 1.0
|
| 32 |
+
cols.append(world.probe_delta(direction, cfg.probe_magnitude,
|
| 33 |
+
cfg.probe_items, rng=rng))
|
| 34 |
+
H = np.stack(cols, axis=1) # (n, d)
|
| 35 |
+
|
| 36 |
+
Psi_blocks, groups, expert_maps = [], [], []
|
| 37 |
+
col0 = 0
|
| 38 |
+
for e in range(E):
|
| 39 |
+
He = H[:, e * R:(e + 1) * R] # (n, R)
|
| 40 |
+
U, S, Vt = np.linalg.svd(He, full_matrices=False)
|
| 41 |
+
r = min(R, (S > 1e-8).sum())
|
| 42 |
+
r = max(r, 1)
|
| 43 |
+
Psi_blocks.append(U[:, :r])
|
| 44 |
+
groups.append(np.arange(col0, col0 + r))
|
| 45 |
+
expert_maps.append((S[:r], Vt[:r]))
|
| 46 |
+
col0 += r
|
| 47 |
+
Psi = np.concatenate(Psi_blocks, axis=1)
|
| 48 |
+
return Dictionary(Psi=Psi, groups=groups, H=H, expert_maps=expert_maps)
|
| 49 |
+
|
| 50 |
+
|
| 51 |
+
def pilot_linearity_check(world, cand, dictionary, cfg, rng=None):
|
| 52 |
+
"""Trust-region pilot (Assumption 1): compare probe-superposition
|
| 53 |
+
prediction H beta with a small paired evaluation on random slices.
|
| 54 |
+
Returns (relative_residual, passed)."""
|
| 55 |
+
rng = rng or np.random.default_rng(cfg.seed + 7)
|
| 56 |
+
pred = dictionary.H @ cand.beta
|
| 57 |
+
# test where the superposition prediction claims signal exists: the
|
| 58 |
+
# top-|pred| slices; a linearity violation shows up exactly there.
|
| 59 |
+
k = min(world.n, 40)
|
| 60 |
+
idx = np.argsort(np.abs(pred))[::-1][:k]
|
| 61 |
+
obs = np.array([world.paired_scores(cand, int(j), cfg.pilot_items,
|
| 62 |
+
rng=rng).mean() for j in idx])
|
| 63 |
+
denom = max(np.linalg.norm(obs), np.linalg.norm(pred[idx])) + 1e-9
|
| 64 |
+
rel = np.linalg.norm(obs - pred[idx]) / denom
|
| 65 |
+
return rel, rel <= cfg.trust_region_residual
|
spectra_rsi/recovery.py
ADDED
|
@@ -0,0 +1,88 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Weighted sparse-group recovery (Eq. recovery) via FISTA, plus
|
| 2 |
+
debiased refit and bootstrap support stability.
|
| 3 |
+
|
| 4 |
+
min_x 0.5 ||y - M x||^2 + lambda1 ||x||_1 + lambda_g sum_e w_e ||x_Ge||_2
|
| 5 |
+
"""
|
| 6 |
+
import numpy as np
|
| 7 |
+
|
| 8 |
+
|
| 9 |
+
def _soft(x, t):
|
| 10 |
+
return np.sign(x) * np.maximum(np.abs(x) - t, 0.0)
|
| 11 |
+
|
| 12 |
+
|
| 13 |
+
def _group_soft(x, groups, thresholds):
|
| 14 |
+
out = x.copy()
|
| 15 |
+
for g, t in zip(groups, thresholds):
|
| 16 |
+
nrm = np.linalg.norm(x[g])
|
| 17 |
+
out[g] = 0.0 if nrm <= t else (1 - t / nrm) * x[g]
|
| 18 |
+
return out
|
| 19 |
+
|
| 20 |
+
|
| 21 |
+
def _lipschitz(M, iters=60, rng=None):
|
| 22 |
+
rng = rng or np.random.default_rng(0)
|
| 23 |
+
v = rng.normal(size=M.shape[1])
|
| 24 |
+
v /= np.linalg.norm(v)
|
| 25 |
+
for _ in range(iters):
|
| 26 |
+
v = M.T @ (M @ v)
|
| 27 |
+
nv = np.linalg.norm(v)
|
| 28 |
+
if nv < 1e-18:
|
| 29 |
+
return 1.0
|
| 30 |
+
v /= nv
|
| 31 |
+
return nv
|
| 32 |
+
|
| 33 |
+
|
| 34 |
+
def sparse_group_recover(y, M, groups, group_w, lambda_l1, lambda_group,
|
| 35 |
+
n_iters=600):
|
| 36 |
+
"""FISTA with elementwise + group soft-thresholding prox."""
|
| 37 |
+
p = M.shape[1]
|
| 38 |
+
L = _lipschitz(M) + 1e-9
|
| 39 |
+
step = 1.0 / L
|
| 40 |
+
x = np.zeros(p)
|
| 41 |
+
z = x.copy()
|
| 42 |
+
t = 1.0
|
| 43 |
+
for _ in range(n_iters):
|
| 44 |
+
grad = M.T @ (M @ z - y)
|
| 45 |
+
u = z - step * grad
|
| 46 |
+
u = _soft(u, step * lambda_l1)
|
| 47 |
+
u = _group_soft(u, groups, [step * lambda_group * w for w in group_w])
|
| 48 |
+
t_next = 0.5 * (1 + np.sqrt(1 + 4 * t * t))
|
| 49 |
+
z = u + ((t - 1) / t_next) * (u - x)
|
| 50 |
+
x, t = u, t_next
|
| 51 |
+
return x
|
| 52 |
+
|
| 53 |
+
|
| 54 |
+
def debias(y, M, x_hat, tol=1e-8):
|
| 55 |
+
"""Least-squares refit restricted to the selected support (Sec. 5.1)."""
|
| 56 |
+
support = np.nonzero(np.abs(x_hat) > tol)[0]
|
| 57 |
+
if support.size == 0:
|
| 58 |
+
return x_hat, support
|
| 59 |
+
Ms = M[:, support]
|
| 60 |
+
coef, *_ = np.linalg.lstsq(Ms, y, rcond=None)
|
| 61 |
+
out = np.zeros_like(x_hat)
|
| 62 |
+
out[support] = coef
|
| 63 |
+
return out, support
|
| 64 |
+
|
| 65 |
+
|
| 66 |
+
def bootstrap_support(y, M, groups, group_w, lambda_l1, lambda_group,
|
| 67 |
+
reps=60, n_iters=250, rng=None):
|
| 68 |
+
"""Row-resampling bootstrap: per-group selection frequencies (Sec. 8.1)."""
|
| 69 |
+
rng = rng or np.random.default_rng(0)
|
| 70 |
+
m = len(y)
|
| 71 |
+
freq = np.zeros(len(groups))
|
| 72 |
+
for _ in range(reps):
|
| 73 |
+
idx = rng.integers(0, m, m)
|
| 74 |
+
xb = sparse_group_recover(y[idx], M[idx], groups, group_w,
|
| 75 |
+
lambda_l1, lambda_group, n_iters=n_iters)
|
| 76 |
+
for e, g in enumerate(groups):
|
| 77 |
+
if np.linalg.norm(xb[g]) > 1e-8:
|
| 78 |
+
freq[e] += 1
|
| 79 |
+
return freq / reps
|
| 80 |
+
|
| 81 |
+
|
| 82 |
+
def active_groups(x_hat, groups, rel_tol=0.05):
|
| 83 |
+
"""Groups carrying more than rel_tol of the maximum group energy."""
|
| 84 |
+
energies = np.array([np.linalg.norm(x_hat[g]) for g in groups])
|
| 85 |
+
if energies.max() <= 0:
|
| 86 |
+
return [], energies
|
| 87 |
+
keep = np.nonzero(energies > rel_tol * energies.max())[0]
|
| 88 |
+
return list(keep), energies
|
spectra_rsi/sketch.py
ADDED
|
@@ -0,0 +1,83 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Paired randomized sketches and measurement design (Secs. 3.1, 6).
|
| 2 |
+
|
| 3 |
+
Rows are sparse, signed, variance-normalized task weights. Measurements are
|
| 4 |
+
weighted sums of task-level paired mean differences (realizable composites),
|
| 5 |
+
with Neyman item allocation N_ij ~ |a_ij| sigma_j.
|
| 6 |
+
"""
|
| 7 |
+
import numpy as np
|
| 8 |
+
|
| 9 |
+
|
| 10 |
+
def design_matrix(m, n, row_density, sigma, rng, group_slices=None,
|
| 11 |
+
focus_groups=None, explore_fraction=0.15):
|
| 12 |
+
"""Sparse Rademacher design, variance-normalized (Sec. 6.1).
|
| 13 |
+
|
| 14 |
+
If `focus_groups` is given (stage II), row support concentrates on the
|
| 15 |
+
nominated groups' slices with an exploration floor elsewhere (Sec. 6.2).
|
| 16 |
+
"""
|
| 17 |
+
A = np.zeros((m, n))
|
| 18 |
+
if focus_groups is not None and group_slices is not None:
|
| 19 |
+
focus_mask = np.zeros(n, dtype=bool)
|
| 20 |
+
for g in focus_groups:
|
| 21 |
+
focus_mask[group_slices[g]] = True
|
| 22 |
+
p_focus = row_density
|
| 23 |
+
p_off = row_density * explore_fraction
|
| 24 |
+
probs = np.where(focus_mask, p_focus, p_off)
|
| 25 |
+
else:
|
| 26 |
+
probs = np.full(n, row_density)
|
| 27 |
+
|
| 28 |
+
for i in range(m):
|
| 29 |
+
s = rng.random(n) < probs
|
| 30 |
+
if not s.any():
|
| 31 |
+
s[rng.integers(n)] = True
|
| 32 |
+
xi = rng.choice([-1.0, 1.0], size=n)
|
| 33 |
+
row = np.where(s, xi / np.maximum(sigma, 1e-9), 0.0)
|
| 34 |
+
norm = np.sqrt((row ** 2).sum())
|
| 35 |
+
A[i] = row / max(norm, 1e-12)
|
| 36 |
+
return A
|
| 37 |
+
|
| 38 |
+
|
| 39 |
+
def neyman_allocation(A_row, sigma, budget):
|
| 40 |
+
"""N_ij ~ |a_ij| sigma_j over the row support, integer allocation >= 2."""
|
| 41 |
+
support = np.nonzero(A_row)[0]
|
| 42 |
+
w = np.abs(A_row[support]) * sigma[support]
|
| 43 |
+
w = w / w.sum()
|
| 44 |
+
N = np.maximum(2, np.floor(budget * w).astype(int))
|
| 45 |
+
return support, N
|
| 46 |
+
|
| 47 |
+
|
| 48 |
+
def paired_sketch(world, cand, A, sigma, items_per_row, rng):
|
| 49 |
+
"""Collect y_i = sum_j a_ij * mean paired difference on slice j.
|
| 50 |
+
|
| 51 |
+
Returns (y, total_items, per_slice_means) where per_slice_means caches
|
| 52 |
+
reusable slice estimates for residual monitoring.
|
| 53 |
+
"""
|
| 54 |
+
m = A.shape[0]
|
| 55 |
+
y = np.zeros(m)
|
| 56 |
+
total_items = 0
|
| 57 |
+
slice_sums = {}
|
| 58 |
+
slice_counts = {}
|
| 59 |
+
for i in range(m):
|
| 60 |
+
support, N = neyman_allocation(A[i], sigma, items_per_row)
|
| 61 |
+
acc = 0.0
|
| 62 |
+
for j, nj in zip(support, N):
|
| 63 |
+
diffs = world.paired_scores(cand, int(j), int(nj), rng=rng)
|
| 64 |
+
mean_j = diffs.mean()
|
| 65 |
+
acc += A[i, j] * mean_j
|
| 66 |
+
total_items += int(nj)
|
| 67 |
+
slice_sums[j] = slice_sums.get(j, 0.0) + diffs.sum()
|
| 68 |
+
slice_counts[j] = slice_counts.get(j, 0) + int(nj)
|
| 69 |
+
y[i] = acc
|
| 70 |
+
per_slice = {j: slice_sums[j] / slice_counts[j] for j in slice_sums}
|
| 71 |
+
return y, total_items, per_slice
|
| 72 |
+
|
| 73 |
+
|
| 74 |
+
def group_slice_map(dictionary, top_frac=0.25):
|
| 75 |
+
"""Map each dictionary group to the capability slices it loads on most,
|
| 76 |
+
used to focus stage-II measurement rows."""
|
| 77 |
+
n = dictionary.Psi.shape[0]
|
| 78 |
+
out = []
|
| 79 |
+
k = max(1, int(n * top_frac / max(1, len(dictionary.groups))))
|
| 80 |
+
for g in dictionary.groups:
|
| 81 |
+
load = np.abs(dictionary.Psi[:, g]).sum(axis=1)
|
| 82 |
+
out.append(np.argsort(load)[::-1][:max(k, 8)])
|
| 83 |
+
return out
|
spectra_rsi/world.py
ADDED
|
@@ -0,0 +1,171 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Synthetic ground-truth world.
|
| 2 |
+
|
| 3 |
+
Implements the paper's problem formulation (Sec. 3) with a known
|
| 4 |
+
capability-response operator so that reconstruction, attribution, and gating
|
| 5 |
+
can be validated against ground truth.
|
| 6 |
+
|
| 7 |
+
The world exposes exactly the interface a real evaluation harness would:
|
| 8 |
+
paired scoring of (base, candidate) on sampled items with common decoding
|
| 9 |
+
randomness. Everything downstream (probes, sketches, gate) only uses this
|
| 10 |
+
interface, so a real LLM harness can be dropped in behind it.
|
| 11 |
+
"""
|
| 12 |
+
from dataclasses import dataclass, field
|
| 13 |
+
import numpy as np
|
| 14 |
+
|
| 15 |
+
|
| 16 |
+
@dataclass
|
| 17 |
+
class CandidateUpdate:
|
| 18 |
+
"""A frozen candidate with optional off-dictionary capability residual."""
|
| 19 |
+
beta: np.ndarray # (d,) coefficients over stacked expert directions
|
| 20 |
+
label: str = "candidate"
|
| 21 |
+
# ground-truth bookkeeping (for benchmarking only; never read by the loop)
|
| 22 |
+
true_support: set = field(default_factory=set) # expert indices actually perturbed
|
| 23 |
+
# Optional capability-space component not represented by expert coefficients.
|
| 24 |
+
# Used only for synthetic out-of-dictionary benchmark interventions.
|
| 25 |
+
residual_delta: np.ndarray | None = None
|
| 26 |
+
|
| 27 |
+
|
| 28 |
+
class SyntheticWorld:
|
| 29 |
+
"""Ground truth: Delta c = H* beta + nonlinearity + heteroscedastic noise.
|
| 30 |
+
|
| 31 |
+
- H* (n x d): true interventional response matrix, block-structured so
|
| 32 |
+
that each expert's directions load mainly on a band of slices (this is
|
| 33 |
+
the group-compressibility assumption made explicit and controllable).
|
| 34 |
+
- Paired scoring with common random numbers: base and candidate item
|
| 35 |
+
scores share a noise component with correlation `cnr_rho`, so paired
|
| 36 |
+
differences have variance reduced by (1 - rho) relative to independent
|
| 37 |
+
evaluation (Prop. 1 in the paper).
|
| 38 |
+
"""
|
| 39 |
+
|
| 40 |
+
def __init__(self, n_slices, n_experts, rank_per_expert, seed=0,
|
| 41 |
+
band_overlap=0.15, offband_leak=0.05, nonlin=0.0,
|
| 42 |
+
noise_scale=0.25, cnr_rho=0.85):
|
| 43 |
+
self.n = n_slices
|
| 44 |
+
self.E = n_experts
|
| 45 |
+
self.R = rank_per_expert
|
| 46 |
+
self.d = n_experts * rank_per_expert
|
| 47 |
+
self.nonlin = nonlin
|
| 48 |
+
self.cnr_rho = cnr_rho
|
| 49 |
+
rng = np.random.default_rng(seed)
|
| 50 |
+
self._rng = rng
|
| 51 |
+
|
| 52 |
+
# block-structured true response matrix
|
| 53 |
+
band = n_slices // n_experts
|
| 54 |
+
H = np.zeros((self.n, self.d))
|
| 55 |
+
for e in range(n_experts):
|
| 56 |
+
lo, hi = e * band, (e + 1) * band
|
| 57 |
+
# widen band with overlap
|
| 58 |
+
pad = int(band * band_overlap)
|
| 59 |
+
lo2, hi2 = max(0, lo - pad), min(self.n, hi + pad)
|
| 60 |
+
for r in range(rank_per_expert):
|
| 61 |
+
col = e * rank_per_expert + r
|
| 62 |
+
v = np.zeros(self.n)
|
| 63 |
+
# in-band responses are positive: pushing an expert "up"
|
| 64 |
+
# improves the capabilities it owns (sign semantics), while
|
| 65 |
+
# off-band leakage stays signed and diffuse.
|
| 66 |
+
v[lo2:hi2] = np.abs(rng.normal(0, 1, hi2 - lo2))
|
| 67 |
+
v += offband_leak * rng.normal(0, 1, self.n)
|
| 68 |
+
H[:, col] = v / np.linalg.norm(v)
|
| 69 |
+
self.H_true = H
|
| 70 |
+
|
| 71 |
+
# per-slice score noise scale (heteroscedastic)
|
| 72 |
+
self.sigma = noise_scale * (0.5 + rng.random(self.n))
|
| 73 |
+
|
| 74 |
+
# current deployed parameter state in expert-coefficient coordinates
|
| 75 |
+
self.theta_offset = np.zeros(self.d)
|
| 76 |
+
|
| 77 |
+
# ------------------------------------------------------------------ #
|
| 78 |
+
# ground truth (benchmark use only)
|
| 79 |
+
# ------------------------------------------------------------------ #
|
| 80 |
+
def true_delta(self, cand: CandidateUpdate) -> np.ndarray:
|
| 81 |
+
lin = self.H_true @ cand.beta
|
| 82 |
+
if self.nonlin > 0:
|
| 83 |
+
# quadratic saturation: response bends for large updates
|
| 84 |
+
lin = lin - self.nonlin * np.sign(lin) * lin ** 2
|
| 85 |
+
if cand.residual_delta is not None:
|
| 86 |
+
residual = np.asarray(cand.residual_delta, dtype=float)
|
| 87 |
+
if residual.shape != (self.n,):
|
| 88 |
+
raise ValueError(
|
| 89 |
+
f"residual_delta must have shape ({self.n},), "
|
| 90 |
+
f"got {residual.shape}"
|
| 91 |
+
)
|
| 92 |
+
lin = lin + residual
|
| 93 |
+
return lin
|
| 94 |
+
|
| 95 |
+
# ------------------------------------------------------------------ #
|
| 96 |
+
# evaluation interface (what a real harness would implement)
|
| 97 |
+
# ------------------------------------------------------------------ #
|
| 98 |
+
def paired_scores(self, cand: CandidateUpdate, slice_idx: int, n_items: int,
|
| 99 |
+
rng=None) -> np.ndarray:
|
| 100 |
+
"""Paired per-item score differences d_l for one capability slice,
|
| 101 |
+
evaluated with common random numbers (shared item + decoding noise)."""
|
| 102 |
+
rng = rng or self._rng
|
| 103 |
+
mu = self.true_delta(cand)[slice_idx]
|
| 104 |
+
s = self.sigma[slice_idx]
|
| 105 |
+
shared = rng.normal(0, s, n_items)
|
| 106 |
+
e_base = np.sqrt(1 - self.cnr_rho) * rng.normal(0, s, n_items)
|
| 107 |
+
e_cand = np.sqrt(1 - self.cnr_rho) * rng.normal(0, s, n_items)
|
| 108 |
+
base = shared + e_base
|
| 109 |
+
candv = mu + shared + e_cand
|
| 110 |
+
return candv - base # variance ~ 2 (1 - rho) s^2
|
| 111 |
+
|
| 112 |
+
def probe_delta(self, direction: np.ndarray, magnitude: float,
|
| 113 |
+
n_items_per_slice: int, rng=None) -> np.ndarray:
|
| 114 |
+
"""Central-difference fingerprint (Eq. fingerprint): evaluate
|
| 115 |
+
+h and -h probes on ALL slices with finite item budget."""
|
| 116 |
+
rng = rng or self._rng
|
| 117 |
+
plus = CandidateUpdate(beta=magnitude * direction)
|
| 118 |
+
minus = CandidateUpdate(beta=-magnitude * direction)
|
| 119 |
+
mu_p = self.true_delta(plus)
|
| 120 |
+
mu_m = self.true_delta(minus)
|
| 121 |
+
noise = (self.sigma / np.sqrt(max(1, n_items_per_slice))
|
| 122 |
+
* np.sqrt(2 * (1 - self.cnr_rho)))
|
| 123 |
+
obs_p = mu_p + noise * rng.normal(0, 1, self.n)
|
| 124 |
+
obs_m = mu_m + noise * rng.normal(0, 1, self.n)
|
| 125 |
+
return (obs_p - obs_m) / (2 * magnitude)
|
| 126 |
+
|
| 127 |
+
# ------------------------------------------------------------------ #
|
| 128 |
+
# candidate factory helpers (benchmark interventions, Sec. 12.2)
|
| 129 |
+
# ------------------------------------------------------------------ #
|
| 130 |
+
def make_candidate(self, kind="single_gain", experts=None, scale=0.5,
|
| 131 |
+
rng=None) -> CandidateUpdate:
|
| 132 |
+
rng = rng or self._rng
|
| 133 |
+
beta = np.zeros(self.d)
|
| 134 |
+
R = self.R
|
| 135 |
+
if experts is None:
|
| 136 |
+
experts = [int(rng.integers(self.E))]
|
| 137 |
+
if kind == "single_gain":
|
| 138 |
+
for e in experts:
|
| 139 |
+
beta[e * R:(e + 1) * R] = scale * np.abs(rng.normal(0.8, 0.2, R))
|
| 140 |
+
elif kind == "single_regression":
|
| 141 |
+
for e in experts:
|
| 142 |
+
beta[e * R:(e + 1) * R] = -scale * np.abs(rng.normal(0.8, 0.2, R))
|
| 143 |
+
elif kind == "canceling_mixture":
|
| 144 |
+
e1, e2 = experts[0], experts[-1]
|
| 145 |
+
beta[e1 * R:(e1 + 1) * R] = scale
|
| 146 |
+
beta[e2 * R:(e2 + 1) * R] = -scale
|
| 147 |
+
experts = [e1, e2]
|
| 148 |
+
elif kind == "broad_noncompressible":
|
| 149 |
+
beta = scale * rng.normal(0, 0.3, self.d)
|
| 150 |
+
experts = list(range(self.E))
|
| 151 |
+
elif kind == "off_dictionary":
|
| 152 |
+
# Construct a capability-space change orthogonal to the true
|
| 153 |
+
# expert-response span. No expert coefficient vector can
|
| 154 |
+
# represent this residual.
|
| 155 |
+
raw = rng.normal(0, 1, self.n)
|
| 156 |
+
coef, *_ = np.linalg.lstsq(self.H_true, raw, rcond=None)
|
| 157 |
+
residual = raw - self.H_true @ coef
|
| 158 |
+
norm = np.linalg.norm(residual)
|
| 159 |
+
if norm <= 1e-12:
|
| 160 |
+
raise RuntimeError("Failed to construct off-dictionary residual")
|
| 161 |
+
residual = scale * residual / norm
|
| 162 |
+
experts = []
|
| 163 |
+
return CandidateUpdate(
|
| 164 |
+
beta=beta,
|
| 165 |
+
label=kind,
|
| 166 |
+
true_support=set(),
|
| 167 |
+
residual_delta=residual,
|
| 168 |
+
)
|
| 169 |
+
else:
|
| 170 |
+
raise ValueError(kind)
|
| 171 |
+
return CandidateUpdate(beta=beta, label=kind, true_support=set(experts))
|
tests/__init__.py
ADDED
|
File without changes
|
tests/test_gate.py
ADDED
|
@@ -0,0 +1,70 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""E-process gate: type-I control under adversarial optional stopping,
|
| 2 |
+
and power under genuine non-inferiority."""
|
| 3 |
+
import numpy as np
|
| 4 |
+
from spectra_rsi.gate import EProcess
|
| 5 |
+
|
| 6 |
+
|
| 7 |
+
def test_type1_error_under_optional_stopping():
|
| 8 |
+
"""Adversary peeks after every observation and stops at the most
|
| 9 |
+
favorable moment. False rejection of a TRUE null (mean == -tau, i.e.
|
| 10 |
+
exactly materially regressed) must stay below alpha."""
|
| 11 |
+
alpha, tau = 0.05, 0.05
|
| 12 |
+
rng = np.random.default_rng(42)
|
| 13 |
+
n_sims, horizon = 400, 2000
|
| 14 |
+
false_rejects = 0
|
| 15 |
+
for _ in range(n_sims):
|
| 16 |
+
ep = EProcess(tau=tau, alpha=alpha, bound=1.0)
|
| 17 |
+
rejected = False
|
| 18 |
+
for _ in range(horizon // 50):
|
| 19 |
+
obs = np.clip(rng.normal(-tau, 0.3, 50), -1, 1) # H0 boundary
|
| 20 |
+
ep.update(obs)
|
| 21 |
+
if ep.rejects_null: # adversarial peeking
|
| 22 |
+
rejected = True
|
| 23 |
+
break
|
| 24 |
+
false_rejects += rejected
|
| 25 |
+
rate = false_rejects / n_sims
|
| 26 |
+
assert rate <= alpha + 2 * np.sqrt(alpha * (1 - alpha) / n_sims), rate
|
| 27 |
+
|
| 28 |
+
|
| 29 |
+
def test_power_under_true_improvement():
|
| 30 |
+
"""A genuinely improved anchor should be certified with high probability."""
|
| 31 |
+
alpha, tau = 0.05, 0.02
|
| 32 |
+
rng = np.random.default_rng(7)
|
| 33 |
+
certified = 0
|
| 34 |
+
n_sims = 100
|
| 35 |
+
for _ in range(n_sims):
|
| 36 |
+
ep = EProcess(tau=tau, alpha=alpha, bound=1.0)
|
| 37 |
+
for _ in range(60):
|
| 38 |
+
obs = np.clip(rng.normal(+0.08, 0.25, 50), -1, 1)
|
| 39 |
+
ep.update(obs)
|
| 40 |
+
if ep.rejects_null:
|
| 41 |
+
break
|
| 42 |
+
certified += ep.rejects_null
|
| 43 |
+
assert certified / n_sims > 0.9
|
| 44 |
+
|
| 45 |
+
|
| 46 |
+
def test_anchor_gate_never_exceeds_global_budget():
|
| 47 |
+
"""AnchorGate must never consume more than max_items globally."""
|
| 48 |
+
from spectra_rsi.gate import AnchorGate
|
| 49 |
+
|
| 50 |
+
class DummyWorld:
|
| 51 |
+
def paired_scores(self, cand, j, batch, rng=None):
|
| 52 |
+
# Keep anchors unresolved long enough to exhaust the budget.
|
| 53 |
+
return np.full(batch, -0.02)
|
| 54 |
+
|
| 55 |
+
gate = AnchorGate(
|
| 56 |
+
anchor_slices=[0, 1, 2, 3],
|
| 57 |
+
tau=0.02,
|
| 58 |
+
alpha_total=0.05,
|
| 59 |
+
)
|
| 60 |
+
|
| 61 |
+
_, _, used = gate.run(
|
| 62 |
+
DummyWorld(),
|
| 63 |
+
cand=None,
|
| 64 |
+
batch=100,
|
| 65 |
+
max_items=250,
|
| 66 |
+
rng=np.random.default_rng(123),
|
| 67 |
+
)
|
| 68 |
+
|
| 69 |
+
assert used <= 250
|
| 70 |
+
assert used == 200
|
tests/test_loop.py
ADDED
|
@@ -0,0 +1,114 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""End-to-end loop: attribution, gating, and trust-region fallback."""
|
| 2 |
+
import numpy as np
|
| 3 |
+
import pytest
|
| 4 |
+
from spectra_rsi import SyntheticWorld, SpectraConfig, SpectraRSILoop
|
| 5 |
+
from spectra_rsi.metrics import support_f1, normalized_delta_error
|
| 6 |
+
|
| 7 |
+
|
| 8 |
+
@pytest.fixture(scope="module")
|
| 9 |
+
def setup(tmp_path_factory):
|
| 10 |
+
cfg = SpectraConfig(n_slices=200, n_experts=8, rank_per_expert=2,
|
| 11 |
+
m_coarse=40, m_focused=60, items_per_row=150,
|
| 12 |
+
bootstrap_reps=20, fista_iters=400,
|
| 13 |
+
audit_dir=str(tmp_path_factory.mktemp("audit")))
|
| 14 |
+
world = SyntheticWorld(cfg.n_slices, cfg.n_experts, cfg.rank_per_expert,
|
| 15 |
+
seed=11)
|
| 16 |
+
loop = SpectraRSILoop(world, cfg)
|
| 17 |
+
return world, cfg, loop
|
| 18 |
+
|
| 19 |
+
|
| 20 |
+
def test_single_expert_attribution_and_acceptance(setup):
|
| 21 |
+
world, cfg, loop = setup
|
| 22 |
+
cand = world.make_candidate("single_gain", experts=[3], scale=0.4,
|
| 23 |
+
rng=np.random.default_rng(2))
|
| 24 |
+
rep = loop.run_iteration(cand)
|
| 25 |
+
assert rep.pilot_passed
|
| 26 |
+
assert 3 in rep.recovered_experts
|
| 27 |
+
f1 = support_f1(rep.recovered_experts, cand.true_support)
|
| 28 |
+
assert f1 >= 0.5
|
| 29 |
+
err = normalized_delta_error(rep.delta_hat, world.true_delta(cand))
|
| 30 |
+
assert err < 0.8 # smoke-test budget; tightens with more items
|
| 31 |
+
assert rep.gate_decision == "accept"
|
| 32 |
+
|
| 33 |
+
|
| 34 |
+
def test_regression_is_not_accepted(setup):
|
| 35 |
+
"""Construct a candidate that provably regresses a protected anchor:
|
| 36 |
+
beta chosen opposite to the anchor's true response row, so
|
| 37 |
+
Delta c(anchor) = -scale * ||H*(anchor,:)||^2 < -tau."""
|
| 38 |
+
from spectra_rsi.world import CandidateUpdate
|
| 39 |
+
world, cfg, loop = setup
|
| 40 |
+
anchor = int(loop.anchor_slices[0])
|
| 41 |
+
row = world.H_true[anchor] # ground truth (test only)
|
| 42 |
+
scale = 3.0 * cfg.tau_margin / max(np.dot(row, row), 1e-9)
|
| 43 |
+
beta = -scale * row
|
| 44 |
+
true_drop = world.true_delta(CandidateUpdate(beta=beta))[anchor]
|
| 45 |
+
assert true_drop < -cfg.tau_margin # sanity: really regressed
|
| 46 |
+
cand = CandidateUpdate(beta=beta, label="anchor_regression")
|
| 47 |
+
rep = loop.run_iteration(cand)
|
| 48 |
+
if rep.pilot_passed:
|
| 49 |
+
assert rep.gate_decision in ("rollback", "quarantine")
|
| 50 |
+
else:
|
| 51 |
+
assert rep.dense_fallback
|
| 52 |
+
|
| 53 |
+
|
| 54 |
+
def test_out_of_trust_region_falls_back(setup):
|
| 55 |
+
world, cfg, loop = setup
|
| 56 |
+
nl_world = SyntheticWorld(cfg.n_slices, cfg.n_experts,
|
| 57 |
+
cfg.rank_per_expert, seed=13, nonlin=0.9)
|
| 58 |
+
nl_loop = SpectraRSILoop(nl_world, cfg)
|
| 59 |
+
cand = nl_world.make_candidate("single_gain", experts=[1], scale=5.0,
|
| 60 |
+
rng=np.random.default_rng(4))
|
| 61 |
+
rep = nl_loop.run_iteration(cand)
|
| 62 |
+
assert (not rep.pilot_passed and rep.dense_fallback) or rep.pilot_passed
|
| 63 |
+
|
| 64 |
+
|
| 65 |
+
def test_audit_records_written(setup):
|
| 66 |
+
import glob, os
|
| 67 |
+
world, cfg, loop = setup
|
| 68 |
+
files = glob.glob(os.path.join(cfg.audit_dir, "iter_*.json"))
|
| 69 |
+
assert len(files) >= 1
|
| 70 |
+
|
| 71 |
+
|
| 72 |
+
def test_off_dictionary_candidate_triggers_dense_fallback():
|
| 73 |
+
"""A capability change outside the expert-response span must fall back."""
|
| 74 |
+
cfg = SpectraConfig(
|
| 75 |
+
n_slices=256,
|
| 76 |
+
n_experts=8,
|
| 77 |
+
rank_per_expert=2,
|
| 78 |
+
pilot_items=120,
|
| 79 |
+
seed=7,
|
| 80 |
+
)
|
| 81 |
+
|
| 82 |
+
world = SyntheticWorld(
|
| 83 |
+
cfg.n_slices,
|
| 84 |
+
cfg.n_experts,
|
| 85 |
+
cfg.rank_per_expert,
|
| 86 |
+
seed=cfg.seed,
|
| 87 |
+
offband_leak=0.01,
|
| 88 |
+
)
|
| 89 |
+
|
| 90 |
+
rng = np.random.default_rng(99)
|
| 91 |
+
cand = world.make_candidate(
|
| 92 |
+
"off_dictionary",
|
| 93 |
+
scale=0.4,
|
| 94 |
+
rng=rng,
|
| 95 |
+
)
|
| 96 |
+
|
| 97 |
+
# Verify the synthetic residual is genuinely outside span(H_true).
|
| 98 |
+
residual = cand.residual_delta
|
| 99 |
+
coef, *_ = np.linalg.lstsq(world.H_true, residual, rcond=None)
|
| 100 |
+
projection = world.H_true @ coef
|
| 101 |
+
relative_projection = (
|
| 102 |
+
np.linalg.norm(projection) / np.linalg.norm(residual)
|
| 103 |
+
)
|
| 104 |
+
|
| 105 |
+
assert relative_projection < 1e-10
|
| 106 |
+
assert cand.true_support == set()
|
| 107 |
+
|
| 108 |
+
# The trust-region pilot should reject the sparse model and fall back.
|
| 109 |
+
rep = SpectraRSILoop(world, cfg).run_iteration(cand)
|
| 110 |
+
|
| 111 |
+
assert rep.pilot_residual > cfg.trust_region_residual
|
| 112 |
+
assert rep.dense_fallback is True
|
| 113 |
+
assert rep.items_sense == 0
|
| 114 |
+
assert rep.items_anchor == 0
|
tests/test_recovery.py
ADDED
|
@@ -0,0 +1,20 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Sparse-group recovery: exact-support recovery on clean synthetic data."""
|
| 2 |
+
import numpy as np
|
| 3 |
+
from spectra_rsi.recovery import sparse_group_recover, debias, active_groups
|
| 4 |
+
|
| 5 |
+
|
| 6 |
+
def test_group_recovery_identifies_support():
|
| 7 |
+
rng = np.random.default_rng(0)
|
| 8 |
+
p, n_groups = 32, 8
|
| 9 |
+
groups = [np.arange(i * 4, (i + 1) * 4) for i in range(n_groups)]
|
| 10 |
+
x_true = np.zeros(p)
|
| 11 |
+
x_true[groups[2]] = rng.normal(0, 1, 4)
|
| 12 |
+
x_true[groups[5]] = rng.normal(0, 1, 4)
|
| 13 |
+
M = rng.normal(0, 1 / np.sqrt(60), (60, p))
|
| 14 |
+
y = M @ x_true + 0.01 * rng.normal(size=60)
|
| 15 |
+
gw = np.ones(n_groups)
|
| 16 |
+
x_hat = sparse_group_recover(y, M, groups, gw, 1e-3, 5e-3, n_iters=800)
|
| 17 |
+
x_deb, _ = debias(y, M, x_hat)
|
| 18 |
+
act, _ = active_groups(x_deb, groups, rel_tol=0.1)
|
| 19 |
+
assert set(act) == {2, 5}
|
| 20 |
+
assert np.linalg.norm(x_deb - x_true) / np.linalg.norm(x_true) < 0.1
|
tests/test_sketch.py
ADDED
|
@@ -0,0 +1,39 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Paired sketch: unbiasedness and CRN variance reduction (Prop. 1)."""
|
| 2 |
+
import numpy as np
|
| 3 |
+
import pytest
|
| 4 |
+
from spectra_rsi import SyntheticWorld, SpectraConfig
|
| 5 |
+
from spectra_rsi.sketch import design_matrix, paired_sketch
|
| 6 |
+
|
| 7 |
+
|
| 8 |
+
@pytest.fixture
|
| 9 |
+
def world():
|
| 10 |
+
return SyntheticWorld(n_slices=100, n_experts=4, rank_per_expert=2, seed=3)
|
| 11 |
+
|
| 12 |
+
|
| 13 |
+
def test_paired_sketch_unbiased(world):
|
| 14 |
+
rng = np.random.default_rng(0)
|
| 15 |
+
cand = world.make_candidate("single_gain", experts=[1], scale=0.5, rng=rng)
|
| 16 |
+
A = design_matrix(20, world.n, 0.2, world.sigma, rng)
|
| 17 |
+
true = A @ world.true_delta(cand)
|
| 18 |
+
reps = 40
|
| 19 |
+
est = np.zeros((reps, 20))
|
| 20 |
+
for r in range(reps):
|
| 21 |
+
y, _, _ = paired_sketch(world, cand, A, world.sigma, 150, rng)
|
| 22 |
+
est[r] = y
|
| 23 |
+
bias = np.abs(est.mean(axis=0) - true)
|
| 24 |
+
se = est.std(axis=0) / np.sqrt(reps)
|
| 25 |
+
# bias within 4 standard errors on at least 95% of rows
|
| 26 |
+
assert np.mean(bias < 4 * se + 1e-3) > 0.9
|
| 27 |
+
|
| 28 |
+
|
| 29 |
+
def test_crn_variance_reduction():
|
| 30 |
+
"""Common random numbers must shrink paired-difference variance."""
|
| 31 |
+
w_crn = SyntheticWorld(100, 4, 2, seed=5, cnr_rho=0.9)
|
| 32 |
+
w_ind = SyntheticWorld(100, 4, 2, seed=5, cnr_rho=0.0)
|
| 33 |
+
rng = np.random.default_rng(1)
|
| 34 |
+
cand_c = w_crn.make_candidate("single_gain", experts=[0], rng=rng)
|
| 35 |
+
cand_i = w_ind.make_candidate("single_gain", experts=[0],
|
| 36 |
+
rng=np.random.default_rng(1))
|
| 37 |
+
var_crn = np.var(w_crn.paired_scores(cand_c, 3, 5000))
|
| 38 |
+
var_ind = np.var(w_ind.paired_scores(cand_i, 3, 5000))
|
| 39 |
+
assert var_crn < 0.5 * var_ind
|