# SPECTRA-RSI Result Contract Every benchmark invocation writes one self-contained JSON document. Raw benchmark JSON files are the primary experimental records and should not be hand-edited. ## Core result fields Each result records the benchmark identifier, candidate type, complete configuration, system metadata, wall time, dense-fallback status, pilot residual, and probe/sensing/anchor item counts. When structured recovery completes, it also records the gate decision, support F1, normalized delta error, regression recall, true expert support, nominated experts, recovered experts, and bootstrap selection frequencies. A dense fallback can legitimately omit recovery and gate fields because structured sensing and protected-anchor evaluation were not executed. ## Reproducibility configuration The `config` object records `n_slices`, `n_experts`, `rank`, `m_coarse`, `m_focused`, `items_per_row`, `bootstrap_reps`, `lambda_l1`, `lambda_group`, `seed`, `candidate_seed`, `candidate`, `scale`, and `output`. Do not compare benchmark results without checking these settings. ## System metadata The `system` object may record hostname, platform, CUDA visibility, PyTorch version/status, GPU count and names, NVIDIA driver, and the mechanism used for GPU detection. PyTorch is not required by the current NumPy benchmark. When PyTorch is unavailable on an NVIDIA system, the runner attempts GPU detection with `nvidia-smi`. A `torch_probe_error` such as `No module named torch` therefore does not by itself indicate benchmark failure. ## Canonical results `results/validated_results.txt` defines the canonical validated release set. Generate the validated leaderboard with `python benchmarks/aggregate_results.py results --manifest results/validated_results.txt --out results/validated_leaderboard.csv`. Other JSON files under `results/` may be exploratory, diagnostic, pre-fix, or parameter-tuning runs. Preserve them for provenance, but do not present them as canonical release results unless deliberately added to the validated manifest. ## Contributor submissions Preserve the raw JSON generated by the benchmark. Also report execution details that cannot be detected automatically, especially accelerator model/count, CPU, RAM, accelerator memory when known, driver/runtime versions, execution environment, command used, and any backend modifications. Do not submit only screenshots or manually transcribed metrics. Failed runs, dense fallbacks, false accepts, unstable recovery, memory limits, and performance regressions are valid benchmark evidence and should be retained rather than discarded.