# OpenSysOne results Training and evaluation completed on 17 September 2026. The selected Qwen3-4B model uses rank-8 adapters and a scalar decision head. Its weights are Spark B step 1,500, retained unchanged at expanded branch step 0. Selection used validation only; temperature was fitted on a separate 510-decision calibration split. | Reserved evaluation | Decisions | Selected accuracy | Pretrained verifier | | --- | ---: | ---: | ---: | | Original four-family test | 2,042 | **92.90%** | 84.48% | | Social IQA family holdout | 768 | **72.92%** | 70.31% | Paired 95% bootstrap accuracy gains are **+8.42 pp [6.85, 9.89]** on the test and **+2.60 pp [0.13, 5.34]** on the holdout. The latter covers one untrained fine-tuning family; pretrained exposure is unknown. These are decision-scoring results, not general-intelligence scores or a measured comparison with hosted Jev. On the separate matched 320-decision profile, selected/base-verifier/joint-label accuracy was **89.06% / 80.94% / 86.25%**. Warm four-choice latency with a 768-token state was **3.710 / 3.177 / 0.818 seconds** on the same idle Spark in FP32. The selected scorer was 11–17% slower than the per-option verifier across the measured workloads; shared-prefix reuse and merged adapters remain future experiments. The [complete report](results/report.md) includes calibration, per-family results, uncertainty, all timing cells, CSV/JSON data and charts. The [profiling protocol](docs/research/profiling-protocol.md) explains scope; the [full results history](docs/operations/results-history.md) retains earlier smoke experiments, failures and raw evidence links. Operational provenance and remaining service controls are in [HANDOVER.md](HANDOVER.md).