File size: 1,732 Bytes
294f8ea 2d5c26a 294f8ea 2d5c26a 294f8ea 2d5c26a 294f8ea | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 | # OpenSysOne results
Training and evaluation completed on 17 September 2026. The selected Qwen3-4B
model uses rank-8 adapters and a scalar decision head. Its weights are Spark B
step 1,500, retained unchanged at expanded branch step 0. Selection used validation
only; temperature was fitted on a separate 510-decision calibration split.
| Reserved evaluation | Decisions | Selected accuracy | Pretrained verifier |
| --- | ---: | ---: | ---: |
| Original four-family test | 2,042 | **92.90%** | 84.48% |
| Social IQA family holdout | 768 | **72.92%** | 70.31% |
Paired 95% bootstrap accuracy gains are **+8.42 pp [6.85, 9.89]** on the test and
**+2.60 pp [0.13, 5.34]** on the holdout. The latter covers one untrained fine-tuning
family; pretrained exposure is unknown. These are decision-scoring results, not
general-intelligence scores or a measured comparison with hosted Jev.
On the separate matched 320-decision profile, selected/base-verifier/joint-label
accuracy was **89.06% / 80.94% / 86.25%**. Warm four-choice latency with a 768-token
state was **3.710 / 3.177 / 0.818 seconds** on the same idle Spark in FP32. The
selected scorer was 11–17% slower than the per-option verifier across the measured
workloads; shared-prefix reuse and merged adapters remain future experiments.
The [complete report](results/report.md) includes
calibration, per-family results, uncertainty, all timing cells, CSV/JSON data and
charts. The [profiling protocol](docs/research/profiling-protocol.md) explains scope;
the [full results history](docs/operations/results-history.md) retains earlier
smoke experiments, failures and raw evidence links. Operational provenance and
remaining service controls are in [HANDOVER.md](HANDOVER.md).
|