File size: 1,732 Bytes
294f8ea
2d5c26a
294f8ea
 
 
 
2d5c26a
294f8ea
2d5c26a
294f8ea
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
# OpenSysOne results

Training and evaluation completed on 17 September 2026. The selected Qwen3-4B
model uses rank-8 adapters and a scalar decision head. Its weights are Spark B
step 1,500, retained unchanged at expanded branch step 0. Selection used validation
only; temperature was fitted on a separate 510-decision calibration split.

| Reserved evaluation | Decisions | Selected accuracy | Pretrained verifier |
| --- | ---: | ---: | ---: |
| Original four-family test | 2,042 | **92.90%** | 84.48% |
| Social IQA family holdout | 768 | **72.92%** | 70.31% |

Paired 95% bootstrap accuracy gains are **+8.42 pp [6.85, 9.89]** on the test and
**+2.60 pp [0.13, 5.34]** on the holdout. The latter covers one untrained fine-tuning
family; pretrained exposure is unknown. These are decision-scoring results, not
general-intelligence scores or a measured comparison with hosted Jev.

On the separate matched 320-decision profile, selected/base-verifier/joint-label
accuracy was **89.06% / 80.94% / 86.25%**. Warm four-choice latency with a 768-token
state was **3.710 / 3.177 / 0.818 seconds** on the same idle Spark in FP32. The
selected scorer was 11–17% slower than the per-option verifier across the measured
workloads; shared-prefix reuse and merged adapters remain future experiments.

The [complete report](results/report.md) includes
calibration, per-family results, uncertainty, all timing cells, CSV/JSON data and
charts. The [profiling protocol](docs/research/profiling-protocol.md) explains scope;
the [full results history](docs/operations/results-history.md) retains earlier
smoke experiments, failures and raw evidence links. Operational provenance and
remaining service controls are in [HANDOVER.md](HANDOVER.md).