opensysone / source /RESULTS.md
andyshu's picture
Organize verified OpenSysOne publication payload
294f8ea verified
|
Raw History Blame Contribute Delete
1.73 kB

OpenSysOne results

Training and evaluation completed on 17 September 2026. The selected Qwen3-4B model uses rank-8 adapters and a scalar decision head. Its weights are Spark B step 1,500, retained unchanged at expanded branch step 0. Selection used validation only; temperature was fitted on a separate 510-decision calibration split.

Reserved evaluation Decisions Selected accuracy Pretrained verifier
Original four-family test 2,042 92.90% 84.48%
Social IQA family holdout 768 72.92% 70.31%

Paired 95% bootstrap accuracy gains are +8.42 pp [6.85, 9.89] on the test and +2.60 pp [0.13, 5.34] on the holdout. The latter covers one untrained fine-tuning family; pretrained exposure is unknown. These are decision-scoring results, not general-intelligence scores or a measured comparison with hosted Jev.

On the separate matched 320-decision profile, selected/base-verifier/joint-label accuracy was 89.06% / 80.94% / 86.25%. Warm four-choice latency with a 768-token state was 3.710 / 3.177 / 0.818 seconds on the same idle Spark in FP32. The selected scorer was 11–17% slower than the per-option verifier across the measured workloads; shared-prefix reuse and merged adapters remain future experiments.

The complete report includes calibration, per-family results, uncertainty, all timing cells, CSV/JSON data and charts. The profiling protocol explains scope; the full results history retains earlier smoke experiments, failures and raw evidence links. Operational provenance and remaining service controls are in HANDOVER.md.