Download source/RESULTS.md from andyshu/opensysone: direct link, hf CLI and curl.
- Browser
- Download file 1.73 kB
-
https://huggingface.co/andyshu/opensysone/resolve/main/source/RESULTS.md
- Command line
-
hf download hf://andyshu/opensysone/source/RESULTS.md
-
curl -L -o RESULTS.md https://huggingface.co/andyshu/opensysone/resolve/main/source/RESULTS.md
OpenSysOne results
Training and evaluation completed on 17 September 2026. The selected Qwen3-4B model uses rank-8 adapters and a scalar decision head. Its weights are Spark B step 1,500, retained unchanged at expanded branch step 0. Selection used validation only; temperature was fitted on a separate 510-decision calibration split.
| Reserved evaluation | Decisions | Selected accuracy | Pretrained verifier |
|---|---|---|---|
| Original four-family test | 2,042 | 92.90% | 84.48% |
| Social IQA family holdout | 768 | 72.92% | 70.31% |
Paired 95% bootstrap accuracy gains are +8.42 pp [6.85, 9.89] on the test and +2.60 pp [0.13, 5.34] on the holdout. The latter covers one untrained fine-tuning family; pretrained exposure is unknown. These are decision-scoring results, not general-intelligence scores or a measured comparison with hosted Jev.
On the separate matched 320-decision profile, selected/base-verifier/joint-label accuracy was 89.06% / 80.94% / 86.25%. Warm four-choice latency with a 768-token state was 3.710 / 3.177 / 0.818 seconds on the same idle Spark in FP32. The selected scorer was 11–17% slower than the per-option verifier across the measured workloads; shared-prefix reuse and merged adapters remain future experiments.
The complete report includes calibration, per-family results, uncertainty, all timing cells, CSV/JSON data and charts. The profiling protocol explains scope; the full results history retains earlier smoke experiments, failures and raw evidence links. Operational provenance and remaining service controls are in HANDOVER.md.