--- license: unknown pipeline_tag: text-classification base_model: - Qwen/Qwen3-4B-Instruct-2507 - Qwen/Qwen3.5-2B tags: - opensysone - decision-scoring - lora - experimental --- # OpenSysOne **Inspired by [Jev](https://typesafe.ai/), TypeSafe.ai's System One model.** Credit goes to the TypeSafe team for inspiring this project's exploration of structured decisions with probabilities. OpenSysOne is an independent experimental implementation; API compatibility does not establish Jev equivalence. The **completed 4B release** scores a state, question and explicit candidate answers, returning probabilities over those choices. Training, separate calibration, final evaluation and local API verification completed on 17 September 2026. Start with the [model and reconstruction notes](model/README.md), [results report](results/report.md), or [publication guide](docs/README.md). The calibrated artifact is [model/model.pt](model/model.pt). It contains custom OpenSysOne adapter/head weights and metadata. The pinned Qwen3-4B-Instruct-2507 base is required separately; this is not a standalone Transformers model or a standard PEFT adapter package. ## Measured results The full comparison uses the unchanged pretrained yes/no verifier, with a separate temperature fitted for each model. Intervals are paired 95% source-group bootstrap intervals for selected minus base accuracy. | Evaluation | Decisions | Selected | Base verifier | Accuracy gain (95% interval) | | --- | ---: | ---: | ---: | ---: | | Known-family test | 2,042 | 92.90% | 84.48% | +8.42 pp [6.85, 9.89] | | Social IQA family holdout | 768 | 72.92% | 70.31% | +2.60 pp [0.13, 5.34] | On a separate matched 320-decision profile, selected accuracy was 89.06%, versus 80.94% for the base verifier and 86.25% for a base model using one constrained answer-label token. The selected scorer was **slower on all 12 profiled workloads**: 1.11–1.17× the verifier latency and 2.18–15.58× the label baseline latency. These are warm, serial FP32 measurements on one GB10, not concurrent-serving throughput or comparisons with generated reasoning. See [tables and charts](results/). ## Model and limits The release uses rank-8 additive adapters and a scalar head: 16,517,633 trainable parameters. The selected expanded branch's step 0 retains the refinement parent's step-1,500 weights. The expanded branch's later step 159 was not selected; its post-selection diagnostic gains are reported separately. One temperature, 1.745822, was fitted on 510 separate known-family calibration examples. Social IQA calibration remains limited: its top-label ECE is 8.30%. No general intelligence, Jev-level quality or universal calibration claim follows. Four-choice Banking77 is not the full 77-label task; benchmark grouping does not rule out base-model pretraining overlap. The 1,024-token inference limit includes the complete formatted candidate prompt, and longer inputs are rejected. ## Files and provenance - [model/](model/README.md): calibrated artifact, hash and pinned base requirements. - [docs/](docs/README.md): layout and [reproduction guide](docs/reproduce.md). - [source/](source/): complete committed project source, tests and usage guides. - [results/](results/): final metrics, profiling report, tables and charts. - [archive/](archive/README.md): index to preserved experiment history. The original pointers remain authoritative: [FINAL_MODEL.json](FINAL_MODEL.json) identifies the calibrated release, [PROFILE_RESULTS.json](PROFILE_RESULTS.json) identifies verified profiling and wrap-up evidence, and [CURRENT_SNAPSHOT.json](CURRENT_SNAPSHOT.json) identifies the earlier training backup. Historical payload paths and hashes are preserved. The existing `license: unknown` metadata is unchanged. Base-model and dataset licenses remain separate; see the source's [original data provenance](source/results/public-decisions-v1-manifest.json) and [expanded data provenance](source/results/20260917-expanded-data/dataset-manifest.json). Base weights and credentials are excluded. Hosted Jev calls require separate authentication and were not exercised in this evaluation.