# OpenSysOne An experimental natural-language decision scorer built by tuning Qwen3-4B. Give it context, a question and possible answers; it returns a probability for each answer. The project includes a browser playground and a Jev-compatible API. **Credit:** OpenSysOne is inspired by [Jev](https://typesafe.ai/), TypeSafe.ai's System One model for structured decisions with probabilities. Credit goes to the TypeSafe team for motivating this independent experimental implementation. **Training and evaluation are complete.** Start with the [accuracy and speed report](results/report.md), [browser playground guide](docs/usage/playground.md), or [API guide](docs/usage/jev-api.md). The calibrated model and publication files are available in [andyshu/opensysone on Hugging Face](https://huggingface.co/andyshu/opensysone). ## Results | Reserved evaluation | Examples | Trained 4B | Pretrained per-option verifier | | --- | ---: | ---: | ---: | | Original four-family test | 2,042 | **92.90%** | 84.48% | | Social IQA family holdout | 768 | **72.92%** | 70.31% | Current inference has an accuracy–speed tradeoff. On one idle Spark in FP32, one four-choice question with a 768-token state takes **3.710 s** for the trained scorer, **3.177 s** for the pretrained per-option verifier and **0.818 s** for a pretrained joint answer-label prompt. On a separate matched 320-example sample, the trained and joint-label methods score **89.06%** and **86.25%**. These warm measurements exclude model loading and HTTP. See the report for all workloads, calibration scores, confidence intervals and limits. The selected weights are Spark B step 1,500, retained identically at expanded branch step 0. A separate 510-example calibration split fits temperature 1.745822. The model uses custom rank-8 adapters and a scalar head; it requires the pinned pretrained base and the project's reconstruction code. It is not a generic Transformers or PEFT checkpoint. Social IQA was excluded from our fine-tuning; exposure during base-model pretraining is unknown. ## Repository layout | Location | Purpose | | --- | --- | | [docs/](docs/README.md) | Usage guides, research notes and operational records | | [results/](results/README.md) | Final report and charts, plus preserved historical evidence | | [examples/](examples/) | Sample API request | | [web/](web/) | Browser playground assets | | [scripts/](scripts/) | Training, evaluation, verification and archival tools | | [tests/](tests/) | CPU and harness checks | | `experiment.py`, `training_model.py`, `selection.py` | Training, reconstruction and checkpoint selection | | `jev_harness.py`, `playground.py` | API and interactive model comparison | | `smoke_train.py`, `decision_model.py`, `smoke_data.py` | Original 0.5B smoke and shared-prefix experiments | ## Try it on the existing GX10 installation The GUI uses port **7466**. On the Mac, run `ssh -N -L 7466:127.0.0.1:7466 gx10` and open **http://localhost:7466**. Select **Qwen3 4B · Selected** for the calibrated release. The local API is on loopback port **18081**: ```bash curl --fail-with-body http://127.0.0.1:18081/v1/systemone \ -H 'Content-Type: application/json' --data-binary @examples/jev_request.json ``` For another machine, read the [reconstruction requirements](https://huggingface.co/andyshu/opensysone/blob/main/model/README.md) and [reproduction guide](https://huggingface.co/andyshu/opensysone/blob/main/docs/reproduce.md). The checkpoint records its pinned local base-model path; downloading the adapter alone is insufficient. Hosted Jev requests require `TYPESAFE_API_KEY`; no hosted Jev benchmark is claimed. ## Development and continuation The existing isolated environment on GX10 can run the CPU suite: ```bash ~/ai/envs/opensysone/bin/python -m unittest discover -s tests -v ``` Read [HANDOVER.md](HANDOVER.md), [PLAN.md](PLAN.md) and [RESULTS.md](RESULTS.md) before continuing experiments. Model weights, optimizer state and runtime logs live outside this source tree under `~/ai/models/opensysone` and `~/ai/opensysone/runs`. Dataset/source pins and checkpoint provenance accompany the archived results. No additional training is active or scheduled by this publication cleanup.