Download source/README.md from andyshu/opensysone: direct link, hf CLI and curl.
- Browser
- Download file 4.22 kB
-
https://huggingface.co/andyshu/opensysone/resolve/main/source/README.md
- Command line
-
hf download hf://andyshu/opensysone/source/README.md
-
curl -L -o README.md https://huggingface.co/andyshu/opensysone/resolve/main/source/README.md
OpenSysOne
An experimental natural-language decision scorer built by tuning Qwen3-4B. Give it context, a question and possible answers; it returns a probability for each answer. The project includes a browser playground and a Jev-compatible API.
Credit: OpenSysOne is inspired by Jev, TypeSafe.ai's System One model for structured decisions with probabilities. Credit goes to the TypeSafe team for motivating this independent experimental implementation.
Training and evaluation are complete. Start with the accuracy and speed report, browser playground guide, or API guide. The calibrated model and publication files are available in andyshu/opensysone on Hugging Face.
Results
| Reserved evaluation | Examples | Trained 4B | Pretrained per-option verifier |
|---|---|---|---|
| Original four-family test | 2,042 | 92.90% | 84.48% |
| Social IQA family holdout | 768 | 72.92% | 70.31% |
Current inference has an accuracy–speed tradeoff. On one idle Spark in FP32, one four-choice question with a 768-token state takes 3.710 s for the trained scorer, 3.177 s for the pretrained per-option verifier and 0.818 s for a pretrained joint answer-label prompt. On a separate matched 320-example sample, the trained and joint-label methods score 89.06% and 86.25%. These warm measurements exclude model loading and HTTP. See the report for all workloads, calibration scores, confidence intervals and limits.
The selected weights are Spark B step 1,500, retained identically at expanded branch step 0. A separate 510-example calibration split fits temperature 1.745822. The model uses custom rank-8 adapters and a scalar head; it requires the pinned pretrained base and the project's reconstruction code. It is not a generic Transformers or PEFT checkpoint. Social IQA was excluded from our fine-tuning; exposure during base-model pretraining is unknown.
Repository layout
| Location | Purpose |
|---|---|
| docs/ | Usage guides, research notes and operational records |
| results/ | Final report and charts, plus preserved historical evidence |
| examples/ | Sample API request |
| web/ | Browser playground assets |
| scripts/ | Training, evaluation, verification and archival tools |
| tests/ | CPU and harness checks |
experiment.py, training_model.py, selection.py |
Training, reconstruction and checkpoint selection |
jev_harness.py, playground.py |
API and interactive model comparison |
smoke_train.py, decision_model.py, smoke_data.py |
Original 0.5B smoke and shared-prefix experiments |
Try it on the existing GX10 installation
The GUI uses port 7466. On the Mac, run ssh -N -L 7466:127.0.0.1:7466 gx10
and open http://localhost:7466. Select Qwen3 4B · Selected for the calibrated
release. The local API is on loopback port 18081:
curl --fail-with-body http://127.0.0.1:18081/v1/systemone \
-H 'Content-Type: application/json' --data-binary @examples/jev_request.json
For another machine, read the reconstruction requirements
and reproduction guide. The checkpoint records
its pinned local base-model path; downloading the adapter alone is insufficient.
Hosted Jev requests require TYPESAFE_API_KEY; no hosted Jev benchmark is claimed.
Development and continuation
The existing isolated environment on GX10 can run the CPU suite:
~/ai/envs/opensysone/bin/python -m unittest discover -s tests -v
Read HANDOVER.md, PLAN.md and RESULTS.md
before continuing experiments. Model weights, optimizer state and runtime logs
live outside this source tree under ~/ai/models/opensysone and
~/ai/opensysone/runs. Dataset/source pins and checkpoint provenance accompany
the archived results. No additional training is active or scheduled by this
publication cleanup.