opensysone / source /README.md
andyshu's picture
Organize verified OpenSysOne publication payload
75c3d90 verified
|
Raw History Blame Contribute Delete
4.22 kB

OpenSysOne

An experimental natural-language decision scorer built by tuning Qwen3-4B. Give it context, a question and possible answers; it returns a probability for each answer. The project includes a browser playground and a Jev-compatible API.

Credit: OpenSysOne is inspired by Jev, TypeSafe.ai's System One model for structured decisions with probabilities. Credit goes to the TypeSafe team for motivating this independent experimental implementation.

Training and evaluation are complete. Start with the accuracy and speed report, browser playground guide, or API guide. The calibrated model and publication files are available in andyshu/opensysone on Hugging Face.

Results

Reserved evaluation Examples Trained 4B Pretrained per-option verifier
Original four-family test 2,042 92.90% 84.48%
Social IQA family holdout 768 72.92% 70.31%

Current inference has an accuracy–speed tradeoff. On one idle Spark in FP32, one four-choice question with a 768-token state takes 3.710 s for the trained scorer, 3.177 s for the pretrained per-option verifier and 0.818 s for a pretrained joint answer-label prompt. On a separate matched 320-example sample, the trained and joint-label methods score 89.06% and 86.25%. These warm measurements exclude model loading and HTTP. See the report for all workloads, calibration scores, confidence intervals and limits.

The selected weights are Spark B step 1,500, retained identically at expanded branch step 0. A separate 510-example calibration split fits temperature 1.745822. The model uses custom rank-8 adapters and a scalar head; it requires the pinned pretrained base and the project's reconstruction code. It is not a generic Transformers or PEFT checkpoint. Social IQA was excluded from our fine-tuning; exposure during base-model pretraining is unknown.

Repository layout

Location Purpose
docs/ Usage guides, research notes and operational records
results/ Final report and charts, plus preserved historical evidence
examples/ Sample API request
web/ Browser playground assets
scripts/ Training, evaluation, verification and archival tools
tests/ CPU and harness checks
experiment.py, training_model.py, selection.py Training, reconstruction and checkpoint selection
jev_harness.py, playground.py API and interactive model comparison
smoke_train.py, decision_model.py, smoke_data.py Original 0.5B smoke and shared-prefix experiments

Try it on the existing GX10 installation

The GUI uses port 7466. On the Mac, run ssh -N -L 7466:127.0.0.1:7466 gx10 and open http://localhost:7466. Select Qwen3 4B · Selected for the calibrated release. The local API is on loopback port 18081:

curl --fail-with-body http://127.0.0.1:18081/v1/systemone \
  -H 'Content-Type: application/json' --data-binary @examples/jev_request.json

For another machine, read the reconstruction requirements and reproduction guide. The checkpoint records its pinned local base-model path; downloading the adapter alone is insufficient. Hosted Jev requests require TYPESAFE_API_KEY; no hosted Jev benchmark is claimed.

Development and continuation

The existing isolated environment on GX10 can run the CPU suite:

~/ai/envs/opensysone/bin/python -m unittest discover -s tests -v

Read HANDOVER.md, PLAN.md and RESULTS.md before continuing experiments. Model weights, optimizer state and runtime logs live outside this source tree under ~/ai/models/opensysone and ~/ai/opensysone/runs. Dataset/source pins and checkpoint provenance accompany the archived results. No additional training is active or scheduled by this publication cleanup.