opensysone / source /docs /usage /jev-api.md
andyshu's picture
Organize verified OpenSysOne publication payload
294f8ea verified
|
Raw History Blame Contribute Delete
4.78 kB

Use OpenSysOne and hosted Jev with the same request

For an interactive text-and-options interface, use the browser playground. It serves frozen training snapshots through this scoring code. Training is complete; the selected calibrated 4B model is the default, and the final local API is running on loopback port 18081.

The harness implements TypeSafe's documented POST /v1/systemone request and answer shapes for choice, score and noul. See the official API reference and example request. It supports local trained scoring, hosted Jev calls, and a comparison of both responses and elapsed request times. Agreement is not a quality benchmark.

On GX10, use the isolated environment:

cd /home/andy/projects/opensysone
OPENSYSONE_PYTHON=/home/andy/ai/envs/opensysone/bin/python

The completed campaign's calibrated checkpoint is recorded in /home/andy/ai/opensysone/deploy/current.json. Substitute its model path below. Only load trusted project checkpoints; they contain serialized Python state.

$OPENSYSONE_PYTHON jev_harness.py --backend local \
  --checkpoint /absolute/path/to/model.pt --request examples/jev_request.json

$OPENSYSONE_PYTHON jev_harness.py --backend serve \
  --checkpoint /absolute/path/to/model.pt --port 18081

The campaign starts that loopback server automatically after successful calibration, untouched evaluation and a trained-model inference check. GET /health reports the loaded checkpoint and calibration status; GET /v1/models reports the actual local model name. Send requests with:

curl --fail-with-body http://127.0.0.1:18081/v1/systemone \
  -H 'Content-Type: application/json' --data-binary @examples/jev_request.json

The local model accepts opensysone, its actual model name, or jev-latest as a client compatibility alias. Its response always identifies OpenSysOne. Choice answers include the selected key and probabilities; score answers include the probability-weighted rubric index and legend; noul answers give the probability of the positive criterion. State and instructions can be strings, objects or arrays. Shared harness limits are 64 questions, 255 alternatives per question and 512 candidate branches per request. Each complete chat-formatted candidate must fit the configured token limit; oversized inputs are rejected without truncation. The default is the checkpoint's training limit. --max-tokens 1024 allows a separately tested larger inference context, and the campaign uses that limit. /health reports the active limit. Longer-context correctness is a wiring check, not evidence of task generalization at that length.

Local confidence is explicitly 1 - entropy(probabilities) / log(choice_count). This measures concentration and does not claim TypeSafe uses that formula, or that concentration proves calibrated correctness. A single temperature is fitted on separate known-family calibration data; calibration on new task families remains a measured question. Input-token usage counts every processed candidate branch, including repeated state tokens. Local scoring generates no answer tokens.

For hosted calls, configure TYPESAFE_API_KEY securely in the process environment. The default URL is https://api.typesafe.ai; TYPESAFE_BASE_URL can override it. The harness does not read or print secrets from files. Hosted-only use needs just Python's standard library, so it can also run on the Mac.

python3 jev_harness.py --backend jev --request examples/jev_request.json

$OPENSYSONE_PYTHON jev_harness.py --backend compare \
  --checkpoint /absolute/path/to/model.pt --request examples/jev_request.json

The hosted client validates responses, uses bounded retries for rate limits and transient failures, and refuses credential forwarding across redirects. No hosted call is made by local or serve modes. Optional OPENSYSONE_API_KEY protects the local server with bearer authentication; it also applies to health requests. The real-model verification uses an authenticated ephemeral loopback server. Use --timeout 300 for large HTTP requests if the default 60-second timeout is too short. Hundreds of alternatives are processed in bounded chunks and can take much longer than a small routing request.

For access from the Mac, use an SSH tunnel instead of opening a network listener:

ssh -N -L 18081:127.0.0.1:18081 gx10

Inspect and stop the fleet coordinator or its resulting API with scripts/fleet_campaign.py --campaign <fleet-run> --status or --stop. Individual training jobs use scripts/campaign_status.py. See handover.md for exact run paths, deadline, source revisions and restart commands.