opensysone / source /docs /usage /jev-api.md
andyshu's picture
Organize verified OpenSysOne publication payload
294f8ea verified
|
Raw History Blame Contribute Delete
4.78 kB
# Use OpenSysOne and hosted Jev with the same request
For an interactive text-and-options interface, use the
[browser playground](playground.md). It serves frozen training snapshots through
this scoring code. Training is complete; the selected calibrated 4B model is the
default, and the final local API is running on loopback port **18081**.
The harness implements TypeSafe's documented `POST /v1/systemone` request and
answer shapes for `choice`, `score` and `noul`. See the
[official API reference](https://docs.typesafe.ai/api) and
[example request](../../examples/jev_request.json). It supports local trained scoring,
hosted Jev calls, and a comparison of both responses and elapsed request times.
Agreement is not a quality benchmark.
On GX10, use the isolated environment:
```bash
cd /home/andy/projects/opensysone
OPENSYSONE_PYTHON=/home/andy/ai/envs/opensysone/bin/python
```
The completed campaign's calibrated checkpoint is recorded in
`/home/andy/ai/opensysone/deploy/current.json`. Substitute its `model` path below.
Only load trusted project checkpoints; they contain serialized Python state.
```bash
$OPENSYSONE_PYTHON jev_harness.py --backend local \
--checkpoint /absolute/path/to/model.pt --request examples/jev_request.json
$OPENSYSONE_PYTHON jev_harness.py --backend serve \
--checkpoint /absolute/path/to/model.pt --port 18081
```
The campaign starts that loopback server automatically after successful
calibration, untouched evaluation and a trained-model inference check.
`GET /health` reports the loaded checkpoint and calibration status;
`GET /v1/models` reports the actual local model name. Send requests with:
```bash
curl --fail-with-body http://127.0.0.1:18081/v1/systemone \
-H 'Content-Type: application/json' --data-binary @examples/jev_request.json
```
The local model accepts `opensysone`, its actual model name, or `jev-latest` as a
client compatibility alias. Its response always identifies OpenSysOne. Choice
answers include the selected key and probabilities; score answers include the
probability-weighted rubric index and legend; noul answers give the probability
of the positive criterion. State and instructions can be strings, objects or
arrays. Shared harness limits are 64 questions, 255 alternatives per question and 512
candidate branches per request. Each complete chat-formatted candidate must fit
the configured token limit; oversized inputs are rejected without truncation.
The default is the checkpoint's training limit. `--max-tokens 1024` allows a
separately tested larger inference context, and the campaign uses that limit.
`/health` reports the active limit. Longer-context correctness is a wiring check,
not evidence of task generalization at that length.
Local confidence is explicitly `1 - entropy(probabilities) / log(choice_count)`.
This measures concentration and does not claim TypeSafe uses that formula, or
that concentration proves calibrated correctness. A single temperature is fitted
on separate known-family calibration data; calibration on new task families
remains a measured question. Input-token usage counts every processed candidate
branch, including repeated state tokens. Local scoring generates no answer tokens.
For hosted calls, configure `TYPESAFE_API_KEY` securely in the process environment.
The default URL is `https://api.typesafe.ai`; `TYPESAFE_BASE_URL` can override it.
The harness does not read or print secrets from files. Hosted-only use needs just
Python's standard library, so it can also run on the Mac.
```bash
python3 jev_harness.py --backend jev --request examples/jev_request.json
$OPENSYSONE_PYTHON jev_harness.py --backend compare \
--checkpoint /absolute/path/to/model.pt --request examples/jev_request.json
```
The hosted client validates responses, uses bounded retries for rate limits and
transient failures, and refuses credential forwarding across redirects. No
hosted call is made by local or serve modes. Optional `OPENSYSONE_API_KEY` protects
the local server with bearer authentication; it also applies to health requests.
The real-model verification uses an authenticated ephemeral loopback server.
Use `--timeout 300` for large HTTP requests if the default 60-second timeout is
too short. Hundreds of alternatives are processed in bounded chunks and can
take much longer than a small routing request.
For access from the Mac, use an SSH tunnel instead of opening a network listener:
```bash
ssh -N -L 18081:127.0.0.1:18081 gx10
```
Inspect and stop the fleet coordinator or its resulting API with
`scripts/fleet_campaign.py --campaign <fleet-run> --status` or `--stop`.
Individual training jobs use `scripts/campaign_status.py`. See
[handover.md](../operations/handover.md) for exact run paths,
deadline, source revisions and restart commands.