opensysone / source /PLAYGROUND.md
andyshu's picture
Back up verified OpenSysOne training snapshot and pinned source
e9d0e73 verified
|
Raw History Blame
4.8 kB
# Try the trained models
The playground accepts context, a question and one option per line, then shows a
probability for each option. It uses real frozen OpenSysOne checkpoints through
the existing scoring harness. Pick another model to try the same decision again.
On your Mac, leave this tunnel running:
```bash
ssh -N -L 7466:127.0.0.1:7466 gx10
```
Open **http://localhost:7466**. If you are using a browser directly on GX10, no
tunnel is needed. Port 18081 remains reserved for the final evaluated model API.
Paste your text, adjust the question and enter at least two distinct options.
Click **Get probabilities**, or press **Command/Ctrl + Enter**. The examples are
invented inputs you can edit. Copy JSON exports the actual result to your clipboard.
Editing an input invalidates the previous result so it cannot be confused with a
new question. Inputs are not saved to browser storage or request logs.
The available snapshots are:
| Model | Selected step | Calibration |
| --- | ---: | --- |
| Qwen3 4B · Main | 2,500 | Uncalibrated |
| Qwen3 4B · Lower rate | 2,500 | Uncalibrated |
| Qwen3.5 2B | 2,000 | Uncalibrated |
Probabilities are normalized over the supplied options. Adding/removing an option
changes the question being scored. An uncalibrated 90% output is not a demonstrated
90% success rate. These fixed snapshots remain stable while training continues.
Each complete context/question/option prompt must fit the verified **1,024-token**
inference limit, including chat formatting. The server reports an error for longer
inputs instead of silently truncating them. Up to 255 distinct options are allowed;
large lists take longer. The first request after selecting another model also
loads its weights. One request runs at a time; other requests receive a busy reply.
## Runtime and controls
The current server is running as PID **1469393**, source **`2d0ff79`**, with pending
exit status. Its immutable runtime is
`/home/andy/ai/opensysone/runs/20260917T034059Z-playground-port7466`.
The runtime directory is recorded in `~/ai/opensysone/runs/LAST_PLAYGROUND`.
Its `models.json` lists the fixed snapshots, and its launch/verification evidence
records the process, source revision, log and numerical checks. `HANDOVER.md`
records the active process. Inspect the listener and status with:
```bash
curl --fail http://127.0.0.1:7466/api/status
curl --fail http://127.0.0.1:7466/api/models
```
Use the existing isolated environment from this project directory to start it:
```bash
~/ai/envs/opensysone/bin/python playground.py \
--models /home/andy/ai/opensysone/runs/20260917T034059Z-playground-port7466/models.json \
--port 7466 --device cuda --max-tokens 1024
```
Before restarting, verify and stop the exact recorded playground process with
SIGTERM. This is separate from the training jobs, fleet coordinator and final-model
uploader. The server binds only to loopback; it requires no firewall or service
configuration changes. No frontend dependency installation is required.
The server loads one model at a time, checks frozen checkpoint hashes and model
metadata, and releases a previous model before loading a different one. The same
24 GiB available-memory check, 16 GiB CUDA allocation cap, FP32 scoring and OOM
adjustment 0 apply. GPU inference shares compute with GX10 training; frequent or
large requests can slow that training. Idle browser tabs perform no model inference.
The runtime supports `--device cpu` as an alternative, with latency depending on
the model and input. Check current memory and GPU processes before any model load.
The frontend is in `web/`, the server is `playground.py`, and the input/scoring
contract is inherited from `jev_harness.py`. See [JEV_HARNESS.md](JEV_HARNESS.md)
for the underlying API and hosted Jev integration.
## Verification
Seven backend tests pass with
`python3 -m unittest discover -s tests -p 'test_playground.py'`.
Real Chromium checks and screenshots are in
[`results/20260917-playground`](results/20260917-playground/). They score invented
examples with all three actual models; reserved test/holdout data remains untouched.
The main-model warm example took 1.53 seconds; first-load/model-switch requests
were 5.08–9.10 seconds. These single observations are not latency percentiles.
To repeat the real browser check using existing GX10 browser tooling:
```bash
PLAYWRIGHT_MODULE=/home/andy/projects/ultraviolet/node_modules/playwright-core \
CHROMIUM_PATH=/home/andy/.cache/ms-playwright/chromium-1234/chrome-linux/chrome \
node scripts/verify_playground.cjs http://127.0.0.1:7466 \
/home/andy/ai/opensysone/runs/NEW-UNUSED-playground-browser-check
```
That command performs real inference and changes the resident GUI model during
its checks. It does not train or alter checkpoints.