Download source/docs/usage/playground.md from andyshu/opensysone: direct link, hf CLI and curl.
- Browser
- Download file 7.21 kB
-
https://huggingface.co/andyshu/opensysone/resolve/main/source/docs/usage/playground.md
- Command line
-
hf download hf://andyshu/opensysone/source/docs/usage/playground.md
-
curl -L -o playground.md https://huggingface.co/andyshu/opensysone/resolve/main/source/docs/usage/playground.md
Try the trained models
The playground accepts context, a question and one option per line, then shows a probability for each option. It uses real frozen OpenSysOne checkpoints through the existing scoring harness. Pick another model to try the same decision again.
On your Mac, leave this tunnel running:
ssh -N -L 7466:127.0.0.1:7466 gx10
Open http://localhost:7466. If you are using a browser directly on GX10, no tunnel is needed. The final evaluated model API is running on loopback port 18081.
Paste your text, adjust the question and enter at least two distinct options. Click Get probabilities, or press Command/Ctrl + Enter. The examples are invented inputs you can edit. Copy JSON exports the actual result to your clipboard. Editing an input invalidates the previous result so it cannot be confused with a new question. Inputs are not saved to browser storage or request logs.
The workspace fits the browser window. Desktop shows the editor and probabilities side by side. Phones and short windows use Context, Choices, and Results tabs, preserving your inputs when you switch. Arrow keys, Home and End navigate the tabs. Scoring opens Results; validation brings the affected field into view. Long input text and answer lists scroll inside their own areas, with actions remaining visible. Refresh an existing browser tab to load layout changes.
The available snapshots are:
| Model | Selected step | Calibration |
|---|---|---|
| Qwen3 4B · Selected (default) | Spark B 1,500, retained at expanded branch 0 | Separate 510-decision calibration |
| Qwen3 4B · Main | 2,500 | Uncalibrated |
| Qwen3 4B · Lower rate | 2,500 | Uncalibrated |
| Qwen3.5 2B | 2,000 | Uncalibrated |
Probabilities are normalized over the supplied options. Adding/removing an option changes the question being scored. An uncalibrated 90% output is not a demonstrated 90% success rate. The default selected model uses temperature 1.745822 fitted on separate calibration data; calibration on arbitrary new tasks is unproven. Training has stopped. The other three entries are preserved historical snapshots.
Each complete context/question/option prompt must fit the verified 1,024-token inference limit, including chat formatting. The server reports an error for longer inputs instead of silently truncating them. Up to 255 distinct options are allowed; large lists take longer. The first request after selecting another model also loads its weights. One request runs at a time; other requests receive a busy reply.
Runtime and controls
The current server is running as PID 1674635, supervised by 1674634, backend
source 00c80dd, with pending exit status while serving. Its runtime is
/home/andy/ai/opensysone/runs/20260917T081925Z-playground-selected.
The runtime directory is recorded in ~/ai/opensysone/runs/LAST_PLAYGROUND.
Its models.json lists the fixed snapshots, and its launch/verification evidence
records the process, source revision, log and numerical checks. handover.md
records the active process. Inspect the listener and status with:
curl --fail http://127.0.0.1:7466/api/status
curl --fail http://127.0.0.1:7466/api/models
Use the existing isolated environment from this project directory to start it:
~/ai/envs/opensysone/bin/python playground.py \
--models /home/andy/ai/opensysone/runs/20260917T081925Z-playground-selected/models.json \
--port 7466 --device cuda --max-tokens 1024
Before restarting, verify the command in launch.json and stop the exact recorded
playground process with SIGTERM. Its wrapper records the exit in state.json and
exit_code; serving output is in run.log. This is separate from the fleet coordinator and final-model
uploader. The server binds only to loopback; it requires no firewall or service
configuration changes. No frontend dependency installation is required.
Static assets are read afresh on each request; frontend-current.json in the
runtime records the current frontend revision and verified served asset hashes.
The server loads one model at a time, checks frozen checkpoint hashes and model
metadata, and releases a previous model before loading a different one. The same
24 GiB available-memory check, 16 GiB CUDA allocation cap, FP32 scoring and OOM
adjustment 0 apply. Final accuracy evaluation and speed profiling are complete.
Inference requests consume GX10 compute while they run; idle browser tabs perform
no model inference. The recorded speed profile used an otherwise idle Spark.
The runtime supports --device cpu as an alternative, with latency depending on
the model and input. Check current memory and GPU processes before any model load.
The frontend is in web/, the server is playground.py, and the input/scoring
contract is inherited from jev_harness.py. See jev-api.md
for the underlying API and hosted Jev integration.
Verification
The selected-model refresh passed real-browser scoring for all four entries on
2026-09-17 at approximately 08:20 UTC, including normalized probabilities, the
calibrated badge, responsive layout and input/error/copy behavior. Runtime evidence
and screenshots are in 20260917T081925Z-playground-selected/browser-check under
the run root. The earlier screenshots and timings below are historical.
Seven backend tests pass with
python3 -m unittest discover -s tests -p 'test_playground.py'.
Real Chromium checks and screenshots are in
results/20260917-playground. They score invented
examples with all three actual models; reserved test/holdout data remains untouched.
The main-model warm example took 1.53 seconds; first-load/model-switch requests
were 5.08–9.10 seconds. These single observations are not latency percentiles.
To repeat the real browser check using existing GX10 browser tooling:
PLAYWRIGHT_MODULE=/home/andy/projects/ultraviolet/node_modules/playwright-core \
CHROMIUM_PATH=/home/andy/.cache/ms-playwright/chromium-1234/chrome-linux/chrome \
node scripts/verify_playground.cjs http://127.0.0.1:7466 \
/home/andy/ai/opensysone/runs/NEW-UNUSED-playground-browser-check
That command performs real inference and changes the resident GUI model during its checks. It does not train or alter checkpoints.
For responsive layout checks without model inference, use the isolated fixture API:
PLAYWRIGHT_MODULE=/home/andy/projects/ultraviolet/node_modules/playwright-core \
CHROMIUM_PATH=/home/andy/.cache/ms-playwright/chromium-1234/chrome-linux/chrome \
node scripts/verify_playground_layout.cjs \
/home/andy/ai/opensysone/runs/NEW-UNUSED-playground-layout-check
Its probabilities are test fixtures, not model output. It checks viewport fit,
control visibility, tab navigation, input preservation, long result lists and
error recovery across desktop, phone and landscape sizes.
The relayout passed 13 viewport sizes, including 320×568 and short 390×360
windows, with at least one readable line in every visible input. Reports and
screenshots are in results/20260917-playground-layout.