opensysone / source /docs /usage /playground.md
andyshu's picture
Organize verified OpenSysOne publication payload
294f8ea verified
|
Raw History Blame Contribute Delete
7.21 kB

Try the trained models

The playground accepts context, a question and one option per line, then shows a probability for each option. It uses real frozen OpenSysOne checkpoints through the existing scoring harness. Pick another model to try the same decision again.

On your Mac, leave this tunnel running:

ssh -N -L 7466:127.0.0.1:7466 gx10

Open http://localhost:7466. If you are using a browser directly on GX10, no tunnel is needed. The final evaluated model API is running on loopback port 18081.

Paste your text, adjust the question and enter at least two distinct options. Click Get probabilities, or press Command/Ctrl + Enter. The examples are invented inputs you can edit. Copy JSON exports the actual result to your clipboard. Editing an input invalidates the previous result so it cannot be confused with a new question. Inputs are not saved to browser storage or request logs.

The workspace fits the browser window. Desktop shows the editor and probabilities side by side. Phones and short windows use Context, Choices, and Results tabs, preserving your inputs when you switch. Arrow keys, Home and End navigate the tabs. Scoring opens Results; validation brings the affected field into view. Long input text and answer lists scroll inside their own areas, with actions remaining visible. Refresh an existing browser tab to load layout changes.

The available snapshots are:

Model Selected step Calibration
Qwen3 4B · Selected (default) Spark B 1,500, retained at expanded branch 0 Separate 510-decision calibration
Qwen3 4B · Main 2,500 Uncalibrated
Qwen3 4B · Lower rate 2,500 Uncalibrated
Qwen3.5 2B 2,000 Uncalibrated

Probabilities are normalized over the supplied options. Adding/removing an option changes the question being scored. An uncalibrated 90% output is not a demonstrated 90% success rate. The default selected model uses temperature 1.745822 fitted on separate calibration data; calibration on arbitrary new tasks is unproven. Training has stopped. The other three entries are preserved historical snapshots.

Each complete context/question/option prompt must fit the verified 1,024-token inference limit, including chat formatting. The server reports an error for longer inputs instead of silently truncating them. Up to 255 distinct options are allowed; large lists take longer. The first request after selecting another model also loads its weights. One request runs at a time; other requests receive a busy reply.

Runtime and controls

The current server is running as PID 1674635, supervised by 1674634, backend source 00c80dd, with pending exit status while serving. Its runtime is /home/andy/ai/opensysone/runs/20260917T081925Z-playground-selected.

The runtime directory is recorded in ~/ai/opensysone/runs/LAST_PLAYGROUND. Its models.json lists the fixed snapshots, and its launch/verification evidence records the process, source revision, log and numerical checks. handover.md records the active process. Inspect the listener and status with:

curl --fail http://127.0.0.1:7466/api/status
curl --fail http://127.0.0.1:7466/api/models

Use the existing isolated environment from this project directory to start it:

~/ai/envs/opensysone/bin/python playground.py \
  --models /home/andy/ai/opensysone/runs/20260917T081925Z-playground-selected/models.json \
  --port 7466 --device cuda --max-tokens 1024

Before restarting, verify the command in launch.json and stop the exact recorded playground process with SIGTERM. Its wrapper records the exit in state.json and exit_code; serving output is in run.log. This is separate from the fleet coordinator and final-model uploader. The server binds only to loopback; it requires no firewall or service configuration changes. No frontend dependency installation is required. Static assets are read afresh on each request; frontend-current.json in the runtime records the current frontend revision and verified served asset hashes.

The server loads one model at a time, checks frozen checkpoint hashes and model metadata, and releases a previous model before loading a different one. The same 24 GiB available-memory check, 16 GiB CUDA allocation cap, FP32 scoring and OOM adjustment 0 apply. Final accuracy evaluation and speed profiling are complete. Inference requests consume GX10 compute while they run; idle browser tabs perform no model inference. The recorded speed profile used an otherwise idle Spark. The runtime supports --device cpu as an alternative, with latency depending on the model and input. Check current memory and GPU processes before any model load.

The frontend is in web/, the server is playground.py, and the input/scoring contract is inherited from jev_harness.py. See jev-api.md for the underlying API and hosted Jev integration.

Verification

The selected-model refresh passed real-browser scoring for all four entries on 2026-09-17 at approximately 08:20 UTC, including normalized probabilities, the calibrated badge, responsive layout and input/error/copy behavior. Runtime evidence and screenshots are in 20260917T081925Z-playground-selected/browser-check under the run root. The earlier screenshots and timings below are historical.

Seven backend tests pass with python3 -m unittest discover -s tests -p 'test_playground.py'. Real Chromium checks and screenshots are in results/20260917-playground. They score invented examples with all three actual models; reserved test/holdout data remains untouched. The main-model warm example took 1.53 seconds; first-load/model-switch requests were 5.08–9.10 seconds. These single observations are not latency percentiles.

To repeat the real browser check using existing GX10 browser tooling:

PLAYWRIGHT_MODULE=/home/andy/projects/ultraviolet/node_modules/playwright-core \
CHROMIUM_PATH=/home/andy/.cache/ms-playwright/chromium-1234/chrome-linux/chrome \
  node scripts/verify_playground.cjs http://127.0.0.1:7466 \
  /home/andy/ai/opensysone/runs/NEW-UNUSED-playground-browser-check

That command performs real inference and changes the resident GUI model during its checks. It does not train or alter checkpoints.

For responsive layout checks without model inference, use the isolated fixture API:

PLAYWRIGHT_MODULE=/home/andy/projects/ultraviolet/node_modules/playwright-core \
CHROMIUM_PATH=/home/andy/.cache/ms-playwright/chromium-1234/chrome-linux/chrome \
  node scripts/verify_playground_layout.cjs \
  /home/andy/ai/opensysone/runs/NEW-UNUSED-playground-layout-check

Its probabilities are test fixtures, not model output. It checks viewport fit, control visibility, tab navigation, input preservation, long result lists and error recovery across desktop, phone and landscape sizes. The relayout passed 13 viewport sizes, including 320×568 and short 390×360 windows, with at least one readable line in every visible input. Reports and screenshots are in results/20260917-playground-layout.