# Try the trained models The playground accepts context, a question and one option per line, then shows a probability for each option. It uses real frozen OpenSysOne checkpoints through the existing scoring harness. Pick another model to try the same decision again. On your Mac, leave this tunnel running: ```bash ssh -N -L 7466:127.0.0.1:7466 gx10 ``` Open **http://localhost:7466**. If you are using a browser directly on GX10, no tunnel is needed. The final evaluated model API is running on loopback port 18081. Paste your text, adjust the question and enter at least two distinct options. Click **Get probabilities**, or press **Command/Ctrl + Enter**. The examples are invented inputs you can edit. Copy JSON exports the actual result to your clipboard. Editing an input invalidates the previous result so it cannot be confused with a new question. Inputs are not saved to browser storage or request logs. The workspace fits the browser window. Desktop shows the editor and probabilities side by side. Phones and short windows use **Context**, **Choices**, and **Results** tabs, preserving your inputs when you switch. Arrow keys, Home and End navigate the tabs. Scoring opens Results; validation brings the affected field into view. Long input text and answer lists scroll inside their own areas, with actions remaining visible. Refresh an existing browser tab to load layout changes. The available snapshots are: | Model | Selected step | Calibration | | --- | ---: | --- | | Qwen3 4B · Selected (default) | Spark B 1,500, retained at expanded branch 0 | Separate 510-decision calibration | | Qwen3 4B · Main | 2,500 | Uncalibrated | | Qwen3 4B · Lower rate | 2,500 | Uncalibrated | | Qwen3.5 2B | 2,000 | Uncalibrated | Probabilities are normalized over the supplied options. Adding/removing an option changes the question being scored. An uncalibrated 90% output is not a demonstrated 90% success rate. The default selected model uses temperature **1.745822** fitted on separate calibration data; calibration on arbitrary new tasks is unproven. Training has stopped. The other three entries are preserved historical snapshots. Each complete context/question/option prompt must fit the verified **1,024-token** inference limit, including chat formatting. The server reports an error for longer inputs instead of silently truncating them. Up to 255 distinct options are allowed; large lists take longer. The first request after selecting another model also loads its weights. One request runs at a time; other requests receive a busy reply. ## Runtime and controls The current server is running as PID **1674635**, supervised by **1674634**, backend source **`00c80dd`**, with pending exit status while serving. Its runtime is `/home/andy/ai/opensysone/runs/20260917T081925Z-playground-selected`. The runtime directory is recorded in `~/ai/opensysone/runs/LAST_PLAYGROUND`. Its `models.json` lists the fixed snapshots, and its launch/verification evidence records the process, source revision, log and numerical checks. [handover.md](../operations/handover.md) records the active process. Inspect the listener and status with: ```bash curl --fail http://127.0.0.1:7466/api/status curl --fail http://127.0.0.1:7466/api/models ``` Use the existing isolated environment from this project directory to start it: ```bash ~/ai/envs/opensysone/bin/python playground.py \ --models /home/andy/ai/opensysone/runs/20260917T081925Z-playground-selected/models.json \ --port 7466 --device cuda --max-tokens 1024 ``` Before restarting, verify the command in `launch.json` and stop the exact recorded playground process with SIGTERM. Its wrapper records the exit in `state.json` and `exit_code`; serving output is in `run.log`. This is separate from the fleet coordinator and final-model uploader. The server binds only to loopback; it requires no firewall or service configuration changes. No frontend dependency installation is required. Static assets are read afresh on each request; `frontend-current.json` in the runtime records the current frontend revision and verified served asset hashes. The server loads one model at a time, checks frozen checkpoint hashes and model metadata, and releases a previous model before loading a different one. The same 24 GiB available-memory check, 16 GiB CUDA allocation cap, FP32 scoring and OOM adjustment 0 apply. Final accuracy evaluation and speed profiling are complete. Inference requests consume GX10 compute while they run; idle browser tabs perform no model inference. The recorded speed profile used an otherwise idle Spark. The runtime supports `--device cpu` as an alternative, with latency depending on the model and input. Check current memory and GPU processes before any model load. The frontend is in `web/`, the server is `playground.py`, and the input/scoring contract is inherited from `jev_harness.py`. See [jev-api.md](jev-api.md) for the underlying API and hosted Jev integration. ## Verification The selected-model refresh passed real-browser scoring for all four entries on 2026-09-17 at approximately 08:20 UTC, including normalized probabilities, the calibrated badge, responsive layout and input/error/copy behavior. Runtime evidence and screenshots are in `20260917T081925Z-playground-selected/browser-check` under the run root. The earlier screenshots and timings below are historical. Seven backend tests pass with `python3 -m unittest discover -s tests -p 'test_playground.py'`. Real Chromium checks and screenshots are in [`results/20260917-playground`](../../results/20260917-playground). They score invented examples with all three actual models; reserved test/holdout data remains untouched. The main-model warm example took 1.53 seconds; first-load/model-switch requests were 5.08–9.10 seconds. These single observations are not latency percentiles. To repeat the real browser check using existing GX10 browser tooling: ```bash PLAYWRIGHT_MODULE=/home/andy/projects/ultraviolet/node_modules/playwright-core \ CHROMIUM_PATH=/home/andy/.cache/ms-playwright/chromium-1234/chrome-linux/chrome \ node scripts/verify_playground.cjs http://127.0.0.1:7466 \ /home/andy/ai/opensysone/runs/NEW-UNUSED-playground-browser-check ``` That command performs real inference and changes the resident GUI model during its checks. It does not train or alter checkpoints. For responsive layout checks without model inference, use the isolated fixture API: ```bash PLAYWRIGHT_MODULE=/home/andy/projects/ultraviolet/node_modules/playwright-core \ CHROMIUM_PATH=/home/andy/.cache/ms-playwright/chromium-1234/chrome-linux/chrome \ node scripts/verify_playground_layout.cjs \ /home/andy/ai/opensysone/runs/NEW-UNUSED-playground-layout-check ``` Its probabilities are test fixtures, not model output. It checks viewport fit, control visibility, tab navigation, input preservation, long result lists and error recovery across desktop, phone and landscape sizes. The relayout passed 13 viewport sizes, including 320×568 and short 390×360 windows, with at least one readable line in every visible input. Reports and screenshots are in [`results/20260917-playground-layout`](../../results/20260917-playground-layout).