# Reproduce the published scorer The published artifact is a custom adapter/head checkpoint requiring a pinned local base and the supplied project code. These instructions describe the evaluated Linux/GB10 setup and its current path constraints, not a portable one-command installation. ## Retrieve and verify Download the published tree with your authorized Hugging Face client, retaining the sibling `source/` and `model/` directories. For immutable evidence, retrieve the path in [FINAL_MODEL.json](../FINAL_MODEL.json) at its recorded `payload_commit` and verify its model and manifest hashes. The front model must match the same bytes. From the downloaded repository root: ```bash sha256sum model/model.pt cd source ``` Expected model SHA-256: ```text e270e3da905604d97bf5a8f380ea308133403d1c4790a5c012cb1c12e9b6f348 ``` Keep `source/` intact. Scripts import modules relative to its root, examples and web assets use that layout, and evaluation's data signature hashes `training_model.py` relative to the working directory. Run the following commands from `source/`. ## Environment and pinned base The completed [evaluation manifest](../final/20260916T194403396250Z-fleet/evaluation/manifest.json) records NVIDIA GB10, CUDA 13.0 and these installed packages: | Package | Recorded version | | --- | --- | | torch | `2.11.0+cu130` | | transformers | `5.15.0` | | pyarrow | `25.0.1` | | numpy | `2.5.2` | These are measured environment identifiers, not a claim that the same CUDA build is available on every platform. The evaluated isolated interpreter is `/home/andy/ai/envs/opensysone/bin/python`. On another machine, create an isolated compatible environment and verify it against the recorded evidence before claiming reproduction. No dependency version should be inferred from the model card alone. The base is `Qwen/Qwen3-4B-Instruct-2507` at revision `cdbee75f17c01a7cc42f958dc650907174af0554`. On the recorded `/home/andy` account, the existing CPU-only downloader retrieves that pin and writes its provenance: ```bash /home/andy/ai/envs/opensysone/bin/python scripts/download_candidate.py \ --model Qwen/Qwen3-4B-Instruct-2507 ``` It writes under the invoking user's home. The artifact loader specifically expects `/home/andy/ai/models/opensysone/Qwen3-4B-Instruct-2507-cdbee75f`, including `opensysone-provenance.json`. Another home directory requires arranging the pinned base at that recorded location; the current loader has no base-path override. Do not modify the released checkpoint to change its paths. See [model reconstruction notes](../model/README.md). ## Local scoring and API Before loading a model, inspect available memory and existing GPU jobs: ```bash free -b nvidia-smi --query-compute-apps=pid,process_name,used_memory --format=csv ``` The harness checks for at least 24 GiB currently available host memory, restores OOM adjustment 0 and applies a 16 GiB CUDA allocation cap. The verified path uses FP32. The example is an invented request, not a benchmark measurement: ```bash /home/andy/ai/envs/opensysone/bin/python jev_harness.py \ --backend local --checkpoint ../model/model.pt \ --request examples/jev_request.json --device cuda --max-tokens 1024 ``` To run the same scorer as a loopback API on an unused local port: ```bash /home/andy/ai/envs/opensysone/bin/python jev_harness.py \ --backend serve --checkpoint ../model/model.pt \ --device cuda --max-tokens 1024 --port 18081 ``` In another terminal: ```bash curl --fail http://127.0.0.1:18081/health ``` The server binds to `127.0.0.1`; it does not expose a public endpoint. Read the [API guide](../source/docs/usage/jev-api.md) for request shape, optional local authentication and hosted Jev comparison. Hosted Jev needs a separate credential and was not exercised in the published evaluation. The [playground guide](../source/docs/usage/playground.md) covers the browser interface and its fixed checkpoint catalog. Existing machine-specific run paths in usage guides are operational records, not files downloaded with the model. The 1,024-token limit applies separately to each complete chat-formatted candidate prompt. Excess-length input is rejected rather than truncated. The artifact's default training limit is 512, so preserve `--max-tokens 1024` for the documented inference configuration. Scalar calibration is applied by the harness. ## Reproducing evidence The [final report](../results/report.md) separates the full 2,042-decision test and 768-decision Social IQA holdout from the matched 320-decision speed-profile sample and 383 expansion diagnostics. It reports the baseline definition, repeat counts, exact sample sizes and confidence-interval direction. Use [PROFILE_RESULTS.json](../PROFILE_RESULTS.json) and its immutable manifest to retrieve the frozen profiling protocol, requests, raw predictions, timings and source archives. [CURRENT_SNAPSHOT.json](../CURRENT_SNAPSHOT.json) records earlier training backups; it is not the final selected model. Historical dataset and source manifests record the exact hashes required for retraining or evaluation. For resume, keep each `checkpoint.pt`, sibling `best.pt` and matching validation evidence together. The calibrated `model.pt` is an inference artifact without optimizer state. Replay requires those complete evidence bundles and recorded source revisions; the convenient current `source/` view alone is not a substitute for the frozen training/evaluation provenance. Do not select a new checkpoint or fit temperatures using the published test or holdout results. No new inference is required to read the existing report and integrity manifests.