# Use Lex Use the official TypeSafe SDK or curl to send a shared context and named typed questions to a SystemOne-compatible endpoint. The example asks a routing question and an urgency question about the same delivery request. ## Endpoint setup Replace `https://your-decision-endpoint.example` with an endpoint configured to expose `Decision-1.0-Lex-0.6B`. Set the `DECISION_API_KEY` environment variable to the key issued by that endpoint's operator. The sized model name below is a deployment alias that the operator must configure. The Hugging Face repository distributes weights and local inference code; it does not provision an HTTP service or issue API keys. TypeSafe does not host these Decision weights. Installing the official SDK supplies a client for a compatible service, not a model deployment. These are request examples, not recorded model predictions. No example probabilities or performance results are implied. ## Official Python SDK ```bash pip install typesafe-sdk ``` ```python import os from typesafe_sdk import TypeSafeClient, Choice, Noul client = TypeSafeClient( api_key=os.environ["DECISION_API_KEY"], base_url="https://your-decision-endpoint.example", model="Decision-1.0-Lex-0.6B", ) questions = { "route": Choice(instructions="Which team should handle this request?", criteria={"delivery": "Damaged or missing parcels", "billing": "Payments and invoices"}), "urgent": Noul(instructions="Does the customer request action today?"), } response = client.system_one(state="The parcel arrived damaged. Please send a replacement today.", questions=questions) print(response.choices["route"].choice, response.nouls["urgent"].noul) ``` `response.choices["route"].choice` is the selected candidate ID; `response.nouls["urgent"].noul` is the probability that the condition holds. The application decides how to act on the answers. ## Equivalent curl request The model, state, question IDs, instructions and candidate order are identical to the SDK example. ```bash curl -X POST https://your-decision-endpoint.example/v1/systemone \ -H "Authorization: Bearer $DECISION_API_KEY" \ -H "Content-Type: application/json" \ --data '{ "model": "Decision-1.0-Lex-0.6B", "state": "The parcel arrived damaged. Please send a replacement today.", "questions": { "route": {"type": "choice", "instructions": "Which team should handle this request?", "criteria": {"delivery": "Damaged or missing parcels", "billing": "Payments and invoices"}}, "urgent": {"type": "noul", "instructions": "Does the customer request action today?"} } }' ``` See the [official Python usage guide](https://docs.typesafe.ai/sdk/python/usage) and [HTTP request/response reference](https://docs.typesafe.ai/api). ## Several contexts, the same questions Keep the `questions` map and submit another `state` to `client.system_one`. Each request evaluates both questions against its own context. A larger question map expresses more decisions about that context; service concurrency, request limits and scheduling depend on the deployment. The bundled native implementations also support multi-context batching. This is a local inference capability; it does not imply that a deployment exposes a batch HTTP route. All complete-input limits still apply to each rendered state/question/candidate sequence. ## Deployment names and native compatibility The HTTP examples use the configured alias `Decision-1.0-Lex-0.6B`. The bundled local model retains the stable internal identifier `Decision-1.0-Lex`. A deployment maps its public alias to that native model; adding a size suffix to the Hub repository does not change the native identifier or its accepted aliases. Existing local SystemOne CLI examples continue to use the unsized internal identifier. The native code and weights are unchanged. ## Local native installation Download the complete public release with the Hugging Face CLI. No access approval or login is required: ```bash hf download llm-semantic-router/Decision-1.0-Lex-0.6B --local-dir Decision-1.0-Lex-0.6B cd Decision-1.0-Lex-0.6B ``` Use `--local-dir` to materialize ordinary files for the native loader. Use the compatible ROCm/Python environment described below. This release bundles the verified native under `native/`, the public Python API. Run the following commands from the complete distribution root in a compatible environment. `PACKAGE_MANIFEST.json` records the exact payload; do not add files inside `native/` because its loader checks the complete file roster. Use an existing compatible AMD ROCm environment. The verified source pins Transformers 4.57.6; tested companion versions were Python 3.12.13, tokenizers 0.22.2 and safetensors 0.8.0. Actual validation used a ROCm PyTorch 2.12 development build, not a promised generic wheel installation. Choose a matching supported ROCm/PyTorch installation for your host; this release does not supply an installer or a CPU/NVIDIA inference path. Do not upgrade an active environment in place. ## System One: parallel typed questions For local execution, the existing native CLI remains available from the downloaded repository: ```bash ROCR_VISIBLE_DEVICES=0 python systemone.py \ --input examples/system-one.json --output answers.json --batching auto ``` The bundled wrapper accepts up to 128 questions per request and combines independent contexts with up to 512 total decisions. Each complete state/question/candidate sequence must fit 1,024 tokens; validation happens before the first forward. Default physical batching is B8; `auto` opts into the released homogeneous, padding-aware B32 scheduler. These local limits are not a promise about another endpoint's limits. [Native fields and fine-tuning conversion](SYSTEM_ONE.md). ## Native records The original native JSONL interface remains available for applications that need logits, explicit candidate IDs or arbitrary ordered Score values. ```bash export PYTHONPATH="$PWD${PYTHONPATH:+:$PYTHONPATH}" export PYTHONDONTWRITEBYTECODE=1 export HF_HUB_OFFLINE=1 TRANSFORMERS_OFFLINE=1 TOKENIZERS_PARALLELISM=false ROCR_VISIBLE_DEVICES=0 python infer.py \ --input examples/decisions.jsonl --output predictions.jsonl --batch-size 8 ``` The three examples illustrate the Choice, Noul and Score interfaces. They are synthetic interface examples, not evaluated model predictions. Each input line has exactly `id`, `state_text`, and `question`; do not attach labels. Choice uses `options`, Score uses ordered `levels`, and Noul uses the fixed no/yes pair with optional `false_criterion`/`true_criterion`. IDs identify outputs and do not become model tokens. The Python API returns original record/candidate identities, logits and probabilities in the supplied order. Choice includes the chosen candidate ID. Noul includes the yes probability; it does not emit a boolean decision. An application can choose first-argmax (an exact no/yes tie selects no) or `p_yes >= 0.5` (tie selects yes), but should declare that choice. Score includes `expected_value` over the supplied values and `score`, the expected ordinal index from 0 to K−1. Neither is automatically a calibrated business utility. All state/question/candidate text and markers count toward the complete 1024 limit. `predict_1k` checks every input before the first forward and rejects overflow without truncation. Batch sizes 1–8 are supported by this wrapper. Use it rather than directly invoking the native research-capacity entrypoint. A compatible fine-tuned export can be used with `--native --manifest-sha256 `. Only the pinned runtime revision and three-path architecture are accepted. Downloaded HF symlinks must be materialized into real files before loading. The manifest hash is an integrity check, not a reason to execute arbitrary untrusted Python. ### Default typed scheduling The default `SystemOne` path groups complete, admitted questions by decision type in physical batches of eight, then restores the original request and question order. This works for many questions over one state and questions across multiple contexts. No API changes or application-side sorting are needed. The optional `batching="auto"` policy is unchanged. [Measured mixed-question scaling](PERFORMANCE.md).