Download USAGE.md from vllm-sr/Decision-1.0-Nox-4B: direct link, hf CLI and curl.
- Browser
- Download file 5.44 kB
-
https://huggingface.co/vllm-sr/Decision-1.0-Nox-4B/resolve/46505c737a45cbe2c4ac4eee38e4f94cb520e4ed/USAGE.md
- Command line
-
hf download hf://vllm-sr/Decision-1.0-Nox-4B@46505c737a45cbe2c4ac4eee38e4f94cb520e4ed/USAGE.md
-
curl -L -o USAGE.md https://huggingface.co/vllm-sr/Decision-1.0-Nox-4B/resolve/46505c737a45cbe2c4ac4eee38e4f94cb520e4ed/USAGE.md
Load Decision locally
This package loads the exported Sol/Nox decoder bundles and delegates to their frozen numerical engine. It does not use text generation or alter the trained prompt, candidate readout, temperature, or probability semantics.
Install from a downloaded model repository containing this pyproject.toml and src/decision/:
python -m pip install .
# Optional, only for resolving models through Hugging Face:
python -m pip install '.[hub]'
There is no requirement to install an identically named package from PyPI. These commands install this repository's wrapper. They deliberately do not replace your GPU PyTorch installation. The full inference environment is recorded in the model's runtime.json: the qualified run used PyTorch 2.12.0+git6bbd260 / ROCm 7.2.53211, Transformers 5.17.0, FLA 0.5.2, Triton 3.7.1, tokenizers 0.23.2 and safetensors 0.8.0. The recorded PyTorch build is not promised to exist on ordinary PyPI. Use the public digest-pinned build recipe in RUNTIME.md; installation of this lightweight wrapper alone is not installation of that GPU runtime.
Local loading is offline and requires no Hub client or credentials:
from decision import DecisionModel
model = DecisionModel.from_pretrained(
"./Decision-1.0-Sol", local_files_only=True, device="cuda:0"
)
result = model.decide(
state="The customer reports a duplicate invoice charge and asks for a refund.",
questions={
"destination": {
"type": "choice",
"instructions": "Choose the team that handles this request.",
"criteria": {
"billing": "Invoices, payments, refunds and duplicate charges",
"technical": "Product errors and troubleshooting",
},
}
},
)
print(result)
The saved model-card example runner includes Choice, Noul and Score together, compares its actual output with the direct frozen engine, and checks that an overflowing input raises an error:
decision-example ./Decision-1.0-Sol --local-files-only --output actual-example.json
# Or: python -m decision.example ...
The model card must publish output from that model's real run; this document invents no prediction. To load a cached or remote Hub snapshot, pass the published repository ID as the first argument to DecisionModel.from_pretrained and its full commit SHA as revision. local_files_only=True restricts this lookup to already cached files. A local directory always bypasses Hub lookup. The model loads Python inference source from its verified bundle, like other local model packages; choose a repository you trust.
Interface
| Type | Input criteria | Answer |
|---|---|---|
| Choice | Ordered mapping of 2–255 external IDs to complete descriptions | probabilities, selected choice ID, confidence |
| Noul | Optional mapping containing only false / true descriptions |
noul: P(true); a hard judgment uses >= 0.5 |
| Score | Ordered list of 2–10 rubric descriptions | probabilities over string indices, expected level index score, legend, confidence |
Response shape is {"model": name, "answers": {question_name: answer}, "usage": {"input_tokens": total, "scored_questions": count}}. Choice ties select the earliest candidate in insertion order. Score returns the expected ordinal index, not an arbitrary supplied numeric value; this adapter does not implement a supplied-values extension. Noul 0.5 is interpreted as true. Confidence is (K * max(p) - 1)/(K - 1), clipped to [0,1], and is not a claimed reproduction of Jev's confidence statistic. Shipped temperature calibration does not make every confidence value a correctness guarantee.
Question names are preserved as opaque bookkeeping IDs; candidate IDs and descriptions use the frozen renderer. Native objects are deterministically serialized with sorted JSON keys; strings preserve their contents. Do not interpret object/string field-order differences as identical token inputs. Non-finite or non-JSON inputs are rejected.
The bound is 16,384 tokens per complete question, including its state, instructions, all candidates and readout suffix. Questions are separate sequences, grouped in fixed batches of eight; the state is repeated for each question and counted repeatedly in usage.input_tokens. All questions are encoded before any forward pass. If any exceeds the bound, the whole call raises ValueError without truncation or partial answers. A lower max_length can be chosen at load time; a higher limit is rejected. This native wrapper does not impose the Studio's separate 16-question UI limit.
Runtime boundary
The qualified device is an AMD ROCm GPU accessed as cuda:0 in PyTorch. Backbone parameters remain BF16, the candidate head FP32, with FLA Gated DeltaNet and SDPA. CPU and MPS inference are not implemented by the current engine and are rejected explicitly. Other GPU/runtime combinations require validation. The loader checks manifest hashes and the recorded runtime before loading. allow_unvalidated_runtime=True downgrades runtime differences to an explicit warning for experiments; it does not silently substitute a claimed validated environment or bypass unsupported CPU execution.
This package is a loading wrapper, not a trainer, HTTP service or new model selection policy. No prediction logs or private local paths are required by the bundle.