Download README.md from vllm-sr/Decision-1.0-Nox-4B: direct link, hf CLI and curl.
- Browser
- Download file 4.67 kB
-
https://huggingface.co/vllm-sr/Decision-1.0-Nox-4B/resolve/46505c737a45cbe2c4ac4eee38e4f94cb520e4ed/README.md
- Command line
-
hf download hf://vllm-sr/Decision-1.0-Nox-4B@46505c737a45cbe2c4ac4eee38e4f94cb520e4ed/README.md
-
curl -L -o README.md https://huggingface.co/vllm-sr/Decision-1.0-Nox-4B/resolve/46505c737a45cbe2c4ac4eee38e4f94cb520e4ed/README.md
license: apache-2.0
language:
- en
- zh
base_model: Qwen/Qwen3.5-4B
base_model_relation: finetune
tags:
- decision-model
- classification
- qwen3_5
- custom-code
- pytorch
- rocm
Decision-1.0-Nox
Nox, Latin for night.
Your move. Give Nox a state, questions and possible answers. It returns typed decisions and probabilities, with labels defined at runtime.
4.208B parameters · 16K complete-question budget · English / Chinese evaluated · Apache 2.0
| Type | Use it for | Output |
|---|---|---|
| Choice | Route a request or choose among 2–255 actions. | Selected ID + distribution |
| Noul | Check a condition against supplied evidence. | P(true) |
| Score | Apply 2–10 ordered rubric descriptions. | Expected index + distribution |
Measured capability
72.84% overall accuracy across 3,766 scored decisions and 54 tasks. Nox leads Kev-4B by 2.75 percentage points on this decision-focused comparison; reading and transfer remain opportunities to improve.
| Model | Size | Decisions | Composition | Reading | Inference | Transfer | Overall |
|---|---|---|---|---|---|---|---|
| Jev | — | 79.10 | 66.38 | 94.53 | 89.79 | 87.19 | 81.05 |
| Lux | 9B | 83.21 | 51.88 | 90.31 | 90.83 | 77.44 | 76.72 |
| Nox | 4B | 83.00 | 51.79 | 79.06 | 86.25 | 67.97 | 72.84 |
| Kev-9B | 9B | 76.75 | 45.75 | 86.72 | 83.54 | 79.25 | 71.89 |
| Kev-4B | 4B | 71.90 | 48.54 | 81.88 | 84.58 | 76.10 | 70.09 |
| Qwen3.5-9B | 9B | 73.91 | 44.62 | 89.84 | 79.58 | 73.23 | 69.73 |
| Decider | 2B | 64.01 | 46.58 | 92.03 | 84.38 | 69.31 | 67.71 |
| Qwen3.5-4B | 4B | 69.89 | 43.33 | 87.97 | 79.79 | 68.83 | 67.29 |
| Sol | 2B | 73.75 | 46.08 | 76.56 | 84.17 | 57.07 | 66.32 |
| Kev-0.8B | 0.8B | 60.14 | 42.29 | 67.81 | 68.75 | 61.19 | 58.28 |
| Qwen3.5-2B | 2B | 57.12 | 39.00 | 73.75 | 72.29 | 56.31 | 57.24 |
| Laya · English | 0.421B | 56.54 | 35.33 | 51.41 | 63.75 | 53.06 | 51.03 |
| Laya · Multilingual | 0.322B | 47.25 | 38.92 | 50.78 | 57.29 | 47.13 | 47.19 |
Accuracy (%). Overall weights: Decisions 30%, Composition 25%, Reading 15%, Inference 15%, Transfer 15%. These outcome-informed product-priority weights were chosen after observing results; reweighting is not a training improvement. Bold marks Decision-family cells above every external open or untuned reference for that metric, excluding Jev and the other Decision models.
All 54 tasks · Order, missing-evidence and calibration diagnostics · Methods and uncertainty
More questions, one request
Distinct Choice questions at a fixed 499 input tokens per question. Thirty measurements per point across six independently loaded processes on an otherwise idle AMD gfx942 GPU. Python latency includes tokenization and inference; loading and network are excluded. p50, p95 and memory.
Try it
Download the current model with hf download llm-semantic-router/Decision-1.0-Nox --local-dir decision-model, then follow ROCm setup. In that container, with the model mounted at /model:
from decision import DecisionModel
from decision.example import REQUEST
model = DecisionModel.from_pretrained("/model", local_files_only=True)
print(model.decide(**REQUEST)["answers"])
Tested SystemOne request and output · Install and typed API guide
The complete state, question and candidates must fit 16,384 tokens; overflow is rejected. The bundled normalization profile loads automatically. AMD gfx942 is validated; CPU/MPS are unsupported and NVIDIA is unqualified. Use a fresh Python process when switching profiles.
Architecture
A causal Qwen3.5 text backbone combines gated linear and full attention. A shared candidate head reads candidate endpoints and the final query vector. Each question uses one forward pass; questions run independently in batches of eight.
Candidate head · Vector architecture · Inference code
Adapted from Qwen3.5-4B. It evaluates supplied evidence without live retrieval; confidence does not guarantee correctness. License · Attributions.



