--- license: apache-2.0 language: - en - zh base_model: Qwen/Qwen3.5-4B base_model_relation: finetune tags: - decision-model - classification - qwen3_5 - custom-code - pytorch - rocm --- # Decision-1.0-Nox *Nox, Latin for night.* **Your move.** Give Nox a state, questions and possible answers. It returns typed decisions and probabilities, with labels defined at runtime. **4.208B parameters · 16K complete-question budget · English / Chinese evaluated · Apache 2.0** [Decision family](https://huggingface.co/collections/llm-semantic-router/decision-10-6ab12177bd0002394d8409f9) | Type | Use it for | Output | |---|---|---| | **Choice** | Route a request or choose among 2–255 actions. | Selected ID + distribution | | **Noul** | Check a condition against supplied evidence. | P(true) | | **Score** | Apply 2–10 ordered rubric descriptions. | Expected index + distribution | ## Measured capability **72.84% overall accuracy** across 3,766 scored decisions and 54 tasks. Nox leads Kev-4B by **2.75 percentage points** on this decision-focused comparison; reading and transfer remain opportunities to improve. | Model | Size | Decisions | Composition | Reading | Inference | Transfer | Overall | |---|---:|---:|---:|---:|---:|---:|---:| | Jev | — | 79.10 | 66.38 | 94.53 | 89.79 | 87.19 | 81.05 | | Lux | 9B | **83.21** | **51.88** | 90.31 | **90.83** | 77.44 | **76.72** | | Nox | 4B | **83.00** | **51.79** | 79.06 | **86.25** | 67.97 | **72.84** | | Kev-9B | 9B | 76.75 | 45.75 | 86.72 | 83.54 | 79.25 | 71.89 | | Kev-4B | 4B | 71.90 | 48.54 | 81.88 | 84.58 | 76.10 | 70.09 | | Qwen3.5-9B | 9B | 73.91 | 44.62 | 89.84 | 79.58 | 73.23 | 69.73 | | Decider | 2B | 64.01 | 46.58 | 92.03 | 84.38 | 69.31 | 67.71 | | Qwen3.5-4B | 4B | 69.89 | 43.33 | 87.97 | 79.79 | 68.83 | 67.29 | | Sol | 2B | 73.75 | 46.08 | 76.56 | 84.17 | 57.07 | 66.32 | | Kev-0.8B | 0.8B | 60.14 | 42.29 | 67.81 | 68.75 | 61.19 | 58.28 | | Qwen3.5-2B | 2B | 57.12 | 39.00 | 73.75 | 72.29 | 56.31 | 57.24 | | Laya · English | 0.421B | 56.54 | 35.33 | 51.41 | 63.75 | 53.06 | 51.03 | | Laya · Multilingual | 0.322B | 47.25 | 38.92 | 50.78 | 57.29 | 47.13 | 47.19 | Accuracy (%). Overall weights: Decisions **30%**, Composition **25%**, Reading **15%**, Inference **15%**, Transfer **15%**. These outcome-informed product-priority weights were chosen after observing results; reweighting is not a training improvement. Bold marks Decision-family cells above every external open or untuned reference for that metric, excluding Jev and the other Decision models. ![Decision model ranking](assets/decision-ranking.png) ![Capability matrix](assets/decision-matrix.png) [All 54 tasks](TASKS.md) · [Order, missing-evidence and calibration diagnostics](DIAGNOSTICS.md) · [Methods and uncertainty](EVALUATION.md) ## More questions, one request ![Question-count latency](assets/decision-question-scaling.png) Distinct Choice questions at a fixed **499 input tokens per question**. Thirty measurements per point across six independently loaded processes on an otherwise idle AMD gfx942 GPU. Python latency includes tokenization and inference; loading and network are excluded. [p50, p95 and memory](QUESTION-SCALING.md). ## Try it Download the current model with `hf download llm-semantic-router/Decision-1.0-Nox --local-dir decision-model`, then follow [ROCm setup](RUNTIME.md). In that container, with the model mounted at `/model`: ```python from decision import DecisionModel from decision.example import REQUEST model = DecisionModel.from_pretrained("/model", local_files_only=True) print(model.decide(**REQUEST)["answers"]) ``` [Tested SystemOne request and output](model-card-example.json) · [Install and typed API guide](USAGE.md) The complete state, question and candidates must fit 16,384 tokens; overflow is rejected. The bundled normalization profile loads automatically. AMD gfx942 is validated; CPU/MPS are unsupported and NVIDIA is unqualified. Use a fresh Python process when switching profiles. ## Architecture ![Decision decoder architecture](assets/architecture.png) A causal Qwen3.5 text backbone combines gated linear and full attention. A shared candidate head reads candidate endpoints and the final query vector. Each question uses one forward pass; questions run independently in batches of eight. [Candidate head](assets/readout.png) · [Vector architecture](assets/architecture.svg) · [Inference code](code/decision_model.py) Adapted from [Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B). It evaluates supplied evidence without live retrieval; confidence does not guarantee correctness. [License](LICENSE) · [Attributions](ATTRIBUTIONS.md).