Decision-1.0-Nox-4B / README.md
Xunzhuo's picture
Update public comparison roster; preserve model weights and benchmark scores
46505c7 verified
|
Raw History Blame
4.67 kB
metadata
license: apache-2.0
language:
  - en
  - zh
base_model: Qwen/Qwen3.5-4B
base_model_relation: finetune
tags:
  - decision-model
  - classification
  - qwen3_5
  - custom-code
  - pytorch
  - rocm

Decision-1.0-Nox

Nox, Latin for night.

Your move. Give Nox a state, questions and possible answers. It returns typed decisions and probabilities, with labels defined at runtime.

4.208B parameters · 16K complete-question budget · English / Chinese evaluated · Apache 2.0

Decision family

Type Use it for Output
Choice Route a request or choose among 2–255 actions. Selected ID + distribution
Noul Check a condition against supplied evidence. P(true)
Score Apply 2–10 ordered rubric descriptions. Expected index + distribution

Measured capability

72.84% overall accuracy across 3,766 scored decisions and 54 tasks. Nox leads Kev-4B by 2.75 percentage points on this decision-focused comparison; reading and transfer remain opportunities to improve.

Model Size Decisions Composition Reading Inference Transfer Overall
Jev — 79.10 66.38 94.53 89.79 87.19 81.05
Lux 9B 83.21 51.88 90.31 90.83 77.44 76.72
Nox 4B 83.00 51.79 79.06 86.25 67.97 72.84
Kev-9B 9B 76.75 45.75 86.72 83.54 79.25 71.89
Kev-4B 4B 71.90 48.54 81.88 84.58 76.10 70.09
Qwen3.5-9B 9B 73.91 44.62 89.84 79.58 73.23 69.73
Decider 2B 64.01 46.58 92.03 84.38 69.31 67.71
Qwen3.5-4B 4B 69.89 43.33 87.97 79.79 68.83 67.29
Sol 2B 73.75 46.08 76.56 84.17 57.07 66.32
Kev-0.8B 0.8B 60.14 42.29 67.81 68.75 61.19 58.28
Qwen3.5-2B 2B 57.12 39.00 73.75 72.29 56.31 57.24
Laya · English 0.421B 56.54 35.33 51.41 63.75 53.06 51.03
Laya · Multilingual 0.322B 47.25 38.92 50.78 57.29 47.13 47.19

Accuracy (%). Overall weights: Decisions 30%, Composition 25%, Reading 15%, Inference 15%, Transfer 15%. These outcome-informed product-priority weights were chosen after observing results; reweighting is not a training improvement. Bold marks Decision-family cells above every external open or untuned reference for that metric, excluding Jev and the other Decision models.

Decision model ranking

Capability matrix

All 54 tasks · Order, missing-evidence and calibration diagnostics · Methods and uncertainty

More questions, one request

Question-count latency

Distinct Choice questions at a fixed 499 input tokens per question. Thirty measurements per point across six independently loaded processes on an otherwise idle AMD gfx942 GPU. Python latency includes tokenization and inference; loading and network are excluded. p50, p95 and memory.

Try it

Download the current model with hf download llm-semantic-router/Decision-1.0-Nox --local-dir decision-model, then follow ROCm setup. In that container, with the model mounted at /model:

from decision import DecisionModel
from decision.example import REQUEST

model = DecisionModel.from_pretrained("/model", local_files_only=True)
print(model.decide(**REQUEST)["answers"])

Tested SystemOne request and output · Install and typed API guide

The complete state, question and candidates must fit 16,384 tokens; overflow is rejected. The bundled normalization profile loads automatically. AMD gfx942 is validated; CPU/MPS are unsupported and NVIDIA is unqualified. Use a fresh Python process when switching profiles.

Architecture

Decision decoder architecture

A causal Qwen3.5 text backbone combines gated linear and full attention. A shared candidate head reads candidate endpoints and the final query vector. Each question uses one forward pass; questions run independently in batches of eight.

Candidate head · Vector architecture · Inference code

Adapted from Qwen3.5-4B. It evaluates supplied evidence without live retrieval; confidence does not guarantee correctness. License · Attributions.