Decision-1.0-Sol-2B
Sol, Latin for sun.
Give Sol a state, questions and possible answers. It returns decisions and probabilities with labels defined at runtime.
| Type | Use it for | Output |
|---|---|---|
| Choice | Route a request or choose among 2–255 actions. | Selected ID + distribution |
| Noul | Check a condition against supplied evidence. | P(true) |
| Score | Apply 2–10 ordered rubric descriptions. | Expected index + distribution |
Measured capability
66.32% weighted accuracy across 3,766 decisions and 54 tasks. Compare decision, reading and transfer capabilities in the complete results below.
| Model | Size | Decisions | Composition | Reading | Inference | Transfer | Overall |
|---|---|---|---|---|---|---|---|
| Sol-2B | 2B | 73.75 | 46.08 | 76.56 | 84.17 | 57.07 | 66.32 |
| Lux-9B | 9B | 84.38 | 52.75 | 90.16 | 91.46 | 77.72 | 77.40 |
| Nox-4B | 4B | 83.00 | 51.79 | 79.06 | 86.25 | 69.60 | 73.09 |
| Kev-9B | 9B | 76.75 | 45.75 | 86.72 | 83.54 | 79.25 | 71.89 |
| Kev-4B | 4B | 71.90 | 48.54 | 81.88 | 84.58 | 76.10 | 70.09 |
| Qwen3.5-9B | 9B | 73.91 | 44.62 | 89.84 | 79.58 | 73.23 | 69.73 |
| Decider | 2B | 64.01 | 46.58 | 92.03 | 84.38 | 69.31 | 67.71 |
| Qwen3.5-4B | 4B | 69.89 | 43.33 | 87.97 | 79.79 | 68.83 | 67.29 |
| Eos-0.8B | 0.8B | 65.94 | 46.04 | 70.31 | 81.67 | 52.01 | 61.89 |
| Kev-0.8B | 0.8B | 60.14 | 42.29 | 67.81 | 68.75 | 61.19 | 58.28 |
| Qwen3.5-2B | 2B | 57.12 | 39.00 | 73.75 | 72.29 | 56.31 | 57.24 |
| Kai-0.6B | 0.6B | 57.96 | 40.83 | 54.69 | 69.79 | 48.37 | 53.52 |
| Laya · English | 0.421B | 56.54 | 35.33 | 51.41 | 63.75 | 53.06 | 51.03 |
| Laya · Multilingual | 0.322B | 47.25 | 38.92 | 50.78 | 57.29 | 47.13 | 47.19 |
| Jev | — | 79.10 | 66.38 | 94.53 | 89.79 | 87.19 | 81.05 |
Accuracy (%). Overall weights: Decisions 30%, Composition 25%, Reading 15%, Inference 15%, Transfer 15%. These outcome-informed product-priority weights were chosen after observing results. Bold marks a Decision-family cell above every external open or untuned reference; Jev and other Decision models are excluded.
All 54 tasks · Probability, order and missing-evidence diagnostics · Methods and uncertainty
More questions, measured
Distinct Choice questions at a fixed 499 input tokens per question. Thirty measurements per point across six independently loaded processes on an otherwise idle AMD gfx942 GPU. Python request latency includes tokenization and inference; loading and network are excluded. p50, p95 and memory.
Download the complete model repository
hf download llm-semantic-router/Decision-1.0-Sol-2B --local-dir Decision-1.0-Sol-2B
This downloads the complete model release. The root config.json lists the backbone, tokenizer, decision head, and calibration files.
Use with 🤗 Transformers
The repository includes its inference code, so stock Transformers can download and run the complete model locally with trust_remote_code=True. system_one takes and returns the same System One request and response bodies as the Decision runtime; nothing is generated.
pip install "transformers>=5.17" torch safetensors huggingface_hub
from transformers import AutoModel
model = AutoModel.from_pretrained("llm-semantic-router/Decision-1.0-Sol-2B", trust_remote_code=True)
response = model.system_one(
state="Customer requests a refund.",
questions={
"route": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {"billing": "Payments and refunds", "technical": "Product faults"},
},
"priority": {
"type": "score",
"instructions": "How urgent is this request?",
"criteria": ["Not urgent", "Soon", "Today", "Immediately"],
},
},
)
print(response["answers"]["route"]["choice"], response["answers"]["priority"]["score"])
pipeline("decision", model="llm-semantic-router/Decision-1.0-Sol-2B", trust_remote_code=True) accepts the same request body. A malformed question is answered with an invalid_question error. The model loads on the first GPU when one is visible, otherwise on the CPU (pass device="cpu" or device="cuda:0" to choose). On a GPU the backbone runs in BF16 with an FP32 decision head; on the CPU everything runs in FP32. A complete question, its candidates and the state are limited to 16,384 tokens; if a question is longer, every question of the request is answered with a max_length_exceeded error and nothing is truncated. flash-linear-attention speeds up the linear-attention layers on a GPU.
Serve with vLLM Semantic Router
This repository contains model data and its Transformers loading code. Use the vLLM Semantic Router Decision runtime to load llm-semantic-router/Decision-1.0-Sol-2B and serve Choice, Noul, and Score requests. The serving implementation and its dependencies live in vLLM Semantic Router; this release bundles no serving code. The weights do not depend on a particular accelerator; which hardware can serve them is decided by the runtime. For local inference without a server, see Use with 🤗 Transformers above.
After configuring a compatible Decision endpoint, send a SystemOne request (replace the placeholder URL and key):
curl -X POST 'https://your-decision-endpoint.example/v1/systemone' \
-H 'Authorization: Bearer YOUR_ENDPOINT_API_KEY' \
-H 'Content-Type: application/json' \
--data-raw '{"model":"Decision-1.0-Sol-2B","state":"Customer requests a refund.","questions":{"route":{"type":"choice","instructions":"Which team should handle this?","criteria":{"billing":"Payments and refunds","technical":"Product faults"}}}}'
The Hugging Face repository is a model download, not a hosted inference endpoint.
The published model's complete state, question, and candidates have a 16,384-token input limit. See the evaluation scope for measured conditions.
Architecture
A causal Qwen3.5 text backbone combines gated linear and full attention. A shared candidate head reads candidate endpoints and the final query vector. The serving runtime schedules questions according to available hardware and request load.
Candidate head · Vector architecture
Adapted from Qwen3.5-2B. It evaluates supplied evidence without live retrieval; confidence does not guarantee correctness. License · Attributions.
- Downloads last month
- 82




