Instructions to use vllm-sr/Decision-2.0-Sol-2B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use vllm-sr/Decision-2.0-Sol-2B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="vllm-sr/Decision-2.0-Sol-2B", trust_remote_code=True)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("vllm-sr/Decision-2.0-Sol-2B", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Decision-2.0-Sol-2B
Decision-2.0-Sol-2B is the 2B model of Decision 2.0, the decision models of vLLM Semantic Router. Give it an input (text or JSON) and the questions you need answered: pick one of several options, say yes or no, or rate on a scale. It answers them all at once and returns a probability for every answer, without generating text.
| Parameters | 1.88B |
| Context length | 16,384 tokens |
| Decision types | Choice · Yes / No · Score |
| License | Apache-2.0 |
Highlights
- Top JevArena score of its size: 52.1, ahead of the 4 other same-size models compared.
- Ahead of Decision 1.0 Sol: +6.3 on JevArena and +4.2 on the Jev Decision Index.
- Speed: a median of 7.2 ms per single-question request on a single GPU.
- Many questions, one pass: Choice, Yes / No and Score questions about the same input are answered together in one forward pass, with a probability for every option.
Quickstart
pip install "transformers>=5.17" torch safetensors
import json
from transformers import AutoModel
model = AutoModel.from_pretrained("vllm-sr/Decision-2.0-Sol-2B", trust_remote_code=True)
result = model.system_one(
state="The order arrived damaged yesterday. The customer has a receipt and asks for a replacement today.",
questions={
"route": {
"type": "choice",
"instructions": "Which team should handle this request?",
"criteria": {
"returns": "Refunds, replacements and damaged deliveries",
"billing": "Payments, invoices and charges",
"technical": "Product setup and faults"
}
},
"receipt": {
"type": "noul",
"instructions": "Does the customer have a receipt?"
},
"urgency": {
"type": "score",
"instructions": "How urgent is this request?",
"criteria": [
"Routine",
"Soon",
"Today"
]
}
},
)
print(json.dumps(result["answers"], indent=2))
# Or as a pipeline:
# transformers.pipeline("decision", model="vllm-sr/Decision-2.0-Sol-2B", trust_remote_code=True)(state=..., questions=...)
Evaluation
| Model | JevArena ↑ | Human-labelled transfer ↑ | Jev Decision Index ↑ |
|---|---|---|---|
| Decision-2.0-Sol-2B | 52.1 | 51.3 | 29.5 |
| Decider 2B | 49.5 | 42.0 | — |
| This-That 1.2 | 46.1 | 40.5 | — |
| Decision 1.0 Sol | 45.8 | 49.3 | 25.3 |
| Bosun v3.1 1.7B | 42.1 | 38.0 | — |
JevArena
Every model answers the same frozen prompts, scored the same way; missing or invalid answers count as errors. Human-labelled transfer is the median macro-F1 over 15 human-labelled tasks (×100).
Jev Decision Index
Decision 2.0: independent reproduction with the official 0.2.1 kit on the released weights; others: public board snapshot, 2026-09-28. Training data audited at row level against all Index test items.
License
Apache-2.0 (LICENSE).
Citation
@misc{decision_2_0_sol_2b_2026,
title = {{Decision-2.0-Sol-2B}: A Decision 2.0 Model for Structured Decisions},
author = {{vLLM Semantic Router Team}},
year = {2026},
howpublished = {\url{https://huggingface.co/vllm-sr/Decision-2.0-Sol-2B}}
}
- Downloads last month
- 72




