Feature Extraction
Transformers
Safetensors
decision2
decision-model
classification
system-one
custom_code
Instructions to use vllm-sr/Decision-2.0-Sol-2B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use vllm-sr/Decision-2.0-Sol-2B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="vllm-sr/Decision-2.0-Sol-2B", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("vllm-sr/Decision-2.0-Sol-2B", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Decision 2.0 package Decision-2.0-Sol-2B (release) from 088818218df33b6c2dca39badffde47962125d41-src_training_decision2
Browse files- MODEL_MANIFEST.json +7 -7
- README.md +5 -5
MODEL_MANIFEST.json
CHANGED
|
@@ -2,7 +2,7 @@
|
|
| 2 |
"base": null,
|
| 3 |
"builder": {
|
| 4 |
"module_sha256": "7f74839fb11762d66fa4cdc9657e7dfa8f4041bb20293dedbf69d54640fe0f0d",
|
| 5 |
-
"source_commit": "
|
| 6 |
},
|
| 7 |
"calibration": null,
|
| 8 |
"card": {
|
|
@@ -15,11 +15,11 @@
|
|
| 15 |
"assets/jevarena.png": "0d97e2ecf10414cf211a37e8f8f16fbebddbc1dc5a85f4a243b1ce029d489081"
|
| 16 |
},
|
| 17 |
"index_sha256": "97a95897fb13b81addf033dee7984d53e996ddb8baad5076f706a0b7c752e22e",
|
| 18 |
-
"readme_sha256": "
|
| 19 |
},
|
| 20 |
"files_sha256": {
|
| 21 |
"LICENSE": "bbedc3fda3305820b977265f01b8619d87570a6739de3a5582c3464840f1e57a",
|
| 22 |
-
"README.md": "
|
| 23 |
"assets/banner.png": "2363db1a11ea3d8c61656a078c2b2cdf10623177e54ca3ac5a48b22ccb0be724",
|
| 24 |
"assets/index-areas.png": "3670b08f2c186c03a78028ad0c0c9ab07a4475ffb33b81ce88daea600b1a0d37",
|
| 25 |
"assets/index-pareto.png": "bbc3b1c70755ff97ed5943bfa65f5b1c1e24041b0923c102610104182db5ecba",
|
|
@@ -71,12 +71,12 @@
|
|
| 71 |
{
|
| 72 |
"component": "Decision-2.0-Sol-2B weights, decision head, package runtime, card and artwork",
|
| 73 |
"licence": "apache-2.0",
|
| 74 |
-
"source": "
|
| 75 |
},
|
| 76 |
{
|
| 77 |
"component": "Decision 1.0 Sol-2B weights (direct weight origin, fully fine-tuned)",
|
| 78 |
"licence": "apache-2.0",
|
| 79 |
-
"source": "
|
| 80 |
},
|
| 81 |
{
|
| 82 |
"component": "Qwen3.5-2B text backbone and tokenizer (upstream of Decision 1.0 Sol)",
|
|
@@ -101,7 +101,7 @@
|
|
| 101 |
"model_name": "Decision-2.0-Sol-2B",
|
| 102 |
"origin": {
|
| 103 |
"relation": "finetune",
|
| 104 |
-
"repo_id": "
|
| 105 |
"revision": "ce0c018a28de16d6639b1cd203b761bf643b89e6",
|
| 106 |
"summary": "Every weight of Decision 1.0 Sol was fine-tuned (nothing frozen, no adapter). The release is the uniform average of two seeds of that fine-tune, trained with the previous Decision 2.0 Sol 2B's answer probabilities as soft targets (self-distillation) and with extra copies of released multilingual rows that keep the released multilingual token share. Decision 1.0 Sol is itself a text-only fine-tune of [Qwen/Qwen3.5-2B](https://huggingface.co/Qwen/Qwen3.5-2B) at `15852e8c16360a2fea060d615a32b45270f8a8fc` (Apache-2.0), whose text backbone and tokenizer this model inherits; the Qwen3.5 vision tower is not part of it."
|
| 107 |
},
|
|
@@ -172,7 +172,7 @@
|
|
| 172 |
"5.18.0"
|
| 173 |
]
|
| 174 |
},
|
| 175 |
-
"repo_id": "
|
| 176 |
"runtime": {
|
| 177 |
"equivalence": "decision2/qwen.py loads this full checkpoint with the training/model sources vendored from the scored run's own runner mirror (5b246b110), whose SHA-256 equal the scored adapter sources (checked at build time), and applies the per-item batching, BF16-backbone / FP32-head execution, raw probabilities (temperature 1; no calibration file) and answer normalization of v2.dec.infer_dec, which for this non-residual checkpoint wraps the same DecisionModel without extra readouts. The scored checkpoint (8c8e98e3) stored every tensor in FP32; this package (v2.release.bf16_copy) stores its Linear projection matrices in BF16 exactly as BF16 autocast rounds them and every other tensor bit for bit in FP32, and the runtime holds the backbone's BF16-exact Linear weights in BF16, the values BF16 autocast multiplies with. The runtime, the Transformers remote code (AutoModel with trust_remote_code) and the forward token budget are the current revision's. Checked on one GPU, in the scored image with a copy of the persisted Triton autotune cache of the formal run, against the sealed formal predictions of every scored prompt (typed-final 1,600, css15 6,547, public231 231) by release.sh --parity before and after the download, and AutoModel against the native runtime on every scored prompt.",
|
| 178 |
"requirements": {
|
|
|
|
| 2 |
"base": null,
|
| 3 |
"builder": {
|
| 4 |
"module_sha256": "7f74839fb11762d66fa4cdc9657e7dfa8f4041bb20293dedbf69d54640fe0f0d",
|
| 5 |
+
"source_commit": "088818218df33b6c2dca39badffde47962125d41"
|
| 6 |
},
|
| 7 |
"calibration": null,
|
| 8 |
"card": {
|
|
|
|
| 15 |
"assets/jevarena.png": "0d97e2ecf10414cf211a37e8f8f16fbebddbc1dc5a85f4a243b1ce029d489081"
|
| 16 |
},
|
| 17 |
"index_sha256": "97a95897fb13b81addf033dee7984d53e996ddb8baad5076f706a0b7c752e22e",
|
| 18 |
+
"readme_sha256": "0c547477c690faabd9e066c0fa51fa98e9d5388e225bea4703e3e5a9dfe08c24"
|
| 19 |
},
|
| 20 |
"files_sha256": {
|
| 21 |
"LICENSE": "bbedc3fda3305820b977265f01b8619d87570a6739de3a5582c3464840f1e57a",
|
| 22 |
+
"README.md": "0c547477c690faabd9e066c0fa51fa98e9d5388e225bea4703e3e5a9dfe08c24",
|
| 23 |
"assets/banner.png": "2363db1a11ea3d8c61656a078c2b2cdf10623177e54ca3ac5a48b22ccb0be724",
|
| 24 |
"assets/index-areas.png": "3670b08f2c186c03a78028ad0c0c9ab07a4475ffb33b81ce88daea600b1a0d37",
|
| 25 |
"assets/index-pareto.png": "bbc3b1c70755ff97ed5943bfa65f5b1c1e24041b0923c102610104182db5ecba",
|
|
|
|
| 71 |
{
|
| 72 |
"component": "Decision-2.0-Sol-2B weights, decision head, package runtime, card and artwork",
|
| 73 |
"licence": "apache-2.0",
|
| 74 |
+
"source": "vllm-sr/Decision-2.0-Sol-2B"
|
| 75 |
},
|
| 76 |
{
|
| 77 |
"component": "Decision 1.0 Sol-2B weights (direct weight origin, fully fine-tuned)",
|
| 78 |
"licence": "apache-2.0",
|
| 79 |
+
"source": "vllm-sr/Decision-1.0-Sol-2B@ce0c018a"
|
| 80 |
},
|
| 81 |
{
|
| 82 |
"component": "Qwen3.5-2B text backbone and tokenizer (upstream of Decision 1.0 Sol)",
|
|
|
|
| 101 |
"model_name": "Decision-2.0-Sol-2B",
|
| 102 |
"origin": {
|
| 103 |
"relation": "finetune",
|
| 104 |
+
"repo_id": "vllm-sr/Decision-1.0-Sol-2B",
|
| 105 |
"revision": "ce0c018a28de16d6639b1cd203b761bf643b89e6",
|
| 106 |
"summary": "Every weight of Decision 1.0 Sol was fine-tuned (nothing frozen, no adapter). The release is the uniform average of two seeds of that fine-tune, trained with the previous Decision 2.0 Sol 2B's answer probabilities as soft targets (self-distillation) and with extra copies of released multilingual rows that keep the released multilingual token share. Decision 1.0 Sol is itself a text-only fine-tune of [Qwen/Qwen3.5-2B](https://huggingface.co/Qwen/Qwen3.5-2B) at `15852e8c16360a2fea060d615a32b45270f8a8fc` (Apache-2.0), whose text backbone and tokenizer this model inherits; the Qwen3.5 vision tower is not part of it."
|
| 107 |
},
|
|
|
|
| 172 |
"5.18.0"
|
| 173 |
]
|
| 174 |
},
|
| 175 |
+
"repo_id": "vllm-sr/Decision-2.0-Sol-2B",
|
| 176 |
"runtime": {
|
| 177 |
"equivalence": "decision2/qwen.py loads this full checkpoint with the training/model sources vendored from the scored run's own runner mirror (5b246b110), whose SHA-256 equal the scored adapter sources (checked at build time), and applies the per-item batching, BF16-backbone / FP32-head execution, raw probabilities (temperature 1; no calibration file) and answer normalization of v2.dec.infer_dec, which for this non-residual checkpoint wraps the same DecisionModel without extra readouts. The scored checkpoint (8c8e98e3) stored every tensor in FP32; this package (v2.release.bf16_copy) stores its Linear projection matrices in BF16 exactly as BF16 autocast rounds them and every other tensor bit for bit in FP32, and the runtime holds the backbone's BF16-exact Linear weights in BF16, the values BF16 autocast multiplies with. The runtime, the Transformers remote code (AutoModel with trust_remote_code) and the forward token budget are the current revision's. Checked on one GPU, in the scored image with a copy of the persisted Triton autotune cache of the formal run, against the sealed formal predictions of every scored prompt (typed-final 1,600, css15 6,547, public231 231) by release.sh --parity before and after the download, and AutoModel against the native runtime on every scored prompt.",
|
| 178 |
"requirements": {
|
README.md
CHANGED
|
@@ -1,6 +1,6 @@
|
|
| 1 |
---
|
| 2 |
license: apache-2.0
|
| 3 |
-
base_model:
|
| 4 |
base_model_relation: finetune
|
| 5 |
library_name: transformers
|
| 6 |
tags:
|
|
@@ -14,7 +14,7 @@ tags:
|
|
| 14 |
|
| 15 |
# Decision-2.0-Sol-2B
|
| 16 |
|
| 17 |
-
**Decision-2.0-Sol-2B** is the 2B model of [Decision 2.0](https://huggingface.co/collections/
|
| 18 |
|
| 19 |
| | |
|
| 20 |
| --- | --- |
|
|
@@ -41,7 +41,7 @@ import json
|
|
| 41 |
|
| 42 |
from transformers import AutoModel
|
| 43 |
|
| 44 |
-
model = AutoModel.from_pretrained("
|
| 45 |
result = model.system_one(
|
| 46 |
state="The order arrived damaged yesterday. The customer has a receipt and asks for a replacement today.",
|
| 47 |
questions={
|
|
@@ -72,7 +72,7 @@ result = model.system_one(
|
|
| 72 |
print(json.dumps(result["answers"], indent=2))
|
| 73 |
|
| 74 |
# Or as a pipeline:
|
| 75 |
-
# transformers.pipeline("decision", model="
|
| 76 |
```
|
| 77 |
|
| 78 |
## Evaluation
|
|
@@ -112,6 +112,6 @@ Apache-2.0 ([LICENSE](LICENSE)).
|
|
| 112 |
title = {{Decision-2.0-Sol-2B}: A Decision 2.0 Model for Structured Decisions},
|
| 113 |
author = {{vLLM Semantic Router Team}},
|
| 114 |
year = {2026},
|
| 115 |
-
howpublished = {\url{https://huggingface.co/
|
| 116 |
}
|
| 117 |
```
|
|
|
|
| 1 |
---
|
| 2 |
license: apache-2.0
|
| 3 |
+
base_model: vllm-sr/Decision-1.0-Sol-2B
|
| 4 |
base_model_relation: finetune
|
| 5 |
library_name: transformers
|
| 6 |
tags:
|
|
|
|
| 14 |
|
| 15 |
# Decision-2.0-Sol-2B
|
| 16 |
|
| 17 |
+
**Decision-2.0-Sol-2B** is the 2B model of [Decision 2.0](https://huggingface.co/collections/vllm-sr/decision-20-6ab7cf7bdfb506bf8269cb00), the decision models of [vLLM Semantic Router](https://github.com/vllm-project/semantic-router). Give it an input (text or JSON) and the questions you need answered: pick one of several options, say yes or no, or rate on a scale. It answers them all at once and returns a probability for every answer, without generating text.
|
| 18 |
|
| 19 |
| | |
|
| 20 |
| --- | --- |
|
|
|
|
| 41 |
|
| 42 |
from transformers import AutoModel
|
| 43 |
|
| 44 |
+
model = AutoModel.from_pretrained("vllm-sr/Decision-2.0-Sol-2B", trust_remote_code=True)
|
| 45 |
result = model.system_one(
|
| 46 |
state="The order arrived damaged yesterday. The customer has a receipt and asks for a replacement today.",
|
| 47 |
questions={
|
|
|
|
| 72 |
print(json.dumps(result["answers"], indent=2))
|
| 73 |
|
| 74 |
# Or as a pipeline:
|
| 75 |
+
# transformers.pipeline("decision", model="vllm-sr/Decision-2.0-Sol-2B", trust_remote_code=True)(state=..., questions=...)
|
| 76 |
```
|
| 77 |
|
| 78 |
## Evaluation
|
|
|
|
| 112 |
title = {{Decision-2.0-Sol-2B}: A Decision 2.0 Model for Structured Decisions},
|
| 113 |
author = {{vLLM Semantic Router Team}},
|
| 114 |
year = {2026},
|
| 115 |
+
howpublished = {\url{https://huggingface.co/vllm-sr/Decision-2.0-Sol-2B}}
|
| 116 |
}
|
| 117 |
```
|