Xunzhuo commited on
Commit
951e7f7
·
verified ·
1 Parent(s): 1b7c47e

Decision 2.0 package Decision-2.0-Sol-2B (release) from 088818218df33b6c2dca39badffde47962125d41-src_training_decision2

Browse files
Files changed (2) hide show
  1. MODEL_MANIFEST.json +7 -7
  2. README.md +5 -5
MODEL_MANIFEST.json CHANGED
@@ -2,7 +2,7 @@
2
  "base": null,
3
  "builder": {
4
  "module_sha256": "7f74839fb11762d66fa4cdc9657e7dfa8f4041bb20293dedbf69d54640fe0f0d",
5
- "source_commit": "807acc299469e887b63a794830d3e547786a132f"
6
  },
7
  "calibration": null,
8
  "card": {
@@ -15,11 +15,11 @@
15
  "assets/jevarena.png": "0d97e2ecf10414cf211a37e8f8f16fbebddbc1dc5a85f4a243b1ce029d489081"
16
  },
17
  "index_sha256": "97a95897fb13b81addf033dee7984d53e996ddb8baad5076f706a0b7c752e22e",
18
- "readme_sha256": "7b91af2a48bf47e9ef67c774a98720e062948cf4f8c5c8957a6a3c491b9ec360"
19
  },
20
  "files_sha256": {
21
  "LICENSE": "bbedc3fda3305820b977265f01b8619d87570a6739de3a5582c3464840f1e57a",
22
- "README.md": "7b91af2a48bf47e9ef67c774a98720e062948cf4f8c5c8957a6a3c491b9ec360",
23
  "assets/banner.png": "2363db1a11ea3d8c61656a078c2b2cdf10623177e54ca3ac5a48b22ccb0be724",
24
  "assets/index-areas.png": "3670b08f2c186c03a78028ad0c0c9ab07a4475ffb33b81ce88daea600b1a0d37",
25
  "assets/index-pareto.png": "bbc3b1c70755ff97ed5943bfa65f5b1c1e24041b0923c102610104182db5ecba",
@@ -71,12 +71,12 @@
71
  {
72
  "component": "Decision-2.0-Sol-2B weights, decision head, package runtime, card and artwork",
73
  "licence": "apache-2.0",
74
- "source": "llm-semantic-router/Decision-2.0-Sol-2B"
75
  },
76
  {
77
  "component": "Decision 1.0 Sol-2B weights (direct weight origin, fully fine-tuned)",
78
  "licence": "apache-2.0",
79
- "source": "llm-semantic-router/Decision-1.0-Sol-2B@ce0c018a"
80
  },
81
  {
82
  "component": "Qwen3.5-2B text backbone and tokenizer (upstream of Decision 1.0 Sol)",
@@ -101,7 +101,7 @@
101
  "model_name": "Decision-2.0-Sol-2B",
102
  "origin": {
103
  "relation": "finetune",
104
- "repo_id": "llm-semantic-router/Decision-1.0-Sol-2B",
105
  "revision": "ce0c018a28de16d6639b1cd203b761bf643b89e6",
106
  "summary": "Every weight of Decision 1.0 Sol was fine-tuned (nothing frozen, no adapter). The release is the uniform average of two seeds of that fine-tune, trained with the previous Decision 2.0 Sol 2B's answer probabilities as soft targets (self-distillation) and with extra copies of released multilingual rows that keep the released multilingual token share. Decision 1.0 Sol is itself a text-only fine-tune of [Qwen/Qwen3.5-2B](https://huggingface.co/Qwen/Qwen3.5-2B) at `15852e8c16360a2fea060d615a32b45270f8a8fc` (Apache-2.0), whose text backbone and tokenizer this model inherits; the Qwen3.5 vision tower is not part of it."
107
  },
@@ -172,7 +172,7 @@
172
  "5.18.0"
173
  ]
174
  },
175
- "repo_id": "llm-semantic-router/Decision-2.0-Sol-2B",
176
  "runtime": {
177
  "equivalence": "decision2/qwen.py loads this full checkpoint with the training/model sources vendored from the scored run's own runner mirror (5b246b110), whose SHA-256 equal the scored adapter sources (checked at build time), and applies the per-item batching, BF16-backbone / FP32-head execution, raw probabilities (temperature 1; no calibration file) and answer normalization of v2.dec.infer_dec, which for this non-residual checkpoint wraps the same DecisionModel without extra readouts. The scored checkpoint (8c8e98e3) stored every tensor in FP32; this package (v2.release.bf16_copy) stores its Linear projection matrices in BF16 exactly as BF16 autocast rounds them and every other tensor bit for bit in FP32, and the runtime holds the backbone's BF16-exact Linear weights in BF16, the values BF16 autocast multiplies with. The runtime, the Transformers remote code (AutoModel with trust_remote_code) and the forward token budget are the current revision's. Checked on one GPU, in the scored image with a copy of the persisted Triton autotune cache of the formal run, against the sealed formal predictions of every scored prompt (typed-final 1,600, css15 6,547, public231 231) by release.sh --parity before and after the download, and AutoModel against the native runtime on every scored prompt.",
178
  "requirements": {
 
2
  "base": null,
3
  "builder": {
4
  "module_sha256": "7f74839fb11762d66fa4cdc9657e7dfa8f4041bb20293dedbf69d54640fe0f0d",
5
+ "source_commit": "088818218df33b6c2dca39badffde47962125d41"
6
  },
7
  "calibration": null,
8
  "card": {
 
15
  "assets/jevarena.png": "0d97e2ecf10414cf211a37e8f8f16fbebddbc1dc5a85f4a243b1ce029d489081"
16
  },
17
  "index_sha256": "97a95897fb13b81addf033dee7984d53e996ddb8baad5076f706a0b7c752e22e",
18
+ "readme_sha256": "0c547477c690faabd9e066c0fa51fa98e9d5388e225bea4703e3e5a9dfe08c24"
19
  },
20
  "files_sha256": {
21
  "LICENSE": "bbedc3fda3305820b977265f01b8619d87570a6739de3a5582c3464840f1e57a",
22
+ "README.md": "0c547477c690faabd9e066c0fa51fa98e9d5388e225bea4703e3e5a9dfe08c24",
23
  "assets/banner.png": "2363db1a11ea3d8c61656a078c2b2cdf10623177e54ca3ac5a48b22ccb0be724",
24
  "assets/index-areas.png": "3670b08f2c186c03a78028ad0c0c9ab07a4475ffb33b81ce88daea600b1a0d37",
25
  "assets/index-pareto.png": "bbc3b1c70755ff97ed5943bfa65f5b1c1e24041b0923c102610104182db5ecba",
 
71
  {
72
  "component": "Decision-2.0-Sol-2B weights, decision head, package runtime, card and artwork",
73
  "licence": "apache-2.0",
74
+ "source": "vllm-sr/Decision-2.0-Sol-2B"
75
  },
76
  {
77
  "component": "Decision 1.0 Sol-2B weights (direct weight origin, fully fine-tuned)",
78
  "licence": "apache-2.0",
79
+ "source": "vllm-sr/Decision-1.0-Sol-2B@ce0c018a"
80
  },
81
  {
82
  "component": "Qwen3.5-2B text backbone and tokenizer (upstream of Decision 1.0 Sol)",
 
101
  "model_name": "Decision-2.0-Sol-2B",
102
  "origin": {
103
  "relation": "finetune",
104
+ "repo_id": "vllm-sr/Decision-1.0-Sol-2B",
105
  "revision": "ce0c018a28de16d6639b1cd203b761bf643b89e6",
106
  "summary": "Every weight of Decision 1.0 Sol was fine-tuned (nothing frozen, no adapter). The release is the uniform average of two seeds of that fine-tune, trained with the previous Decision 2.0 Sol 2B's answer probabilities as soft targets (self-distillation) and with extra copies of released multilingual rows that keep the released multilingual token share. Decision 1.0 Sol is itself a text-only fine-tune of [Qwen/Qwen3.5-2B](https://huggingface.co/Qwen/Qwen3.5-2B) at `15852e8c16360a2fea060d615a32b45270f8a8fc` (Apache-2.0), whose text backbone and tokenizer this model inherits; the Qwen3.5 vision tower is not part of it."
107
  },
 
172
  "5.18.0"
173
  ]
174
  },
175
+ "repo_id": "vllm-sr/Decision-2.0-Sol-2B",
176
  "runtime": {
177
  "equivalence": "decision2/qwen.py loads this full checkpoint with the training/model sources vendored from the scored run's own runner mirror (5b246b110), whose SHA-256 equal the scored adapter sources (checked at build time), and applies the per-item batching, BF16-backbone / FP32-head execution, raw probabilities (temperature 1; no calibration file) and answer normalization of v2.dec.infer_dec, which for this non-residual checkpoint wraps the same DecisionModel without extra readouts. The scored checkpoint (8c8e98e3) stored every tensor in FP32; this package (v2.release.bf16_copy) stores its Linear projection matrices in BF16 exactly as BF16 autocast rounds them and every other tensor bit for bit in FP32, and the runtime holds the backbone's BF16-exact Linear weights in BF16, the values BF16 autocast multiplies with. The runtime, the Transformers remote code (AutoModel with trust_remote_code) and the forward token budget are the current revision's. Checked on one GPU, in the scored image with a copy of the persisted Triton autotune cache of the formal run, against the sealed formal predictions of every scored prompt (typed-final 1,600, css15 6,547, public231 231) by release.sh --parity before and after the download, and AutoModel against the native runtime on every scored prompt.",
178
  "requirements": {
README.md CHANGED
@@ -1,6 +1,6 @@
1
  ---
2
  license: apache-2.0
3
- base_model: llm-semantic-router/Decision-1.0-Sol-2B
4
  base_model_relation: finetune
5
  library_name: transformers
6
  tags:
@@ -14,7 +14,7 @@ tags:
14
 
15
  # Decision-2.0-Sol-2B
16
 
17
- **Decision-2.0-Sol-2B** is the 2B model of [Decision 2.0](https://huggingface.co/collections/llm-semantic-router/decision-20-6ab7cf7bdfb506bf8269cb00), the decision models of [vLLM Semantic Router](https://github.com/vllm-project/semantic-router). Give it an input (text or JSON) and the questions you need answered: pick one of several options, say yes or no, or rate on a scale. It answers them all at once and returns a probability for every answer, without generating text.
18
 
19
  | | |
20
  | --- | --- |
@@ -41,7 +41,7 @@ import json
41
 
42
  from transformers import AutoModel
43
 
44
- model = AutoModel.from_pretrained("llm-semantic-router/Decision-2.0-Sol-2B", trust_remote_code=True)
45
  result = model.system_one(
46
  state="The order arrived damaged yesterday. The customer has a receipt and asks for a replacement today.",
47
  questions={
@@ -72,7 +72,7 @@ result = model.system_one(
72
  print(json.dumps(result["answers"], indent=2))
73
 
74
  # Or as a pipeline:
75
- # transformers.pipeline("decision", model="llm-semantic-router/Decision-2.0-Sol-2B", trust_remote_code=True)(state=..., questions=...)
76
  ```
77
 
78
  ## Evaluation
@@ -112,6 +112,6 @@ Apache-2.0 ([LICENSE](LICENSE)).
112
  title = {{Decision-2.0-Sol-2B}: A Decision 2.0 Model for Structured Decisions},
113
  author = {{vLLM Semantic Router Team}},
114
  year = {2026},
115
- howpublished = {\url{https://huggingface.co/llm-semantic-router/Decision-2.0-Sol-2B}}
116
  }
117
  ```
 
1
  ---
2
  license: apache-2.0
3
+ base_model: vllm-sr/Decision-1.0-Sol-2B
4
  base_model_relation: finetune
5
  library_name: transformers
6
  tags:
 
14
 
15
  # Decision-2.0-Sol-2B
16
 
17
+ **Decision-2.0-Sol-2B** is the 2B model of [Decision 2.0](https://huggingface.co/collections/vllm-sr/decision-20-6ab7cf7bdfb506bf8269cb00), the decision models of [vLLM Semantic Router](https://github.com/vllm-project/semantic-router). Give it an input (text or JSON) and the questions you need answered: pick one of several options, say yes or no, or rate on a scale. It answers them all at once and returns a probability for every answer, without generating text.
18
 
19
  | | |
20
  | --- | --- |
 
41
 
42
  from transformers import AutoModel
43
 
44
+ model = AutoModel.from_pretrained("vllm-sr/Decision-2.0-Sol-2B", trust_remote_code=True)
45
  result = model.system_one(
46
  state="The order arrived damaged yesterday. The customer has a receipt and asks for a replacement today.",
47
  questions={
 
72
  print(json.dumps(result["answers"], indent=2))
73
 
74
  # Or as a pipeline:
75
+ # transformers.pipeline("decision", model="vllm-sr/Decision-2.0-Sol-2B", trust_remote_code=True)(state=..., questions=...)
76
  ```
77
 
78
  ## Evaluation
 
112
  title = {{Decision-2.0-Sol-2B}: A Decision 2.0 Model for Structured Decisions},
113
  author = {{vLLM Semantic Router Team}},
114
  year = {2026},
115
+ howpublished = {\url{https://huggingface.co/vllm-sr/Decision-2.0-Sol-2B}}
116
  }
117
  ```