Xunzhuo commited on
Commit
c9bad4f
·
verified ·
1 Parent(s): 43d6004

Clarify model card and add measured question-count latency

Browse files
QUESTION-SCALING.md ADDED
@@ -0,0 +1,53 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Fixed-question latency scaling
2
+
3
+ A fixed English state (819 characters), identical question text and a four-choice menu are repeated Q times in one API request. Native rendered rows, including state and question, contain 309 tokens for Sol/Nox, 228 for Laya and 250 for Decider. Tokenizers and templates differ. Every model uses the same physical AMD gfx942 GPU and matches its quality/runtime fingerprint. Native question packing is preserved, with no cross-request cache or concurrent-client load.
4
+
5
+ Points show the empirical median of 30 timed requests, from three blocks with randomized model order. Each point receives three warmups and ten measured repetitions per block. Rendering, tokenization, transfers, model forward and answer assembly are included; loading, warmup and network are excluded. Lines connect measured points for readability and do not represent measured intermediate values. The horizontal axis is logarithmic in base 2.
6
+
7
+ Q1/Q8/Q32 come from the original GPU-quiet round; Q2/Q4/Q16 are an additive round under the same frozen setup. A repeated Q1 bridge is reported separately and never pooled or selected. This is local request scaling, not maximum server throughput. Laya uses its English branch for these English inputs.
8
+
9
+ Base-model appendix retains the original measured Q1/Q8/Q32 points only.
10
+
11
+ The machine exposes 261824 MiB of device memory. Parameters, native precision, shipped temperature and model/runtime fingerprints are unchanged from the quality benchmark. No throughput saturation claim is made. p95 is an empirical percentile, not a confidence interval.
12
+
13
+ Full aggregate evidence: [question-scaling.json](metrics/question-scaling.json).
14
+
15
+ | Model | Questions | Tokens / question | p50 (ms) | p95 (ms) | Samples | Round |
16
+ |---|---:|---:|---:|---:|---:|---|
17
+ | Sol 2B | 1 | 309 | 21.04 | 21.41 | 30 | original |
18
+ | Sol 2B | 2 | 309 | 21.48 | 21.81 | 30 | additive |
19
+ | Sol 2B | 4 | 309 | 23.47 | 31.89 | 30 | additive |
20
+ | Sol 2B | 8 | 309 | 37.29 | 37.51 | 30 | original |
21
+ | Sol 2B | 16 | 309 | 73.88 | 74.46 | 30 | additive |
22
+ | Sol 2B | 32 | 309 | 147.50 | 147.91 | 30 | original |
23
+ | Nox 4B | 1 | 309 | 27.37 | 28.95 | 30 | original |
24
+ | Nox 4B | 2 | 309 | 28.17 | 28.54 | 30 | additive |
25
+ | Nox 4B | 4 | 309 | 40.89 | 41.04 | 30 | additive |
26
+ | Nox 4B | 8 | 309 | 69.25 | 69.84 | 30 | original |
27
+ | Nox 4B | 16 | 309 | 137.64 | 137.82 | 30 | additive |
28
+ | Nox 4B | 32 | 309 | 275.65 | 278.26 | 30 | original |
29
+ | Laya | 1 | 228 | 11.54 | 12.21 | 30 | original |
30
+ | Laya | 2 | 228 | 12.60 | 13.02 | 30 | additive |
31
+ | Laya | 4 | 228 | 13.37 | 13.68 | 30 | additive |
32
+ | Laya | 8 | 228 | 14.56 | 15.02 | 30 | original |
33
+ | Laya | 16 | 228 | 22.47 | 22.80 | 30 | additive |
34
+ | Laya | 32 | 228 | 37.89 | 38.30 | 30 | original |
35
+ | Decider | 1 | 250 | 23.35 | 23.91 | 30 | original |
36
+ | Decider | 2 | 250 | 23.52 | 23.89 | 30 | additive |
37
+ | Decider | 4 | 250 | 24.05 | 24.31 | 30 | additive |
38
+ | Decider | 8 | 250 | 31.59 | 31.77 | 30 | original |
39
+ | Decider | 16 | 250 | 53.12 | 53.32 | 30 | additive |
40
+ | Decider | 32 | 250 | 96.67 | 97.02 | 30 | original |
41
+
42
+ ## Nonpooled Q1 bridge
43
+
44
+ Flag threshold: absolute median difference greater than max(10% of the original Q1 median, 2 ms). Original Q1 remains the plotted point in every case; no normalization, replacement or pooling.
45
+
46
+ | Model | Original p50 (ms) | Bridge p50 (ms) | Difference (ms) | Flag |
47
+ |---|---:|---:|---:|---|
48
+ | Sol 2B | 21.04 | 20.87 | -0.17 | No |
49
+ | Nox 4B | 27.37 | 27.16 | -0.21 | No |
50
+ | Laya | 11.54 | 12.14 | +0.60 | No |
51
+ | Decider | 23.35 | 23.18 | -0.17 | No |
52
+
53
+ All bridge checks are below the predeclared threshold.
README.md CHANGED
@@ -24,95 +24,76 @@ datasets:
24
 
25
  # Decision-1.0-Nox
26
 
27
- **Your move. State in. Decisions out.**
28
 
29
- **A capable decoder for runtime-defined decisions.** Nox turns your state and questions into choices, yes/no judgments and rubric scores in one numerical forward pass per question.
30
 
31
- **4.208B parameters · 16,384-token complete-question budget · English / Chinese evaluated · Apache 2.0**
32
 
33
- | Choose | Judge | Score |
34
- |---|---|---|
35
- | Select an action from 2–255 candidates you provide. | Return P(true) for a condition applied to the supplied state. | Apply 2–10 ordered criteria and return the expected index. |
36
 
37
- Each response preserves your question names and candidate IDs. There is no explanatory token generation and no fixed application label set.
 
 
 
 
 
 
 
 
 
 
38
 
39
  ## Measured capability
40
 
41
- Nox reaches **79.32%** overall: **+22.31 points over Laya**, **+15.31 over Decider**, and **+9.43 over its untuned 4B parent**. Jev scores 79.10% on the same panel; similar averages do not establish equivalent capability.
42
 
43
  | Model | Overall accuracy ↑ | Choice ↑ | Noul ↑ | Score ↑ |
44
  | --- | ---: | ---: | ---: | ---: |
45
- | Nox · 4B | 79.32 | 76.06 | 90.62 | 87.50 |
46
- | Jev 1.13.0 | 79.10 | 72.61 | 100.00 | 100.00 |
47
  | Qwen3.5 · 4B, untuned | 69.89 | 68.48 | 62.50 | 90.62 |
48
  | Sol · 2B | 66.25 | 71.41 | 42.19 | 43.75 |
49
  | Decider · 2B | 64.01 | 60.77 | 71.88 | 84.38 |
50
  | Qwen3.5 · 2B, untuned | 57.12 | 56.38 | 37.50 | 81.25 |
51
- | Laya · EN/ML routed | 57.01 | 64.10 | 43.75 | 18.75 |
52
 
53
- Accuracy (%), higher is better. Overall is the equal-weight mean across **10 families / 880 questions / 432 semantic groups**, not the average of the three type columns. Models receive the same frozen requests; no benchmark-specific adaptation. Laya uses its fixed EN/ML router and reports 112 truncated core inputs. Untuned Qwen uses the exact post-trained parent and a frozen LM-head adapter. [Methods, uncertainty and all input-support details](EVALUATION.md).
54
 
55
- ![Decision model ranking with 95% confidence intervals](assets/decision-quality.png)
56
 
57
- ![Accuracy across all ten measured task families](assets/decision-capabilities.png)
58
 
59
- Relational composition, state tracking, probability calibration and the separate native-capability probe remain important gaps, including against Jev. These models judge supplied evidence; they do not browse for current facts or retrieve private information. Use an explicit unknown option when the evidence may be insufficient.
60
 
61
- ## Inference cost
62
 
63
- | Choice workload | Tokens / question | p50 ms | p95 ms | Peak GiB |
64
- | --- | ---: | ---: | ---: | ---: |
65
- | Short · 1 question / 4 options | 309 | 27.37 | 28.95 | 8.06 |
66
- | Short · 8 questions / 4 options | 309 | 69.25 | 69.84 | 8.35 |
67
- | Long · 1 question / 4 options | 8863 | 278.96 | 296.37 | 9.25 |
68
- | 1 question / 255 options | 7870 | 248.31 | 249.58 | 9.10 |
69
 
70
- Measured on one AMD gfx942 GPU (261824 MiB), BF16 backbone / FP32 head, with the exact quality-tested runtime. Local Python-request latency includes rendering, tokenization, transfer, forward and answer assembly; loading, warmup and network are excluded. Thirty measurements per cell across three randomized blocks. Q8 means eight questions in one request, not eight concurrent clients. Memory is peak PyTorch allocation including weights. [All models, typed workloads and measurement details](TIMING.md).
71
 
72
  ## Try it
73
 
74
- With the Hugging Face CLI and Docker installed, download this release and follow the tested [public ROCm runtime recipe](RUNTIME.md). It builds from a public digest-pinned image and installs the lightweight `decision-local` wrapper included here. No API key is needed for local inference.
75
-
76
- ```bash
77
- hf download llm-semantic-router/Decision-1.0-Nox --revision v1.0 --local-dir decision-model
78
- cd decision-model
79
- docker build --pull -f Dockerfile.runtime -t decision-runtime:1.0 .
80
- mkdir -p runtime-output
81
- docker run --rm --device=/dev/kfd --device=/dev/dri --group-add video --ipc=host \
82
- -v "$PWD":/model:ro -v "$PWD/runtime-output":/output \
83
- decision-runtime:1.0 \
84
- python3 -m decision.example /model --local-files-only --output /output/example.json
85
- ```
86
-
87
- This executes the packaged three-question billing example, checks the wrapper against the direct numerical engine, and verifies explicit overflow rejection. The recorded output selected **billing**, returned **P(refund requested) = 1.000000**, and scored urgency **0.999770** on the three-level 0–2 rubric. [Exact request and full measured response](model-card-example.json).
88
-
89
- From Python inside that runtime:
90
 
91
  ```python
92
  from decision import DecisionModel
93
  from decision.example import REQUEST
94
 
95
  model = DecisionModel.from_pretrained("/model", local_files_only=True)
96
- result = model.decide(**REQUEST)
97
- print(result["answers"])
98
  ```
99
 
100
- The complete state, instructions and candidates must fit the token budget for **every question**; overflow rejects the call instead of truncating. Score returns an expected ordinal index, not arbitrary supplied numerical values. [Interface and loading reference](USAGE.md).
101
-
102
- ## Architecture
103
-
104
- ![Actual Decision decoder backbone and candidate readout](assets/architecture.png)
105
-
106
- The text-only Qwen3.5 backbone combines gated linear-attention blocks with full causal attention. A shared head reads candidate endpoints together with a final query vector, then returns a masked softmax over the current candidates. The backbone runs in BF16 and the small decision head in FP32. Questions execute independently in batches of eight; this release does not cache a shared state prefix across questions.
107
 
108
- ![Shared candidate-pointer head and typed outputs](assets/readout.png)
109
 
110
- Editable [architecture SVG](assets/architecture.svg), [readout SVG](assets/readout.svg), and [vector PDF](assets/architecture-atlas.pdf). [Exact architecture and inference source](code/decision_model.py).
111
 
112
- ## Deployment and scope
113
 
114
- The packaged runtime is validated on an AMD GPU with the gfx942 architecture and ROCm. CPU/MPS inference is not implemented, and NVIDIA compatibility has not been qualified. [Runtime, dependency pins and reproduction](RUNTIME.md).
115
 
116
- This is a general decision checkpoint adapted from [Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B); it is not a chat generator. English and Chinese were measured on the declared suite; other languages and application-specific reliability require evaluation. Probabilities can be overconfident, and the concentration-based confidence field is not a correctness guarantee.
117
 
118
- Weights and code are distributed under [Apache 2.0](LICENSE), with [upstream and training-data attribution](ATTRIBUTIONS.md). Explore the [Decision collection](https://huggingface.co/collections/llm-semantic-router/decision-10-6ab12177bd0002394d8409f9).
 
24
 
25
  # Decision-1.0-Nox
26
 
27
+ *Nox, Latin for night.*
28
 
29
+ **Your move.**
30
 
31
+ A capable decoder for decisions defined by you. Give Nox a state, questions and possible answers. It returns choices, yes/no judgments and rubric scores with probability distributions—one forward pass per question.
32
 
33
+ **4.208B parameters · 16K complete-question budget · English / Chinese evaluated · Apache 2.0**
 
 
34
 
35
+ [Decision family](https://huggingface.co/collections/llm-semantic-router/decision-10-6ab12177bd0002394d8409f9) · [Meet Sol](https://huggingface.co/llm-semantic-router/Decision-1.0-Sol)
36
+
37
+ ## Three ways to decide
38
+
39
+ | Type | Use it for | Output |
40
+ | --- | --- | --- |
41
+ | **Choice** | Route a request; select an action from 2–255 candidates. | Selected ID + distribution |
42
+ | **Noul** | Check a condition against available evidence. | P(true) |
43
+ | **Score** | Apply 2–10 ordered rubric descriptions. | Expected index + distribution |
44
+
45
+ Your question names and candidate IDs are preserved. Labels are defined at runtime.
46
 
47
  ## Measured capability
48
 
49
+ **79.32% overall**: +22.31 points over Laya, +15.31 over Decider, and +9.43 over its untuned parent. Jev scores 79.10%; similar averages do not imply equivalent capability.
50
 
51
  | Model | Overall accuracy ↑ | Choice ↑ | Noul ↑ | Score ↑ |
52
  | --- | ---: | ---: | ---: | ---: |
53
+ | Nox · 4B | **79.32** | **76.06** | 90.62 | 87.50 |
54
+ | Jev 1.13.0 | 79.10 | 72.61 | **100.00** | **100.00** |
55
  | Qwen3.5 · 4B, untuned | 69.89 | 68.48 | 62.50 | 90.62 |
56
  | Sol · 2B | 66.25 | 71.41 | 42.19 | 43.75 |
57
  | Decider · 2B | 64.01 | 60.77 | 71.88 | 84.38 |
58
  | Qwen3.5 · 2B, untuned | 57.12 | 56.38 | 37.50 | 81.25 |
59
+ | Laya · EN/ML | 57.01 | 64.10 | 43.75 | 18.75 |
60
 
61
+ Accuracy (%). Overall averages **10 families / 880 questions** equally; type columns pool their questions. Same requests for every model. Untuned means the post-trained parent before decision adaptation. Laya's native interface truncated 112 inputs. [Methods and uncertainty](EVALUATION.md).
62
 
63
+ ![Decision quality ranking](assets/decision-quality.png)
64
 
65
+ ![Capability comparison across ten families](assets/decision-capabilities.png)
66
 
67
+ Evidence reasoning, probability calibration and the separate native-capability probe remain behind Jev.
68
 
69
+ ## More questions, measured
70
 
71
+ ![Latency as questions per request increase](assets/decision-question-scaling.png)
 
 
 
 
 
72
 
73
+ Same state, same question length, four choices; only the number of questions changes. Each model keeps its native interface (Nox: 309 tokens/question). Median of 30 requests per point on one AMD gfx942 GPU; local Python request time includes tokenization and inference, excluding loading and network. [p95, drift checks and full methods](QUESTION-SCALING.md).
74
 
75
  ## Try it
76
 
77
+ Download the fixed release with `hf download llm-semantic-router/Decision-1.0-Nox --revision v1.0 --local-dir decision-model`, then follow the [ROCm setup](RUNTIME.md). Inside that container, with the model mounted at `/model`:
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
78
 
79
  ```python
80
  from decision import DecisionModel
81
  from decision.example import REQUEST
82
 
83
  model = DecisionModel.from_pretrained("/model", local_files_only=True)
84
+ print(model.decide(**REQUEST)["answers"])
 
85
  ```
86
 
87
+ [Actual request and measured output](model-card-example.json) · [Install and API guide](USAGE.md)
 
 
 
 
 
 
88
 
89
+ The complete state, question and candidates must fit 16,384 tokens. Overflow is rejected. Score returns an expected ordinal index. Runtime: AMD gfx942 validated; CPU/MPS unsupported; NVIDIA unqualified.
90
 
91
+ ## Architecture
92
 
93
+ ![Decision decoder architecture](assets/architecture.png)
94
 
95
+ A causal Qwen3.5 text backbone combines gated linear attention with full attention. A shared candidate head reads candidate endpoints and the final query vector. Questions run independently in batches of eight.
96
 
97
+ [Candidate head](assets/readout.png) · [Editable SVG](assets/architecture.svg) · [Vector atlas](assets/architecture-atlas.pdf) · [Inference code](code/decision_model.py)
98
 
99
+ Adapted from [Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B). It evaluates supplied evidence, without live retrieval; confidence is not a correctness guarantee. [Apache 2.0](https://huggingface.co/llm-semantic-router/Decision-1.0-Nox/blob/main/LICENSE) · [Attributions](ATTRIBUTIONS.md).
assets/decision-capabilities.pdf CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:8cd467055c05d5c7616e776174ed0e3e953e71672558ea346e347e8663b5e5ce
3
- size 106372
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:1314aae927ea0aa598e62a80138e2ea9bfd8e851bc6faa91504d5aa691426e24
3
+ size 45170
assets/decision-capabilities.png CHANGED

Git LFS Details

  • SHA256: ec86958b569c84c9ed8fcd8180095f88bbc3fe8f8ded70f685b08bd6f59f2ef6
  • Pointer size: 131 Bytes
  • Size of remote file: 483 kB

Git LFS Details

  • SHA256: ed88c09a355c8583ac4d7b31bd4ff93b7be8725fde436f3d76e1754ae92a5962
  • Pointer size: 131 Bytes
  • Size of remote file: 276 kB
assets/decision-capabilities.svg CHANGED
assets/decision-quality.pdf CHANGED
Binary files a/assets/decision-quality.pdf and b/assets/decision-quality.pdf differ
 
assets/decision-quality.png CHANGED

Git LFS Details

  • SHA256: 50af70f9120e1f13024a18b264d9471c0f441464f33188306e5f5e0dc0cc4955
  • Pointer size: 131 Bytes
  • Size of remote file: 383 kB

Git LFS Details

  • SHA256: dbf0e32d3e7846f94b8d6fd224eaf54677cd1ecb140f1671fb1b80f2a6d25935
  • Pointer size: 131 Bytes
  • Size of remote file: 176 kB
assets/decision-quality.svg CHANGED
assets/decision-question-scaling.pdf ADDED
Binary file (25.8 kB). View file
 
assets/decision-question-scaling.png ADDED
assets/decision-question-scaling.svg ADDED
metrics/question-scaling.json ADDED
@@ -0,0 +1,327 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "scope": "Fixed same question text and819-characterstate, K4; actual request latency, no concurrent serving claim",
3
+ "original_timing_sha256": "a0e5f1c94277b31a922ae1e1a708ce70a2f03b5e49996a688112fc1e9697615a",
4
+ "additive_protocol_sha256": "391114fe1d4bd3641957a28920a14347418d4e27b8ff495e743627a73aef0135",
5
+ "model_runtime_keys": {
6
+ "Decision-1.0-Sol": "aa382311fc7d9249e78ae4459d9238d10b1b2e02a7b454949be4ec87c01c971d",
7
+ "Decision-1.0-Nox": "34be384b21efb03b65b90d28463bdab1fb180438a9c7538dbf0fadb0e049c230",
8
+ "laya-routed": "4c6fdee5192ac14c2842de9fef374aa6d2fe636dbd593ad182a75156422e663f",
9
+ "decider": "417e15690883f8700b02327aacb6ef2f39c5a14737ee081163c239183e5f27eb",
10
+ "base-2b": "4cfa0428c0f0332610cc36e7a92c97d1c039b08961f1c16aea4d552165097ea5",
11
+ "base-4b": "2edfae75b3a649b147bd49096006f70e5eb29515cb8fa134e79b9255f198c278"
12
+ },
13
+ "points": {
14
+ "Decision-1.0-Sol": [
15
+ {
16
+ "q": 1,
17
+ "p50_ms": 21.03651,
18
+ "p95_ms": 21.4133929,
19
+ "round": "original",
20
+ "samples": 30,
21
+ "tokens_per_native_row": 309
22
+ },
23
+ {
24
+ "q": 2,
25
+ "p50_ms": 21.4783635,
26
+ "p95_ms": 21.808785,
27
+ "round": "additive",
28
+ "samples": 30,
29
+ "tokens_per_native_row": 309
30
+ },
31
+ {
32
+ "q": 4,
33
+ "p50_ms": 23.473785,
34
+ "p95_ms": 31.8870149,
35
+ "round": "additive",
36
+ "samples": 30,
37
+ "tokens_per_native_row": 309
38
+ },
39
+ {
40
+ "q": 8,
41
+ "p50_ms": 37.294218,
42
+ "p95_ms": 37.50742325,
43
+ "round": "original",
44
+ "samples": 30,
45
+ "tokens_per_native_row": 309
46
+ },
47
+ {
48
+ "q": 16,
49
+ "p50_ms": 73.88406900000001,
50
+ "p95_ms": 74.45837850000001,
51
+ "round": "additive",
52
+ "samples": 30,
53
+ "tokens_per_native_row": 309
54
+ },
55
+ {
56
+ "q": 32,
57
+ "p50_ms": 147.50142499999998,
58
+ "p95_ms": 147.90909245,
59
+ "round": "original",
60
+ "samples": 30,
61
+ "tokens_per_native_row": 309
62
+ }
63
+ ],
64
+ "Decision-1.0-Nox": [
65
+ {
66
+ "q": 1,
67
+ "p50_ms": 27.374125,
68
+ "p95_ms": 28.94694755,
69
+ "round": "original",
70
+ "samples": 30,
71
+ "tokens_per_native_row": 309
72
+ },
73
+ {
74
+ "q": 2,
75
+ "p50_ms": 28.1707225,
76
+ "p95_ms": 28.5395078,
77
+ "round": "additive",
78
+ "samples": 30,
79
+ "tokens_per_native_row": 309
80
+ },
81
+ {
82
+ "q": 4,
83
+ "p50_ms": 40.8942485,
84
+ "p95_ms": 41.03577595,
85
+ "round": "additive",
86
+ "samples": 30,
87
+ "tokens_per_native_row": 309
88
+ },
89
+ {
90
+ "q": 8,
91
+ "p50_ms": 69.247188,
92
+ "p95_ms": 69.84472105,
93
+ "round": "original",
94
+ "samples": 30,
95
+ "tokens_per_native_row": 309
96
+ },
97
+ {
98
+ "q": 16,
99
+ "p50_ms": 137.64264550000001,
100
+ "p95_ms": 137.81669829999998,
101
+ "round": "additive",
102
+ "samples": 30,
103
+ "tokens_per_native_row": 309
104
+ },
105
+ {
106
+ "q": 32,
107
+ "p50_ms": 275.653244,
108
+ "p95_ms": 278.2649851,
109
+ "round": "original",
110
+ "samples": 30,
111
+ "tokens_per_native_row": 309
112
+ }
113
+ ],
114
+ "laya-routed": [
115
+ {
116
+ "q": 1,
117
+ "p50_ms": 11.539303,
118
+ "p95_ms": 12.21025405,
119
+ "round": "original",
120
+ "samples": 30,
121
+ "tokens_per_native_row": 228
122
+ },
123
+ {
124
+ "q": 2,
125
+ "p50_ms": 12.602103499999998,
126
+ "p95_ms": 13.020027549999998,
127
+ "round": "additive",
128
+ "samples": 30,
129
+ "tokens_per_native_row": 228
130
+ },
131
+ {
132
+ "q": 4,
133
+ "p50_ms": 13.366638,
134
+ "p95_ms": 13.68192385,
135
+ "round": "additive",
136
+ "samples": 30,
137
+ "tokens_per_native_row": 228
138
+ },
139
+ {
140
+ "q": 8,
141
+ "p50_ms": 14.559828,
142
+ "p95_ms": 15.0213171,
143
+ "round": "original",
144
+ "samples": 30,
145
+ "tokens_per_native_row": 228
146
+ },
147
+ {
148
+ "q": 16,
149
+ "p50_ms": 22.471257,
150
+ "p95_ms": 22.797051449999998,
151
+ "round": "additive",
152
+ "samples": 30,
153
+ "tokens_per_native_row": 228
154
+ },
155
+ {
156
+ "q": 32,
157
+ "p50_ms": 37.885172999999995,
158
+ "p95_ms": 38.3012614,
159
+ "round": "original",
160
+ "samples": 30,
161
+ "tokens_per_native_row": 228
162
+ }
163
+ ],
164
+ "decider": [
165
+ {
166
+ "q": 1,
167
+ "p50_ms": 23.3526305,
168
+ "p95_ms": 23.906231950000002,
169
+ "round": "original",
170
+ "samples": 30,
171
+ "tokens_per_native_row": 250
172
+ },
173
+ {
174
+ "q": 2,
175
+ "p50_ms": 23.515689000000002,
176
+ "p95_ms": 23.8890495,
177
+ "round": "additive",
178
+ "samples": 30,
179
+ "tokens_per_native_row": 250
180
+ },
181
+ {
182
+ "q": 4,
183
+ "p50_ms": 24.045369,
184
+ "p95_ms": 24.310676,
185
+ "round": "additive",
186
+ "samples": 30,
187
+ "tokens_per_native_row": 250
188
+ },
189
+ {
190
+ "q": 8,
191
+ "p50_ms": 31.591898999999998,
192
+ "p95_ms": 31.7650492,
193
+ "round": "original",
194
+ "samples": 30,
195
+ "tokens_per_native_row": 250
196
+ },
197
+ {
198
+ "q": 16,
199
+ "p50_ms": 53.1165,
200
+ "p95_ms": 53.321273049999995,
201
+ "round": "additive",
202
+ "samples": 30,
203
+ "tokens_per_native_row": 250
204
+ },
205
+ {
206
+ "q": 32,
207
+ "p50_ms": 96.6671255,
208
+ "p95_ms": 97.02467545,
209
+ "round": "original",
210
+ "samples": 30,
211
+ "tokens_per_native_row": 250
212
+ }
213
+ ],
214
+ "base-2b": [
215
+ {
216
+ "q": 1,
217
+ "p50_ms": 20.508525,
218
+ "p95_ms": 20.714382399999998,
219
+ "round": "original",
220
+ "samples": 30,
221
+ "tokens_per_native_row": 296
222
+ },
223
+ {
224
+ "q": 8,
225
+ "p50_ms": 37.320352,
226
+ "p95_ms": 37.61330755,
227
+ "round": "original",
228
+ "samples": 30,
229
+ "tokens_per_native_row": 296
230
+ },
231
+ {
232
+ "q": 32,
233
+ "p50_ms": 146.8754105,
234
+ "p95_ms": 147.15678465000002,
235
+ "round": "original",
236
+ "samples": 30,
237
+ "tokens_per_native_row": 296
238
+ }
239
+ ],
240
+ "base-4b": [
241
+ {
242
+ "q": 1,
243
+ "p50_ms": 27.469792499999997,
244
+ "p95_ms": 28.755623,
245
+ "round": "original",
246
+ "samples": 30,
247
+ "tokens_per_native_row": 296
248
+ },
249
+ {
250
+ "q": 8,
251
+ "p50_ms": 69.02867599999999,
252
+ "p95_ms": 69.45783845,
253
+ "round": "original",
254
+ "samples": 30,
255
+ "tokens_per_native_row": 296
256
+ },
257
+ {
258
+ "q": 32,
259
+ "p50_ms": 274.68847300000004,
260
+ "p95_ms": 275.19135435000004,
261
+ "round": "original",
262
+ "samples": 30,
263
+ "tokens_per_native_row": 296
264
+ }
265
+ ]
266
+ },
267
+ "Q1_bridge": {
268
+ "Decision-1.0-Sol": {
269
+ "original_p50_ms": 21.03651,
270
+ "bridge_p50_ms": 20.8673795,
271
+ "delta_ms": -0.1691305000000014,
272
+ "delta_fraction": -0.008039855470322821,
273
+ "flag": false,
274
+ "pooled_or_selected": false
275
+ },
276
+ "Decision-1.0-Nox": {
277
+ "original_p50_ms": 27.374125,
278
+ "bridge_p50_ms": 27.164138,
279
+ "delta_ms": -0.20998699999999815,
280
+ "delta_fraction": -0.007671003182750047,
281
+ "flag": false,
282
+ "pooled_or_selected": false
283
+ },
284
+ "laya-routed": {
285
+ "original_p50_ms": 11.539303,
286
+ "bridge_p50_ms": 12.141239,
287
+ "delta_ms": 0.6019360000000002,
288
+ "delta_fraction": 0.05216398252130139,
289
+ "flag": false,
290
+ "pooled_or_selected": false
291
+ },
292
+ "decider": {
293
+ "original_p50_ms": 23.3526305,
294
+ "bridge_p50_ms": 23.177763,
295
+ "delta_ms": -0.17486750000000129,
296
+ "delta_fraction": -0.007488128585771192,
297
+ "flag": false,
298
+ "pooled_or_selected": false
299
+ }
300
+ },
301
+ "all_bridge_drift_checks_pass": true,
302
+ "raw_additive_measurement_sha256": {
303
+ "block-0/Decision-1.0-Sol/measurements.jsonl": "e186d2eb157df15eca30befa39b92c02ebb47288dffbca82d601f4ffe6ab8a6e",
304
+ "block-1/Decision-1.0-Sol/measurements.jsonl": "715ef039bcd8a35bc75d25d94adc64dab02cfdb37697c9f0274690a06321f9e2",
305
+ "block-2/Decision-1.0-Sol/measurements.jsonl": "333007726ded345cf37281b6df4afee351cd9bf099ca59a807daeba272b4bf6e",
306
+ "block-0/Decision-1.0-Nox/measurements.jsonl": "6770f8d60ea448ee6c02df901f90eea489f0e8d198c245211ae8520ddfe11c86",
307
+ "block-1/Decision-1.0-Nox/measurements.jsonl": "3893b592c8af7ae163453891cbd1373c2edb5289db4d27cfd8ab1276e6b415b4",
308
+ "block-2/Decision-1.0-Nox/measurements.jsonl": "b0a73b6436323613b27b09f3b609a640ad9d1f00317a1161dff65f4444520521",
309
+ "block-0/laya-routed/measurements.jsonl": "2cf05c1438bb62d7510c00e00f4da2e134472661635c420fe58cba4b8bd4dd59",
310
+ "block-1/laya-routed/measurements.jsonl": "18582d373500d0826e4fa03f23f24d15c4f2a506b99f837723e07d19ca44fb95",
311
+ "block-2/laya-routed/measurements.jsonl": "20e699577a1fe5edccb3fb1dfa47b491930ffc3875cf7b62a2f9409881b73aca",
312
+ "block-0/decider/measurements.jsonl": "0116ea8a843afe0496762442dd0f3ef05c7b3639664bde2d6de47071d67463ea",
313
+ "block-1/decider/measurements.jsonl": "2d99b9ef948b16cf2c2b39d971c79ae15b05819ce4f1a89c1c85e22534384b9a",
314
+ "block-2/decider/measurements.jsonl": "a8a84c27bac9d3050f75ea65c97b35c8ba7c515b9c7ed45f77f7374ad0b277c1"
315
+ },
316
+ "logo_sha256": "d83cf7878c3839f3085feb8c02334217ad2fef0a615b805e1726b292552c46e8",
317
+ "logo_modified": false,
318
+ "renderer_sha256": "cd9ccebdcce07cfda8aa301be71ab8c23a748d9e1b79c020873322d2f085c0b0",
319
+ "figure_footer_annotations": false,
320
+ "canvas_width_px": 792,
321
+ "designed_card_width_px": 600,
322
+ "axis_legend_effective_css_px_at_600": 12.152777777777779,
323
+ "display_label_aliases": {
324
+ "Laya": "Fixed upstream EN/ML routed Laya; English branch used for these English inputs",
325
+ "Decider": "Mapika/decider-2b"
326
+ }
327
+ }
release-manifest.json CHANGED
@@ -1,9 +1,9 @@
1
  {
2
  "format": "decision-public-release-v1",
3
- "status": "assembled-not-published",
4
  "bundle_manifest_sha256": "adf5d88146d4f4a7fdcf0b265aa0c804a6116537c1e92bd620e40dfdf6803550",
5
- "readiness_sha256": "d5b4d11c20f886845577a43314802bbeef51f7fa4f2c1f438c4236323f4182d0",
6
- "model_card_sha256": "57c4022b8a683383db0e5e5270827029764d83ee49d742f712da22a88fa7f133",
7
  "repo_id": "llm-semantic-router/Decision-1.0-Nox",
8
  "assembly_script_sha256": "0ca32612da99c16b9801149d69f375b157e51b6e426b0a74efa3efb6bcb9f792",
9
  "original_bundle_manifest_preserved": true,
@@ -34,6 +34,11 @@
34
  "bytes": 11544,
35
  "sha256": "bbedc3fda3305820b977265f01b8619d87570a6739de3a5582c3464840f1e57a"
36
  },
 
 
 
 
 
37
  {
38
  "file": "QWEN-LICENSE",
39
  "bytes": 11544,
@@ -41,8 +46,8 @@
41
  },
42
  {
43
  "file": "README.md",
44
- "bytes": 6981,
45
- "sha256": "57c4022b8a683383db0e5e5270827029764d83ee49d742f712da22a88fa7f133"
46
  },
47
  {
48
  "file": "RUNTIME.md",
@@ -81,18 +86,18 @@
81
  },
82
  {
83
  "file": "assets/decision-capabilities.pdf",
84
- "bytes": 106372,
85
- "sha256": "8cd467055c05d5c7616e776174ed0e3e953e71672558ea346e347e8663b5e5ce"
86
  },
87
  {
88
  "file": "assets/decision-capabilities.png",
89
- "bytes": 483212,
90
- "sha256": "ec86958b569c84c9ed8fcd8180095f88bbc3fe8f8ded70f685b08bd6f59f2ef6"
91
  },
92
  {
93
  "file": "assets/decision-capabilities.svg",
94
- "bytes": 153471,
95
- "sha256": "e3b03c292b420b039b972756ae40e93e73c0610672a9bca5d795736faeb7f697"
96
  },
97
  {
98
  "file": "assets/decision-family-header.png",
@@ -106,18 +111,33 @@
106
  },
107
  {
108
  "file": "assets/decision-quality.pdf",
109
- "bytes": 92352,
110
- "sha256": "403a02af4ba6292d27fe2863dc568a026278002fbb58082befed37ac55f31ce4"
111
  },
112
  {
113
  "file": "assets/decision-quality.png",
114
- "bytes": 383098,
115
- "sha256": "50af70f9120e1f13024a18b264d9471c0f441464f33188306e5f5e0dc0cc4955"
116
  },
117
  {
118
  "file": "assets/decision-quality.svg",
119
- "bytes": 111684,
120
- "sha256": "b430df49b3c74fa0cd8b7e39ab1a7f6768295feaa423213816fa6ce8f2fde8a2"
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
121
  },
122
  {
123
  "file": "assets/readout.pdf",
@@ -204,6 +224,11 @@
204
  "bytes": 1214948,
205
  "sha256": "b2181b7636f0ddace6af33a45010bfb4e62fb32f987f4abe2d2a4d58b3a007e4"
206
  },
 
 
 
 
 
207
  {
208
  "file": "metrics/timing-aggregate.json",
209
  "bytes": 148812,
@@ -277,7 +302,7 @@
277
  "FIGURE-NOTICES.md": "copy",
278
  "LICENSE": "copy",
279
  "QWEN-LICENSE": "copy",
280
- "README.md": "copy",
281
  "RUNTIME.md": "copy",
282
  "TIMING.md": "copy",
283
  "USAGE.md": "copy",
@@ -285,14 +310,14 @@
285
  "assets/architecture.pdf": "copy",
286
  "assets/architecture.png": "copy",
287
  "assets/architecture.svg": "copy",
288
- "assets/decision-capabilities.pdf": "copy",
289
- "assets/decision-capabilities.png": "copy",
290
- "assets/decision-capabilities.svg": "copy",
291
  "assets/decision-family-header.png": "copy",
292
  "assets/decision-mark.png": "copy",
293
- "assets/decision-quality.pdf": "copy",
294
- "assets/decision-quality.png": "copy",
295
- "assets/decision-quality.svg": "copy",
296
  "assets/readout.pdf": "copy",
297
  "assets/readout.png": "copy",
298
  "assets/readout.svg": "copy",
@@ -322,7 +347,15 @@
322
  "src/decision/model.py": "copy",
323
  "temperature.json": "copy",
324
  "tokenizer.json": "copy",
325
- "tokenizer_config.json": "copy"
 
 
 
 
 
326
  },
327
- "scope": "Inference weights, original numerical code, calibrated runtime metadata, wrapper, model card, license, attribution and approved assets only."
 
 
 
328
  }
 
1
  {
2
  "format": "decision-public-release-v1",
3
+ "status": "documentation-refresh",
4
  "bundle_manifest_sha256": "adf5d88146d4f4a7fdcf0b265aa0c804a6116537c1e92bd620e40dfdf6803550",
5
+ "readiness_sha256": "461241c90fbf15ff07859b8709999559dffa13ef9b0f66bd674b44166a9b3f88",
6
+ "model_card_sha256": "d273cd46bb50341ac0f7a16a2a6241632dcf82f34d9b91ca3d80e562e4888ecc",
7
  "repo_id": "llm-semantic-router/Decision-1.0-Nox",
8
  "assembly_script_sha256": "0ca32612da99c16b9801149d69f375b157e51b6e426b0a74efa3efb6bcb9f792",
9
  "original_bundle_manifest_preserved": true,
 
34
  "bytes": 11544,
35
  "sha256": "bbedc3fda3305820b977265f01b8619d87570a6739de3a5582c3464840f1e57a"
36
  },
37
+ {
38
+ "file": "QUESTION-SCALING.md",
39
+ "bytes": 3608,
40
+ "sha256": "ad72203f319d9344ffa917b6d12ae34fd976c71c39fc715fa862dd23685bbbf1"
41
+ },
42
  {
43
  "file": "QWEN-LICENSE",
44
  "bytes": 11544,
 
46
  },
47
  {
48
  "file": "README.md",
49
+ "bytes": 4512,
50
+ "sha256": "d273cd46bb50341ac0f7a16a2a6241632dcf82f34d9b91ca3d80e562e4888ecc"
51
  },
52
  {
53
  "file": "RUNTIME.md",
 
86
  },
87
  {
88
  "file": "assets/decision-capabilities.pdf",
89
+ "bytes": 45170,
90
+ "sha256": "1314aae927ea0aa598e62a80138e2ea9bfd8e851bc6faa91504d5aa691426e24"
91
  },
92
  {
93
  "file": "assets/decision-capabilities.png",
94
+ "bytes": 276453,
95
+ "sha256": "ed88c09a355c8583ac4d7b31bd4ff93b7be8725fde436f3d76e1754ae92a5962"
96
  },
97
  {
98
  "file": "assets/decision-capabilities.svg",
99
+ "bytes": 69169,
100
+ "sha256": "cae4bf2c1aa13e8d6dc94c435eed1fd8bde1e523fbfa83a0b7dde5ba37650022"
101
  },
102
  {
103
  "file": "assets/decision-family-header.png",
 
111
  },
112
  {
113
  "file": "assets/decision-quality.pdf",
114
+ "bytes": 37484,
115
+ "sha256": "aec3769c803f3b87668332e2a6d036cab1bd65b21c8ebbd919fbb79f1637290f"
116
  },
117
  {
118
  "file": "assets/decision-quality.png",
119
+ "bytes": 175813,
120
+ "sha256": "dbf0e32d3e7846f94b8d6fd224eaf54677cd1ecb140f1671fb1b80f2a6d25935"
121
  },
122
  {
123
  "file": "assets/decision-quality.svg",
124
+ "bytes": 37293,
125
+ "sha256": "e324b1446ac18beeefd6da1d3d4c2ec4db579b7d989390663c1ab5e43bf76841"
126
+ },
127
+ {
128
+ "file": "assets/decision-question-scaling.pdf",
129
+ "bytes": 25845,
130
+ "sha256": "e1e6606982eed8280db78072e30f7cf3ca27ceaabdd69d1253371f6c9d6fafc0"
131
+ },
132
+ {
133
+ "file": "assets/decision-question-scaling.png",
134
+ "bytes": 54575,
135
+ "sha256": "83c68fe79fd316c1640e56e85e5e1424bc6b145a7baf83ea174c11237118cb55"
136
+ },
137
+ {
138
+ "file": "assets/decision-question-scaling.svg",
139
+ "bytes": 19837,
140
+ "sha256": "93de52c74cfed97fe660bd79619a9e41eeeeb03d0a9f082b4033d8cb919deeb0"
141
  },
142
  {
143
  "file": "assets/readout.pdf",
 
224
  "bytes": 1214948,
225
  "sha256": "b2181b7636f0ddace6af33a45010bfb4e62fb32f987f4abe2d2a4d58b3a007e4"
226
  },
227
+ {
228
+ "file": "metrics/question-scaling.json",
229
+ "bytes": 9629,
230
+ "sha256": "4708635c6f26cd0209dafb4e45af6921ed22091e1996bbeb5861c98214ff5886"
231
+ },
232
  {
233
  "file": "metrics/timing-aggregate.json",
234
  "bytes": 148812,
 
302
  "FIGURE-NOTICES.md": "copy",
303
  "LICENSE": "copy",
304
  "QWEN-LICENSE": "copy",
305
+ "README.md": "documentation-copy",
306
  "RUNTIME.md": "copy",
307
  "TIMING.md": "copy",
308
  "USAGE.md": "copy",
 
310
  "assets/architecture.pdf": "copy",
311
  "assets/architecture.png": "copy",
312
  "assets/architecture.svg": "copy",
313
+ "assets/decision-capabilities.pdf": "documentation-copy",
314
+ "assets/decision-capabilities.png": "documentation-copy",
315
+ "assets/decision-capabilities.svg": "documentation-copy",
316
  "assets/decision-family-header.png": "copy",
317
  "assets/decision-mark.png": "copy",
318
+ "assets/decision-quality.pdf": "documentation-copy",
319
+ "assets/decision-quality.png": "documentation-copy",
320
+ "assets/decision-quality.svg": "documentation-copy",
321
  "assets/readout.pdf": "copy",
322
  "assets/readout.png": "copy",
323
  "assets/readout.svg": "copy",
 
347
  "src/decision/model.py": "copy",
348
  "temperature.json": "copy",
349
  "tokenizer.json": "copy",
350
+ "tokenizer_config.json": "copy",
351
+ "QUESTION-SCALING.md": "documentation-copy",
352
+ "metrics/question-scaling.json": "documentation-copy",
353
+ "assets/decision-question-scaling.pdf": "documentation-copy",
354
+ "assets/decision-question-scaling.png": "documentation-copy",
355
+ "assets/decision-question-scaling.svg": "documentation-copy"
356
  },
357
+ "scope": "Inference weights, original numerical code, calibrated runtime metadata, wrapper, model card, license, attribution and approved assets only.",
358
+ "previous_release_manifest_sha256": "dc5ca2bcd8caa6199f5c9bbb71d6bdff604a77f1b568a651c81c6fdb190fb042",
359
+ "documentation_refresh_script_sha256": "faf40a7508c73266e2337acac18d05d5fe1a119cee13768d41947b4369a44afc",
360
+ "inference_files_unchanged": true
361
  }