diff --git a/.gitattributes b/.gitattributes
index e02cc0fd205ad99a2897c1d1fc744fa02cf08426..d5f3fa5c0c85bcd687438d45b22cbfbe6fbb80da 100644
--- a/.gitattributes
+++ b/.gitattributes
@@ -44,3 +44,12 @@ assets/decision-quality.png filter=lfs diff=lfs merge=lfs -text
assets/readout.pdf filter=lfs diff=lfs merge=lfs -text
assets/readout.png filter=lfs diff=lfs merge=lfs -text
tokenizer.json filter=lfs diff=lfs merge=lfs -text
+assets/decision-expanded-old_core-600px.png filter=lfs diff=lfs merge=lfs -text
+assets/decision-expanded-old_core.png filter=lfs diff=lfs merge=lfs -text
+assets/decision-expanded-overview.png filter=lfs diff=lfs merge=lfs -text
+assets/decision-expanded-ranking.png filter=lfs diff=lfs merge=lfs -text
+assets/decision-expanded-v3_core-600px.png filter=lfs diff=lfs merge=lfs -text
+assets/decision-expanded-v3_core.png filter=lfs diff=lfs merge=lfs -text
+assets/decision-expanded-v4.png filter=lfs diff=lfs merge=lfs -text
+assets/decision-expanded-v5.png filter=lfs diff=lfs merge=lfs -text
+assets/decision-question-scaling.png filter=lfs diff=lfs merge=lfs -text
diff --git a/ATTRIBUTIONS.md b/ATTRIBUTIONS.md
index 04c64ff35e515f5604cf31d6f66a22e8fd55aa4e..2410992a66f3020e95629019fa4d3aac4a1ffe3f 100644
--- a/ATTRIBUTIONS.md
+++ b/ATTRIBUTIONS.md
@@ -23,3 +23,13 @@ MultiNLI is by Adina Williams, Nikita Nangia and Samuel R. Bowman, *A Broad-Cove
- [MultiNLI paper](https://aclanthology.org/N18-1101/).
MASSIVE and SLURP remain excluded from custom training, checkpoint selection and calibration. MASSIVE is used only for evaluation. Official Jev outputs are never used as training labels. The initial-release training histories above remain historical; the subsequent selected model and evaluation provenance identify the actual update.
+
+## Natural-language decision adaptation
+
+This update adds 8,000 human-annotated training examples to 16,000 retained decision examples: 4,000 Cosmos QA reading questions, 2,000 SQuAD 2.0 answerability judgments, and 2,000 SNLI inference pairs. The selected natural-data checkpoints complete one pass of this 24,000-example mixture. Human source labels are preserved; SQuAD answerability uses its supplied impossible/answerable annotation, and each source is converted to the model's decision interface. Source-parent groups and near duplicates are separated between custom training, selection and calibration. No official Jev output supplies a training label.
+
+- **Cosmos QA**, by Lifu Huang, Ronan Le Bras, Chandra Bhagavatula and Yejin Choi. Data from the [author repository at the pinned revision](https://github.com/wilburOne/cosmosqa/tree/b6eb99cca4e2a51dd28a9a6f562534872d851639). The [official AllenAI dataset card](https://huggingface.co/datasets/allenai/cosmos_qa/blob/28d9d5e2aae025e73e11177891a88dba51190013/README.md) records the author-confirmed [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/) license.
+- **SQuAD 2.0**, by Pranav Rajpurkar, Robin Jia and Percy Liang. The official [SQuAD project](https://rajpurkar.github.io/SQuAD-explorer/) distributes the dataset under [CC BY-SA 4.0](https://creativecommons.org/licenses/by-sa/4.0/). The adaptation here is answerability classification, not the official span-extraction benchmark.
+- **SNLI 1.0**, by Samuel R. Bowman, Gabor Angeli, Christopher Potts and Christopher D. Manning. The official [Stanford Natural Language Inference project](https://nlp.stanford.edu/projects/snli/) and release README identify [CC BY-SA 4.0](https://creativecommons.org/licenses/by-sa/4.0/) for the corpus. Original entailment, neutral and contradiction labels supply the three decision alternatives.
+
+These dataset licenses govern their respective source material; they are not replaced by the model package's Apache 2.0 license. Source corpus text, transformed training records and individual evaluation predictions are not redistributed in this model package. New evaluation panels, including supplied-fact QASC questions, are described separately in EVALUATION.md; evaluation data are not used for checkpoint selection or temperature fitting. As with other public datasets, exclusion from this custom training does not establish absence from upstream pretraining.
diff --git a/EVALUATION.md b/EVALUATION.md
index 7bbf3d8b06732f75bd7af2e87c3a487275a1faa3..6b1a0f062f6285fc44bbd083d43bc4c261e949aa 100644
--- a/EVALUATION.md
+++ b/EVALUATION.md
@@ -1,93 +1,92 @@
-# Nox: measured quality update
-
-The primary score averages 20 task-family accuracies equally: one half the original 880-question panel and one half the independently constructed 880-question confirmation panel. The original panel was already observed during development. The added panel was excluded from custom training, selection and calibration; candidate weights, prompt and temperature were fixed before its inference. Once unsealed for the first candidate, it remains an observed regression panel. A later release does not make it fresh again.
-
-| Model | Overall accuracy ↑ | Choice ↑ | Noul ↑ | Score ↑ |
-| --- | ---: | ---: | ---: | ---: |
-| Jev · 1.13.0 | **72.74** | **71.77** | **80.36** | **68.06** |
-| Nox · 4B · v1.1 | 67.05 | 70.55 | 60.27 | 55.56 |
-| Qwen3.5 · 4B · untuned | 56.61 | 60.06 | 53.57 | 52.08 |
-| Sol · 2B · v1.0 | 55.58 | 61.93 | 50.89 | 32.64 |
-| Decider · 2B | 55.30 | 57.26 | 58.93 | 50.00 |
-| Qwen3.5 · 2B · untuned | 48.06 | 50.36 | 45.09 | 47.22 |
-| Laya · EN/ML | 47.38 | 53.02 | 52.23 | 19.44 |
-
-All values are percentages. Choice, Noul and Score columns pool correct/total questions of that type across both panels; their mean is not the overall family-macro score. Score accuracy tests the highest-probability rubric level; expected-value error is reported separately. Missing or invalid predictions count as incorrect. Confidence intervals use 10,000 paired cluster-bootstrap resamples. The two language views of each MASSIVE source utterance share draws within menu strata; synthetic counterfactual blocks remain intact. The original and confirmation panels are resampled independently, then combined with equal weight.
-
-## Change from the released weights
-
-| Panel | Accuracy change (points) | Paired 95% interval |
-| --- | ---: | ---: |
-| Old | +3.53 | [+0.37, +6.59] |
-| Fresh | +5.12 | [+2.21, +8.08] |
-| Joint | +4.33 | [+2.17, +6.46] |
-
-Release eligibility follows the prospectively fixed positive joint point improvement plus complete, finite, valid candidate execution. It does not require every family to improve or the confidence interval to exclude zero. An eligible update is not evidence of universal superiority.
-
-## Native API coverage
-
-Each separate native panel has 68 requests and 220 requested answers, including repeated packing and isolation probes. These answers are descriptive diagnostics, not 220 independent questions and not part of the primary score. Unsupported requests are reported with the full denominator rather than relabeled as wrong semantic predictions.
-
-| Model | Original: valid / 220 | Original: correct / 220 | New: valid / 220 | New: correct / 220 |
-| --- | ---: | ---: | ---: | ---: |
-| Nox · 4B | 220 | 200 | 220 | 138 |
-| Sol · 2B | 220 | 116 | 220 | 116 |
-| Jev · 1.13.0 | 220 | 203 | 220 | 168 |
-| Laya · EN/ML | 208 | 88 | 208 | 98 |
-| Decider · 2B | 220 | 111 | 220 | 96 |
-| Qwen3.5 · 2B · untuned | 196 | 86 | 196 | 71 |
-| Qwen3.5 · 4B · untuned | 196 | 112 | 196 | 104 |
-
-## Scope and interpretation
-
-All comparators receive the same frozen requests through documented native or fixed adaptation interfaces. The untuned baselines are post-trained Qwen3.5 parents before decision adaptation, evaluated with a fixed LM-head decision adapter. Laya uses fixed English/multilingual routing. Jev is the official API; server-side truncation and internal model architecture cannot be independently inspected.
-
-| Model | Original reported truncations | New reported truncations |
-| --- | ---: | ---: |
-| Nox · 4B | 0 | 0 |
-| Sol · 2B | 0 | 0 |
-| Jev · 1.13.0 | 0 | 0 |
-| Laya · EN/ML | 112 | 40 |
-| Decider · 2B | 0 | 0 |
-| Qwen3.5 · 2B · untuned | 0 | 0 |
-| Qwen3.5 · 4B · untuned | 0 | 0 |
-
-Zero reported truncations for a remote service is not proof of complete server-side input processing. The Decision candidate additionally passes explicit complete-token and overflow checks.
-
-The confirmation panel includes unseen MASSIVE English/Chinese intents and eight independently generated decision families: temporal exclusion, constraint assignment, transaction recovery, conflicting-rule reasoning, multiset reconciliation, record identity, capacity constraints and service-loss scoring. Training, selection and calibration exclude MASSIVE and SLURP. Exact overlap was zero; three moderate lexical near-pairs with training were disclosed and retained. Gold labels were not supplied by Jev.
-
-MASSIVE uses the [official 1.1 archive](https://amazon-massive-nlu-dataset.s3.amazonaws.com/amazon-massive-dataset-1.1.tar.gz), CC BY 4.0. Exposure during base-model pretraining is unknown. Synthetic English/Chinese views share symbolic JSON state; they do not establish natural multilingual document comprehension.
-
-The [complete aggregate measurements](quality-metrics.json) retain every family, raw and calibrated probability diagnostics, native coverage, failure counts, Brier/NLL and ordinal metrics. Proper probability scores are conditional on valid probability coverage; zero probability on the true answer produces explicit infinite NLL rather than clipping. Per-family regressions remain visible in the capability matrix.
-
-The question-scaling figure retains measurements from v1.0 weights and is labeled accordingly. Architecture and API behavior are unchanged; these measurements are not a timing claim for the updated weights.
-
-Evaluation policy SHA-256: `359de2b6b81be91a66c6c538b3655bed06f4728357328cd318b1a765e64565cd`. Scoring receipt SHA-256: `927c92121c2df534f2ba12523e466ecebd8bd8936f949c7792de3e747c92a47c`.
-
-## Human-source confirmation (V4)
-
-Nox improves from **75.47% to 78.44%** on this separately constructed, human-annotated panel: **+2.97 percentage points**, paired 95% CI **[+0.77, +5.32]**. Reading comprehension remains a clear gap: Decider reaches **92.03%**, the untuned 4B model with its fixed LM-head adapter **87.97%**, and Jev **94.53%**. This supplemental result does not change the fixed V3 publication rule.
-
-The score weights 160 balanced BoolQ questions at 50% and 160 parallel Belebele questions in each of English and Chinese at 25% each. All **480 requested examples** are included; the weighted score differs from raw correct/480. Bold marks the best result in each score column.
-
-| Model | Weighted % | BoolQ % (correct) | Belebele EN % (correct) | Belebele ZH % (correct) | Raw correct |
-|---|---:|---:|---:|---:|---:|
-| Nox (this update) | 78.44 | 87.50 (140/160) | 71.25 (114/160) | 67.50 (108/160) | 362/480 |
-| Sol v1.0 | 69.38 | 74.38 (119/160) | 66.25 (106/160) | 62.50 (100/160) | 325/480 |
-| Jev 1.13.0 | **94.53** | **92.50 (148/160)** | **96.88 (155/160)** | **96.25 (154/160)** | 457/480 |
-| Laya EN/ML routed | 51.25 | 69.38 (111/160) | 38.12 (61/160) | 28.12 (45/160) | 217/480 |
-| Decider 2B | 92.03 | 91.88 (147/160) | 93.12 (149/160) | 91.25 (146/160) | 442/480 |
-| Qwen3.5 2B + LM-head adapter | 73.75 | 65.00 (104/160) | 83.75 (134/160) | 81.25 (130/160) | 368/480 |
-| Qwen3.5 4B + LM-head adapter | 87.97 | 83.75 (134/160) | 93.12 (149/160) | 91.25 (146/160) | 429/480 |
-
-Nox's absolute weighted 95% CI is **[73.99%, 82.65%]**, compared with **[70.71%, 80.02%]** for Nox v1.0. Its BoolQ false and true recalls are both **87.50%** (70/80 each). Machine-readable results include all model intervals and Noul/Choice NLL and Brier scores. Probabilities and released calibration are unchanged; zero assigned gold probability yields infinite NLL without epsilon clipping.
-
-All seven models returned 480/480 valid outputs. Complete inputs were verified for all six local models, with zero truncations; retention inside Jev's closed API is unknown. Sol is the immutable v1.0 reference. Qwen baselines use the same starting model revisions with a fixed LM-head readout, not randomly initialized decision heads. These are quality measurements, not latency comparisons.
-
-Intervals use 10,000 paired passage-cluster bootstrap draws. BoolQ has 160 passage parents; Belebele has 144, with every selected question and both translations from each passage resampled together. All models share the same draws.
-
-[BoolQ validation](https://huggingface.co/datasets/google/boolq/tree/35b264d03638db9f4ce671b711558bf7ff0f80d5) (CC BY-SA 3.0) and [Belebele test](https://huggingface.co/datasets/facebook/belebele/tree/7899cdfa4e1e0d733fd77c848e2c273cb1d32be2) (CC BY-SA 4.0) were excluded from custom TRAIN/SELECT/CAL. Original human labels and English/Chinese alignment were independently checked; this was not new human adjudication. The lexical audit found no exact or >=0.5 character-5-gram Jaccard overlaps across 51 data/panel files, but cannot establish semantic independence or unknown upstream-pretraining exposure. BoolQ lacks article-title metadata, so grouping is by passage, not article.
-
-The candidate was frozen before the panel's first global unseal on 2026-09-21. The panel is now observed regression for subsequent unfrozen candidates. It covers evidence yes/no and four-choice comprehension, not ordinal Score or all decision capabilities. Aggregate data and SHA-256 evidence accompany this section.
-
-[Reading benchmark aggregate measurements](metrics/natural-confirmation.json).
+# Evaluation
+
+Four complete panels, shared requests and recorded native decisions. Failures and rejections remain in the requested denominator. The four-panel mean uses 25% each for general decisions, compositional tasks, natural reading, and reading/inference. Within-panel source/family weights are retained. Intervals use the frozen paired source-cluster procedure; they are reported without introducing another release gate.
+
+| Model | Version | General decisions | Compositional tasks | Natural reading | Reading and inference | Mean |
+|---|---|---:|---:|---:|---:|---:|
+| Jev | 1.13.0 | 79.10 | **66.38** | **94.53** | **89.79** | **82.45** |
+| Nox | v1.2 · this release | **83.21** | 51.08 | 78.12 | 84.38 | 74.20 |
+| Nox | v1.1 · previous | 82.85 | 51.25 | 78.44 | 79.17 | 72.93 |
+| Decider | 2B | 64.01 | 46.58 | 92.03 | 84.38 | 71.75 |
+| Qwen3.5 | 4B · untuned | 69.89 | 43.33 | 87.97 | 79.79 | 70.25 |
+| Qwen3.5 | 2B · untuned | 57.12 | 39.00 | 73.75 | 72.29 | 60.54 |
+| Laya | Upstream default | 57.01 | 37.75 | 51.25 | 63.75 | 52.44 |
+
+Accuracy (%). Mean weights each panel equally; each panel retains its frozen family/source weights.
+Bold marks the exact highest score in a column, including exact ties. Rounded visual ties can differ.
+
+### General decisions
+
+| Task | Jev · 1.13.0 | Nox · v1.2 · this release | Nox · v1.1 · previous | Decider · 2B | Qwen3.5 · 4B · untuned | Qwen3.5 · 2B · untuned | Laya · Upstream default |
+|---|---:|---:|---:|---:|---:|---:|---:|
+| ag news | 85.16 | 85.16 | 83.59 | 86.72 | 84.38 | 80.47 | **91.41** |
+| boolean constraints | **100.00** | 93.75 | 93.75 | 71.88 | 62.50 | 37.50 | 43.75 |
+| dbpedia 14 | 96.43 | 96.43 | 96.43 | **98.21** | 97.32 | 92.86 | 83.93 |
+| natural intents | **100.00** | **100.00** | **100.00** | **100.00** | **100.00** | 96.88 | 85.94 |
+| option carrier | 30.21 | **100.00** | **100.00** | 8.33 | 72.92 | 9.38 | 95.83 |
+| ordinal rubric | **100.00** | 89.06 | 89.06 | 84.38 | 90.62 | 81.25 | 18.75 |
+| relational composition | **56.25** | 54.17 | 55.21 | 54.17 | 51.04 | 48.96 | 25.00 |
+| scoped evidence | **89.58** | 78.12 | 75.00 | 47.92 | 51.04 | 36.46 | 37.50 |
+| state tracking | 33.33 | **35.42** | **35.42** | 29.17 | 28.12 | 25.00 | 23.96 |
+| unknown rejection | **100.00** | **100.00** | **100.00** | 59.38 | 60.94 | 62.50 | 64.06 |
+
+### Compositional tasks
+
+| Task | Jev · 1.13.0 | Nox · v1.2 · this release | Nox · v1.1 · previous | Decider · 2B | Qwen3.5 · 4B · untuned | Qwen3.5 · 2B · untuned | Laya · Upstream default |
+|---|---:|---:|---:|---:|---:|---:|---:|
+| canonical record identity | **76.25** | 50.00 | 50.00 | 56.25 | 50.00 | 46.25 | 47.50 |
+| capacitated assignment | **68.75** | 41.25 | 43.75 | 51.25 | 50.00 | 50.00 | 63.75 |
+| constraint assignment | **55.00** | 41.25 | 42.50 | 23.75 | 23.75 | 21.25 | 22.50 |
+| massive en | **91.67** | 90.83 | **91.67** | 90.00 | **91.67** | 70.83 | 74.17 |
+| massive zh | **88.33** | 87.50 | **88.33** | **88.33** | 86.67 | 74.17 | 73.33 |
+| multiset reconciliation | **58.75** | 28.75 | 31.25 | 23.75 | 27.50 | 27.50 | 26.25 |
+| ordinal service loss | **42.50** | 31.25 | 28.75 | 22.50 | 21.25 | 20.00 | 20.00 |
+| paraconsistent rule closure | **66.25** | 33.75 | 35.00 | 26.25 | 25.00 | 27.50 | 18.75 |
+| temporal exclusion | 41.25 | **45.00** | 41.25 | 42.50 | 31.25 | 28.75 | 12.50 |
+| transaction recovery | **75.00** | 61.25 | 60.00 | 41.25 | 26.25 | 23.75 | 18.75 |
+
+### Natural reading
+
+| Task | Jev · 1.13.0 | Nox · v1.2 · this release | Nox · v1.1 · previous | Decider · 2B | Qwen3.5 · 4B · untuned | Qwen3.5 · 2B · untuned | Laya · Upstream default |
+|---|---:|---:|---:|---:|---:|---:|---:|
+| boolq | **92.50** | 85.62 | 87.50 | 91.88 | 83.75 | 65.00 | 69.38 |
+| belebele en | **96.88** | 71.88 | 71.25 | 93.12 | 93.12 | 83.75 | 38.12 |
+| belebele zh | **96.25** | 69.38 | 67.50 | 91.25 | 91.25 | 81.25 | 28.12 |
+
+### Reading and inference
+
+| Task | Jev · 1.13.0 | Nox · v1.2 · this release | Nox · v1.1 · previous | Decider · 2B | Qwen3.5 · 4B · untuned | Qwen3.5 · 2B · untuned | Laya · Upstream default |
+|---|---:|---:|---:|---:|---:|---:|---:|
+| cosmos qa | **86.67** | 68.33 | 57.50 | 66.67 | 60.83 | 56.67 | 30.00 |
+| squad2 answerability | **90.83** | 85.00 | 75.83 | 84.17 | 81.67 | 82.50 | 57.50 |
+| snli | 82.50 | 86.67 | 85.83 | **90.83** | 80.00 | 64.17 | 72.50 |
+| qasc | **99.17** | 97.50 | 97.50 | 95.83 | 96.67 | 85.83 | 95.00 |
+
+### Probability quality
+
+| Model | Version | Weighted Brier ↓ | Valid probability rows |
+|---|---|---:|---:|
+| Jev | 1.13.0 | **0.2287** | 2,720 / 2,720 |
+| Nox | v1.2 · this release | 0.3416 | 2,720 / 2,720 |
+| Nox | v1.1 · previous | 0.3630 | 2,720 / 2,720 |
+| Decider | 2B | 0.3535 | 2,720 / 2,720 |
+| Qwen3.5 | 4B · untuned | 0.3996 | 2,720 / 2,720 |
+| Qwen3.5 | 2B · untuned | 0.5130 | 2,720 / 2,720 |
+| Laya | Upstream default | 0.5903 | 2,720 / 2,720 |
+
+Missing Brier means probability coverage was incomplete; no supported-only average is substituted.
+The two native contract supplements are reported separately and are outside this quality mean.
+
+
+
+
+
+
+
+
+
+## Scope
+
+These panels include previously observed regression sets. They are not all pristine holdouts. The untuned Qwen models use their frozen chat/LM-head adapters. Laya uses its unmodified upstream default language router (no explicit language override); Decider retains its native adapter. Documented native truncation is retained. Jev is a closed service snapshot, not a reproducible weight release. Full native-interface supplements are outside the quality mean.
+
+[Exact aggregate statistics](metrics/expanded-quality.json) · [Probability and adapter coverage](metrics/comparator-coverage.json) · [Measured latency](QUESTION-SCALING.md) · [Material provenance](metrics/materials-provenance.json)
diff --git a/NORMALIZATION_RUNTIME.md b/NORMALIZATION_RUNTIME.md
new file mode 100644
index 0000000000000000000000000000000000000000..d9a89dd1133c5f07ec94f3a3278018639c3a6c61
--- /dev/null
+++ b/NORMALIZATION_RUNTIME.md
@@ -0,0 +1,5 @@
+# Validated normalization runtime
+
+This Nox bundle installs its recorded FLA normalization profile through the default public entrypoint. It requires the pinned ROCm runtime on gfx942 and a fresh process. The profile covers dimension 128, BF16 input/output, FP32 reciprocal norms, and buckets 1–64 for the actual 32 normalized value heads, batch size at most 8 and complete inputs at most 16,384 tokens. Unknown keys fail closed. Weights, tokenizer, prompt, readout and temperature remain unchanged from the source export.
+
+No private compilation cache is required. Other FLA kernels retain normal runtime behavior. Numerical validation of this newly bound bundle is still required. Earlier latency measurements do not measure this runtime or candidate.
diff --git a/QUESTION-SCALING.md b/QUESTION-SCALING.md
index 8409dd27fa34cd1dcab1cdb59fac07da68321f39..dbdcdcd2fe6cbcce30fe90f95abdb210cf9bd3aa 100644
--- a/QUESTION-SCALING.md
+++ b/QUESTION-SCALING.md
@@ -1,55 +1,64 @@
-# Fixed-question latency scaling — v1.0
-
-These measurements use the initial Sol/Nox v1.0 weights. They are retained as versioned evidence and do not measure the updated weights.
-
-A fixed English state (819 characters), identical question text and a four-choice menu are repeated Q times in one API request. Native rendered rows, including state and question, contain 309 tokens for Sol/Nox, 228 for Laya and 250 for Decider. Tokenizers and templates differ. Every model uses the same physical AMD gfx942 GPU and matches its quality/runtime fingerprint. Native question packing is preserved, with no cross-request cache or concurrent-client load.
-
-Points show the empirical median of 30 timed requests, from three blocks with randomized model order. Each point receives three warmups and ten measured repetitions per block. Rendering, tokenization, transfers, model forward and answer assembly are included; loading, warmup and network are excluded. Lines connect measured points for readability and do not represent measured intermediate values. The horizontal axis is logarithmic in base 2.
-
-Q1/Q8/Q32 come from the original GPU-quiet round; Q2/Q4/Q16 are an additive round under the same frozen setup. A repeated Q1 bridge is reported separately and never pooled or selected. This is local request scaling, not maximum server throughput. Laya uses its English branch for these English inputs.
-
-Base-model appendix retains the original measured Q1/Q8/Q32 points only.
-
-The machine exposes 261824 MiB of device memory. Parameters, native precision, shipped temperature and model/runtime fingerprints are those of the original v1.0 quality benchmark. No throughput saturation claim is made. p95 is an empirical percentile, not a confidence interval.
-
-Full aggregate evidence: [question-scaling.json](metrics/question-scaling.json).
-
-| Model | Questions | Tokens / question | p50 (ms) | p95 (ms) | Samples | Round |
-|---|---:|---:|---:|---:|---:|---|
-| Sol 2B | 1 | 309 | 21.04 | 21.41 | 30 | original |
-| Sol 2B | 2 | 309 | 21.48 | 21.81 | 30 | additive |
-| Sol 2B | 4 | 309 | 23.47 | 31.89 | 30 | additive |
-| Sol 2B | 8 | 309 | 37.29 | 37.51 | 30 | original |
-| Sol 2B | 16 | 309 | 73.88 | 74.46 | 30 | additive |
-| Sol 2B | 32 | 309 | 147.50 | 147.91 | 30 | original |
-| Nox 4B | 1 | 309 | 27.37 | 28.95 | 30 | original |
-| Nox 4B | 2 | 309 | 28.17 | 28.54 | 30 | additive |
-| Nox 4B | 4 | 309 | 40.89 | 41.04 | 30 | additive |
-| Nox 4B | 8 | 309 | 69.25 | 69.84 | 30 | original |
-| Nox 4B | 16 | 309 | 137.64 | 137.82 | 30 | additive |
-| Nox 4B | 32 | 309 | 275.65 | 278.26 | 30 | original |
-| Laya | 1 | 228 | 11.54 | 12.21 | 30 | original |
-| Laya | 2 | 228 | 12.60 | 13.02 | 30 | additive |
-| Laya | 4 | 228 | 13.37 | 13.68 | 30 | additive |
-| Laya | 8 | 228 | 14.56 | 15.02 | 30 | original |
-| Laya | 16 | 228 | 22.47 | 22.80 | 30 | additive |
-| Laya | 32 | 228 | 37.89 | 38.30 | 30 | original |
-| Decider | 1 | 250 | 23.35 | 23.91 | 30 | original |
-| Decider | 2 | 250 | 23.52 | 23.89 | 30 | additive |
-| Decider | 4 | 250 | 24.05 | 24.31 | 30 | additive |
-| Decider | 8 | 250 | 31.59 | 31.77 | 30 | original |
-| Decider | 16 | 250 | 53.12 | 53.32 | 30 | additive |
-| Decider | 32 | 250 | 96.67 | 97.02 | 30 | original |
-
-## Nonpooled Q1 bridge
-
-Flag threshold: absolute median difference greater than max(10% of the original Q1 median, 2 ms). Original Q1 remains the plotted point in every case; no normalization, replacement or pooling.
-
-| Model | Original p50 (ms) | Bridge p50 (ms) | Difference (ms) | Flag |
-|---|---:|---:|---:|---|
-| Sol 2B | 21.04 | 20.87 | -0.17 | No |
-| Nox 4B | 27.37 | 27.16 | -0.21 | No |
-| Laya | 11.54 | 12.14 | +0.60 | No |
-| Decider | 23.35 | 23.18 | -0.17 | No |
-
-All bridge checks are below the predeclared threshold.
+# Request latency
+
+Fresh matched measurements of this candidate and its own published predecessor.
+
+The same physical AMD gfx942 GPU runs three independently loaded blocks per model in a counterbalanced order. Each of 18 fixed Choice/Noul/Score cases has three warmups and ten measurements per block. The tables pool all 30 measurements per point. Loading and network time are excluded; tokenization and inference are included. No cross-request prefix cache is enabled. These are sequential requests, not concurrent-service throughput.
+
+
+
+## Choice
+
+| Model | Q | p50 ms ↓ | p95 ms ↓ | Input tokens |
+|---|---:|---:|---:|---:|
+| Nox · v1.2 · this release | 1 | 32.108 | 33.047 | 309 |
+| Nox · v1.1 · previous | 1 | **28.976** | **30.240** | 309 |
+| Nox · v1.2 · this release | 2 | 33.025 | 34.047 | 618 |
+| Nox · v1.1 · previous | 2 | **29.337** | **29.953** | 618 |
+| Nox · v1.2 · this release | 4 | **41.706** | **42.005** | 1236 |
+| Nox · v1.1 · previous | 4 | 42.041 | 42.413 | 1236 |
+| Nox · v1.2 · this release | 8 | **69.483** | **70.719** | 2472 |
+| Nox · v1.1 · previous | 8 | 70.106 | 70.850 | 2472 |
+| Nox · v1.2 · this release | 16 | **141.299** | **141.921** | 4944 |
+| Nox · v1.1 · previous | 16 | 141.369 | 142.338 | 4944 |
+| Nox · v1.2 · this release | 32 | **282.326** | **284.856** | 9888 |
+| Nox · v1.1 · previous | 32 | 283.681 | 286.013 | 9888 |
+
+## Noul
+
+| Model | Q | p50 ms ↓ | p95 ms ↓ | Input tokens |
+|---|---:|---:|---:|---:|
+| Nox · v1.2 · this release | 1 | 32.282 | 32.977 | 262 |
+| Nox · v1.1 · previous | 1 | **28.993** | **29.371** | 262 |
+| Nox · v1.2 · this release | 2 | 32.573 | 33.978 | 524 |
+| Nox · v1.1 · previous | 2 | **29.196** | **30.459** | 524 |
+| Nox · v1.2 · this release | 4 | **38.967** | **39.277** | 1048 |
+| Nox · v1.1 · previous | 4 | 39.373 | 39.573 | 1048 |
+| Nox · v1.2 · this release | 8 | 64.497 | **65.068** | 2096 |
+| Nox · v1.1 · previous | 8 | **64.382** | 65.297 | 2096 |
+| Nox · v1.2 · this release | 16 | **129.421** | **130.165** | 4192 |
+| Nox · v1.1 · previous | 16 | 129.926 | 130.926 | 4192 |
+| Nox · v1.2 · this release | 32 | 260.115 | **261.017** | 8384 |
+| Nox · v1.1 · previous | 32 | **259.700** | 262.418 | 8384 |
+
+## Score
+
+| Model | Q | p50 ms ↓ | p95 ms ↓ | Input tokens |
+|---|---:|---:|---:|---:|
+| Nox · v1.2 · this release | 1 | 32.601 | 33.154 | 307 |
+| Nox · v1.1 · previous | 1 | **28.816** | **29.119** | 307 |
+| Nox · v1.2 · this release | 2 | 33.221 | 33.625 | 614 |
+| Nox · v1.1 · previous | 2 | **29.339** | **29.777** | 614 |
+| Nox · v1.2 · this release | 4 | **41.776** | 42.379 | 1228 |
+| Nox · v1.1 · previous | 4 | 41.965 | **42.220** | 1228 |
+| Nox · v1.2 · this release | 8 | **69.535** | **70.140** | 2456 |
+| Nox · v1.1 · previous | 8 | 69.583 | 70.672 | 2456 |
+| Nox · v1.2 · this release | 16 | **140.447** | **141.969** | 4912 |
+| Nox · v1.1 · previous | 16 | 141.124 | 142.791 | 4912 |
+| Nox · v1.2 · this release | 32 | **281.576** | **284.902** | 9824 |
+| Nox · v1.1 · previous | 32 | 283.684 | 286.097 | 9824 |
+
+Choice six-point geometric-mean latency ratio (candidate / predecessor): **1.0337**.
+
+The candidate has 3.37% higher measured latency by this summary. This is a descriptive result for these six request sizes, not a claim of universal speedup.
+
+The runtime/profile and temperature belong to these exact measured bundles. Older release timing is not reused.
diff --git a/README.md b/README.md
index c450fdb566400b75d96871ce36baea53a73e4f69..8d485977771209af7f0ab93fc0608805ed0a624a 100644
--- a/README.md
+++ b/README.md
@@ -12,13 +12,6 @@ tags:
- custom-code
- pytorch
- rocm
-- choice
-- noul
-- scoring
-datasets:
-- PolyAI/banking77
-- clinc/clinc_oos
-- nyu-mll/multi_nli
---

@@ -27,69 +20,47 @@ datasets:
*Nox, Latin for night.*
-**Your move.**
-
-A capable decoder for decisions defined by you. Give Nox a state, questions and possible answers. It returns choices, yes/no judgments and rubric scores with probability distributions—one forward pass per question.
+**Your move.** Give Nox a state, questions and possible answers. It returns decisions and probabilities with labels defined at runtime.
**4.208B parameters · 16K complete-question budget · English / Chinese evaluated · Apache 2.0**
-[Decision family](https://huggingface.co/collections/llm-semantic-router/decision-10-6ab12177bd0002394d8409f9) · [Meet Sol](https://huggingface.co/llm-semantic-router/Decision-1.0-Sol)
-
-## Three ways to decide
+[Decision family](https://huggingface.co/collections/llm-semantic-router/decision-10-6ab12177bd0002394d8409f9)
| Type | Use it for | Output |
-| --- | --- | --- |
-| **Choice** | Route a request; select an action from 2–255 candidates. | Selected ID + distribution |
-| **Noul** | Check a condition against available evidence. | P(true) |
+|---|---|---|
+| **Choice** | Route a request or choose among 2–255 actions. | Selected ID + distribution |
+| **Noul** | Check a condition against supplied evidence. | P(true) |
| **Score** | Apply 2–10 ordered rubric descriptions. | Expected index + distribution |
-Your question names and candidate IDs are preserved. Labels are defined at runtime.
-
## Measured capability
-**67.05% overall**, up 4.33 points from Nox v1.0 across 20 equally weighted capabilities.
-
-| Model | Overall accuracy ↑ | Choice ↑ | Noul ↑ | Score ↑ |
-| --- | ---: | ---: | ---: | ---: |
-| Jev · 1.13.0 | **72.74** | **71.77** | **80.36** | **68.06** |
-| Nox · 4B · v1.1 | 67.05 | 70.55 | 60.27 | 55.56 |
-| Qwen3.5 · 4B · untuned | 56.61 | 60.06 | 53.57 | 52.08 |
-| Sol · 2B · v1.0 | 55.58 | 61.93 | 50.89 | 32.64 |
-| Decider · 2B | 55.30 | 57.26 | 58.93 | 50.00 |
-| Qwen3.5 · 2B · untuned | 48.06 | 50.36 | 45.09 | 47.22 |
-| Laya · EN/ML | 47.38 | 53.02 | 52.23 | 19.44 |
-
-Accuracy (%). Overall is the equal-weight mean of **20 task families**, covering **1,760 questions**; type columns pool their questions. Same requests for every model. Laya's native interface reported 152 truncated inputs. [Methods, coverage and uncertainty](EVALUATION.md).
-
-
-
-
+**74.20% four-panel mean**, +1.27 points versus its published predecessor.
-### Reading comprehension
+| Model | Mean accuracy ↑ | Decisions | Composition | Reading | Inference |
+|---|---:|---:|---:|---:|---:|
+| Jev · 1.13.0 | **82.45** | 79.10 | **66.38** | **94.53** | **89.79** |
+| Nox · v1.2 · this release | 74.20 | **83.21** | 51.08 | 78.12 | 84.38 |
+| Nox · v1.1 · previous | 72.93 | 82.85 | 51.25 | 78.44 | 79.17 |
+| Decider · 2B | 71.75 | 64.01 | 46.58 | 92.03 | 84.38 |
+| Qwen3.5 · 4B · untuned | 70.25 | 69.89 | 43.33 | 87.97 | 79.79 |
+| Qwen3.5 · 2B · untuned | 60.54 | 57.12 | 39.00 | 73.75 | 72.29 |
+| Laya · Upstream default | 52.44 | 57.01 | 37.75 | 51.25 | 63.75 |
-On a separate **480-question human-annotated reading benchmark**, Nox improves **75.47% → 78.44%**. Jev, Decider and the untuned 4B parent remain stronger on this task.
+Accuracy (%), **2,720 decisions**. Each panel contributes one quarter; its original family/source weights are retained. Bold marks each column's exact highest score. [Methods, uncertainty and all task results](EVALUATION.md).
-| Model | Weighted accuracy (%) ↑ |
-| --- | ---: |
-| Jev 1.13.0 | **94.53** |
-| Decider 2B | 92.03 |
-| Qwen3.5 4B + LM-head adapter | 87.97 |
-| Nox · 4B · v1.1 | 78.44 |
-| Qwen3.5 2B + LM-head adapter | 73.75 |
-| Sol · 2B · v1.0 | 69.38 |
-| Laya EN/ML routed | 51.25 |
+
-Weights: 50% BoolQ, 25% Belebele English, 25% Belebele Chinese. [Exact counts, uncertainty and source limitations](EVALUATION.md#human-source-confirmation-v4).
+
## More questions, measured
-
+
-**Measured with v1.0 weights.** Same state, same question length, four choices; only the number of questions changes. Each model keeps its native interface (Nox: 309 tokens/question). Median of 30 requests per point on one AMD gfx942 GPU; local Python request time includes tokenization and inference, excluding loading and network. [p95, drift checks and full methods](QUESTION-SCALING.md).
+Same inputs and physical AMD gfx942 GPU; 30 measured requests per point across three blocks. Python request latency includes tokenization and inference, excluding loading and network. [p95 and all three native types](QUESTION-SCALING.md).
## Try it
-Download the fixed release with `hf download llm-semantic-router/Decision-1.0-Nox --revision v1.1 --local-dir decision-model`, then follow the [ROCm setup](RUNTIME.md). Inside that container, with the model mounted at `/model`:
+Download `hf download llm-semantic-router/Decision-1.0-Nox --revision v1.2 --local-dir decision-model`, then follow [ROCm setup](RUNTIME.md). In that container, with the model mounted at `/model`:
```python
from decision import DecisionModel
@@ -99,16 +70,16 @@ model = DecisionModel.from_pretrained("/model", local_files_only=True)
print(model.decide(**REQUEST)["answers"])
```
-[Actual request and measured output](model-card-example.json) · [Install and API guide](USAGE.md)
+[Tested request and output](model-card-example.json) · [Install and API guide](USAGE.md)
-The complete state, question and candidates must fit 16,384 tokens. Overflow is rejected. Score returns an expected ordinal index. Runtime: AMD gfx942 validated; CPU/MPS unsupported; NVIDIA unqualified.
+The complete state, question and candidates must fit 16,384 tokens; overflow is rejected. The bundled normalization profile loads automatically. AMD gfx942 is validated; CPU/MPS are unsupported and NVIDIA is unqualified. Use a fresh Python process when switching profiles.
## Architecture

-A causal Qwen3.5 text backbone combines gated linear attention with full attention. A shared candidate head reads candidate endpoints and the final query vector. Questions run independently in batches of eight.
+A causal Qwen3.5 text backbone combines gated linear and full attention. A shared candidate head reads candidate endpoints and the final query vector. Each question uses one forward pass; questions run independently in batches of eight.
-[Candidate head](assets/readout.png) · [Editable SVG](assets/architecture.svg) · [Vector atlas](assets/architecture-atlas.pdf) · [Inference code](code/decision_model.py)
+[Candidate head](assets/readout.png) · [Vector architecture](assets/architecture.svg) · [Inference code](code/decision_model.py)
-Adapted from [Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B). It evaluates supplied evidence, without live retrieval; confidence is not a correctness guarantee. [Apache 2.0](https://huggingface.co/llm-semantic-router/Decision-1.0-Nox/blob/main/LICENSE) · [Attributions](ATTRIBUTIONS.md).
+Adapted from [Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B). It evaluates supplied evidence without live retrieval; confidence does not guarantee correctness. [License](LICENSE) · [Attributions](ATTRIBUTIONS.md).
diff --git a/RUNTIME.md b/RUNTIME.md
index db2642566b6503f2bab14dcae15a95e18f55c7ff..2106cc049a44e2ec2493aa5b84148b32e716492d 100644
--- a/RUNTIME.md
+++ b/RUNTIME.md
@@ -2,7 +2,7 @@
The package has a public, digest-pinned installation path. `Dockerfile.runtime` starts from `vllm/vllm-openai-rocm@sha256:1fd21abe66455b4df5a2e83629e97cdcc9d58913b16052d8118b92b239792339`, adds two hash-checked FLA wheels, and installs this repository's loading wrapper. It keeps the base image's ROCm PyTorch and Triton builds. The vLLM server is not used by Decision inference.
-**Validation boundary:** the public registry manifest, base-image ancestry, package metadata and critical PyTorch binary hashes have been checked. The recipe built successfully and passed CPU imports and real AMD ROCm gfx942 GPU inference for both initial v1.0 bundles. On the packaged three-question Choice/Noul/Score example, its complete responses matched the qualified research runtime exactly; the wrapper matched the direct engine and rejected an oversized complete input. The initial v1.0 build evidence is in `runtime-build-provenance.json`. Updated weights retain the same installation recipe and are separately checked during release; their actual request and output are in `model-card-example.json`. This example establishes a working public installation path; it is not a full rerun of the quality or timing benchmark. Published benchmark results use the qualified runtime in each bundle's `runtime.json`.
+**Validation boundary:** the public registry manifest, base-image ancestry, package metadata and critical PyTorch binary hashes have been checked. The recipe built successfully and passed CPU imports and real AMD ROCm gfx942 GPU GPU inference for both released bundles. On the packaged three-question Choice/Noul/Score example, its complete responses matched the qualified research runtime exactly; the wrapper matched the direct engine and rejected an oversized complete input. Evidence is in `runtime-build-provenance.json`. This example establishes a working public installation path; it is not a full rerun of the quality or timing benchmark. Published benchmark results use the qualified runtime in each bundle's `runtime.json`.
## Build and run
diff --git a/RUNTIME_BINDING.json b/RUNTIME_BINDING.json
new file mode 100644
index 0000000000000000000000000000000000000000..02627d018524d3d6a21799cb8eb3a8489f9ce2af
--- /dev/null
+++ b/RUNTIME_BINDING.json
@@ -0,0 +1,76 @@
+{
+ "source_bundle_manifest_sha256": "6d402d6d55e734851cb0d6417d6f3b416b25ae807c3a0d6dcf12e40acb6ac546",
+ "profile_sha256": "be32858d15233e0a3fbee0e4257fb02be0b3439deee4eb9c3f61151df7b73850",
+ "profile_validation_receipt_sha256": "7a66474325b5575fdd15a33e5e0340c19e809bdcc22adf4cbb45baa0e1b32693",
+ "unchanged_files": [
+ {
+ "file": "backbone/config.json",
+ "bytes": 1978,
+ "sha256": "ae3a463b32e95b6cc207a7af4f1defb4195f388eb6f9ff19d2690b73d4966953"
+ },
+ {
+ "file": "backbone/model-00001-of-00003.safetensors",
+ "bytes": 3991295368,
+ "sha256": "5f9c4bc396605b551b5983a1699de8e96755fafdb6afdb228a5cc4d9b438a5da"
+ },
+ {
+ "file": "backbone/model-00002-of-00003.safetensors",
+ "bytes": 3979828128,
+ "sha256": "656cc757ba03f87cedc0e4c88237db1ae8215686f9b6477334a34dc0a32136db"
+ },
+ {
+ "file": "backbone/model-00003-of-00003.safetensors",
+ "bytes": 440425856,
+ "sha256": "0a30a07b6efbaa37d016008848c1fa6dadcb46ae451a0e5e4b927ac9b5771804"
+ },
+ {
+ "file": "backbone/model.safetensors.index.json",
+ "bytes": 33047,
+ "sha256": "1602d52e38d81586af85bc4ce29ce082c5fc5877c763b1ebcab7545320016599"
+ },
+ {
+ "file": "chat_template.jinja",
+ "bytes": 7756,
+ "sha256": "a4aee8afcf2e0711942cf848899be66016f8d14a889ff9ede07bca099c28f715"
+ },
+ {
+ "file": "code/decision_model.py",
+ "bytes": 10114,
+ "sha256": "d3e28489c09f3bd7130e2d43d92e0b5c4a08e25b09b21303904defb0ff1c3646"
+ },
+ {
+ "file": "code/predict.py",
+ "bytes": 3164,
+ "sha256": "02352e8385ab47157b6459910da54d962e5da4bb4940d571faedc86bc5da9aee"
+ },
+ {
+ "file": "decision_config.json",
+ "bytes": 753,
+ "sha256": "443a9b3f191a8387915c606de8303700fc1a069fe7c8ad46a0eba5528a2549a8"
+ },
+ {
+ "file": "decision_head.safetensors",
+ "bytes": 10529624,
+ "sha256": "8b8e342445035e503b4963e16c5f7a4fef95c08368e27eace312145fc3660772"
+ },
+ {
+ "file": "temperature.json",
+ "bytes": 6095,
+ "sha256": "1d652e75f49b5c232d1e742b189c1c826ad52ced1a931a124e29e65e7d7fb569"
+ },
+ {
+ "file": "tokenizer.json",
+ "bytes": 19989325,
+ "sha256": "06b9509352d2af50381ab2247e083b80d32d5c0aba91c272ca9ff729b6a0e523"
+ },
+ {
+ "file": "tokenizer_config.json",
+ "bytes": 1123,
+ "sha256": "bee8eba30f0eb4af73c0fe2cd06d0f89b657d7819941c438157ec42f7c80ea87"
+ }
+ ],
+ "training_selection_calibration_unchanged": true,
+ "recomputed_or_recalibrated": false,
+ "private_compiled_cache_required": false,
+ "new_bundle_offline_proof_required": true
+}
diff --git a/SOURCE_BUNDLE_MANIFEST.json b/SOURCE_BUNDLE_MANIFEST.json
new file mode 100644
index 0000000000000000000000000000000000000000..e93c76ac366e0f8c4a748927663c5917adc30475
--- /dev/null
+++ b/SOURCE_BUNDLE_MANIFEST.json
@@ -0,0 +1,173 @@
+{
+ "format": "research-pointer-bundle-v1",
+ "status": "candidate-export-awaiting-independent-reload-and-quality-gates",
+ "files": [
+ {
+ "file": "backbone/config.json",
+ "bytes": 1978,
+ "sha256": "ae3a463b32e95b6cc207a7af4f1defb4195f388eb6f9ff19d2690b73d4966953"
+ },
+ {
+ "file": "backbone/model-00001-of-00003.safetensors",
+ "bytes": 3991295368,
+ "sha256": "5f9c4bc396605b551b5983a1699de8e96755fafdb6afdb228a5cc4d9b438a5da"
+ },
+ {
+ "file": "backbone/model-00002-of-00003.safetensors",
+ "bytes": 3979828128,
+ "sha256": "656cc757ba03f87cedc0e4c88237db1ae8215686f9b6477334a34dc0a32136db"
+ },
+ {
+ "file": "backbone/model-00003-of-00003.safetensors",
+ "bytes": 440425856,
+ "sha256": "0a30a07b6efbaa37d016008848c1fa6dadcb46ae451a0e5e4b927ac9b5771804"
+ },
+ {
+ "file": "backbone/model.safetensors.index.json",
+ "bytes": 33047,
+ "sha256": "1602d52e38d81586af85bc4ce29ce082c5fc5877c763b1ebcab7545320016599"
+ },
+ {
+ "file": "chat_template.jinja",
+ "bytes": 7756,
+ "sha256": "a4aee8afcf2e0711942cf848899be66016f8d14a889ff9ede07bca099c28f715"
+ },
+ {
+ "file": "code/decision_api.py",
+ "bytes": 6724,
+ "sha256": "147b2fec32cbbbbb1b92cf2a19bb887f9d945e4974e141ca7ea93b8c542f5b21"
+ },
+ {
+ "file": "code/decision_model.py",
+ "bytes": 10114,
+ "sha256": "d3e28489c09f3bd7130e2d43d92e0b5c4a08e25b09b21303904defb0ff1c3646"
+ },
+ {
+ "file": "code/predict.py",
+ "bytes": 3164,
+ "sha256": "02352e8385ab47157b6459910da54d962e5da4bb4940d571faedc86bc5da9aee"
+ },
+ {
+ "file": "decision_config.json",
+ "bytes": 753,
+ "sha256": "443a9b3f191a8387915c606de8303700fc1a069fe7c8ad46a0eba5528a2549a8"
+ },
+ {
+ "file": "decision_head.safetensors",
+ "bytes": 10529624,
+ "sha256": "8b8e342445035e503b4963e16c5f7a4fef95c08368e27eace312145fc3660772"
+ },
+ {
+ "file": "runtime.json",
+ "bytes": 378,
+ "sha256": "c5d3521358b2817f4e56ea8150c5c612139b4e2b2c5bb09f512d9a7c0b5298b9"
+ },
+ {
+ "file": "temperature.json",
+ "bytes": 6095,
+ "sha256": "1d652e75f49b5c232d1e742b189c1c826ad52ced1a931a124e29e65e7d7fb569"
+ },
+ {
+ "file": "tokenizer.json",
+ "bytes": 19989325,
+ "sha256": "06b9509352d2af50381ab2247e083b80d32d5c0aba91c272ca9ff729b6a0e523"
+ },
+ {
+ "file": "tokenizer_config.json",
+ "bytes": 1123,
+ "sha256": "bee8eba30f0eb4af73c0fe2cd06d0f89b657d7819941c438157ec42f7c80ea87"
+ }
+ ],
+ "tensors": [
+ {
+ "file": "backbone/model-00001-of-00003.safetensors",
+ "elements": 1995638528,
+ "elements_by_dtype": {
+ "BF16": 1995638528
+ }
+ },
+ {
+ "file": "backbone/model-00002-of-00003.safetensors",
+ "elements": 1989900928,
+ "elements_by_dtype": {
+ "BF16": 1989900928
+ }
+ },
+ {
+ "file": "backbone/model-00003-of-00003.safetensors",
+ "elements": 220211840,
+ "elements_by_dtype": {
+ "BF16": 220211840
+ }
+ },
+ {
+ "file": "decision_head.safetensors",
+ "elements": 2632192,
+ "elements_by_dtype": {
+ "F32": 2632192
+ }
+ }
+ ],
+ "source_checkpoint_files": [
+ {
+ "file": "backbone/config.json",
+ "bytes": 1977,
+ "sha256": "a5ed4156fda05f0f9149c66964d6165916754e7355488a8a07d9b0398acdbdb9"
+ },
+ {
+ "file": "backbone/model-00001-of-00005.safetensors",
+ "bytes": 3992269864,
+ "sha256": "8adb3dbf9db30b3ecd1f29ce6a853bc6d7e3f465857471cfe579c570537250f8"
+ },
+ {
+ "file": "backbone/model-00002-of-00005.safetensors",
+ "bytes": 3990302320,
+ "sha256": "8aa80bed34a466a7208ad177d5e5742773e4140c55cba69911fcc805633afa5e"
+ },
+ {
+ "file": "backbone/model-00003-of-00005.safetensors",
+ "bytes": 3937871784,
+ "sha256": "9af681e845c99e0d933dc3fa643271ee9f5bb382f98454ceb41eb1faf0979f87"
+ },
+ {
+ "file": "backbone/model-00004-of-00005.safetensors",
+ "bytes": 3926578128,
+ "sha256": "34de5c39c2338e1c7110beaf2834c856f3830411002893f7946045c0a083cc92"
+ },
+ {
+ "file": "backbone/model-00005-of-00005.safetensors",
+ "bytes": 976029360,
+ "sha256": "ce20083dd8deb786b14fe1126a12cdbf678a8a3c25367a6676c0ca55d9ae7980"
+ },
+ {
+ "file": "decision_config.json",
+ "bytes": 1110,
+ "sha256": "a12bd49b459771f22d22f29c59fe3cb6779c52a5d39b40cdbe5e247c5181b3fc"
+ },
+ {
+ "file": "decision_head.safetensors",
+ "bytes": 10529624,
+ "sha256": "8b8e342445035e503b4963e16c5f7a4fef95c08368e27eace312145fc3660772"
+ },
+ {
+ "file": "tokenizer.json",
+ "bytes": 19989325,
+ "sha256": "06b9509352d2af50381ab2247e083b80d32d5c0aba91c272ca9ff729b6a0e523"
+ },
+ {
+ "file": "tokenizer_config.json",
+ "bytes": 1123,
+ "sha256": "bee8eba30f0eb4af73c0fe2cd06d0f89b657d7819941c438157ec42f7c80ea87"
+ }
+ ],
+ "source_model_code_sha256": "d3e28489c09f3bd7130e2d43d92e0b5c4a08e25b09b21303904defb0ff1c3646",
+ "source_api_code_sha256": "147b2fec32cbbbbb1b92cf2a19bb887f9d945e4974e141ca7ea93b8c542f5b21",
+ "dev_sha256": "45b4cd46acc4be53b95c7ed8f333b3533972296a4d0557c7b8c48c9cab6ced61",
+ "production_predictions_sha256": "cf0b1a17008bddc2e3e18289866993770c41d4a5af5f6f4219508bb34fcabefa",
+ "temperature_sha256": "1d652e75f49b5c232d1e742b189c1c826ad52ced1a931a124e29e65e7d7fb569",
+ "runtime_sha256": "c5d3521358b2817f4e56ea8150c5c612139b4e2b2c5bb09f512d9a7c0b5298b9",
+ "production_batch_size": 8,
+ "input_length_limit": 16384,
+ "original_checkpoint_name": "winner",
+ "no_publication_performed": true
+}
diff --git a/assets/architecture-atlas.pdf b/assets/architecture-atlas.pdf
deleted file mode 100644
index 0026e78ffed2d4b4be209ad5646d3166dbdb985e..0000000000000000000000000000000000000000
--- a/assets/architecture-atlas.pdf
+++ /dev/null
@@ -1,3 +0,0 @@
-version https://git-lfs.github.com/spec/v1
-oid sha256:d1582ac115aa6a40129b6221e15a197e33d20847e03d69396f02fdab82771376
-size 306242
diff --git a/assets/architecture.pdf b/assets/architecture.pdf
deleted file mode 100644
index 3d94346a23104039a8ca5d04c5cd214ba203d0fc..0000000000000000000000000000000000000000
--- a/assets/architecture.pdf
+++ /dev/null
@@ -1,3 +0,0 @@
-version https://git-lfs.github.com/spec/v1
-oid sha256:ee87f997a0aa4343c2ec66038b7526361e0a3c0ae02b383f8928aac9dd9bd74c
-size 183111
diff --git a/assets/decision-capabilities.pdf b/assets/decision-capabilities.pdf
deleted file mode 100644
index 7ef95dd53592ddac69407a4f74d08702931dff7a..0000000000000000000000000000000000000000
--- a/assets/decision-capabilities.pdf
+++ /dev/null
@@ -1,3 +0,0 @@
-version https://git-lfs.github.com/spec/v1
-oid sha256:1db9c9fcad30a0721692eb2143149001686f37e1ab9639395aa40886185f0e1f
-size 51584
diff --git a/assets/decision-capabilities.png b/assets/decision-capabilities.png
deleted file mode 100644
index 86f86ab1291a6d8ac78b3cd5fa21286fb50d0d75..0000000000000000000000000000000000000000
--- a/assets/decision-capabilities.png
+++ /dev/null
@@ -1,3 +0,0 @@
-version https://git-lfs.github.com/spec/v1
-oid sha256:bf584ab8924fe6f1e4623a5d402cac6fbfde96f12eaefa8c6b650ede9d95f21b
-size 451771
diff --git a/assets/decision-capabilities.svg b/assets/decision-capabilities.svg
deleted file mode 100644
index e4e29a43b37904d523aaa65fb33b6920704ad1c6..0000000000000000000000000000000000000000
--- a/assets/decision-capabilities.svg
+++ /dev/null
@@ -1,2533 +0,0 @@
-
-
-
diff --git a/assets/decision-expanded-old_core-600px.png b/assets/decision-expanded-old_core-600px.png
new file mode 100644
index 0000000000000000000000000000000000000000..7e3d59c6c97fad779ae2883b578a4bfb43e3dd77
--- /dev/null
+++ b/assets/decision-expanded-old_core-600px.png
@@ -0,0 +1,3 @@
+version https://git-lfs.github.com/spec/v1
+oid sha256:519923f36d62df99c6769731416d70413c9f5ec3991ba43a75de0714767a12cc
+size 140155
diff --git a/assets/decision-expanded-old_core.pdf b/assets/decision-expanded-old_core.pdf
new file mode 100644
index 0000000000000000000000000000000000000000..36ff27c92f667f8a821dea19ee62a1f22381fd0f
Binary files /dev/null and b/assets/decision-expanded-old_core.pdf differ
diff --git a/assets/decision-expanded-old_core.png b/assets/decision-expanded-old_core.png
new file mode 100644
index 0000000000000000000000000000000000000000..c2987afc600b1f1d5d402a46a7f362d49fad5dac
--- /dev/null
+++ b/assets/decision-expanded-old_core.png
@@ -0,0 +1,3 @@
+version https://git-lfs.github.com/spec/v1
+oid sha256:c59e464a07320a13a182ec841fef0a826eed777d03264260d3a3d6bcf99a2819
+size 261350
diff --git a/assets/decision-expanded-old_core.svg b/assets/decision-expanded-old_core.svg
new file mode 100644
index 0000000000000000000000000000000000000000..27c9c63c48354d64653349997b2ff3cf53935c3d
--- /dev/null
+++ b/assets/decision-expanded-old_core.svg
@@ -0,0 +1,1434 @@
+
+
+
diff --git a/assets/decision-expanded-overview-600px.png b/assets/decision-expanded-overview-600px.png
new file mode 100644
index 0000000000000000000000000000000000000000..8aab2201341a316a1bd8bc1cd169c8550d442329
Binary files /dev/null and b/assets/decision-expanded-overview-600px.png differ
diff --git a/assets/decision-expanded-overview.pdf b/assets/decision-expanded-overview.pdf
new file mode 100644
index 0000000000000000000000000000000000000000..944776f1923284c45f1887ede44b7f6f7e59ef77
Binary files /dev/null and b/assets/decision-expanded-overview.pdf differ
diff --git a/assets/decision-expanded-overview.png b/assets/decision-expanded-overview.png
new file mode 100644
index 0000000000000000000000000000000000000000..0af17bbdacd38de51acb504ce2f241408f4d48e3
--- /dev/null
+++ b/assets/decision-expanded-overview.png
@@ -0,0 +1,3 @@
+version https://git-lfs.github.com/spec/v1
+oid sha256:fa9972beb8e9081bb5a8c59e487e69f31567470fd4af6b0d6d36bdb82d60cdf1
+size 167930
diff --git a/assets/decision-expanded-overview.svg b/assets/decision-expanded-overview.svg
new file mode 100644
index 0000000000000000000000000000000000000000..df15c54a630f18e5d14a1d79506fdbf9de081928
--- /dev/null
+++ b/assets/decision-expanded-overview.svg
@@ -0,0 +1,714 @@
+
+
+
diff --git a/assets/decision-expanded-ranking-600px.png b/assets/decision-expanded-ranking-600px.png
new file mode 100644
index 0000000000000000000000000000000000000000..64ae164f9b9f1314caac05b86445dd4768c8c286
Binary files /dev/null and b/assets/decision-expanded-ranking-600px.png differ
diff --git a/assets/decision-expanded-ranking.pdf b/assets/decision-expanded-ranking.pdf
new file mode 100644
index 0000000000000000000000000000000000000000..c4e3ac0d899d841a520a1ad3ada741933f22a4c0
Binary files /dev/null and b/assets/decision-expanded-ranking.pdf differ
diff --git a/assets/decision-expanded-ranking.png b/assets/decision-expanded-ranking.png
new file mode 100644
index 0000000000000000000000000000000000000000..9211da8f4b64b9902306b60152d9a2a740015a35
--- /dev/null
+++ b/assets/decision-expanded-ranking.png
@@ -0,0 +1,3 @@
+version https://git-lfs.github.com/spec/v1
+oid sha256:038b995974df964403278b2d121bf9b888c52387aeed4e1726ddf5cbbb18f0b3
+size 171628
diff --git a/assets/decision-quality.svg b/assets/decision-expanded-ranking.svg
similarity index 56%
rename from assets/decision-quality.svg
rename to assets/decision-expanded-ranking.svg
index 0d287a906213b6b7f0472f3663f45b0e601a90d8..89911fb6cac9a858d413c1ed111d197b21c606f9 100644
--- a/assets/decision-quality.svg
+++ b/assets/decision-expanded-ranking.svg
@@ -1,16 +1,16 @@
-