Xunzhuo commited on
Commit
3e46eb1
·
verified ·
1 Parent(s): e0da3b2

Release Decision-1.0-Lux

Browse files
This view is limited to 50 files because it contains too many changes.   See raw diff
Files changed (50) hide show
  1. .gitattributes +6 -0
  2. ATTRIBUTIONS.md +19 -0
  3. DIAGNOSTICS.md +108 -0
  4. Dockerfile.runtime +7 -0
  5. EVALUATION.md +42 -0
  6. LICENSE +202 -0
  7. METHODS.md +13 -0
  8. QUESTION-SCALING.md +33 -0
  9. QWEN-LICENSE +202 -0
  10. README.md +105 -0
  11. RUNTIME.md +23 -0
  12. SENSITIVITY.md +18 -0
  13. TASKS.md +97 -0
  14. USAGE.md +29 -0
  15. WEIGHTING.md +18 -0
  16. assets/architecture.pdf +3 -0
  17. assets/architecture.png +3 -0
  18. assets/architecture.svg +157 -0
  19. assets/decision-matrix.pdf +0 -0
  20. assets/decision-matrix.png +3 -0
  21. assets/decision-matrix.svg +1253 -0
  22. assets/decision-question-scaling.pdf +0 -0
  23. assets/decision-question-scaling.png +0 -0
  24. assets/decision-question-scaling.svg +193 -0
  25. assets/decision-ranking.pdf +0 -0
  26. assets/decision-ranking.png +3 -0
  27. assets/decision-ranking.svg +352 -0
  28. assets/readout.png +3 -0
  29. assets/readout.svg +92 -0
  30. backbone/config.json +83 -0
  31. backbone/model-00001-of-00004.safetensors +3 -0
  32. backbone/model-00002-of-00004.safetensors +3 -0
  33. backbone/model-00003-of-00004.safetensors +3 -0
  34. backbone/model-00004-of-00004.safetensors +3 -0
  35. backbone/model.safetensors.index.json +434 -0
  36. bundle-manifest.json +161 -0
  37. chat_template.jinja +154 -0
  38. code/decision_api.py +198 -0
  39. code/decision_model.py +173 -0
  40. code/profile_guard.py +82 -0
  41. code/runtime_profile.py +36 -0
  42. decision_config.json +17 -0
  43. decision_head.safetensors +3 -0
  44. metrics/benchmark.json +0 -0
  45. metrics/evaluation-provenance.json +1019 -0
  46. metrics/question-scaling.json +342 -0
  47. model-card-example.json +134 -0
  48. pyproject.toml +19 -0
  49. release-manifest.json +48 -0
  50. runtime-fla-requirements.lock +4 -0
.gitattributes CHANGED
@@ -33,3 +33,9 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
 
 
 
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ assets/architecture.pdf filter=lfs diff=lfs merge=lfs -text
37
+ assets/architecture.png filter=lfs diff=lfs merge=lfs -text
38
+ assets/decision-matrix.png filter=lfs diff=lfs merge=lfs -text
39
+ assets/decision-ranking.png filter=lfs diff=lfs merge=lfs -text
40
+ assets/readout.png filter=lfs diff=lfs merge=lfs -text
41
+ tokenizer.json filter=lfs diff=lfs merge=lfs -text
ATTRIBUTIONS.md ADDED
@@ -0,0 +1,19 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Attribution
2
+
3
+ Lux adapts the text backbone and tokenizer of [Qwen3.5-9B at the pinned source revision](https://huggingface.co/Qwen/Qwen3.5-9B/tree/c202236235762e1c871ad0ccb60c8ee5ba337b9a), developed by the Qwen team at Alibaba Cloud and released under Apache License 2.0. The original [Qwen license](QWEN-LICENSE) is retained without modification. Lux excludes the vision tower and vocabulary-generation head, adds a shared candidate head, and adapts the text model for typed decisions.
4
+
5
+ The retained decision mixture includes programmatically verified tasks, human-labeled evidence judgments from MultiNLI and replay from BANKING77 and CLINC150. It is a subset of the established training pool; Lux does not inherit the trained Sol or Nox weights. Official Jev outputs are not used as training labels.
6
+
7
+ - **BANKING77**, by Iñigo Casanueva and colleagues, from the [PolyAI source repository](https://github.com/PolyAI-LDN/task-specific-datasets/tree/57ec275d8078af65b7731c2a98be812d844a6d6b/banking_data), CC BY 4.0. Original intent annotations are converted to runtime-defined choices.
8
+ - **CLINC150**, by Stefan Larson and colleagues, from the [CLINC source repository](https://github.com/clinc/oos-eval/tree/828f8093932c8fe6ca7936c3d2e52903b1c523de), CC BY 3.0. Source domain and intent labels are retained in the decision adaptation.
9
+ - **MultiNLI**, by Adina Williams, Nikita Nangia and Samuel R. Bowman, [NAACL 2018](https://aclanthology.org/N18-1101/). The retained subset uses the non-fiction government, slate, telephone and travel genres; fiction is excluded. The [pinned dataset card](https://huggingface.co/datasets/nyu-mll/multi_nli/blob/da70db2af9d09693783c3320c4249840212ee221/README.md) describes the source-specific terms, including the Open American National Corpus's permissive terms.
10
+
11
+ ## Natural-language decision adaptation
12
+
13
+ Lux uses 8,000 human-annotated training examples alongside 16,000 retained decision examples: 4,000 Cosmos QA reading questions, 2,000 SQuAD 2.0 answerability judgments, and 2,000 SNLI inference pairs. The released checkpoint completes one pass of this 24,000-example mixture. Human source labels are preserved; SQuAD answerability uses its supplied impossible/answerable annotation, and each source is converted to the model's decision interface. Source-parent groups and near duplicates are separated between custom training, selection and calibration. No official Jev output supplies a training label.
14
+
15
+ - **Cosmos QA**, by Lifu Huang, Ronan Le Bras, Chandra Bhagavatula and Yejin Choi. Data from the [author repository at the pinned revision](https://github.com/wilburOne/cosmosqa/tree/b6eb99cca4e2a51dd28a9a6f562534872d851639). The [official AllenAI dataset card](https://huggingface.co/datasets/allenai/cosmos_qa/blob/28d9d5e2aae025e73e11177891a88dba51190013/README.md) records the author-confirmed [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/) license.
16
+ - **SQuAD 2.0**, by Pranav Rajpurkar, Robin Jia and Percy Liang. The official [SQuAD project](https://rajpurkar.github.io/SQuAD-explorer/) distributes the dataset under [CC BY-SA 4.0](https://creativecommons.org/licenses/by-sa/4.0/). The adaptation here is answerability classification, not the official span-extraction benchmark.
17
+ - **SNLI 1.0**, by Samuel R. Bowman, Gabor Angeli, Christopher Potts and Christopher D. Manning. The official [Stanford Natural Language Inference project](https://nlp.stanford.edu/projects/snli/) and release README identify [CC BY-SA 4.0](https://creativecommons.org/licenses/by-sa/4.0/) for the corpus. Original entailment, neutral and contradiction labels supply the three decision alternatives.
18
+
19
+ These dataset licenses govern their respective source material; they are not replaced by the model package's Apache 2.0 license. Source corpus text, transformed training records and individual evaluation predictions are not redistributed in this model package. New evaluation panels, including supplied-fact QASC questions, are described separately in EVALUATION.md; evaluation data are not used for checkpoint selection or temperature fitting. As with other public datasets, exclusion from this custom training does not establish absence from upstream pretraining.
DIAGNOSTICS.md ADDED
@@ -0,0 +1,108 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Diagnostics
2
+
3
+ These axes remain separate from headline accuracy. Probability metrics use the shipped distribution; no temperature is fitted on this benchmark. An unavailable full-population metric is shown as —, without replacing it with a successful-only average.
4
+
5
+ ## Probability quality on transfer
6
+
7
+ | Model | Valid / requested | Brier ↓ | NLL ↓ | ECE % ↓ | Coverage at ≤5% error % ↑ | AURC ↓ |
8
+ |---|---:|---:|---:|---:|---:|---:|
9
+ | Jev | 1046/1046 | 0.1912 | 0.5743 | 3.35 | 76.96 | 0.0346 |
10
+ | Lux | 1046/1046 | 0.3058 | 0.6201 | 5.13 | 57.36 | 0.0639 |
11
+ | Nox | 1046/1046 | 0.4350 | 0.9375 | 10.27 | 35.18 | 0.1195 |
12
+ | Kev-9B | 1046/1046 | 0.2948 | 0.6028 | 4.45 | 58.03 | 0.0586 |
13
+ | Kev-4B | 1046/1046 | 0.3145 | 0.6434 | 2.61 | 56.79 | 0.0659 |
14
+ | Qwen3.5-9B | 1046/1046 | 0.3673 | 0.7597 | 9.75 | 47.42 | 0.0872 |
15
+ | Decider | 1046/1046 | 0.4122 | 0.7969 | 8.09 | 35.18 | 0.1172 |
16
+ | Qwen3.5-4B | 1046/1046 | 0.4199 | 0.8186 | 8.98 | 34.70 | 0.1233 |
17
+ | Sol | 1046/1046 | 0.5454 | 1.1094 | 11.25 | 16.16 | 0.2073 |
18
+ | Kev-0.8B | 1046/1046 | 0.4839 | 0.9489 | 2.58 | 21.61 | 0.1782 |
19
+ | Qwen3.5-2B | 1046/1046 | 0.5708 | 1.2327 | 12.30 | 11.76 | 0.2699 |
20
+ | Laya · English | 1046/1046 | 0.5832 | 1.2275 | 12.90 | 0.48 | 0.2780 |
21
+ | Laya · Multilingual | 1046/1046 | 0.6873 | 1.4653 | 21.36 | 0.00 | 0.3460 |
22
+ | Kai | 1046/1046 | 0.6066 | 1.1552 | 7.95 | 1.24 | 0.3343 |
23
+
24
+ Coverage at an error threshold keeps whole confidence-tie groups together. These are empirical observed-sample results, not a deployment error guarantee.
25
+
26
+ ## Option-order robustness
27
+
28
+ | Model | Valid / requested pairs | Both correct % ↑ | Semantic flip % ↓ | Mean half-L1 ↓ |
29
+ |---|---:|---:|---:|---:|
30
+ | Jev | 36/36 | 86.11 | 0.00 | 0.0208 |
31
+ | Lux | 36/36 | 77.78 | 8.33 | 0.0923 |
32
+ | Nox | 36/36 | 58.33 | 25.00 | 0.1049 |
33
+ | Kev-9B | 36/36 | 80.56 | 2.78 | 0.0615 |
34
+ | Kev-4B | 36/36 | 77.78 | 5.56 | 0.0667 |
35
+ | Qwen3.5-9B | 36/36 | 75.00 | 11.11 | 0.1157 |
36
+ | Decider | 36/36 | 83.33 | 11.11 | 0.0701 |
37
+ | Qwen3.5-4B | 36/36 | 77.78 | 13.89 | 0.1390 |
38
+ | Sol | 36/36 | 41.67 | 38.89 | 0.0694 |
39
+ | Kev-0.8B | 36/36 | 55.56 | 16.67 | 0.0808 |
40
+ | Qwen3.5-2B | 36/36 | 55.56 | 30.56 | 0.2163 |
41
+ | Laya · English | 36/36 | 44.44 | 19.44 | 0.0928 |
42
+ | Laya · Multilingual | 36/36 | 50.00 | 22.22 | 0.1491 |
43
+ | Kai | 36/36 | 41.67 | 27.78 | 0.1170 |
44
+
45
+ The 36 paired permutations test the same semantics under changed option order. Stable answers can still be wrong; both-correct rate therefore accompanies flip rate. Refused pairs remain in the both-correct denominator.
46
+
47
+ ## Missing evidence
48
+
49
+ | Model | Valid / requested | Intact/control accuracy % ↑ | Mean max P % ↓ | P≥0.9 share % ↓ | Normalized entropy ↑ | Paired confidence drop pp ↑ |
50
+ |---|---:|---:|---:|---:|---:|---:|
51
+ | Jev | 110/110 | 93.64 | 62.04 | 16.36 | 0.7657 | 30.30 |
52
+ | Lux | 110/110 | 84.55 | 68.81 | 19.09 | 0.6785 | 19.59 |
53
+ | Nox | 110/110 | 72.73 | 78.65 | 27.27 | 0.5072 | 12.64 |
54
+ | Kev-9B | 110/110 | 91.82 | 39.61 | 0.00 | 0.9981 | 53.19 |
55
+ | Kev-4B | 110/110 | 91.82 | 41.47 | 0.00 | 0.9918 | 51.57 |
56
+ | Qwen3.5-9B | 110/110 | 80.91 | 67.62 | 7.27 | 0.7585 | 19.09 |
57
+ | Decider | 110/110 | 69.09 | 68.82 | 16.36 | 0.6772 | 13.59 |
58
+ | Qwen3.5-4B | 110/110 | 72.73 | 59.89 | 0.91 | 0.8120 | 22.35 |
59
+ | Sol | 110/110 | 65.45 | 77.19 | 24.55 | 0.5424 | 5.68 |
60
+ | Kev-0.8B | 110/110 | 87.27 | 42.11 | 0.00 | 0.9796 | 40.34 |
61
+ | Qwen3.5-2B | 110/110 | 51.82 | 65.19 | 9.09 | 0.7370 | 4.68 |
62
+ | Laya · English | 110/110 | 49.09 | 67.96 | 0.00 | 0.7497 | -5.58 |
63
+ | Laya · Multilingual | 110/110 | 37.27 | 75.70 | 29.09 | 0.5307 | -0.49 |
64
+ | Kai | 110/110 | 37.27 | 55.39 | 18.18 | 0.8262 | -0.07 |
65
+
66
+ The 110 unknowable examples have no scored true class and are excluded from accuracy. Confidence is compared with matched evidence-bearing controls. Lower unsupported confidence and a positive evidence-removal confidence drop are desirable; these are not correctness scores.
67
+
68
+ ## Native contract coverage
69
+
70
+ | Model | Original probability rows | Transfer probability rows | Transfer truncated questions |
71
+ |---|---:|---:|---:|
72
+ | Jev | 2720/2720 | 1264/1264 | — / not observable |
73
+ | Lux | 2720/2720 | 1264/1264 | 0 |
74
+ | Nox | 2720/2720 | 1264/1264 | 0 |
75
+ | Kev-9B | 2720/2720 | 1264/1264 | 0 |
76
+ | Kev-4B | 2720/2720 | 1264/1264 | 0 |
77
+ | Qwen3.5-9B | 2720/2720 | 1264/1264 | 0 |
78
+ | Decider | 2720/2720 | 1264/1264 | 0 |
79
+ | Qwen3.5-4B | 2720/2720 | 1264/1264 | 0 |
80
+ | Sol | 2720/2720 | 1264/1264 | 0 |
81
+ | Kev-0.8B | 2720/2720 | 1264/1264 | 0 |
82
+ | Qwen3.5-2B | 2720/2720 | 1264/1264 | 0 |
83
+ | Laya · English | 2720/2720 | 1264/1264 | 34 |
84
+ | Laya · Multilingual | 2720/2720 | 1264/1264 | 14 |
85
+ | Kai | 2720/2720 | 1264/1264 | 0 |
86
+
87
+ Transfer coverage includes all 1,264 questions: clean, unknown-evidence and order variants. Accuracy uses 1,046 clean knowable questions. Native refusals count as incorrect when accuracy applies, and absent probabilities are never invented. Laya retains upstream truncation; Kai retains the shipped complete-request 1,024-token limit. Server-side truncation for Jev cannot be observed.
88
+
89
+ ## Uncertainty
90
+
91
+ | Model | Overall % | 95% component-bootstrap interval |
92
+ |---|---:|---:|
93
+ | Jev | 81.05 | 79.70–82.35 |
94
+ | Lux | 76.72 | 75.35–78.07 |
95
+ | Nox | 72.84 | 71.33–74.31 |
96
+ | Kev-9B | 71.89 | 70.42–73.35 |
97
+ | Kev-4B | 70.09 | 68.45–71.63 |
98
+ | Qwen3.5-9B | 69.73 | 68.27–71.20 |
99
+ | Decider | 67.71 | 66.10–69.34 |
100
+ | Qwen3.5-4B | 67.29 | 65.89–68.69 |
101
+ | Sol | 66.32 | 64.77–67.85 |
102
+ | Kev-0.8B | 58.28 | 56.64–59.89 |
103
+ | Qwen3.5-2B | 57.24 | 55.74–58.76 |
104
+ | Laya · English | 51.03 | 49.43–52.68 |
105
+ | Laya · Multilingual | 47.19 | 45.58–48.82 |
106
+ | Kai | 46.49 | 45.02–47.99 |
107
+
108
+ Intervals use 10,000 paired whole-source-component bootstrap draws. Related source, variant, parent, control and exact-state records stay together; all models receive the same draws. Five panels are sampled independently. These descriptive intervals do not create a positive-confidence-interval release requirement.
Dockerfile.runtime ADDED
@@ -0,0 +1,7 @@
 
 
 
 
 
 
 
 
1
+ FROM vllm/vllm-openai-rocm@sha256:1fd21abe66455b4df5a2e83629e97cdcc9d58913b16052d8118b92b239792339
2
+ COPY runtime-fla-requirements.lock /tmp/runtime-fla-requirements.lock
3
+ RUN python3 -m pip install --no-cache-dir --no-index --no-deps --require-hashes --target /opt/decision-fla -r /tmp/runtime-fla-requirements.lock
4
+ ENV PYTHONPATH=/opt/decision-fla
5
+ WORKDIR /work
6
+ ENTRYPOINT []
7
+ CMD ["/bin/bash"]
EVALUATION.md ADDED
@@ -0,0 +1,42 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Evaluation
2
+
3
+ Lux achieves **76.72%** on the decision-focused benchmark. The headline combines 880 general decisions, 880 compositional tasks, 480 natural-reading questions, 480 reading/inference questions and 1,046 answerable transfer decisions. The exact weights are 30%, 25%, 15%, 15% and 15%, preserving each original panel's family/source weights. Refusals and failures remain in the requested accuracy denominator.
4
+
5
+ **These product weights were chosen after earlier results were observed.** Reweighting changes the product score, not model predictions or trained capability. [Weight sensitivity](WEIGHTING.md) retains both the earlier five-panel weighting and the original four-panel comparison for these same current model predictions.
6
+
7
+ | Model | Size | Decisions | Composition | Reading | Inference | Transfer | Overall |
8
+ |---|---:|---:|---:|---:|---:|---:|---:|
9
+ | Jev | — | 79.10 | 66.38 | 94.53 | 89.79 | 87.19 | 81.05 |
10
+ | Lux | 9B | **83.21** | **51.88** | 90.31 | **90.83** | 77.44 | **76.72** |
11
+ | Nox | 4B | **83.00** | **51.79** | 79.06 | **86.25** | 67.97 | **72.84** |
12
+ | Kev-9B | 9B | 76.75 | 45.75 | 86.72 | 83.54 | 79.25 | 71.89 |
13
+ | Kev-4B | 4B | 71.90 | 48.54 | 81.88 | 84.58 | 76.10 | 70.09 |
14
+ | Qwen3.5-9B | 9B | 73.91 | 44.62 | 89.84 | 79.58 | 73.23 | 69.73 |
15
+ | Decider | 2B | 64.01 | 46.58 | 92.03 | 84.38 | 69.31 | 67.71 |
16
+ | Qwen3.5-4B | 4B | 69.89 | 43.33 | 87.97 | 79.79 | 68.83 | 67.29 |
17
+ | Sol | 2B | 73.75 | 46.08 | 76.56 | 84.17 | 57.07 | 66.32 |
18
+ | Kev-0.8B | 0.8B | 60.14 | 42.29 | 67.81 | 68.75 | 61.19 | 58.28 |
19
+ | Qwen3.5-2B | 2B | 57.12 | 39.00 | 73.75 | 72.29 | 56.31 | 57.24 |
20
+ | Laya · English | 0.421B | 56.54 | 35.33 | 51.41 | 63.75 | 53.06 | 51.03 |
21
+ | Laya · Multilingual | 0.322B | 47.25 | 38.92 | 50.78 | 57.29 | 47.13 | 47.19 |
22
+ | Kai | 0.572B | 42.24 | 39.88 | 55.94 | 54.58 | 48.47 | 46.49 |
23
+
24
+ Models are ordered by unrounded overall accuracy. Bold identifies a Decision-family result strictly higher than every external open reference in that column, excluding other Decision-family models and the closed Jev frontier. Ties are not bold. Sizes identify the source-model tier; Lux's deployed text-plus-head parameter count is 7.941B. Kai is a 0.572B encoder. Laya English and multilingual are evaluated separately.
25
+
26
+ ## What the benchmark measures
27
+
28
+ The original panels cover 27 tasks spanning routing, supplied-evidence decisions, state and rule composition, natural reading and inference. The transfer projection adds source generalization, controlled evidence transformations and answerable counterfactual controls from the audited Kev test suites. Unknown-information probes, option permutation and semantic consistency are separate diagnostics; they are not assigned fabricated correctness labels or silently mixed into the headline.
29
+
30
+ [All 54 task results](TASKS.md) and [robustness diagnostics](DIAGNOSTICS.md) preserve the individual capabilities instead of hiding them behind the five aggregate columns. [Machine-readable benchmark and uncertainty](metrics/benchmark.json) includes the exact component scores, paired comparisons and comparator identities. Probability scores require complete valid distributions; unsupported rows do not receive invented probabilities.
31
+
32
+ Confidence intervals use 10,000 paired bootstrap draws over source components, preserving the frozen within-panel weighting. They describe this finite benchmark and do not establish universal superiority. These suites are observed regression evidence; exclusion from custom training does not prove absence from upstream pretraining.
33
+
34
+ ## Compared inference paths
35
+
36
+ All models receive the same frozen request population through their recorded adapters. Native support limits are retained, including Kai's complete-input admission limit and reference-model option limits. Untuned Qwen models use a fixed pretrained vocabulary-head letter readout at temperature one, without decision-head training or generated reasoning. Lux uses its exported BF16 text backbone, FP32 candidate head and independent-CAL temperature. Jev is the recorded hosted service snapshot, not a locally inspectable architecture.
37
+
38
+ Lux was selected and calibrated before this benchmark was inspected. No test result reselected another checkpoint. The portable public loader reproduced all 1,600 calibration raw-logit vectors exactly in an isolated container without original training-weight access. Exact source, model and runtime identities are retained in the bundle metadata.
39
+
40
+ ## Measured latency
41
+
42
+ [Question-count scaling](QUESTION-SCALING.md) measures Lux's current native local API with fixed input length per question. Its host differs from the earlier Sol/Nox timing host, so the curves are not combined into a cross-hardware speed ranking. Loading and network transport are excluded; tokenization, model execution and response construction are included.
LICENSE ADDED
@@ -0,0 +1,202 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+
2
+ Apache License
3
+ Version 2.0, January 2004
4
+ http://www.apache.org/licenses/
5
+
6
+ TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
7
+
8
+ 1. Definitions.
9
+
10
+ "License" shall mean the terms and conditions for use, reproduction,
11
+ and distribution as defined by Sections 1 through 9 of this document.
12
+
13
+ "Licensor" shall mean the copyright owner or entity authorized by
14
+ the copyright owner that is granting the License.
15
+
16
+ "Legal Entity" shall mean the union of the acting entity and all
17
+ other entities that control, are controlled by, or are under common
18
+ control with that entity. For the purposes of this definition,
19
+ "control" means (i) the power, direct or indirect, to cause the
20
+ direction or management of such entity, whether by contract or
21
+ otherwise, or (ii) ownership of fifty percent (50%) or more of the
22
+ outstanding shares, or (iii) beneficial ownership of such entity.
23
+
24
+ "You" (or "Your") shall mean an individual or Legal Entity
25
+ exercising permissions granted by this License.
26
+
27
+ "Source" form shall mean the preferred form for making modifications,
28
+ including but not limited to software source code, documentation
29
+ source, and configuration files.
30
+
31
+ "Object" form shall mean any form resulting from mechanical
32
+ transformation or translation of a Source form, including but
33
+ not limited to compiled object code, generated documentation,
34
+ and conversions to other media types.
35
+
36
+ "Work" shall mean the work of authorship, whether in Source or
37
+ Object form, made available under the License, as indicated by a
38
+ copyright notice that is included in or attached to the work
39
+ (an example is provided in the Appendix below).
40
+
41
+ "Derivative Works" shall mean any work, whether in Source or Object
42
+ form, that is based on (or derived from) the Work and for which the
43
+ editorial revisions, annotations, elaborations, or other modifications
44
+ represent, as a whole, an original work of authorship. For the purposes
45
+ of this License, Derivative Works shall not include works that remain
46
+ separable from, or merely link (or bind by name) to the interfaces of,
47
+ the Work and Derivative Works thereof.
48
+
49
+ "Contribution" shall mean any work of authorship, including
50
+ the original version of the Work and any modifications or additions
51
+ to that Work or Derivative Works thereof, that is intentionally
52
+ submitted to Licensor for inclusion in the Work by the copyright owner
53
+ or by an individual or Legal Entity authorized to submit on behalf of
54
+ the copyright owner. For the purposes of this definition, "submitted"
55
+ means any form of electronic, verbal, or written communication sent
56
+ to the Licensor or its representatives, including but not limited to
57
+ communication on electronic mailing lists, source code control systems,
58
+ and issue tracking systems that are managed by, or on behalf of, the
59
+ Licensor for the purpose of discussing and improving the Work, but
60
+ excluding communication that is conspicuously marked or otherwise
61
+ designated in writing by the copyright owner as "Not a Contribution."
62
+
63
+ "Contributor" shall mean Licensor and any individual or Legal Entity
64
+ on behalf of whom a Contribution has been received by Licensor and
65
+ subsequently incorporated within the Work.
66
+
67
+ 2. Grant of Copyright License. Subject to the terms and conditions of
68
+ this License, each Contributor hereby grants to You a perpetual,
69
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
70
+ copyright license to reproduce, prepare Derivative Works of,
71
+ publicly display, publicly perform, sublicense, and distribute the
72
+ Work and such Derivative Works in Source or Object form.
73
+
74
+ 3. Grant of Patent License. Subject to the terms and conditions of
75
+ this License, each Contributor hereby grants to You a perpetual,
76
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
77
+ (except as stated in this section) patent license to make, have made,
78
+ use, offer to sell, sell, import, and otherwise transfer the Work,
79
+ where such license applies only to those patent claims licensable
80
+ by such Contributor that are necessarily infringed by their
81
+ Contribution(s) alone or by combination of their Contribution(s)
82
+ with the Work to which such Contribution(s) was submitted. If You
83
+ institute patent litigation against any entity (including a
84
+ cross-claim or counterclaim in a lawsuit) alleging that the Work
85
+ or a Contribution incorporated within the Work constitutes direct
86
+ or contributory patent infringement, then any patent licenses
87
+ granted to You under this License for that Work shall terminate
88
+ as of the date such litigation is filed.
89
+
90
+ 4. Redistribution. You may reproduce and distribute copies of the
91
+ Work or Derivative Works thereof in any medium, with or without
92
+ modifications, and in Source or Object form, provided that You
93
+ meet the following conditions:
94
+
95
+ (a) You must give any other recipients of the Work or
96
+ Derivative Works a copy of this License; and
97
+
98
+ (b) You must cause any modified files to carry prominent notices
99
+ stating that You changed the files; and
100
+
101
+ (c) You must retain, in the Source form of any Derivative Works
102
+ that You distribute, all copyright, patent, trademark, and
103
+ attribution notices from the Source form of the Work,
104
+ excluding those notices that do not pertain to any part of
105
+ the Derivative Works; and
106
+
107
+ (d) If the Work includes a "NOTICE" text file as part of its
108
+ distribution, then any Derivative Works that You distribute must
109
+ include a readable copy of the attribution notices contained
110
+ within such NOTICE file, excluding those notices that do not
111
+ pertain to any part of the Derivative Works, in at least one
112
+ of the following places: within a NOTICE text file distributed
113
+ as part of the Derivative Works; within the Source form or
114
+ documentation, if provided along with the Derivative Works; or,
115
+ within a display generated by the Derivative Works, if and
116
+ wherever such third-party notices normally appear. The contents
117
+ of the NOTICE file are for informational purposes only and
118
+ do not modify the License. You may add Your own attribution
119
+ notices within Derivative Works that You distribute, alongside
120
+ or as an addendum to the NOTICE text from the Work, provided
121
+ that such additional attribution notices cannot be construed
122
+ as modifying the License.
123
+
124
+ You may add Your own copyright statement to Your modifications and
125
+ may provide additional or different license terms and conditions
126
+ for use, reproduction, or distribution of Your modifications, or
127
+ for any such Derivative Works as a whole, provided Your use,
128
+ reproduction, and distribution of the Work otherwise complies with
129
+ the conditions stated in this License.
130
+
131
+ 5. Submission of Contributions. Unless You explicitly state otherwise,
132
+ any Contribution intentionally submitted for inclusion in the Work
133
+ by You to the Licensor shall be under the terms and conditions of
134
+ this License, without any additional terms or conditions.
135
+ Notwithstanding the above, nothing herein shall supersede or modify
136
+ the terms of any separate license agreement you may have executed
137
+ with Licensor regarding such Contributions.
138
+
139
+ 6. Trademarks. This License does not grant permission to use the trade
140
+ names, trademarks, service marks, or product names of the Licensor,
141
+ except as required for reasonable and customary use in describing the
142
+ origin of the Work and reproducing the content of the NOTICE file.
143
+
144
+ 7. Disclaimer of Warranty. Unless required by applicable law or
145
+ agreed to in writing, Licensor provides the Work (and each
146
+ Contributor provides its Contributions) on an "AS IS" BASIS,
147
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
148
+ implied, including, without limitation, any warranties or conditions
149
+ of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
150
+ PARTICULAR PURPOSE. You are solely responsible for determining the
151
+ appropriateness of using or redistributing the Work and assume any
152
+ risks associated with Your exercise of permissions under this License.
153
+
154
+ 8. Limitation of Liability. In no event and under no legal theory,
155
+ whether in tort (including negligence), contract, or otherwise,
156
+ unless required by applicable law (such as deliberate and grossly
157
+ negligent acts) or agreed to in writing, shall any Contributor be
158
+ liable to You for damages, including any direct, indirect, special,
159
+ incidental, or consequential damages of any character arising as a
160
+ result of this License or out of the use or inability to use the
161
+ Work (including but not limited to damages for loss of goodwill,
162
+ work stoppage, computer failure or malfunction, or any and all
163
+ other commercial damages or losses), even if such Contributor
164
+ has been advised of the possibility of such damages.
165
+
166
+ 9. Accepting Warranty or Additional Liability. While redistributing
167
+ the Work or Derivative Works thereof, You may choose to offer,
168
+ and charge a fee for, acceptance of support, warranty, indemnity,
169
+ or other liability obligations and/or rights consistent with this
170
+ License. However, in accepting such obligations, You may act only
171
+ on Your own behalf and on Your sole responsibility, not on behalf
172
+ of any other Contributor, and only if You agree to indemnify,
173
+ defend, and hold each Contributor harmless for any liability
174
+ incurred by, or claims asserted against, such Contributor by reason
175
+ of your accepting any such warranty or additional liability.
176
+
177
+ END OF TERMS AND CONDITIONS
178
+
179
+ APPENDIX: How to apply the Apache License to your work.
180
+
181
+ To apply the Apache License to your work, attach the following
182
+ boilerplate notice, with the fields enclosed by brackets "[]"
183
+ replaced with your own identifying information. (Don't include
184
+ the brackets!) The text should be enclosed in the appropriate
185
+ comment syntax for the file format. We also recommend that a
186
+ file or class name and description of purpose be included on the
187
+ same "printed page" as the copyright notice for easier
188
+ identification within third-party archives.
189
+
190
+ Copyright 2026 Alibaba Cloud
191
+
192
+ Licensed under the Apache License, Version 2.0 (the "License");
193
+ you may not use this file except in compliance with the License.
194
+ You may obtain a copy of the License at
195
+
196
+ http://www.apache.org/licenses/LICENSE-2.0
197
+
198
+ Unless required by applicable law or agreed to in writing, software
199
+ distributed under the License is distributed on an "AS IS" BASIS,
200
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
201
+ See the License for the specific language governing permissions and
202
+ limitations under the License.
METHODS.md ADDED
@@ -0,0 +1,13 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Model details
2
+
3
+ Lux adapts the text backbone and tokenizer of [Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B/tree/c202236235762e1c871ad0ccb60c8ee5ba337b9a). The deployed decision model has **7,940,895,744 parameters**: 7,936,684,544 in the text backbone and 4,211,200 in the shared candidate head. The upstream vision tower, language-generation head and auxiliary prediction components are excluded.
4
+
5
+ The backbone contains 32 causal layers: 24 Gated DeltaNet layers and eight full-attention layers. Hidden width is 4,096 and feed-forward width is 12,288. Full attention uses 16 query heads and four key/value heads; Gated DeltaNet uses 16 key heads and 32 value heads with dimension 128.
6
+
7
+ For each question, the model encodes complete evidence, instructions and candidate descriptions. The decision head combines each contextual candidate-endpoint vector with the final global-query vector through a shared 256-dimensional bilinear/MLP readout. It scores all supplied candidates in one forward pass per question, with a BF16 backbone and FP32 head. Dynamic labels come from the current request.
8
+
9
+ The model was adapted on a fixed 24,000-example mixture of retained decision tasks and human-annotated natural-language judgments. A 100-update head warmup preceded full-parameter cross-entropy training. Two fixed learning rates were compared using separate selection data; the selected model completed 375 full-parameter updates. A single positive temperature was fitted afterward on 1,600 independent calibration examples. Evaluation labels were not used to select the checkpoint or fit its temperature. [Data attribution](ATTRIBUTIONS.md).
10
+
11
+ The exported bundle was reloaded through the default public API in a container that could not access the original training checkpoint, base model or cache. All 1,600 calibration raw-logit vectors matched the production selection/calibration runtime exactly. Mixed Choice/Noul/Score requests and rejection of complete-input overflow were also verified.
12
+
13
+ The decision benchmark reports observed regression evidence across multiple task families. Exclusion from this custom training does not establish absence from upstream pretraining. Accuracy, semantic consistency and probability calibration answer different questions and are reported separately in [evaluation methods](EVALUATION.md).
QUESTION-SCALING.md ADDED
@@ -0,0 +1,33 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Question-count scaling
2
+
3
+ Current Lux public API, one otherwise idle AMD gfx942 GPU (SKU M3250101, PCI device 0x74b9, 274,542,362,624 bytes VRAM). These measurements are specific to this host; no cross-hardware speed ranking is claimed.
4
+
5
+ Each point contains 30 measured requests across six independently loaded processes, with three warmups and five timed calls per cell per process. Input state and each question’s length remain fixed while the number of questions grows. The primary condition uses distinct instructions; repeated instructions are reported separately. Default batching is eight independent questions.
6
+
7
+ The timer starts after pre-request GPU synchronization and includes request validation, rendering, tokenization, model execution, response construction and final GPU synchronization. It excludes model loading, warmup, network transport and result serialization. Private process caches warm all shapes first; autotune configuration hashes remain unchanged during timing.
8
+
9
+ ## Distinct questions · 499 tokens per question
10
+
11
+ | Questions | Median (ms) | p95 (ms) | Peak allocated (GiB) |
12
+ |---:|---:|---:|---:|
13
+ | 1 | 33.17 | 33.52 | 14.95 |
14
+ | 2 | 49.33 | 49.72 | 15.02 |
15
+ | 4 | 82.01 | 82.69 | 15.17 |
16
+ | 8 | 150.41 | 150.92 | 15.46 |
17
+ | 16 | 301.06 | 302.22 | 15.46 |
18
+ | 32 | 600.74 | 605.47 | 15.46 |
19
+
20
+ ## Repeated questions · 309 tokens per question
21
+
22
+ | Questions | Median (ms) | p95 (ms) | Peak allocated (GiB) |
23
+ |---:|---:|---:|---:|
24
+ | 1 | 32.27 | 32.78 | 14.92 |
25
+ | 2 | 36.25 | 36.76 | 14.96 |
26
+ | 4 | 56.30 | 56.55 | 15.06 |
27
+ | 8 | 98.17 | 98.68 | 15.24 |
28
+ | 16 | 195.26 | 196.56 | 15.24 |
29
+ | 32 | 390.27 | 393.29 | 15.24 |
30
+
31
+ Every timed output preserved exact response values and key order; each workload also produced one identical output hash across all six processes. Whole-host GPU process checks observed no unrelated GPU work. Reported peaks are PyTorch allocated memory, not total physical reservation.
32
+
33
+ This is warm local-request latency for two fixed short-input workloads, not concurrent HTTP throughput or long-context performance. [Raw summary and process-block uncertainty](metrics/question-scaling.json).
QWEN-LICENSE ADDED
@@ -0,0 +1,202 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+
2
+ Apache License
3
+ Version 2.0, January 2004
4
+ http://www.apache.org/licenses/
5
+
6
+ TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
7
+
8
+ 1. Definitions.
9
+
10
+ "License" shall mean the terms and conditions for use, reproduction,
11
+ and distribution as defined by Sections 1 through 9 of this document.
12
+
13
+ "Licensor" shall mean the copyright owner or entity authorized by
14
+ the copyright owner that is granting the License.
15
+
16
+ "Legal Entity" shall mean the union of the acting entity and all
17
+ other entities that control, are controlled by, or are under common
18
+ control with that entity. For the purposes of this definition,
19
+ "control" means (i) the power, direct or indirect, to cause the
20
+ direction or management of such entity, whether by contract or
21
+ otherwise, or (ii) ownership of fifty percent (50%) or more of the
22
+ outstanding shares, or (iii) beneficial ownership of such entity.
23
+
24
+ "You" (or "Your") shall mean an individual or Legal Entity
25
+ exercising permissions granted by this License.
26
+
27
+ "Source" form shall mean the preferred form for making modifications,
28
+ including but not limited to software source code, documentation
29
+ source, and configuration files.
30
+
31
+ "Object" form shall mean any form resulting from mechanical
32
+ transformation or translation of a Source form, including but
33
+ not limited to compiled object code, generated documentation,
34
+ and conversions to other media types.
35
+
36
+ "Work" shall mean the work of authorship, whether in Source or
37
+ Object form, made available under the License, as indicated by a
38
+ copyright notice that is included in or attached to the work
39
+ (an example is provided in the Appendix below).
40
+
41
+ "Derivative Works" shall mean any work, whether in Source or Object
42
+ form, that is based on (or derived from) the Work and for which the
43
+ editorial revisions, annotations, elaborations, or other modifications
44
+ represent, as a whole, an original work of authorship. For the purposes
45
+ of this License, Derivative Works shall not include works that remain
46
+ separable from, or merely link (or bind by name) to the interfaces of,
47
+ the Work and Derivative Works thereof.
48
+
49
+ "Contribution" shall mean any work of authorship, including
50
+ the original version of the Work and any modifications or additions
51
+ to that Work or Derivative Works thereof, that is intentionally
52
+ submitted to Licensor for inclusion in the Work by the copyright owner
53
+ or by an individual or Legal Entity authorized to submit on behalf of
54
+ the copyright owner. For the purposes of this definition, "submitted"
55
+ means any form of electronic, verbal, or written communication sent
56
+ to the Licensor or its representatives, including but not limited to
57
+ communication on electronic mailing lists, source code control systems,
58
+ and issue tracking systems that are managed by, or on behalf of, the
59
+ Licensor for the purpose of discussing and improving the Work, but
60
+ excluding communication that is conspicuously marked or otherwise
61
+ designated in writing by the copyright owner as "Not a Contribution."
62
+
63
+ "Contributor" shall mean Licensor and any individual or Legal Entity
64
+ on behalf of whom a Contribution has been received by Licensor and
65
+ subsequently incorporated within the Work.
66
+
67
+ 2. Grant of Copyright License. Subject to the terms and conditions of
68
+ this License, each Contributor hereby grants to You a perpetual,
69
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
70
+ copyright license to reproduce, prepare Derivative Works of,
71
+ publicly display, publicly perform, sublicense, and distribute the
72
+ Work and such Derivative Works in Source or Object form.
73
+
74
+ 3. Grant of Patent License. Subject to the terms and conditions of
75
+ this License, each Contributor hereby grants to You a perpetual,
76
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
77
+ (except as stated in this section) patent license to make, have made,
78
+ use, offer to sell, sell, import, and otherwise transfer the Work,
79
+ where such license applies only to those patent claims licensable
80
+ by such Contributor that are necessarily infringed by their
81
+ Contribution(s) alone or by combination of their Contribution(s)
82
+ with the Work to which such Contribution(s) was submitted. If You
83
+ institute patent litigation against any entity (including a
84
+ cross-claim or counterclaim in a lawsuit) alleging that the Work
85
+ or a Contribution incorporated within the Work constitutes direct
86
+ or contributory patent infringement, then any patent licenses
87
+ granted to You under this License for that Work shall terminate
88
+ as of the date such litigation is filed.
89
+
90
+ 4. Redistribution. You may reproduce and distribute copies of the
91
+ Work or Derivative Works thereof in any medium, with or without
92
+ modifications, and in Source or Object form, provided that You
93
+ meet the following conditions:
94
+
95
+ (a) You must give any other recipients of the Work or
96
+ Derivative Works a copy of this License; and
97
+
98
+ (b) You must cause any modified files to carry prominent notices
99
+ stating that You changed the files; and
100
+
101
+ (c) You must retain, in the Source form of any Derivative Works
102
+ that You distribute, all copyright, patent, trademark, and
103
+ attribution notices from the Source form of the Work,
104
+ excluding those notices that do not pertain to any part of
105
+ the Derivative Works; and
106
+
107
+ (d) If the Work includes a "NOTICE" text file as part of its
108
+ distribution, then any Derivative Works that You distribute must
109
+ include a readable copy of the attribution notices contained
110
+ within such NOTICE file, excluding those notices that do not
111
+ pertain to any part of the Derivative Works, in at least one
112
+ of the following places: within a NOTICE text file distributed
113
+ as part of the Derivative Works; within the Source form or
114
+ documentation, if provided along with the Derivative Works; or,
115
+ within a display generated by the Derivative Works, if and
116
+ wherever such third-party notices normally appear. The contents
117
+ of the NOTICE file are for informational purposes only and
118
+ do not modify the License. You may add Your own attribution
119
+ notices within Derivative Works that You distribute, alongside
120
+ or as an addendum to the NOTICE text from the Work, provided
121
+ that such additional attribution notices cannot be construed
122
+ as modifying the License.
123
+
124
+ You may add Your own copyright statement to Your modifications and
125
+ may provide additional or different license terms and conditions
126
+ for use, reproduction, or distribution of Your modifications, or
127
+ for any such Derivative Works as a whole, provided Your use,
128
+ reproduction, and distribution of the Work otherwise complies with
129
+ the conditions stated in this License.
130
+
131
+ 5. Submission of Contributions. Unless You explicitly state otherwise,
132
+ any Contribution intentionally submitted for inclusion in the Work
133
+ by You to the Licensor shall be under the terms and conditions of
134
+ this License, without any additional terms or conditions.
135
+ Notwithstanding the above, nothing herein shall supersede or modify
136
+ the terms of any separate license agreement you may have executed
137
+ with Licensor regarding such Contributions.
138
+
139
+ 6. Trademarks. This License does not grant permission to use the trade
140
+ names, trademarks, service marks, or product names of the Licensor,
141
+ except as required for reasonable and customary use in describing the
142
+ origin of the Work and reproducing the content of the NOTICE file.
143
+
144
+ 7. Disclaimer of Warranty. Unless required by applicable law or
145
+ agreed to in writing, Licensor provides the Work (and each
146
+ Contributor provides its Contributions) on an "AS IS" BASIS,
147
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
148
+ implied, including, without limitation, any warranties or conditions
149
+ of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
150
+ PARTICULAR PURPOSE. You are solely responsible for determining the
151
+ appropriateness of using or redistributing the Work and assume any
152
+ risks associated with Your exercise of permissions under this License.
153
+
154
+ 8. Limitation of Liability. In no event and under no legal theory,
155
+ whether in tort (including negligence), contract, or otherwise,
156
+ unless required by applicable law (such as deliberate and grossly
157
+ negligent acts) or agreed to in writing, shall any Contributor be
158
+ liable to You for damages, including any direct, indirect, special,
159
+ incidental, or consequential damages of any character arising as a
160
+ result of this License or out of the use or inability to use the
161
+ Work (including but not limited to damages for loss of goodwill,
162
+ work stoppage, computer failure or malfunction, or any and all
163
+ other commercial damages or losses), even if such Contributor
164
+ has been advised of the possibility of such damages.
165
+
166
+ 9. Accepting Warranty or Additional Liability. While redistributing
167
+ the Work or Derivative Works thereof, You may choose to offer,
168
+ and charge a fee for, acceptance of support, warranty, indemnity,
169
+ or other liability obligations and/or rights consistent with this
170
+ License. However, in accepting such obligations, You may act only
171
+ on Your own behalf and on Your sole responsibility, not on behalf
172
+ of any other Contributor, and only if You agree to indemnify,
173
+ defend, and hold each Contributor harmless for any liability
174
+ incurred by, or claims asserted against, such Contributor by reason
175
+ of your accepting any such warranty or additional liability.
176
+
177
+ END OF TERMS AND CONDITIONS
178
+
179
+ APPENDIX: How to apply the Apache License to your work.
180
+
181
+ To apply the Apache License to your work, attach the following
182
+ boilerplate notice, with the fields enclosed by brackets "[]"
183
+ replaced with your own identifying information. (Don't include
184
+ the brackets!) The text should be enclosed in the appropriate
185
+ comment syntax for the file format. We also recommend that a
186
+ file or class name and description of purpose be included on the
187
+ same "printed page" as the copyright notice for easier
188
+ identification within third-party archives.
189
+
190
+ Copyright 2026 Alibaba Cloud
191
+
192
+ Licensed under the Apache License, Version 2.0 (the "License");
193
+ you may not use this file except in compliance with the License.
194
+ You may obtain a copy of the License at
195
+
196
+ http://www.apache.org/licenses/LICENSE-2.0
197
+
198
+ Unless required by applicable law or agreed to in writing, software
199
+ distributed under the License is distributed on an "AS IS" BASIS,
200
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
201
+ See the License for the specific language governing permissions and
202
+ limitations under the License.
README.md ADDED
@@ -0,0 +1,105 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ language:
4
+ - en
5
+ - zh
6
+ base_model: Qwen/Qwen3.5-9B
7
+ base_model_relation: finetune
8
+ tags:
9
+ - decision-model
10
+ - classification
11
+ - qwen3_5
12
+ - custom-code
13
+ - pytorch
14
+ - rocm
15
+ ---
16
+
17
+ # Decision-1.0-Lux
18
+
19
+ *Lux, Latin for light.*
20
+
21
+ **Your move.** Give Lux evidence, questions and possible answers. It returns decisions and probabilities for labels you define at runtime.
22
+
23
+ **Qwen3.5-9B foundation · 16K complete-question budget · English / Chinese evaluated · Apache 2.0**
24
+
25
+ [Decision family](https://huggingface.co/collections/llm-semantic-router/decision-10-6ab12177bd0002394d8409f9)
26
+
27
+ | Type | Use it for | Output |
28
+ |---|---|---|
29
+ | **Choice** | Route a request or choose among 2–255 actions. | Selected ID + probability distribution |
30
+ | **Noul** | Judge a condition against supplied evidence. | P(true) |
31
+ | **Score** | Apply 2–10 ordered rubric descriptions. | Expected index + probability distribution |
32
+
33
+ ## Measured capability
34
+
35
+ **76.72% weighted accuracy** across 3,766 decisions: **+4.83 points over Kev-9B** and **+3.88 over Nox** on the same benchmark.
36
+
37
+ | Model | Size | Decisions | Composition | Reading | Inference | Transfer | Overall |
38
+ |---|---:|---:|---:|---:|---:|---:|---:|
39
+ | Jev | — | 79.10 | 66.38 | 94.53 | 89.79 | 87.19 | 81.05 |
40
+ | Lux | 9B | **83.21** | **51.88** | 90.31 | **90.83** | 77.44 | **76.72** |
41
+ | Nox | 4B | **83.00** | **51.79** | 79.06 | **86.25** | 67.97 | **72.84** |
42
+ | Kev-9B | 9B | 76.75 | 45.75 | 86.72 | 83.54 | 79.25 | 71.89 |
43
+ | Kev-4B | 4B | 71.90 | 48.54 | 81.88 | 84.58 | 76.10 | 70.09 |
44
+ | Qwen3.5-9B | 9B | 73.91 | 44.62 | 89.84 | 79.58 | 73.23 | 69.73 |
45
+ | Decider | 2B | 64.01 | 46.58 | 92.03 | 84.38 | 69.31 | 67.71 |
46
+ | Qwen3.5-4B | 4B | 69.89 | 43.33 | 87.97 | 79.79 | 68.83 | 67.29 |
47
+ | Sol | 2B | 73.75 | 46.08 | 76.56 | 84.17 | 57.07 | 66.32 |
48
+ | Kev-0.8B | 0.8B | 60.14 | 42.29 | 67.81 | 68.75 | 61.19 | 58.28 |
49
+ | Qwen3.5-2B | 2B | 57.12 | 39.00 | 73.75 | 72.29 | 56.31 | 57.24 |
50
+ | Laya · English | 0.421B | 56.54 | 35.33 | 51.41 | 63.75 | 53.06 | 51.03 |
51
+ | Laya · Multilingual | 0.322B | 47.25 | 38.92 | 50.78 | 57.29 | 47.13 | 47.19 |
52
+ | Kai | 0.572B | 42.24 | 39.88 | 55.94 | 54.58 | 48.47 | 46.49 |
53
+
54
+ Accuracy (%), using the same five-panel decision benchmark. General decisions contribute 30%; composition contributes 25%; reading, inference and external transfer each contribute 15%. Bold marks a Decision model strictly above every external open reference in that column; Jev and other Decision models are excluded from this threshold. [Full tasks, uncertainty and comparator identities](EVALUATION.md).
55
+
56
+ ![Decision benchmark ranking](assets/decision-ranking.png)
57
+
58
+ ![Capability matrix](assets/decision-matrix.png)
59
+
60
+ [All 54 tasks](TASKS.md) · [Probability quality, order and missing evidence](DIAGNOSTICS.md)
61
+
62
+ ## Use Lux
63
+
64
+ Download `hf download llm-semantic-router/Decision-1.0-Lux --local-dir decision-model`, then follow [ROCm setup](RUNTIME.md). In that container, with the repository mounted at `/model`:
65
+
66
+ ```python
67
+ from decision import DecisionModel
68
+
69
+ model = DecisionModel.from_pretrained("/model", local_files_only=True)
70
+ result = model.decide(
71
+ state="Customer reports a duplicate charge and asks for a refund.",
72
+ questions={
73
+ "route": {
74
+ "type": "choice",
75
+ "instructions": "Which team should handle this request?",
76
+ "criteria": {"billing": "Payments and refunds", "technical": "Product faults"},
77
+ },
78
+ "refund_requested": {
79
+ "type": "noul",
80
+ "instructions": "Did the customer request a refund?",
81
+ },
82
+ },
83
+ )
84
+ print(result["answers"])
85
+ ```
86
+
87
+ The local API follows the SystemOne `state` + named `questions` format. Questions run independently in batches of eight; no explanation tokens are generated. [API guide](USAGE.md) · [Tested requests and actual outputs](model-card-example.json).
88
+
89
+ ## More questions, measured
90
+
91
+ ![Lux request latency](assets/decision-question-scaling.png)
92
+
93
+ Distinct Choice questions, fixed at 499 input tokens per question. Thirty measured requests per point across six fresh processes on an otherwise idle AMD GPU. Python latency includes tokenization, inference and response construction; loading and network are excluded. [p95, memory and hardware](QUESTION-SCALING.md).
94
+
95
+ ## Architecture
96
+
97
+ ![Lux decoder architecture](assets/architecture.png)
98
+
99
+ A causal Qwen3.5 text backbone combines Gated DeltaNet and full attention. A shared candidate head reads contextual candidate endpoints and the final query vector, producing one probability per supplied answer.
100
+
101
+ [Candidate head](assets/readout.png) · [Vector architecture](assets/architecture.svg) · [Model details](METHODS.md)
102
+
103
+ The full state, question and candidates must fit 16,384 tokens; overflow is rejected. AMD gfx942 is validated; other hardware requires separate qualification. Lux judges supplied evidence without live retrieval, so confidence does not guarantee factual correctness.
104
+
105
+ [License](LICENSE) · [Attributions](ATTRIBUTIONS.md) · [Runtime](RUNTIME.md)
RUNTIME.md ADDED
@@ -0,0 +1,23 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # ROCm runtime
2
+
3
+ Lux was validated on AMD gfx942 using PyTorch `2.12.0+git6bbd260`, HIP `7.2.53211`, Transformers `5.17.0`, Triton `3.7.1`, FLA `0.5.2`, Tokenizers `0.23.2` and Safetensors `0.8.0`. The bundle uses BF16 backbone weights, an FP32 candidate head, SDPA, FLA Gated DeltaNet and reference PyTorch causal convolution.
4
+
5
+ `Dockerfile.runtime` is the exact public-base recipe used to build the qualified runtime. It pins the public `vllm/vllm-openai-rocm` image by digest and installs two hash-verified FLA wheels without replacing PyTorch or Triton. Decision inference does not use the vLLM server. Weights and the local Python package are mounted at runtime.
6
+
7
+ From the downloaded model directory on an AMD Linux host with Docker and ROCm driver support:
8
+
9
+ ```bash
10
+ docker build -f Dockerfile.runtime -t decision-lux-runtime .
11
+ mkdir -p runtime-output
12
+ docker run --rm \
13
+ --device=/dev/kfd --device=/dev/dri --group-add video --ipc=host \
14
+ -e HIP_VISIBLE_DEVICES=0 \
15
+ -e PYTHONPATH=/opt/decision-fla:/model/src \
16
+ -v "$PWD":/model:ro -v "$PWD/runtime-output":/output \
17
+ decision-lux-runtime \
18
+ python -m decision.example /model --local-files-only --output /output/example.json
19
+ ```
20
+
21
+ The bundled profile is installed automatically by `DecisionModel.from_pretrained`. Keep the runtime checks enabled and use a fresh process when switching bound profiles. Runtime versions and numerical kernels can change model probabilities near a decision boundary; a different environment needs separate qualification. CPU and MPS inference are unsupported; NVIDIA has not been qualified for this release.
22
+
23
+ The default public loader was tested in an isolated container with only the exported bundle and small verification inputs mounted. All 1,600 reference calibration raw-logit vectors matched exactly. Mixed Choice/Noul/Score output and rejection of input overflow were also verified. This proves the exported package's inference path; detailed benchmark measurements and their scope are recorded separately.
SENSITIVITY.md ADDED
@@ -0,0 +1,18 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Product weights were changed after earlier results were observed. These are identical model predictions under different weights, not trained-model improvements.
2
+
3
+ | Model | Current 30/25/15/15/15 | Earlier 25/25/15/15/20 | Original four-panel mean |
4
+ |---|---:|---:|---:|
5
+ | Jev | 81.05 | 81.45 | 82.45 |
6
+ | Lux | 76.72 | 76.43 | 79.06 |
7
+ | Nox | 72.84 | 72.09 | 75.03 |
8
+ | Kev-9B | 71.89 | 72.01 | 73.19 |
9
+ | Kev-4B | 70.09 | 70.30 | 71.73 |
10
+ | Qwen3.5-9B | 69.73 | 69.70 | 71.99 |
11
+ | Decider | 67.71 | 67.97 | 71.75 |
12
+ | Qwen3.5-4B | 67.29 | 67.24 | 70.25 |
13
+ | Sol | 66.32 | 65.48 | 70.14 |
14
+ | Kev-0.8B | 58.28 | 58.33 | 59.75 |
15
+ | Qwen3.5-2B | 57.24 | 57.20 | 60.54 |
16
+ | Laya · English | 51.03 | 50.85 | 51.76 |
17
+ | Laya · Multilingual | 47.19 | 47.18 | 48.56 |
18
+ | Kai | 46.49 | 46.80 | 48.16 |
TASKS.md ADDED
@@ -0,0 +1,97 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # All 54 tasks
2
+
3
+ Accuracy (%) on the same requested rows. Bold marks a Decision-family result strictly above every external open or untuned reference; Jev and the other Decision models do not set that threshold. Ties are not bold. Each task uses its full requested denominator. These task rows are diagnostic; the headline retains its declared within-panel weights.
4
+
5
+ <details>
6
+ <summary>Decisions · 10 tasks</summary>
7
+
8
+ | Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual | Kai |
9
+ |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
10
+ | News classification | 128 | 85.16 | 88.28 | 85.16 | 87.50 | 87.50 | 85.94 | 86.72 | 84.38 | 83.59 | 86.72 | 80.47 | 91.41 | 89.84 | 78.12 |
11
+ | Boolean constraints | 64 | 100.00 | 95.31 | 93.75 | 95.31 | 71.88 | 98.44 | 71.88 | 62.50 | 50.00 | 81.25 | 37.50 | 43.75 | 39.06 | 53.12 |
12
+ | Entity classification | 112 | 96.43 | 96.43 | 96.43 | 98.21 | 98.21 | 96.43 | 98.21 | 97.32 | 95.54 | 96.43 | 92.86 | 83.93 | 57.14 | 50.00 |
13
+ | Intent routing | 64 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 96.88 | 81.25 | 81.25 | 81.25 |
14
+ | Evidence placement | 96 | 30.21 | 61.46 | **100.00** | 50.00 | 41.67 | 79.17 | 8.33 | 72.92 | **98.96** | 33.33 | 9.38 | 95.83 | 54.17 | 1.04 |
15
+ | Ordered rubric | 64 | 100.00 | 100.00 | 89.06 | 96.88 | 100.00 | 84.38 | 84.38 | 90.62 | 65.62 | 59.38 | 81.25 | 18.75 | 12.50 | 6.25 |
16
+ | Relation composition | 96 | 56.25 | 61.46 | 52.08 | 51.04 | 62.50 | 36.46 | 54.17 | 51.04 | 37.50 | 44.79 | 48.96 | 25.00 | 35.42 | 51.04 |
17
+ | Scoped evidence | 96 | 89.58 | **98.96** | **78.12** | 65.62 | 47.92 | 55.21 | 47.92 | 51.04 | **79.17** | 10.42 | 36.46 | 37.50 | 36.46 | 29.17 |
18
+ | State tracking | 96 | 33.33 | 33.33 | 35.42 | 38.54 | 34.38 | 31.25 | 29.17 | 28.12 | 27.08 | 31.25 | 25.00 | 23.96 | 29.17 | 30.21 |
19
+ | In / out of menu | 64 | 100.00 | **96.88** | **100.00** | 84.38 | 75.00 | 71.88 | 59.38 | 60.94 | **100.00** | 57.81 | 62.50 | 64.06 | 37.50 | 42.19 |
20
+
21
+ </details>
22
+
23
+ <details>
24
+ <summary>Composition · 10 tasks</summary>
25
+
26
+ | Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual | Kai |
27
+ |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
28
+ | Record identity | 80 | 76.25 | **71.25** | 53.75 | 45.00 | 50.00 | 57.50 | 56.25 | 50.00 | 50.00 | 50.00 | 46.25 | 47.50 | 50.00 | 51.25 |
29
+ | Capacity assignment | 80 | 68.75 | **71.25** | 46.25 | 50.00 | 50.00 | 50.00 | 51.25 | 50.00 | 47.50 | 50.00 | 50.00 | 63.75 | 50.00 | 50.00 |
30
+ | Constraint assignment | 80 | 55.00 | **43.75** | **41.25** | 27.50 | 32.50 | 25.00 | 23.75 | 23.75 | 26.25 | 20.00 | 21.25 | 22.50 | 27.50 | 31.25 |
31
+ | Intent routing · EN | 120 | 91.67 | 90.83 | 91.67 | 86.67 | 87.50 | 88.33 | 90.00 | 91.67 | 90.83 | 83.33 | 70.83 | 74.17 | 70.83 | 72.50 |
32
+ | Intent routing · ZH | 120 | 88.33 | 86.67 | 87.50 | 85.83 | 86.67 | 86.67 | 88.33 | 86.67 | 87.50 | 85.83 | 74.17 | 44.17 | 73.33 | 67.50 |
33
+ | Multiset reconciliation | 80 | 58.75 | 26.25 | 28.75 | 37.50 | 28.75 | 20.00 | 23.75 | 27.50 | 28.75 | 25.00 | 27.50 | 26.25 | 32.50 | 32.50 |
34
+ | Ordered service loss | 80 | 42.50 | 20.00 | **30.00** | 25.00 | 20.00 | 20.00 | 22.50 | 21.25 | 25.00 | 27.50 | 20.00 | 20.00 | 17.50 | 21.25 |
35
+ | Conflicting rule closure | 80 | 66.25 | **36.25** | 33.75 | 27.50 | 33.75 | 27.50 | 26.25 | 25.00 | 25.00 | 27.50 | 27.50 | 18.75 | 20.00 | 23.75 |
36
+ | Temporal exclusion | 80 | 41.25 | 21.25 | 42.50 | 28.75 | 45.00 | 36.25 | 42.50 | 31.25 | 41.25 | 21.25 | 28.75 | 12.50 | 28.75 | 22.50 |
37
+ | Transaction recovery | 80 | 75.00 | 51.25 | **62.50** | 43.75 | 51.25 | 35.00 | 41.25 | 26.25 | 38.75 | 32.50 | 23.75 | 23.75 | 18.75 | 26.25 |
38
+
39
+ </details>
40
+
41
+ <details>
42
+ <summary>Reading · 3 tasks</summary>
43
+
44
+ | Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual | Kai |
45
+ |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
46
+ | Yes / no reading | 160 | 92.50 | 90.62 | 86.25 | 91.25 | 88.75 | 84.38 | 91.88 | 83.75 | 83.12 | 74.38 | 65.00 | 69.38 | 69.38 | 74.38 |
47
+ | Reading · EN | 160 | 96.88 | 91.25 | 73.12 | 83.12 | 75.62 | 98.75 | 93.12 | 93.12 | 70.00 | 63.75 | 83.75 | 38.12 | 36.25 | 39.38 |
48
+ | Reading · ZH | 160 | 96.25 | 88.75 | 70.62 | 81.25 | 74.38 | 91.88 | 91.25 | 91.25 | 70.00 | 58.75 | 81.25 | 28.75 | 28.12 | 35.62 |
49
+
50
+ </details>
51
+
52
+ <details>
53
+ <summary>Inference · 4 tasks</summary>
54
+
55
+ | Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual | Kai |
56
+ |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
57
+ | Contextual reasoning | 120 | 86.67 | **83.33** | 70.83 | 68.33 | 70.83 | 64.17 | 66.67 | 60.83 | 69.17 | 42.50 | 56.67 | 30.00 | 20.00 | 32.50 |
58
+ | Answerability | 120 | 90.83 | **93.33** | **88.33** | 78.33 | 81.67 | 74.17 | 84.17 | 81.67 | 83.33 | 66.67 | 82.50 | 57.50 | 55.00 | 55.83 |
59
+ | Textual entailment | 120 | 82.50 | 90.00 | 89.17 | 90.83 | 89.17 | 81.67 | 90.83 | 80.00 | 88.33 | 77.50 | 64.17 | 72.50 | 72.50 | 34.17 |
60
+ | Scientific inference | 120 | 99.17 | 96.67 | 96.67 | 96.67 | 96.67 | 98.33 | 95.83 | 96.67 | 95.83 | 88.33 | 85.83 | 95.00 | 81.67 | 95.83 |
61
+
62
+ </details>
63
+
64
+ <details>
65
+ <summary>Transfer · 27 tasks</summary>
66
+
67
+ | Task | n | Jev | Lux | Nox | Kev-9B | Kev-4B | Qwen3.5-9B | Decider | Qwen3.5-4B | Sol | Kev-0.8B | Qwen3.5-2B | Laya · English | Laya · Multilingual | Kai |
68
+ |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
69
+ | Buried emotion | 20 | 70.00 | 40.00 | 20.00 | 65.00 | 55.00 | 40.00 | 85.00 | 45.00 | 20.00 | 40.00 | 55.00 | 60.00 | 35.00 | 55.00 |
70
+ | Buried paraphrase | 20 | 95.00 | **85.00** | 75.00 | 80.00 | 75.00 | 80.00 | 75.00 | 75.00 | 70.00 | 60.00 | 70.00 | 55.00 | 70.00 | 35.00 |
71
+ | Buried entailment | 20 | 95.00 | 90.00 | 90.00 | 95.00 | 90.00 | 95.00 | 90.00 | 75.00 | 75.00 | 85.00 | 65.00 | 60.00 | 55.00 | 55.00 |
72
+ | Buried offensive-language detection | 20 | 55.00 | 35.00 | 45.00 | 80.00 | 65.00 | 45.00 | 80.00 | 40.00 | 35.00 | 60.00 | 35.00 | 70.00 | 80.00 | 75.00 |
73
+ | Combined policy conditions | 32 | 96.88 | **90.62** | 75.00 | 68.75 | 78.12 | 53.12 | 53.12 | 59.38 | 78.12 | 40.62 | 62.50 | 62.50 | 43.75 | 34.38 |
74
+ | Policy exceptions | 32 | 100.00 | 96.88 | 96.88 | 87.50 | 96.88 | 59.38 | 56.25 | 75.00 | 71.88 | 65.62 | 40.62 | 43.75 | 37.50 | 62.50 |
75
+ | Policy negation | 32 | 90.62 | 87.50 | 87.50 | 90.62 | 87.50 | 62.50 | 62.50 | 53.12 | 71.88 | 71.88 | 46.88 | 53.12 | 50.00 | 56.25 |
76
+ | Authorization contrast | 40 | 100.00 | 100.00 | 97.50 | 100.00 | 100.00 | 100.00 | 100.00 | 97.50 | 50.00 | 97.50 | 62.50 | 67.50 | 50.00 | 50.00 |
77
+ | Deadline contrast | 40 | 92.50 | 87.50 | 70.00 | 87.50 | 75.00 | 77.50 | 30.00 | 52.50 | 40.00 | 45.00 | 25.00 | 35.00 | 22.50 | 47.50 |
78
+ | Emotion | 80 | 67.50 | 61.25 | 40.00 | 67.50 | 65.00 | 52.50 | 86.25 | 65.00 | 22.50 | 60.00 | 67.50 | 62.50 | 56.25 | 52.50 |
79
+ | MMLU | 80 | 88.75 | **77.50** | 61.25 | 75.00 | 68.75 | 73.75 | 62.50 | 66.25 | 52.50 | 51.25 | 53.75 | 22.50 | 27.50 | 31.25 |
80
+ | MMLU-Pro | 200 | 84.00 | 53.00 | 37.00 | 53.00 | 45.50 | 53.50 | 37.50 | 45.00 | 24.50 | 22.50 | 28.00 | 11.00 | 11.50 | 17.50 |
81
+ | Paraphrase | 80 | 87.50 | **91.25** | 85.00 | 81.25 | 81.25 | 88.75 | 77.50 | 83.75 | 80.00 | 58.75 | 70.00 | 86.25 | 75.00 | 47.50 |
82
+ | Question entailment | 80 | 91.25 | 93.75 | 91.25 | 95.00 | 91.25 | 90.00 | 87.50 | 87.50 | 83.75 | 81.25 | 63.75 | 80.00 | 73.75 | 71.25 |
83
+ | Science questions | 80 | 100.00 | 100.00 | 98.75 | 100.00 | 100.00 | 100.00 | 98.75 | 98.75 | 97.50 | 96.25 | 96.25 | 90.00 | 72.50 | 92.50 |
84
+ | Offensive-language detection | 80 | 76.25 | 73.75 | 75.00 | 86.25 | 85.00 | 83.75 | 88.75 | 77.50 | 75.00 | 72.50 | 83.75 | 81.25 | 82.50 | 78.75 |
85
+ | Evidence control · age eligibility | 10 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 90.00 | 50.00 | 50.00 |
86
+ | Evidence control · authorization | 10 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 80.00 | 100.00 | 50.00 | 100.00 | 50.00 | 70.00 | 50.00 | 50.00 |
87
+ | Evidence control · deadline | 10 | 100.00 | 80.00 | 60.00 | 70.00 | 60.00 | 80.00 | 30.00 | 50.00 | 30.00 | 40.00 | 30.00 | 30.00 | 30.00 | 50.00 |
88
+ | Evidence control · late fee | 10 | 100.00 | 90.00 | 90.00 | 100.00 | 100.00 | 90.00 | 70.00 | 80.00 | 80.00 | 90.00 | 70.00 | 40.00 | 40.00 | 30.00 |
89
+ | Evidence control · quantity limit | 10 | 100.00 | 100.00 | 80.00 | 100.00 | 100.00 | 70.00 | 80.00 | 70.00 | 70.00 | 100.00 | 20.00 | 60.00 | 30.00 | 50.00 |
90
+ | Evidence control · return window | 10 | 50.00 | 60.00 | 40.00 | 70.00 | 70.00 | 60.00 | 60.00 | 40.00 | 40.00 | 80.00 | 50.00 | 60.00 | 40.00 | 50.00 |
91
+ | Evidence control · shipping delay | 10 | 90.00 | 90.00 | 40.00 | 90.00 | 80.00 | 80.00 | 80.00 | 60.00 | 80.00 | 100.00 | 30.00 | 30.00 | 30.00 | 20.00 |
92
+ | Evidence control · sla response | 10 | 100.00 | 70.00 | 40.00 | 100.00 | 100.00 | 50.00 | 30.00 | 50.00 | 40.00 | 100.00 | 50.00 | 50.00 | 30.00 | 20.00 |
93
+ | Evidence control · spend threshold | 10 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 90.00 | 100.00 | 100.00 | 50.00 | 50.00 | 50.00 |
94
+ | Evidence control · volume discount | 10 | 100.00 | 100.00 | 80.00 | 100.00 | 100.00 | 100.00 | 90.00 | 100.00 | 90.00 | 100.00 | 20.00 | 20.00 | 30.00 | 30.00 |
95
+ | Evidence control · warranty claim | 10 | 90.00 | 40.00 | 70.00 | 80.00 | 100.00 | 60.00 | 40.00 | 50.00 | 50.00 | 50.00 | 50.00 | 40.00 | 30.00 | 10.00 |
96
+
97
+ </details>
USAGE.md ADDED
@@ -0,0 +1,29 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Use Lux
2
+
3
+ Lux accepts the SystemOne-style request shape: one `state` and a mapping of named `questions`. Use the bundled Python API in the [qualified ROCm environment](RUNTIME.md).
4
+
5
+ ```python
6
+ from decision import DecisionModel
7
+
8
+ model = DecisionModel.from_pretrained("/model", local_files_only=True)
9
+ request = {
10
+ "state": {"owner": "Lee", "status": "active", "severity": "medium"},
11
+ "questions": {
12
+ "owner": {"type": "choice", "instructions": "Choose the owner.",
13
+ "criteria": {"lee": "Lee", "sam": "Sam"}},
14
+ "active": {"type": "noul", "instructions": "Is the status active?"},
15
+ "severity": {"type": "score", "instructions": "Apply the severity scale.",
16
+ "criteria": ["low", "medium", "high"]},
17
+ },
18
+ }
19
+ result = model.decide(**request)
20
+ print(result["answers"])
21
+ ```
22
+
23
+ The state can be text or JSON-compatible structured data. Arbitrary string question IDs and Choice IDs are preserved in the response. Choice requires 2–255 candidates; Score requires 2–10 descriptions ordered from index zero upward. Noul returns `noul`, the probability of the condition being true; optional `criteria` may define the `false` and `true` meanings.
24
+
25
+ Choice returns `choice`, `probabilities` and `confidence`. Score returns its expected zero-based index as `score`, a distribution and the index-to-description `legend`. Confidence is normalized maximum probability, `(K × max(p) − 1)/(K − 1)`; it is not a claim to reproduce TypeSafe's unpublished confidence statistic.
26
+
27
+ One request can contain many questions. Lux batches independent complete questions eight at a time, preserving request order. Each question reads the complete state and its own candidates. This runtime does not reuse a model-state cache across questions. The 16,384-token limit applies to each complete rendered question, including state, instructions, all candidates and formatting. Oversized inputs raise `ValueError` before a forward pass; they are not silently truncated.
28
+
29
+ Use a fresh Python process when switching bound model profiles. The package validates its runtime and automatically installs its bundled normalization profile. This is a local inference API; an HTTP service can wrap the same request and response objects.
WEIGHTING.md ADDED
@@ -0,0 +1,18 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Product weights were changed after earlier results were observed. These are identical model predictions under different weights, not trained-model improvements.
2
+
3
+ | Model | Current 30/25/15/15/15 | Earlier 25/25/15/15/20 | Original four-panel mean |
4
+ |---|---:|---:|---:|
5
+ | Jev | 81.05 | 81.45 | 82.45 |
6
+ | Lux | 76.72 | 76.43 | 79.06 |
7
+ | Nox | 72.84 | 72.09 | 75.03 |
8
+ | Kev-9B | 71.89 | 72.01 | 73.19 |
9
+ | Kev-4B | 70.09 | 70.30 | 71.73 |
10
+ | Qwen3.5-9B | 69.73 | 69.70 | 71.99 |
11
+ | Decider | 67.71 | 67.97 | 71.75 |
12
+ | Qwen3.5-4B | 67.29 | 67.24 | 70.25 |
13
+ | Sol | 66.32 | 65.48 | 70.14 |
14
+ | Kev-0.8B | 58.28 | 58.33 | 59.75 |
15
+ | Qwen3.5-2B | 57.24 | 57.20 | 60.54 |
16
+ | Laya · English | 51.03 | 50.85 | 51.76 |
17
+ | Laya · Multilingual | 47.19 | 47.18 | 48.56 |
18
+ | Kai | 46.49 | 46.80 | 48.16 |
assets/architecture.pdf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:482fac3bd270e5155e7d7fe61898c113279214222e4c2124b3e495b04b0fe1ea
3
+ size 172410
assets/architecture.png ADDED

Git LFS Details

  • SHA256: 1e8fbe7655e7c1aadb4b2323290cdb20e726c0a42b7914c0a1d5954a6a029931
  • Pointer size: 131 Bytes
  • Size of remote file: 495 kB
assets/architecture.svg ADDED
assets/decision-matrix.pdf ADDED
Binary file (28.6 kB). View file
 
assets/decision-matrix.png ADDED

Git LFS Details

  • SHA256: b6e20d8a2dd7a9bd6cf923e1b99cb1ac8f837532012d5f3e6f0cb245b12f0acc
  • Pointer size: 131 Bytes
  • Size of remote file: 345 kB
assets/decision-matrix.svg ADDED
assets/decision-question-scaling.pdf ADDED
Binary file (18.5 kB). View file
 
assets/decision-question-scaling.png ADDED
assets/decision-question-scaling.svg ADDED
assets/decision-ranking.pdf ADDED
Binary file (24.6 kB). View file
 
assets/decision-ranking.png ADDED

Git LFS Details

  • SHA256: 1bef771a47e7a770c44bd174aaf8c33db3ec8c5dfa20087d5eff15f5a18848e0
  • Pointer size: 131 Bytes
  • Size of remote file: 226 kB
assets/decision-ranking.svg ADDED
assets/readout.png ADDED

Git LFS Details

  • SHA256: fb6b3d1cf149a1a3ccb584eb1c213df73c07bcd328732b4238928e2a2a4af085
  • Pointer size: 131 Bytes
  • Size of remote file: 285 kB
assets/readout.svg ADDED
backbone/config.json ADDED
@@ -0,0 +1,83 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "Qwen3_5TextModel"
4
+ ],
5
+ "attention_bias": false,
6
+ "attention_dropout": 0.0,
7
+ "attn_output_gate": true,
8
+ "bos_token_id": null,
9
+ "dtype": "bfloat16",
10
+ "eos_token_id": 248044,
11
+ "full_attention_interval": 4,
12
+ "head_dim": 256,
13
+ "hidden_act": "silu",
14
+ "hidden_size": 4096,
15
+ "initializer_range": 0.02,
16
+ "intermediate_size": 12288,
17
+ "layer_types": [
18
+ "linear_attention",
19
+ "linear_attention",
20
+ "linear_attention",
21
+ "full_attention",
22
+ "linear_attention",
23
+ "linear_attention",
24
+ "linear_attention",
25
+ "full_attention",
26
+ "linear_attention",
27
+ "linear_attention",
28
+ "linear_attention",
29
+ "full_attention",
30
+ "linear_attention",
31
+ "linear_attention",
32
+ "linear_attention",
33
+ "full_attention",
34
+ "linear_attention",
35
+ "linear_attention",
36
+ "linear_attention",
37
+ "full_attention",
38
+ "linear_attention",
39
+ "linear_attention",
40
+ "linear_attention",
41
+ "full_attention",
42
+ "linear_attention",
43
+ "linear_attention",
44
+ "linear_attention",
45
+ "full_attention",
46
+ "linear_attention",
47
+ "linear_attention",
48
+ "linear_attention",
49
+ "full_attention"
50
+ ],
51
+ "linear_conv_kernel_dim": 4,
52
+ "linear_key_head_dim": 128,
53
+ "linear_num_key_heads": 16,
54
+ "linear_num_value_heads": 32,
55
+ "linear_value_head_dim": 128,
56
+ "mamba_ssm_dtype": "float32",
57
+ "max_position_embeddings": 262144,
58
+ "mlp_only_layers": [],
59
+ "model_type": "qwen3_5_text",
60
+ "mtp_num_hidden_layers": 1,
61
+ "mtp_use_dedicated_embeddings": false,
62
+ "num_attention_heads": 16,
63
+ "num_hidden_layers": 32,
64
+ "num_key_value_heads": 4,
65
+ "pad_token_id": null,
66
+ "partial_rotary_factor": 0.25,
67
+ "rms_norm_eps": 1e-06,
68
+ "rope_parameters": {
69
+ "mrope_interleaved": true,
70
+ "mrope_section": [
71
+ 11,
72
+ 11,
73
+ 10
74
+ ],
75
+ "partial_rotary_factor": 0.25,
76
+ "rope_theta": 10000000,
77
+ "rope_type": "default"
78
+ },
79
+ "tie_word_embeddings": false,
80
+ "transformers_version": "5.17.0",
81
+ "use_cache": false,
82
+ "vocab_size": 248320
83
+ }
backbone/model-00001-of-00004.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:1c64dafca4c6214e64d73f83ec81634773d2be3fdc85fd93f82dd53140e07e2c
3
+ size 3999614440
backbone/model-00002-of-00004.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:55b052d3f60ac4bc7cb0b168cb185899eb5ba64b2b3a98dcef9211e27a89b076
3
+ size 3997271592
backbone/model-00003-of-00004.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:1d8a356f5a18768e4dae02fea0834fee5f136a3052ec8472672f41018ec2b17b
3
+ size 3997288320
backbone/model-00004-of-00004.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:bc97b5ffb4fd5fe2f171e73420dfaf2ccb415e749ba48d5864c12c0a44af9e71
3
+ size 3879241544
backbone/model.safetensors.index.json ADDED
@@ -0,0 +1,434 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "metadata": {
3
+ "total_parameters": 7936684544,
4
+ "total_size": 15873369088
5
+ },
6
+ "weight_map": {
7
+ "embed_tokens.weight": "model-00001-of-00004.safetensors",
8
+ "layers.0.input_layernorm.weight": "model-00001-of-00004.safetensors",
9
+ "layers.0.linear_attn.A_log": "model-00001-of-00004.safetensors",
10
+ "layers.0.linear_attn.conv1d.weight": "model-00001-of-00004.safetensors",
11
+ "layers.0.linear_attn.dt_bias": "model-00001-of-00004.safetensors",
12
+ "layers.0.linear_attn.in_proj_a.weight": "model-00001-of-00004.safetensors",
13
+ "layers.0.linear_attn.in_proj_b.weight": "model-00001-of-00004.safetensors",
14
+ "layers.0.linear_attn.in_proj_qkv.weight": "model-00001-of-00004.safetensors",
15
+ "layers.0.linear_attn.in_proj_z.weight": "model-00001-of-00004.safetensors",
16
+ "layers.0.linear_attn.norm.weight": "model-00001-of-00004.safetensors",
17
+ "layers.0.linear_attn.out_proj.weight": "model-00001-of-00004.safetensors",
18
+ "layers.0.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
19
+ "layers.0.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
20
+ "layers.0.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
21
+ "layers.0.post_attention_layernorm.weight": "model-00001-of-00004.safetensors",
22
+ "layers.1.input_layernorm.weight": "model-00001-of-00004.safetensors",
23
+ "layers.1.linear_attn.A_log": "model-00001-of-00004.safetensors",
24
+ "layers.1.linear_attn.conv1d.weight": "model-00001-of-00004.safetensors",
25
+ "layers.1.linear_attn.dt_bias": "model-00001-of-00004.safetensors",
26
+ "layers.1.linear_attn.in_proj_a.weight": "model-00001-of-00004.safetensors",
27
+ "layers.1.linear_attn.in_proj_b.weight": "model-00001-of-00004.safetensors",
28
+ "layers.1.linear_attn.in_proj_qkv.weight": "model-00001-of-00004.safetensors",
29
+ "layers.1.linear_attn.in_proj_z.weight": "model-00001-of-00004.safetensors",
30
+ "layers.1.linear_attn.norm.weight": "model-00001-of-00004.safetensors",
31
+ "layers.1.linear_attn.out_proj.weight": "model-00001-of-00004.safetensors",
32
+ "layers.1.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
33
+ "layers.1.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
34
+ "layers.1.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
35
+ "layers.1.post_attention_layernorm.weight": "model-00001-of-00004.safetensors",
36
+ "layers.10.input_layernorm.weight": "model-00002-of-00004.safetensors",
37
+ "layers.10.linear_attn.A_log": "model-00002-of-00004.safetensors",
38
+ "layers.10.linear_attn.conv1d.weight": "model-00002-of-00004.safetensors",
39
+ "layers.10.linear_attn.dt_bias": "model-00002-of-00004.safetensors",
40
+ "layers.10.linear_attn.in_proj_a.weight": "model-00002-of-00004.safetensors",
41
+ "layers.10.linear_attn.in_proj_b.weight": "model-00002-of-00004.safetensors",
42
+ "layers.10.linear_attn.in_proj_qkv.weight": "model-00002-of-00004.safetensors",
43
+ "layers.10.linear_attn.in_proj_z.weight": "model-00002-of-00004.safetensors",
44
+ "layers.10.linear_attn.norm.weight": "model-00002-of-00004.safetensors",
45
+ "layers.10.linear_attn.out_proj.weight": "model-00002-of-00004.safetensors",
46
+ "layers.10.mlp.down_proj.weight": "model-00002-of-00004.safetensors",
47
+ "layers.10.mlp.gate_proj.weight": "model-00002-of-00004.safetensors",
48
+ "layers.10.mlp.up_proj.weight": "model-00002-of-00004.safetensors",
49
+ "layers.10.post_attention_layernorm.weight": "model-00002-of-00004.safetensors",
50
+ "layers.11.input_layernorm.weight": "model-00002-of-00004.safetensors",
51
+ "layers.11.mlp.down_proj.weight": "model-00002-of-00004.safetensors",
52
+ "layers.11.mlp.gate_proj.weight": "model-00002-of-00004.safetensors",
53
+ "layers.11.mlp.up_proj.weight": "model-00002-of-00004.safetensors",
54
+ "layers.11.post_attention_layernorm.weight": "model-00002-of-00004.safetensors",
55
+ "layers.11.self_attn.k_norm.weight": "model-00002-of-00004.safetensors",
56
+ "layers.11.self_attn.k_proj.weight": "model-00002-of-00004.safetensors",
57
+ "layers.11.self_attn.o_proj.weight": "model-00002-of-00004.safetensors",
58
+ "layers.11.self_attn.q_norm.weight": "model-00002-of-00004.safetensors",
59
+ "layers.11.self_attn.q_proj.weight": "model-00002-of-00004.safetensors",
60
+ "layers.11.self_attn.v_proj.weight": "model-00002-of-00004.safetensors",
61
+ "layers.12.input_layernorm.weight": "model-00002-of-00004.safetensors",
62
+ "layers.12.linear_attn.A_log": "model-00002-of-00004.safetensors",
63
+ "layers.12.linear_attn.conv1d.weight": "model-00002-of-00004.safetensors",
64
+ "layers.12.linear_attn.dt_bias": "model-00002-of-00004.safetensors",
65
+ "layers.12.linear_attn.in_proj_a.weight": "model-00002-of-00004.safetensors",
66
+ "layers.12.linear_attn.in_proj_b.weight": "model-00002-of-00004.safetensors",
67
+ "layers.12.linear_attn.in_proj_qkv.weight": "model-00002-of-00004.safetensors",
68
+ "layers.12.linear_attn.in_proj_z.weight": "model-00002-of-00004.safetensors",
69
+ "layers.12.linear_attn.norm.weight": "model-00002-of-00004.safetensors",
70
+ "layers.12.linear_attn.out_proj.weight": "model-00002-of-00004.safetensors",
71
+ "layers.12.mlp.down_proj.weight": "model-00002-of-00004.safetensors",
72
+ "layers.12.mlp.gate_proj.weight": "model-00002-of-00004.safetensors",
73
+ "layers.12.mlp.up_proj.weight": "model-00002-of-00004.safetensors",
74
+ "layers.12.post_attention_layernorm.weight": "model-00002-of-00004.safetensors",
75
+ "layers.13.input_layernorm.weight": "model-00002-of-00004.safetensors",
76
+ "layers.13.linear_attn.A_log": "model-00002-of-00004.safetensors",
77
+ "layers.13.linear_attn.conv1d.weight": "model-00002-of-00004.safetensors",
78
+ "layers.13.linear_attn.dt_bias": "model-00002-of-00004.safetensors",
79
+ "layers.13.linear_attn.in_proj_a.weight": "model-00002-of-00004.safetensors",
80
+ "layers.13.linear_attn.in_proj_b.weight": "model-00002-of-00004.safetensors",
81
+ "layers.13.linear_attn.in_proj_qkv.weight": "model-00002-of-00004.safetensors",
82
+ "layers.13.linear_attn.in_proj_z.weight": "model-00002-of-00004.safetensors",
83
+ "layers.13.linear_attn.norm.weight": "model-00002-of-00004.safetensors",
84
+ "layers.13.linear_attn.out_proj.weight": "model-00002-of-00004.safetensors",
85
+ "layers.13.mlp.down_proj.weight": "model-00002-of-00004.safetensors",
86
+ "layers.13.mlp.gate_proj.weight": "model-00002-of-00004.safetensors",
87
+ "layers.13.mlp.up_proj.weight": "model-00003-of-00004.safetensors",
88
+ "layers.13.post_attention_layernorm.weight": "model-00003-of-00004.safetensors",
89
+ "layers.14.input_layernorm.weight": "model-00003-of-00004.safetensors",
90
+ "layers.14.linear_attn.A_log": "model-00003-of-00004.safetensors",
91
+ "layers.14.linear_attn.conv1d.weight": "model-00003-of-00004.safetensors",
92
+ "layers.14.linear_attn.dt_bias": "model-00003-of-00004.safetensors",
93
+ "layers.14.linear_attn.in_proj_a.weight": "model-00003-of-00004.safetensors",
94
+ "layers.14.linear_attn.in_proj_b.weight": "model-00003-of-00004.safetensors",
95
+ "layers.14.linear_attn.in_proj_qkv.weight": "model-00003-of-00004.safetensors",
96
+ "layers.14.linear_attn.in_proj_z.weight": "model-00003-of-00004.safetensors",
97
+ "layers.14.linear_attn.norm.weight": "model-00003-of-00004.safetensors",
98
+ "layers.14.linear_attn.out_proj.weight": "model-00003-of-00004.safetensors",
99
+ "layers.14.mlp.down_proj.weight": "model-00003-of-00004.safetensors",
100
+ "layers.14.mlp.gate_proj.weight": "model-00003-of-00004.safetensors",
101
+ "layers.14.mlp.up_proj.weight": "model-00003-of-00004.safetensors",
102
+ "layers.14.post_attention_layernorm.weight": "model-00003-of-00004.safetensors",
103
+ "layers.15.input_layernorm.weight": "model-00003-of-00004.safetensors",
104
+ "layers.15.mlp.down_proj.weight": "model-00003-of-00004.safetensors",
105
+ "layers.15.mlp.gate_proj.weight": "model-00003-of-00004.safetensors",
106
+ "layers.15.mlp.up_proj.weight": "model-00003-of-00004.safetensors",
107
+ "layers.15.post_attention_layernorm.weight": "model-00003-of-00004.safetensors",
108
+ "layers.15.self_attn.k_norm.weight": "model-00003-of-00004.safetensors",
109
+ "layers.15.self_attn.k_proj.weight": "model-00003-of-00004.safetensors",
110
+ "layers.15.self_attn.o_proj.weight": "model-00003-of-00004.safetensors",
111
+ "layers.15.self_attn.q_norm.weight": "model-00003-of-00004.safetensors",
112
+ "layers.15.self_attn.q_proj.weight": "model-00003-of-00004.safetensors",
113
+ "layers.15.self_attn.v_proj.weight": "model-00003-of-00004.safetensors",
114
+ "layers.16.input_layernorm.weight": "model-00003-of-00004.safetensors",
115
+ "layers.16.linear_attn.A_log": "model-00003-of-00004.safetensors",
116
+ "layers.16.linear_attn.conv1d.weight": "model-00003-of-00004.safetensors",
117
+ "layers.16.linear_attn.dt_bias": "model-00003-of-00004.safetensors",
118
+ "layers.16.linear_attn.in_proj_a.weight": "model-00003-of-00004.safetensors",
119
+ "layers.16.linear_attn.in_proj_b.weight": "model-00003-of-00004.safetensors",
120
+ "layers.16.linear_attn.in_proj_qkv.weight": "model-00003-of-00004.safetensors",
121
+ "layers.16.linear_attn.in_proj_z.weight": "model-00003-of-00004.safetensors",
122
+ "layers.16.linear_attn.norm.weight": "model-00003-of-00004.safetensors",
123
+ "layers.16.linear_attn.out_proj.weight": "model-00003-of-00004.safetensors",
124
+ "layers.16.mlp.down_proj.weight": "model-00003-of-00004.safetensors",
125
+ "layers.16.mlp.gate_proj.weight": "model-00003-of-00004.safetensors",
126
+ "layers.16.mlp.up_proj.weight": "model-00003-of-00004.safetensors",
127
+ "layers.16.post_attention_layernorm.weight": "model-00003-of-00004.safetensors",
128
+ "layers.17.input_layernorm.weight": "model-00003-of-00004.safetensors",
129
+ "layers.17.linear_attn.A_log": "model-00003-of-00004.safetensors",
130
+ "layers.17.linear_attn.conv1d.weight": "model-00003-of-00004.safetensors",
131
+ "layers.17.linear_attn.dt_bias": "model-00003-of-00004.safetensors",
132
+ "layers.17.linear_attn.in_proj_a.weight": "model-00003-of-00004.safetensors",
133
+ "layers.17.linear_attn.in_proj_b.weight": "model-00003-of-00004.safetensors",
134
+ "layers.17.linear_attn.in_proj_qkv.weight": "model-00003-of-00004.safetensors",
135
+ "layers.17.linear_attn.in_proj_z.weight": "model-00003-of-00004.safetensors",
136
+ "layers.17.linear_attn.norm.weight": "model-00003-of-00004.safetensors",
137
+ "layers.17.linear_attn.out_proj.weight": "model-00003-of-00004.safetensors",
138
+ "layers.17.mlp.down_proj.weight": "model-00003-of-00004.safetensors",
139
+ "layers.17.mlp.gate_proj.weight": "model-00003-of-00004.safetensors",
140
+ "layers.17.mlp.up_proj.weight": "model-00003-of-00004.safetensors",
141
+ "layers.17.post_attention_layernorm.weight": "model-00003-of-00004.safetensors",
142
+ "layers.18.input_layernorm.weight": "model-00003-of-00004.safetensors",
143
+ "layers.18.linear_attn.A_log": "model-00003-of-00004.safetensors",
144
+ "layers.18.linear_attn.conv1d.weight": "model-00003-of-00004.safetensors",
145
+ "layers.18.linear_attn.dt_bias": "model-00003-of-00004.safetensors",
146
+ "layers.18.linear_attn.in_proj_a.weight": "model-00003-of-00004.safetensors",
147
+ "layers.18.linear_attn.in_proj_b.weight": "model-00003-of-00004.safetensors",
148
+ "layers.18.linear_attn.in_proj_qkv.weight": "model-00003-of-00004.safetensors",
149
+ "layers.18.linear_attn.in_proj_z.weight": "model-00003-of-00004.safetensors",
150
+ "layers.18.linear_attn.norm.weight": "model-00003-of-00004.safetensors",
151
+ "layers.18.linear_attn.out_proj.weight": "model-00003-of-00004.safetensors",
152
+ "layers.18.mlp.down_proj.weight": "model-00003-of-00004.safetensors",
153
+ "layers.18.mlp.gate_proj.weight": "model-00003-of-00004.safetensors",
154
+ "layers.18.mlp.up_proj.weight": "model-00003-of-00004.safetensors",
155
+ "layers.18.post_attention_layernorm.weight": "model-00003-of-00004.safetensors",
156
+ "layers.19.input_layernorm.weight": "model-00003-of-00004.safetensors",
157
+ "layers.19.mlp.down_proj.weight": "model-00003-of-00004.safetensors",
158
+ "layers.19.mlp.gate_proj.weight": "model-00003-of-00004.safetensors",
159
+ "layers.19.mlp.up_proj.weight": "model-00003-of-00004.safetensors",
160
+ "layers.19.post_attention_layernorm.weight": "model-00003-of-00004.safetensors",
161
+ "layers.19.self_attn.k_norm.weight": "model-00003-of-00004.safetensors",
162
+ "layers.19.self_attn.k_proj.weight": "model-00003-of-00004.safetensors",
163
+ "layers.19.self_attn.o_proj.weight": "model-00003-of-00004.safetensors",
164
+ "layers.19.self_attn.q_norm.weight": "model-00003-of-00004.safetensors",
165
+ "layers.19.self_attn.q_proj.weight": "model-00003-of-00004.safetensors",
166
+ "layers.19.self_attn.v_proj.weight": "model-00003-of-00004.safetensors",
167
+ "layers.2.input_layernorm.weight": "model-00001-of-00004.safetensors",
168
+ "layers.2.linear_attn.A_log": "model-00001-of-00004.safetensors",
169
+ "layers.2.linear_attn.conv1d.weight": "model-00001-of-00004.safetensors",
170
+ "layers.2.linear_attn.dt_bias": "model-00001-of-00004.safetensors",
171
+ "layers.2.linear_attn.in_proj_a.weight": "model-00001-of-00004.safetensors",
172
+ "layers.2.linear_attn.in_proj_b.weight": "model-00001-of-00004.safetensors",
173
+ "layers.2.linear_attn.in_proj_qkv.weight": "model-00001-of-00004.safetensors",
174
+ "layers.2.linear_attn.in_proj_z.weight": "model-00001-of-00004.safetensors",
175
+ "layers.2.linear_attn.norm.weight": "model-00001-of-00004.safetensors",
176
+ "layers.2.linear_attn.out_proj.weight": "model-00001-of-00004.safetensors",
177
+ "layers.2.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
178
+ "layers.2.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
179
+ "layers.2.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
180
+ "layers.2.post_attention_layernorm.weight": "model-00001-of-00004.safetensors",
181
+ "layers.20.input_layernorm.weight": "model-00003-of-00004.safetensors",
182
+ "layers.20.linear_attn.A_log": "model-00003-of-00004.safetensors",
183
+ "layers.20.linear_attn.conv1d.weight": "model-00003-of-00004.safetensors",
184
+ "layers.20.linear_attn.dt_bias": "model-00003-of-00004.safetensors",
185
+ "layers.20.linear_attn.in_proj_a.weight": "model-00003-of-00004.safetensors",
186
+ "layers.20.linear_attn.in_proj_b.weight": "model-00003-of-00004.safetensors",
187
+ "layers.20.linear_attn.in_proj_qkv.weight": "model-00003-of-00004.safetensors",
188
+ "layers.20.linear_attn.in_proj_z.weight": "model-00003-of-00004.safetensors",
189
+ "layers.20.linear_attn.norm.weight": "model-00003-of-00004.safetensors",
190
+ "layers.20.linear_attn.out_proj.weight": "model-00003-of-00004.safetensors",
191
+ "layers.20.mlp.down_proj.weight": "model-00003-of-00004.safetensors",
192
+ "layers.20.mlp.gate_proj.weight": "model-00003-of-00004.safetensors",
193
+ "layers.20.mlp.up_proj.weight": "model-00003-of-00004.safetensors",
194
+ "layers.20.post_attention_layernorm.weight": "model-00003-of-00004.safetensors",
195
+ "layers.21.input_layernorm.weight": "model-00003-of-00004.safetensors",
196
+ "layers.21.linear_attn.A_log": "model-00003-of-00004.safetensors",
197
+ "layers.21.linear_attn.conv1d.weight": "model-00003-of-00004.safetensors",
198
+ "layers.21.linear_attn.dt_bias": "model-00003-of-00004.safetensors",
199
+ "layers.21.linear_attn.in_proj_a.weight": "model-00003-of-00004.safetensors",
200
+ "layers.21.linear_attn.in_proj_b.weight": "model-00003-of-00004.safetensors",
201
+ "layers.21.linear_attn.in_proj_qkv.weight": "model-00003-of-00004.safetensors",
202
+ "layers.21.linear_attn.in_proj_z.weight": "model-00003-of-00004.safetensors",
203
+ "layers.21.linear_attn.norm.weight": "model-00003-of-00004.safetensors",
204
+ "layers.21.linear_attn.out_proj.weight": "model-00003-of-00004.safetensors",
205
+ "layers.21.mlp.down_proj.weight": "model-00003-of-00004.safetensors",
206
+ "layers.21.mlp.gate_proj.weight": "model-00003-of-00004.safetensors",
207
+ "layers.21.mlp.up_proj.weight": "model-00003-of-00004.safetensors",
208
+ "layers.21.post_attention_layernorm.weight": "model-00003-of-00004.safetensors",
209
+ "layers.22.input_layernorm.weight": "model-00003-of-00004.safetensors",
210
+ "layers.22.linear_attn.A_log": "model-00003-of-00004.safetensors",
211
+ "layers.22.linear_attn.conv1d.weight": "model-00003-of-00004.safetensors",
212
+ "layers.22.linear_attn.dt_bias": "model-00003-of-00004.safetensors",
213
+ "layers.22.linear_attn.in_proj_a.weight": "model-00003-of-00004.safetensors",
214
+ "layers.22.linear_attn.in_proj_b.weight": "model-00003-of-00004.safetensors",
215
+ "layers.22.linear_attn.in_proj_qkv.weight": "model-00003-of-00004.safetensors",
216
+ "layers.22.linear_attn.in_proj_z.weight": "model-00003-of-00004.safetensors",
217
+ "layers.22.linear_attn.norm.weight": "model-00003-of-00004.safetensors",
218
+ "layers.22.linear_attn.out_proj.weight": "model-00003-of-00004.safetensors",
219
+ "layers.22.mlp.down_proj.weight": "model-00003-of-00004.safetensors",
220
+ "layers.22.mlp.gate_proj.weight": "model-00003-of-00004.safetensors",
221
+ "layers.22.mlp.up_proj.weight": "model-00003-of-00004.safetensors",
222
+ "layers.22.post_attention_layernorm.weight": "model-00003-of-00004.safetensors",
223
+ "layers.23.input_layernorm.weight": "model-00003-of-00004.safetensors",
224
+ "layers.23.mlp.down_proj.weight": "model-00004-of-00004.safetensors",
225
+ "layers.23.mlp.gate_proj.weight": "model-00004-of-00004.safetensors",
226
+ "layers.23.mlp.up_proj.weight": "model-00004-of-00004.safetensors",
227
+ "layers.23.post_attention_layernorm.weight": "model-00004-of-00004.safetensors",
228
+ "layers.23.self_attn.k_norm.weight": "model-00004-of-00004.safetensors",
229
+ "layers.23.self_attn.k_proj.weight": "model-00004-of-00004.safetensors",
230
+ "layers.23.self_attn.o_proj.weight": "model-00004-of-00004.safetensors",
231
+ "layers.23.self_attn.q_norm.weight": "model-00004-of-00004.safetensors",
232
+ "layers.23.self_attn.q_proj.weight": "model-00004-of-00004.safetensors",
233
+ "layers.23.self_attn.v_proj.weight": "model-00004-of-00004.safetensors",
234
+ "layers.24.input_layernorm.weight": "model-00004-of-00004.safetensors",
235
+ "layers.24.linear_attn.A_log": "model-00004-of-00004.safetensors",
236
+ "layers.24.linear_attn.conv1d.weight": "model-00004-of-00004.safetensors",
237
+ "layers.24.linear_attn.dt_bias": "model-00004-of-00004.safetensors",
238
+ "layers.24.linear_attn.in_proj_a.weight": "model-00004-of-00004.safetensors",
239
+ "layers.24.linear_attn.in_proj_b.weight": "model-00004-of-00004.safetensors",
240
+ "layers.24.linear_attn.in_proj_qkv.weight": "model-00004-of-00004.safetensors",
241
+ "layers.24.linear_attn.in_proj_z.weight": "model-00004-of-00004.safetensors",
242
+ "layers.24.linear_attn.norm.weight": "model-00004-of-00004.safetensors",
243
+ "layers.24.linear_attn.out_proj.weight": "model-00004-of-00004.safetensors",
244
+ "layers.24.mlp.down_proj.weight": "model-00004-of-00004.safetensors",
245
+ "layers.24.mlp.gate_proj.weight": "model-00004-of-00004.safetensors",
246
+ "layers.24.mlp.up_proj.weight": "model-00004-of-00004.safetensors",
247
+ "layers.24.post_attention_layernorm.weight": "model-00004-of-00004.safetensors",
248
+ "layers.25.input_layernorm.weight": "model-00004-of-00004.safetensors",
249
+ "layers.25.linear_attn.A_log": "model-00004-of-00004.safetensors",
250
+ "layers.25.linear_attn.conv1d.weight": "model-00004-of-00004.safetensors",
251
+ "layers.25.linear_attn.dt_bias": "model-00004-of-00004.safetensors",
252
+ "layers.25.linear_attn.in_proj_a.weight": "model-00004-of-00004.safetensors",
253
+ "layers.25.linear_attn.in_proj_b.weight": "model-00004-of-00004.safetensors",
254
+ "layers.25.linear_attn.in_proj_qkv.weight": "model-00004-of-00004.safetensors",
255
+ "layers.25.linear_attn.in_proj_z.weight": "model-00004-of-00004.safetensors",
256
+ "layers.25.linear_attn.norm.weight": "model-00004-of-00004.safetensors",
257
+ "layers.25.linear_attn.out_proj.weight": "model-00004-of-00004.safetensors",
258
+ "layers.25.mlp.down_proj.weight": "model-00004-of-00004.safetensors",
259
+ "layers.25.mlp.gate_proj.weight": "model-00004-of-00004.safetensors",
260
+ "layers.25.mlp.up_proj.weight": "model-00004-of-00004.safetensors",
261
+ "layers.25.post_attention_layernorm.weight": "model-00004-of-00004.safetensors",
262
+ "layers.26.input_layernorm.weight": "model-00004-of-00004.safetensors",
263
+ "layers.26.linear_attn.A_log": "model-00004-of-00004.safetensors",
264
+ "layers.26.linear_attn.conv1d.weight": "model-00004-of-00004.safetensors",
265
+ "layers.26.linear_attn.dt_bias": "model-00004-of-00004.safetensors",
266
+ "layers.26.linear_attn.in_proj_a.weight": "model-00004-of-00004.safetensors",
267
+ "layers.26.linear_attn.in_proj_b.weight": "model-00004-of-00004.safetensors",
268
+ "layers.26.linear_attn.in_proj_qkv.weight": "model-00004-of-00004.safetensors",
269
+ "layers.26.linear_attn.in_proj_z.weight": "model-00004-of-00004.safetensors",
270
+ "layers.26.linear_attn.norm.weight": "model-00004-of-00004.safetensors",
271
+ "layers.26.linear_attn.out_proj.weight": "model-00004-of-00004.safetensors",
272
+ "layers.26.mlp.down_proj.weight": "model-00004-of-00004.safetensors",
273
+ "layers.26.mlp.gate_proj.weight": "model-00004-of-00004.safetensors",
274
+ "layers.26.mlp.up_proj.weight": "model-00004-of-00004.safetensors",
275
+ "layers.26.post_attention_layernorm.weight": "model-00004-of-00004.safetensors",
276
+ "layers.27.input_layernorm.weight": "model-00004-of-00004.safetensors",
277
+ "layers.27.mlp.down_proj.weight": "model-00004-of-00004.safetensors",
278
+ "layers.27.mlp.gate_proj.weight": "model-00004-of-00004.safetensors",
279
+ "layers.27.mlp.up_proj.weight": "model-00004-of-00004.safetensors",
280
+ "layers.27.post_attention_layernorm.weight": "model-00004-of-00004.safetensors",
281
+ "layers.27.self_attn.k_norm.weight": "model-00004-of-00004.safetensors",
282
+ "layers.27.self_attn.k_proj.weight": "model-00004-of-00004.safetensors",
283
+ "layers.27.self_attn.o_proj.weight": "model-00004-of-00004.safetensors",
284
+ "layers.27.self_attn.q_norm.weight": "model-00004-of-00004.safetensors",
285
+ "layers.27.self_attn.q_proj.weight": "model-00004-of-00004.safetensors",
286
+ "layers.27.self_attn.v_proj.weight": "model-00004-of-00004.safetensors",
287
+ "layers.28.input_layernorm.weight": "model-00004-of-00004.safetensors",
288
+ "layers.28.linear_attn.A_log": "model-00004-of-00004.safetensors",
289
+ "layers.28.linear_attn.conv1d.weight": "model-00004-of-00004.safetensors",
290
+ "layers.28.linear_attn.dt_bias": "model-00004-of-00004.safetensors",
291
+ "layers.28.linear_attn.in_proj_a.weight": "model-00004-of-00004.safetensors",
292
+ "layers.28.linear_attn.in_proj_b.weight": "model-00004-of-00004.safetensors",
293
+ "layers.28.linear_attn.in_proj_qkv.weight": "model-00004-of-00004.safetensors",
294
+ "layers.28.linear_attn.in_proj_z.weight": "model-00004-of-00004.safetensors",
295
+ "layers.28.linear_attn.norm.weight": "model-00004-of-00004.safetensors",
296
+ "layers.28.linear_attn.out_proj.weight": "model-00004-of-00004.safetensors",
297
+ "layers.28.mlp.down_proj.weight": "model-00004-of-00004.safetensors",
298
+ "layers.28.mlp.gate_proj.weight": "model-00004-of-00004.safetensors",
299
+ "layers.28.mlp.up_proj.weight": "model-00004-of-00004.safetensors",
300
+ "layers.28.post_attention_layernorm.weight": "model-00004-of-00004.safetensors",
301
+ "layers.29.input_layernorm.weight": "model-00004-of-00004.safetensors",
302
+ "layers.29.linear_attn.A_log": "model-00004-of-00004.safetensors",
303
+ "layers.29.linear_attn.conv1d.weight": "model-00004-of-00004.safetensors",
304
+ "layers.29.linear_attn.dt_bias": "model-00004-of-00004.safetensors",
305
+ "layers.29.linear_attn.in_proj_a.weight": "model-00004-of-00004.safetensors",
306
+ "layers.29.linear_attn.in_proj_b.weight": "model-00004-of-00004.safetensors",
307
+ "layers.29.linear_attn.in_proj_qkv.weight": "model-00004-of-00004.safetensors",
308
+ "layers.29.linear_attn.in_proj_z.weight": "model-00004-of-00004.safetensors",
309
+ "layers.29.linear_attn.norm.weight": "model-00004-of-00004.safetensors",
310
+ "layers.29.linear_attn.out_proj.weight": "model-00004-of-00004.safetensors",
311
+ "layers.29.mlp.down_proj.weight": "model-00004-of-00004.safetensors",
312
+ "layers.29.mlp.gate_proj.weight": "model-00004-of-00004.safetensors",
313
+ "layers.29.mlp.up_proj.weight": "model-00004-of-00004.safetensors",
314
+ "layers.29.post_attention_layernorm.weight": "model-00004-of-00004.safetensors",
315
+ "layers.3.input_layernorm.weight": "model-00001-of-00004.safetensors",
316
+ "layers.3.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
317
+ "layers.3.mlp.gate_proj.weight": "model-00001-of-00004.safetensors",
318
+ "layers.3.mlp.up_proj.weight": "model-00001-of-00004.safetensors",
319
+ "layers.3.post_attention_layernorm.weight": "model-00001-of-00004.safetensors",
320
+ "layers.3.self_attn.k_norm.weight": "model-00001-of-00004.safetensors",
321
+ "layers.3.self_attn.k_proj.weight": "model-00001-of-00004.safetensors",
322
+ "layers.3.self_attn.o_proj.weight": "model-00001-of-00004.safetensors",
323
+ "layers.3.self_attn.q_norm.weight": "model-00001-of-00004.safetensors",
324
+ "layers.3.self_attn.q_proj.weight": "model-00001-of-00004.safetensors",
325
+ "layers.3.self_attn.v_proj.weight": "model-00001-of-00004.safetensors",
326
+ "layers.30.input_layernorm.weight": "model-00004-of-00004.safetensors",
327
+ "layers.30.linear_attn.A_log": "model-00004-of-00004.safetensors",
328
+ "layers.30.linear_attn.conv1d.weight": "model-00004-of-00004.safetensors",
329
+ "layers.30.linear_attn.dt_bias": "model-00004-of-00004.safetensors",
330
+ "layers.30.linear_attn.in_proj_a.weight": "model-00004-of-00004.safetensors",
331
+ "layers.30.linear_attn.in_proj_b.weight": "model-00004-of-00004.safetensors",
332
+ "layers.30.linear_attn.in_proj_qkv.weight": "model-00004-of-00004.safetensors",
333
+ "layers.30.linear_attn.in_proj_z.weight": "model-00004-of-00004.safetensors",
334
+ "layers.30.linear_attn.norm.weight": "model-00004-of-00004.safetensors",
335
+ "layers.30.linear_attn.out_proj.weight": "model-00004-of-00004.safetensors",
336
+ "layers.30.mlp.down_proj.weight": "model-00004-of-00004.safetensors",
337
+ "layers.30.mlp.gate_proj.weight": "model-00004-of-00004.safetensors",
338
+ "layers.30.mlp.up_proj.weight": "model-00004-of-00004.safetensors",
339
+ "layers.30.post_attention_layernorm.weight": "model-00004-of-00004.safetensors",
340
+ "layers.31.input_layernorm.weight": "model-00004-of-00004.safetensors",
341
+ "layers.31.mlp.down_proj.weight": "model-00004-of-00004.safetensors",
342
+ "layers.31.mlp.gate_proj.weight": "model-00004-of-00004.safetensors",
343
+ "layers.31.mlp.up_proj.weight": "model-00004-of-00004.safetensors",
344
+ "layers.31.post_attention_layernorm.weight": "model-00004-of-00004.safetensors",
345
+ "layers.31.self_attn.k_norm.weight": "model-00004-of-00004.safetensors",
346
+ "layers.31.self_attn.k_proj.weight": "model-00004-of-00004.safetensors",
347
+ "layers.31.self_attn.o_proj.weight": "model-00004-of-00004.safetensors",
348
+ "layers.31.self_attn.q_norm.weight": "model-00004-of-00004.safetensors",
349
+ "layers.31.self_attn.q_proj.weight": "model-00004-of-00004.safetensors",
350
+ "layers.31.self_attn.v_proj.weight": "model-00004-of-00004.safetensors",
351
+ "layers.4.input_layernorm.weight": "model-00001-of-00004.safetensors",
352
+ "layers.4.linear_attn.A_log": "model-00001-of-00004.safetensors",
353
+ "layers.4.linear_attn.conv1d.weight": "model-00001-of-00004.safetensors",
354
+ "layers.4.linear_attn.dt_bias": "model-00001-of-00004.safetensors",
355
+ "layers.4.linear_attn.in_proj_a.weight": "model-00001-of-00004.safetensors",
356
+ "layers.4.linear_attn.in_proj_b.weight": "model-00001-of-00004.safetensors",
357
+ "layers.4.linear_attn.in_proj_qkv.weight": "model-00001-of-00004.safetensors",
358
+ "layers.4.linear_attn.in_proj_z.weight": "model-00001-of-00004.safetensors",
359
+ "layers.4.linear_attn.norm.weight": "model-00001-of-00004.safetensors",
360
+ "layers.4.linear_attn.out_proj.weight": "model-00001-of-00004.safetensors",
361
+ "layers.4.mlp.down_proj.weight": "model-00001-of-00004.safetensors",
362
+ "layers.4.mlp.gate_proj.weight": "model-00002-of-00004.safetensors",
363
+ "layers.4.mlp.up_proj.weight": "model-00002-of-00004.safetensors",
364
+ "layers.4.post_attention_layernorm.weight": "model-00002-of-00004.safetensors",
365
+ "layers.5.input_layernorm.weight": "model-00002-of-00004.safetensors",
366
+ "layers.5.linear_attn.A_log": "model-00002-of-00004.safetensors",
367
+ "layers.5.linear_attn.conv1d.weight": "model-00002-of-00004.safetensors",
368
+ "layers.5.linear_attn.dt_bias": "model-00002-of-00004.safetensors",
369
+ "layers.5.linear_attn.in_proj_a.weight": "model-00002-of-00004.safetensors",
370
+ "layers.5.linear_attn.in_proj_b.weight": "model-00002-of-00004.safetensors",
371
+ "layers.5.linear_attn.in_proj_qkv.weight": "model-00002-of-00004.safetensors",
372
+ "layers.5.linear_attn.in_proj_z.weight": "model-00002-of-00004.safetensors",
373
+ "layers.5.linear_attn.norm.weight": "model-00002-of-00004.safetensors",
374
+ "layers.5.linear_attn.out_proj.weight": "model-00002-of-00004.safetensors",
375
+ "layers.5.mlp.down_proj.weight": "model-00002-of-00004.safetensors",
376
+ "layers.5.mlp.gate_proj.weight": "model-00002-of-00004.safetensors",
377
+ "layers.5.mlp.up_proj.weight": "model-00002-of-00004.safetensors",
378
+ "layers.5.post_attention_layernorm.weight": "model-00002-of-00004.safetensors",
379
+ "layers.6.input_layernorm.weight": "model-00002-of-00004.safetensors",
380
+ "layers.6.linear_attn.A_log": "model-00002-of-00004.safetensors",
381
+ "layers.6.linear_attn.conv1d.weight": "model-00002-of-00004.safetensors",
382
+ "layers.6.linear_attn.dt_bias": "model-00002-of-00004.safetensors",
383
+ "layers.6.linear_attn.in_proj_a.weight": "model-00002-of-00004.safetensors",
384
+ "layers.6.linear_attn.in_proj_b.weight": "model-00002-of-00004.safetensors",
385
+ "layers.6.linear_attn.in_proj_qkv.weight": "model-00002-of-00004.safetensors",
386
+ "layers.6.linear_attn.in_proj_z.weight": "model-00002-of-00004.safetensors",
387
+ "layers.6.linear_attn.norm.weight": "model-00002-of-00004.safetensors",
388
+ "layers.6.linear_attn.out_proj.weight": "model-00002-of-00004.safetensors",
389
+ "layers.6.mlp.down_proj.weight": "model-00002-of-00004.safetensors",
390
+ "layers.6.mlp.gate_proj.weight": "model-00002-of-00004.safetensors",
391
+ "layers.6.mlp.up_proj.weight": "model-00002-of-00004.safetensors",
392
+ "layers.6.post_attention_layernorm.weight": "model-00002-of-00004.safetensors",
393
+ "layers.7.input_layernorm.weight": "model-00002-of-00004.safetensors",
394
+ "layers.7.mlp.down_proj.weight": "model-00002-of-00004.safetensors",
395
+ "layers.7.mlp.gate_proj.weight": "model-00002-of-00004.safetensors",
396
+ "layers.7.mlp.up_proj.weight": "model-00002-of-00004.safetensors",
397
+ "layers.7.post_attention_layernorm.weight": "model-00002-of-00004.safetensors",
398
+ "layers.7.self_attn.k_norm.weight": "model-00002-of-00004.safetensors",
399
+ "layers.7.self_attn.k_proj.weight": "model-00002-of-00004.safetensors",
400
+ "layers.7.self_attn.o_proj.weight": "model-00002-of-00004.safetensors",
401
+ "layers.7.self_attn.q_norm.weight": "model-00002-of-00004.safetensors",
402
+ "layers.7.self_attn.q_proj.weight": "model-00002-of-00004.safetensors",
403
+ "layers.7.self_attn.v_proj.weight": "model-00002-of-00004.safetensors",
404
+ "layers.8.input_layernorm.weight": "model-00002-of-00004.safetensors",
405
+ "layers.8.linear_attn.A_log": "model-00002-of-00004.safetensors",
406
+ "layers.8.linear_attn.conv1d.weight": "model-00002-of-00004.safetensors",
407
+ "layers.8.linear_attn.dt_bias": "model-00002-of-00004.safetensors",
408
+ "layers.8.linear_attn.in_proj_a.weight": "model-00002-of-00004.safetensors",
409
+ "layers.8.linear_attn.in_proj_b.weight": "model-00002-of-00004.safetensors",
410
+ "layers.8.linear_attn.in_proj_qkv.weight": "model-00002-of-00004.safetensors",
411
+ "layers.8.linear_attn.in_proj_z.weight": "model-00002-of-00004.safetensors",
412
+ "layers.8.linear_attn.norm.weight": "model-00002-of-00004.safetensors",
413
+ "layers.8.linear_attn.out_proj.weight": "model-00002-of-00004.safetensors",
414
+ "layers.8.mlp.down_proj.weight": "model-00002-of-00004.safetensors",
415
+ "layers.8.mlp.gate_proj.weight": "model-00002-of-00004.safetensors",
416
+ "layers.8.mlp.up_proj.weight": "model-00002-of-00004.safetensors",
417
+ "layers.8.post_attention_layernorm.weight": "model-00002-of-00004.safetensors",
418
+ "layers.9.input_layernorm.weight": "model-00002-of-00004.safetensors",
419
+ "layers.9.linear_attn.A_log": "model-00002-of-00004.safetensors",
420
+ "layers.9.linear_attn.conv1d.weight": "model-00002-of-00004.safetensors",
421
+ "layers.9.linear_attn.dt_bias": "model-00002-of-00004.safetensors",
422
+ "layers.9.linear_attn.in_proj_a.weight": "model-00002-of-00004.safetensors",
423
+ "layers.9.linear_attn.in_proj_b.weight": "model-00002-of-00004.safetensors",
424
+ "layers.9.linear_attn.in_proj_qkv.weight": "model-00002-of-00004.safetensors",
425
+ "layers.9.linear_attn.in_proj_z.weight": "model-00002-of-00004.safetensors",
426
+ "layers.9.linear_attn.norm.weight": "model-00002-of-00004.safetensors",
427
+ "layers.9.linear_attn.out_proj.weight": "model-00002-of-00004.safetensors",
428
+ "layers.9.mlp.down_proj.weight": "model-00002-of-00004.safetensors",
429
+ "layers.9.mlp.gate_proj.weight": "model-00002-of-00004.safetensors",
430
+ "layers.9.mlp.up_proj.weight": "model-00002-of-00004.safetensors",
431
+ "layers.9.post_attention_layernorm.weight": "model-00002-of-00004.safetensors",
432
+ "norm.weight": "model-00004-of-00004.safetensors"
433
+ }
434
+ }
bundle-manifest.json ADDED
@@ -0,0 +1,161 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "format": "research-pointer-bundle-v1",
3
+ "status": "candidate-awaiting-public-reload-and-heldout",
4
+ "files": [
5
+ {
6
+ "file": "backbone/config.json",
7
+ "bytes": 1980,
8
+ "sha256": "4a87e4e7e11a11284a210066648b6d7616fda2872caeadec015aecb26dd9cffe"
9
+ },
10
+ {
11
+ "file": "backbone/model-00001-of-00004.safetensors",
12
+ "bytes": 3999614440,
13
+ "sha256": "1c64dafca4c6214e64d73f83ec81634773d2be3fdc85fd93f82dd53140e07e2c"
14
+ },
15
+ {
16
+ "file": "backbone/model-00002-of-00004.safetensors",
17
+ "bytes": 3997271592,
18
+ "sha256": "55b052d3f60ac4bc7cb0b168cb185899eb5ba64b2b3a98dcef9211e27a89b076"
19
+ },
20
+ {
21
+ "file": "backbone/model-00003-of-00004.safetensors",
22
+ "bytes": 3997288320,
23
+ "sha256": "1d8a356f5a18768e4dae02fea0834fee5f136a3052ec8472672f41018ec2b17b"
24
+ },
25
+ {
26
+ "file": "backbone/model-00004-of-00004.safetensors",
27
+ "bytes": 3879241544,
28
+ "sha256": "bc97b5ffb4fd5fe2f171e73420dfaf2ccb415e749ba48d5864c12c0a44af9e71"
29
+ },
30
+ {
31
+ "file": "backbone/model.safetensors.index.json",
32
+ "bytes": 33048,
33
+ "sha256": "f943816d8882f0acb572029805817240c2310d0d7e5b76fe1ececff63eab0686"
34
+ },
35
+ {
36
+ "file": "chat_template.jinja",
37
+ "bytes": 7756,
38
+ "sha256": "a4aee8afcf2e0711942cf848899be66016f8d14a889ff9ede07bca099c28f715"
39
+ },
40
+ {
41
+ "file": "code/decision_api.py",
42
+ "bytes": 10952,
43
+ "sha256": "273f6f10f22d5a68b8db34cfcbd35407fb43d8030f6d7cd188bcf118dd90a152"
44
+ },
45
+ {
46
+ "file": "code/decision_model.py",
47
+ "bytes": 10114,
48
+ "sha256": "d3e28489c09f3bd7130e2d43d92e0b5c4a08e25b09b21303904defb0ff1c3646"
49
+ },
50
+ {
51
+ "file": "code/profile_guard.py",
52
+ "bytes": 6184,
53
+ "sha256": "d3ceca1abe8f92dd378ecccff076f7a2c5438691d2fd734e9c026388d2188205"
54
+ },
55
+ {
56
+ "file": "code/runtime_profile.py",
57
+ "bytes": 2857,
58
+ "sha256": "a85281716be87788dae554386d434c8fa7dd2f22a3b01b56ee97d22cc0d30740"
59
+ },
60
+ {
61
+ "file": "decision_config.json",
62
+ "bytes": 659,
63
+ "sha256": "656b1717b8cfebad6253e9d1321023323651ad3c80327562b7d1542ec74c23d1"
64
+ },
65
+ {
66
+ "file": "decision_head.safetensors",
67
+ "bytes": 16845656,
68
+ "sha256": "66540f0bea02a117f1c0b5d7b5191735d4cde2a5b6daf1330ecefa6e10497dac"
69
+ },
70
+ {
71
+ "file": "pyproject.toml",
72
+ "bytes": 436,
73
+ "sha256": "135a9516e87ca2fb41ef54a974e29132256b6c2d528a1cbdea1001e28f306946"
74
+ },
75
+ {
76
+ "file": "runtime-profile/l2norm_fwd_kernel.json",
77
+ "bytes": 26330,
78
+ "sha256": "d7ed7c9962a48efdfa76ed695c8bdc0f56afe1397c202c936eb78315e87a26df"
79
+ },
80
+ {
81
+ "file": "runtime-profile/profile.json",
82
+ "bytes": 35237,
83
+ "sha256": "cd73a4df3146d950e900acce4e5cff7ab61f79d866cc86c77035e098ecedcb68"
84
+ },
85
+ {
86
+ "file": "runtime.json",
87
+ "bytes": 991,
88
+ "sha256": "2fffc75c6c681d7056b24ca660f95491f12816d0ed1c4a70ff3dbee3896044bb"
89
+ },
90
+ {
91
+ "file": "src/decision/__init__.py",
92
+ "bytes": 167,
93
+ "sha256": "70de37df98b6fc8e3b9f9d43935ba31a32350496214c11adf8c5fba72c433313"
94
+ },
95
+ {
96
+ "file": "src/decision/example.py",
97
+ "bytes": 3581,
98
+ "sha256": "a54dec885f92c2d38d07ef2333dff51965be52bc119f772cc74200ed314e69b6"
99
+ },
100
+ {
101
+ "file": "src/decision/model.py",
102
+ "bytes": 9415,
103
+ "sha256": "ba240d7493fc29203fe036966f0ab911200cd4a9252b50b05423409977639ee0"
104
+ },
105
+ {
106
+ "file": "src/decision_local.egg-info/PKG-INFO",
107
+ "bytes": 225,
108
+ "sha256": "cc5266721b2e02c5c963b59856d979d619f7d9d687b60bda037c7bad2f22cd1d"
109
+ },
110
+ {
111
+ "file": "src/decision_local.egg-info/SOURCES.txt",
112
+ "bytes": 361,
113
+ "sha256": "741f317b6e9b27453e8f2fbeb76c0d1da326e89de7978fe87ce906ed1e2a68a9"
114
+ },
115
+ {
116
+ "file": "src/decision_local.egg-info/dependency_links.txt",
117
+ "bytes": 1,
118
+ "sha256": "01ba4719c80b6fe911b091a7c05124b64eeece964e09c058ef8f9805daca546b"
119
+ },
120
+ {
121
+ "file": "src/decision_local.egg-info/entry_points.txt",
122
+ "bytes": 59,
123
+ "sha256": "5400ff8d993d37849bc01b702e4674b8d9c4122af514101ace0bcd691eab39fb"
124
+ },
125
+ {
126
+ "file": "src/decision_local.egg-info/requires.txt",
127
+ "bytes": 31,
128
+ "sha256": "b8bf334329a333c3bba85388fbf41377b17305ee71b402c431c1af7445eadc08"
129
+ },
130
+ {
131
+ "file": "src/decision_local.egg-info/top_level.txt",
132
+ "bytes": 9,
133
+ "sha256": "834608781b00d5df58e3ae55a6b15c201aeebacc26ec99a1035fd75b83e16760"
134
+ },
135
+ {
136
+ "file": "temperature.json",
137
+ "bytes": 1138,
138
+ "sha256": "8e04e0b07c179ee81d46b6fc989fc9f8b76dad4c9e87b0ffe0ed9b7f132f575a"
139
+ },
140
+ {
141
+ "file": "tokenizer.json",
142
+ "bytes": 19989325,
143
+ "sha256": "06b9509352d2af50381ab2247e083b80d32d5c0aba91c272ca9ff729b6a0e523"
144
+ },
145
+ {
146
+ "file": "tokenizer_config.json",
147
+ "bytes": 1123,
148
+ "sha256": "bee8eba30f0eb4af73c0fe2cd06d0f89b657d7819941c438157ec42f7c80ea87"
149
+ }
150
+ ],
151
+ "production_batch_size": 8,
152
+ "input_length_limit": 16384,
153
+ "source_model_manifest_sha256": "7c3a8e778011843fd8e2e44b3c641a23fb9d03145a4655b6290e17025b83643a",
154
+ "selection_ranking_sha256": "6fb02c05f33db0dde22b702ec1cd7f10a6419aea50a1ca9d259993a72ba9f42d",
155
+ "calibration_sha256": "d058d6e287b4c17d4adbb7b0bbc2ed09a90efb4d4e880d2bb24871d2b1528bb2",
156
+ "base_revision": "c202236235762e1c871ad0ccb60c8ee5ba337b9a",
157
+ "body_parameters": 7936684544,
158
+ "head_parameters": 4211200,
159
+ "public_wrapper_included": true,
160
+ "normalization_profile_sha256": "cd73a4df3146d950e900acce4e5cff7ab61f79d866cc86c77035e098ecedcb68"
161
+ }
chat_template.jinja ADDED
@@ -0,0 +1,154 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {%- set image_count = namespace(value=0) %}
2
+ {%- set video_count = namespace(value=0) %}
3
+ {%- macro render_content(content, do_vision_count, is_system_content=false) %}
4
+ {%- if content is string %}
5
+ {{- content }}
6
+ {%- elif content is iterable and content is not mapping %}
7
+ {%- for item in content %}
8
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
9
+ {%- if is_system_content %}
10
+ {{- raise_exception('System message cannot contain images.') }}
11
+ {%- endif %}
12
+ {%- if do_vision_count %}
13
+ {%- set image_count.value = image_count.value + 1 %}
14
+ {%- endif %}
15
+ {%- if add_vision_id %}
16
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
17
+ {%- endif %}
18
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
19
+ {%- elif 'video' in item or item.type == 'video' %}
20
+ {%- if is_system_content %}
21
+ {{- raise_exception('System message cannot contain videos.') }}
22
+ {%- endif %}
23
+ {%- if do_vision_count %}
24
+ {%- set video_count.value = video_count.value + 1 %}
25
+ {%- endif %}
26
+ {%- if add_vision_id %}
27
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
28
+ {%- endif %}
29
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
30
+ {%- elif 'text' in item %}
31
+ {{- item.text }}
32
+ {%- else %}
33
+ {{- raise_exception('Unexpected item type in content.') }}
34
+ {%- endif %}
35
+ {%- endfor %}
36
+ {%- elif content is none or content is undefined %}
37
+ {{- '' }}
38
+ {%- else %}
39
+ {{- raise_exception('Unexpected content type.') }}
40
+ {%- endif %}
41
+ {%- endmacro %}
42
+ {%- if not messages %}
43
+ {{- raise_exception('No messages provided.') }}
44
+ {%- endif %}
45
+ {%- if tools and tools is iterable and tools is not mapping %}
46
+ {{- '<|im_start|>system\n' }}
47
+ {{- "# Tools\n\nYou have access to the following functions:\n\n<tools>" }}
48
+ {%- for tool in tools %}
49
+ {{- "\n" }}
50
+ {{- tool | tojson }}
51
+ {%- endfor %}
52
+ {{- "\n</tools>" }}
53
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n<tool_call>\n<function=example_function_name>\n<parameter=example_parameter_1>\nvalue_1\n</parameter>\n<parameter=example_parameter_2>\nThis is the value for the second parameter\nthat can span\nmultiple lines\n</parameter>\n</function>\n</tool_call>\n\n<IMPORTANT>\nReminder:\n- Function calls MUST follow the specified format: an inner <function=...></function> block must be nested within <tool_call></tool_call> XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n</IMPORTANT>' }}
54
+ {%- if messages[0].role == 'system' %}
55
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
56
+ {%- if content %}
57
+ {{- '\n\n' + content }}
58
+ {%- endif %}
59
+ {%- endif %}
60
+ {{- '<|im_end|>\n' }}
61
+ {%- else %}
62
+ {%- if messages[0].role == 'system' %}
63
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
64
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
65
+ {%- endif %}
66
+ {%- endif %}
67
+ {%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
68
+ {%- for message in messages[::-1] %}
69
+ {%- set index = (messages|length - 1) - loop.index0 %}
70
+ {%- if ns.multi_step_tool and message.role == "user" %}
71
+ {%- set content = render_content(message.content, false)|trim %}
72
+ {%- if not(content.startswith('<tool_response>') and content.endswith('</tool_response>')) %}
73
+ {%- set ns.multi_step_tool = false %}
74
+ {%- set ns.last_query_index = index %}
75
+ {%- endif %}
76
+ {%- endif %}
77
+ {%- endfor %}
78
+ {%- if ns.multi_step_tool %}
79
+ {{- raise_exception('No user query found in messages.') }}
80
+ {%- endif %}
81
+ {%- for message in messages %}
82
+ {%- set content = render_content(message.content, true)|trim %}
83
+ {%- if message.role == "system" %}
84
+ {%- if not loop.first %}
85
+ {{- raise_exception('System message must be at the beginning.') }}
86
+ {%- endif %}
87
+ {%- elif message.role == "user" %}
88
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
89
+ {%- elif message.role == "assistant" %}
90
+ {%- set reasoning_content = '' %}
91
+ {%- if message.reasoning_content is string %}
92
+ {%- set reasoning_content = message.reasoning_content %}
93
+ {%- else %}
94
+ {%- if '</think>' in content %}
95
+ {%- set reasoning_content = content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
96
+ {%- set content = content.split('</think>')[-1].lstrip('\n') %}
97
+ {%- endif %}
98
+ {%- endif %}
99
+ {%- set reasoning_content = reasoning_content|trim %}
100
+ {%- if loop.index0 > ns.last_query_index %}
101
+ {{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content + '\n</think>\n\n' + content }}
102
+ {%- else %}
103
+ {{- '<|im_start|>' + message.role + '\n' + content }}
104
+ {%- endif %}
105
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
106
+ {%- for tool_call in message.tool_calls %}
107
+ {%- if tool_call.function is defined %}
108
+ {%- set tool_call = tool_call.function %}
109
+ {%- endif %}
110
+ {%- if loop.first %}
111
+ {%- if content|trim %}
112
+ {{- '\n\n<tool_call>\n<function=' + tool_call.name + '>\n' }}
113
+ {%- else %}
114
+ {{- '<tool_call>\n<function=' + tool_call.name + '>\n' }}
115
+ {%- endif %}
116
+ {%- else %}
117
+ {{- '\n<tool_call>\n<function=' + tool_call.name + '>\n' }}
118
+ {%- endif %}
119
+ {%- if tool_call.arguments is defined %}
120
+ {%- for args_name, args_value in tool_call.arguments|items %}
121
+ {{- '<parameter=' + args_name + '>\n' }}
122
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
123
+ {{- args_value }}
124
+ {{- '\n</parameter>\n' }}
125
+ {%- endfor %}
126
+ {%- endif %}
127
+ {{- '</function>\n</tool_call>' }}
128
+ {%- endfor %}
129
+ {%- endif %}
130
+ {{- '<|im_end|>\n' }}
131
+ {%- elif message.role == "tool" %}
132
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
133
+ {{- '<|im_start|>user' }}
134
+ {%- endif %}
135
+ {{- '\n<tool_response>\n' }}
136
+ {{- content }}
137
+ {{- '\n</tool_response>' }}
138
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
139
+ {{- '<|im_end|>\n' }}
140
+ {%- elif loop.last %}
141
+ {{- '<|im_end|>\n' }}
142
+ {%- endif %}
143
+ {%- else %}
144
+ {{- raise_exception('Unexpected message role.') }}
145
+ {%- endif %}
146
+ {%- endfor %}
147
+ {%- if add_generation_prompt %}
148
+ {{- '<|im_start|>assistant\n' }}
149
+ {%- if enable_thinking is defined and enable_thinking is false %}
150
+ {{- '<think>\n\n</think>\n\n' }}
151
+ {%- else %}
152
+ {{- '<think>\n' }}
153
+ {%- endif %}
154
+ {%- endif %}
code/decision_api.py ADDED
@@ -0,0 +1,198 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Typed local inference adapter for the research decision checkpoints.
2
+
3
+ The response schema resembles TypeSafe's primitives. Confidence uses this
4
+ implementation's documented normalized maximum probability, not a claimed
5
+ reimplementation of TypeSafe's unpublished statistic. No text generation.
6
+ """
7
+ from __future__ import annotations
8
+ import importlib.util
9
+ import math
10
+ from pathlib import Path
11
+
12
+
13
+ def prepare_runtime_profile(checkpoint, device='cuda:0'):
14
+ # This profile is verified before any dependency import can choose kernels.
15
+ import hashlib, json, sys
16
+ root=Path(checkpoint);runtime=json.loads((root/'runtime.json').read_text())
17
+ spec=runtime.get('normalization_profile')
18
+ if spec is None:
19
+ if '_decision_process_normalization_profile_v1' in sys.modules:
20
+ raise RuntimeError('Use separate processes for profiled and unprofiled models')
21
+ return None
22
+ import torch
23
+ target=torch.device(device)
24
+ if target.type!='cuda' or not torch.cuda.is_available():
25
+ raise RuntimeError('The bound profile requires a ROCm CUDA device')
26
+ arch=getattr(torch.cuda.get_device_properties(target),'gcnArchName','').split(':')[0]
27
+ if arch!=spec['validated_arch']:
28
+ raise RuntimeError('Target GPU architecture does not match the bound profile: '+arch)
29
+ relative=Path(spec['loader_file'])
30
+ if relative.is_absolute() or '..' in relative.parts:raise ValueError('Unsafe profile loader path')
31
+ path=root/relative
32
+ if hashlib.sha256(path.read_bytes()).hexdigest()!=spec['loader_sha256']:
33
+ raise ValueError('Bound runtime profile loader changed')
34
+ definition=importlib.util.spec_from_file_location('decision_bundle_runtime_profile',path)
35
+ module=importlib.util.module_from_spec(definition);definition.loader.exec_module(module)
36
+ return module.ensure_profile(root)
37
+
38
+
39
+ def question_row(state, name, question):
40
+ kind=question.get('type')
41
+ if kind not in {'choice','noul','score'}:raise ValueError('Unknown question type')
42
+ if 'instructions' not in question:raise ValueError('instructions is required')
43
+ criteria=question.get('criteria')
44
+ if kind=='noul':
45
+ criteria={} if criteria is None else criteria
46
+ if not isinstance(criteria,dict) or set(criteria)-{'true','false'}:
47
+ raise ValueError('noul criteria may contain only true and false')
48
+ options=[{'key':'false','description':criteria.get('false','The answer to the question is no.')},
49
+ {'key':'true','description':criteria.get('true','The answer to the question is yes.')}]
50
+ elif kind=='score':
51
+ if not isinstance(criteria,list) or not 2<=len(criteria)<=10:
52
+ raise ValueError('score requires an ordered list of 2..10 criteria')
53
+ options=[{'key':str(i),'description':value} for i,value in enumerate(criteria)]
54
+ else:
55
+ if not isinstance(criteria,dict) or not 2<=len(criteria)<=255:
56
+ raise ValueError('choice requires a mapping of 2..255 criteria')
57
+ if not all(isinstance(k,str) for k in criteria):raise ValueError('Choice keys must be strings')
58
+ options=[{'key':key,'description':value} for key,value in criteria.items()]
59
+ # The question name is used for bookkeeping only; encoders never render id.
60
+ return {'id':name,'state':state,'instructions':question['instructions'],
61
+ 'options':options,'task_type':kind,'family':'inference'}
62
+
63
+
64
+ def typed_answer(row, probabilities):
65
+ p=[float(v) for v in probabilities];k=len(row['options'])
66
+ if len(p)!=k or any(not math.isfinite(v) or v<0 for v in p):
67
+ raise ValueError('Invalid probability vector')
68
+ total=sum(p)
69
+ if total<=0 or abs(total-1)>1e-4:raise ValueError('Probabilities must sum to one')
70
+ p=[v/total for v in p];selected=max(range(k),key=p.__getitem__)
71
+ kind=row['task_type']
72
+ if kind=='noul':
73
+ keys=[o['key'] for o in row['options']]
74
+ if set(keys)!={'false','true'}:raise ValueError('Native noul rows require false/true keys')
75
+ return {'type':'noul','noul':p[keys.index('true')]}
76
+ answer={'type':kind,'probabilities':{o['key']:v for o,v in zip(row['options'],p)},
77
+ 'confidence':max(0.,min(1.,(k*max(p)-1)/(k-1)))}
78
+ if kind=='choice':answer['choice']=row['options'][selected]['key']
79
+ else:
80
+ if [o['key'] for o in row['options']] != [str(i) for i in range(k)]:
81
+ raise ValueError('Native score rows require ordered numeric level keys')
82
+ answer['score']=sum(i*v for i,v in enumerate(p))
83
+ answer['legend']={str(i):o['description'] for i,o in enumerate(row['options'])}
84
+ return answer
85
+
86
+
87
+ class DecisionEngine:
88
+ def __init__(self, checkpoint, model_code, *, device='cuda:0', max_length=16384,
89
+ batch_size=8, temperatures=None, model_name='local-decision-research'):
90
+ self.normalization_profile=prepare_runtime_profile(checkpoint, device=device)
91
+ import torch
92
+ path=Path(model_code)/'decision_model.py'
93
+ spec=importlib.util.spec_from_file_location('research_decision_runtime',path)
94
+ module=importlib.util.module_from_spec(spec);spec.loader.exec_module(module)
95
+ model,tokenizer=module.DecisionModel.from_checkpoint(checkpoint,dtype=torch.bfloat16)
96
+ self.model=model.to(device).eval();self.tokenizer=tokenizer;self.module=module
97
+ self.device=device;self.max_length=max_length;self.batch_size=batch_size
98
+ self.temperatures=temperatures or {};self.model_name=model_name
99
+ if batch_size<1 or max_length<1:raise ValueError('Positive batch_size/max_length required')
100
+ if any(not math.isfinite(v) or v<=0 for v in self.temperatures.values()):
101
+ raise ValueError('Temperatures must be finite positive numbers')
102
+
103
+ def predict_rows(self, rows):
104
+ import torch
105
+ encoded=encode_request(rows,self.tokenizer,self.module,self.max_length)
106
+ pad=self.tokenizer.pad_token_id if self.tokenizer.pad_token_id is not None else self.tokenizer.eos_token_id
107
+ records=[]
108
+ with torch.inference_mode():
109
+ for start in range(0,len(rows),self.batch_size):
110
+ items=encoded[start:start+self.batch_size]
111
+ batch={key:value.to(self.device) if torch.is_tensor(value) else value
112
+ for key,value in self.module.collate(items,pad).items()}
113
+ with torch.autocast('cuda',dtype=torch.bfloat16):logits=self.model(**batch)
114
+ # Preserve each row's original float/temperature/softmax math,
115
+ # but defer host synchronization until the complete batch.
116
+ staged=[];transfers=[]
117
+ for row,item,values in zip(rows[start:start+self.batch_size],items,logits):
118
+ k=len(row['options']);values=values[:k].float()
119
+ temperature=self.temperatures.get(row['task_type'],1.)
120
+ probabilities=(values/temperature).softmax(-1)
121
+ staged.append((row,item,k,temperature))
122
+ transfers.extend((values,probabilities))
123
+ host_values=torch.cat(transfers).tolist()
124
+ offset=0
125
+ for row,item,k,temperature in staged:
126
+ values=host_values[offset:offset+k]
127
+ probabilities=host_values[offset+k:offset+2*k]
128
+ offset+=2*k
129
+ answer=typed_answer(row,probabilities)
130
+ prediction=max(range(k),key=probabilities.__getitem__)
131
+ if row['task_type']=='noul':
132
+ chosen='true' if answer['noul']>=.5 else 'false'
133
+ prediction=[o['key'] for o in row['options']].index(chosen)
134
+ rec={'id':row['id'],'status':'ok','prediction':prediction,
135
+ 'probabilities':probabilities,'logits':values,'temperature':temperature,
136
+ 'native_contract':True,'truncated':False,'input_tokens':len(item['ids']),
137
+ 'prompt_sha256':item['prompt_sha256'],'answer':answer}
138
+ if row['task_type']=='noul':rec['native_noul']=answer['noul']
139
+ if row['task_type']=='score':rec['native_score']=answer['score']
140
+ records.append(rec)
141
+ return records
142
+
143
+ def decide(self, state, questions):
144
+ if not isinstance(questions,dict) or not questions:
145
+ raise ValueError('questions must be a nonempty mapping')
146
+ if not all(isinstance(name,str) for name in questions):raise ValueError('Question names must be strings')
147
+ rows=[question_row(state,name,q) for name,q in questions.items()]
148
+ result=self.predict_rows(rows)
149
+ return {'model':self.model_name,'answers':{r['id']:r['answer'] for r in result},
150
+ 'usage':{'input_tokens':sum(r['input_tokens'] for r in result),'scored_questions':len(result)}}
151
+
152
+
153
+ """Experimental request-local exact-segment tokenization.
154
+
155
+ Original encode/segments functions remain authoritative. Batch tokenize exact
156
+ whole segments, never split a BPE prefix at a new boundary. No cross-request
157
+ cache, GPU change, prompt change or change to the eight-row inference groups.
158
+ """
159
+
160
+
161
+ class SegmentLookup:
162
+ def __init__(self, tokenizer, cache):
163
+ self.tokenizer, self.cache = tokenizer, cache
164
+
165
+ def encode(self, text, **kwargs):
166
+ if kwargs == {'add_special_tokens': False} and text in self.cache:
167
+ # Original encode extends its prefix list in place.
168
+ return list(self.cache[text])
169
+ return self.tokenizer.encode(text, **kwargs)
170
+
171
+
172
+ def encode_request(rows, tokenizer, module, max_length=16384,
173
+ max_cached_characters=8_000_000, segment_batch_size=64):
174
+ if max_cached_characters < 0 or segment_batch_size < 1:
175
+ raise ValueError('Invalid tokenizer resource bound')
176
+ unique = {}
177
+ characters = 0
178
+ for row in rows:
179
+ prefix, options, suffix = module.segments(row)
180
+ for segment in (prefix, *options, suffix):
181
+ if segment not in unique:
182
+ unique[segment] = None
183
+ characters += len(segment)
184
+ if characters > max_cached_characters:
185
+ # Preserve the original behavior under the resource cap.
186
+ return [module.encode(r, tokenizer, max_length) for r in rows]
187
+ strings = list(unique)
188
+ for start in range(0, len(strings), segment_batch_size):
189
+ batch = strings[start:start + segment_batch_size]
190
+ result = tokenizer(batch, add_special_tokens=False, padding=False,
191
+ truncation=False, return_attention_mask=False,
192
+ return_token_type_ids=False)['input_ids']
193
+ if len(result) != len(batch):
194
+ raise ValueError('Batch tokenizer output count differs')
195
+ for segment, ids in zip(batch, result):
196
+ unique[segment] = tuple(ids)
197
+ lookup = SegmentLookup(tokenizer, unique)
198
+ return [module.encode(row, lookup, max_length) for row in rows]
code/decision_model.py ADDED
@@ -0,0 +1,173 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Dynamic candidate readout over a causal Qwen3.5 text backbone.
2
+
3
+ Candidate endpoints retain their contextual vectors. A final global-query
4
+ vector can incorporate all options before a shared bilinear + MLP scorer
5
+ scores every candidate. This is a research architecture, not a Jev claim.
6
+ """
7
+ import hashlib
8
+ import json
9
+ import math
10
+ from pathlib import Path
11
+
12
+ import torch
13
+ from torch import nn
14
+ import torch.nn.functional as F
15
+ from safetensors.torch import load_file, save_file
16
+ from transformers import AutoTokenizer, Qwen3_5ForConditionalGeneration
17
+ from transformers.models.qwen3_5.modeling_qwen3_5 import Qwen3_5TextModel
18
+
19
+ PROMPT_VERSION = "structured-segmented-candidate-endpoints-global-query-v2"
20
+ MAX_OPTIONS = 255
21
+
22
+
23
+ def canonical(value):
24
+ return json.dumps(value, ensure_ascii=False, sort_keys=True, separators=(",", ":"))
25
+
26
+
27
+ def payload(value):
28
+ return value if isinstance(value, str) else canonical(value)
29
+
30
+
31
+ def segments(row):
32
+ opts = row["options"]
33
+ if not 2 <= len(opts) <= MAX_OPTIONS:
34
+ raise ValueError(f"{row['id']}: expected 2..255 options")
35
+ if not all(isinstance(o["key"], str) for o in opts):
36
+ raise ValueError("Option keys must be strings")
37
+ if len({o["key"] for o in opts}) != len(opts):
38
+ raise ValueError("Duplicate option keys")
39
+ prefix = f"Context:\n{payload(row['state'])}\n\nTask type: {row.get('task_type', 'choice')}\nQuestion:\n{payload(row['instructions'])}\nOptions:"
40
+ # Tokenize each part separately. This deliberately fixes boundaries and
41
+ # avoids guessing endpoint indices from merged BPE character offsets.
42
+ options = ["\n<option>\n" + canonical({"key": o["key"], "description": o.get("description")}) + "\n</option>" for o in opts]
43
+ suffix = "\n\nSelect the single option best supported by the context and instructions.\nDecision:"
44
+ return prefix, options, suffix
45
+
46
+
47
+ def render(row):
48
+ prefix, opts, suffix = segments(row)
49
+ return prefix + "".join(opts) + suffix
50
+
51
+
52
+ def encode(row, tokenizer, max_length=16384):
53
+ prefix, opts, suffix = segments(row)
54
+ ids = tokenizer.encode(prefix, add_special_tokens=False)
55
+ candidate_positions = []
56
+ for option in opts:
57
+ part = tokenizer.encode(option, add_special_tokens=False)
58
+ if not part:
59
+ raise ValueError("Empty tokenized candidate")
60
+ ids.extend(part)
61
+ candidate_positions.append(len(ids) - 1)
62
+ ids.extend(tokenizer.encode(suffix, add_special_tokens=False))
63
+ if len(ids) > max_length:
64
+ raise ValueError(f"{row['id']}: {len(ids)} tokens exceeds max_length={max_length}; no truncation allowed")
65
+ label = row.get("label", -1)
66
+ if label != -1 and not 0 <= label < len(opts):
67
+ raise ValueError("Invalid label")
68
+ prompt = prefix + "".join(opts) + suffix
69
+ return {"id": row["id"], "ids": ids, "label": label, "nopts": len(opts), "family": row.get("family", "unspecified"), "candidate_positions": candidate_positions, "query_position": len(ids) - 1, "target_probs": row.get("target_probs"), "prompt_sha256": hashlib.sha256(prompt.encode()).hexdigest(), "token_ids_sha256": hashlib.sha256(canonical(ids).encode()).hexdigest(), "segmented_tokenization": True}
70
+
71
+
72
+ def collate(items, pad_id):
73
+ length = ((max(len(x["ids"]) for x in items) + 31) // 32) * 32
74
+ nopts = max(x["nopts"] for x in items)
75
+ ids = torch.full((len(items), length), pad_id, dtype=torch.long)
76
+ mask = torch.zeros_like(ids)
77
+ positions = torch.zeros((len(items), nopts), dtype=torch.long)
78
+ candidate_mask = torch.zeros((len(items), nopts), dtype=torch.bool)
79
+ for i, item in enumerate(items):
80
+ if len(item["candidate_positions"]) != item["nopts"]:
81
+ raise ValueError("Candidate count does not match endpoint count")
82
+ if not all(0 <= p < item["query_position"] < len(item["ids"]) for p in item["candidate_positions"]):
83
+ raise ValueError("Candidate endpoints must precede global query")
84
+ if len(set(item["candidate_positions"])) != item["nopts"]:
85
+ raise ValueError("Duplicate candidate endpoint")
86
+ ids[i, :len(item["ids"])] = torch.tensor(item["ids"])
87
+ mask[i, :len(item["ids"])] = 1
88
+ positions[i, :item["nopts"]] = torch.tensor(item["candidate_positions"])
89
+ candidate_mask[i, :item["nopts"]] = True
90
+ return {"input_ids": ids, "attention_mask": mask, "candidate_positions": positions, "candidate_mask": candidate_mask, "query_positions": torch.tensor([x["query_position"] for x in items]), "labels": torch.tensor([x["label"] for x in items]), "nopts": torch.tensor([x["nopts"] for x in items]), "ids": [x["id"] for x in items], "families": [x["family"] for x in items]}
91
+
92
+
93
+ class CandidateHead(nn.Module):
94
+ def __init__(self, hidden_size, head_dim=256):
95
+ super().__init__()
96
+ self.head_dim = head_dim
97
+ self.candidate_norm = nn.LayerNorm(hidden_size)
98
+ self.query_norm = nn.LayerNorm(hidden_size)
99
+ self.key = nn.Linear(hidden_size, head_dim, bias=False)
100
+ self.query = nn.Linear(hidden_size, head_dim, bias=False)
101
+ self.candidate_mlp = nn.Linear(hidden_size, head_dim, bias=True)
102
+ self.query_mlp = nn.Linear(hidden_size, head_dim, bias=False)
103
+ self.scalar = nn.Linear(head_dim, 1, bias=False)
104
+ nn.init.normal_(self.scalar.weight, mean=0., std=0.01)
105
+
106
+ def forward(self, candidates, query):
107
+ # Keep the small shared head in FP32 even when the backbone uses BF16.
108
+ # The v1 letter head exhibited BF16 ties sensitive to batch padding.
109
+ with torch.autocast(device_type=candidates.device.type, enabled=False):
110
+ c = self.candidate_norm(candidates.float())
111
+ q = self.query_norm(query.float())
112
+ bilinear = (self.key(c) * self.query(q)[:, None, :]).sum(-1) / math.sqrt(self.head_dim)
113
+ interaction = self.scalar(F.gelu(self.candidate_mlp(c) + self.query_mlp(q)[:, None, :])).squeeze(-1)
114
+ return bilinear + interaction
115
+
116
+
117
+ class DecisionModel(nn.Module):
118
+ def __init__(self, backbone, head, metadata):
119
+ super().__init__()
120
+ self.backbone, self.head, self.metadata = backbone, head, metadata
121
+
122
+ @classmethod
123
+ def from_base(cls, path, revision="local", dtype=torch.bfloat16, attention="sdpa", head_dim=256):
124
+ tokenizer = AutoTokenizer.from_pretrained(path, local_files_only=True)
125
+ full, info = Qwen3_5ForConditionalGeneration.from_pretrained(path, dtype=dtype, local_files_only=True, attn_implementation=attention, output_loading_info=True)
126
+ if any(info.get(k) for k in ("missing_keys", "mismatched_keys", "error_msgs")):
127
+ raise RuntimeError(f"Incomplete base loading: {info}")
128
+ backbone = full.model.language_model
129
+ backbone.config.use_cache = False
130
+ head = CandidateHead(backbone.config.hidden_size, head_dim)
131
+ metadata = {"base_revision": revision, "text_parameter_count": sum(p.numel() for p in backbone.parameters()), "prompt_version": PROMPT_VERSION, "attention": attention, "head_dim": head_dim, "max_options": MAX_OPTIONS, "architecture": "contextual-candidate-endpoint-plus-global-query-shared-bilinear-mlp", "head_initialization": "random-shared-content-scorer", "head_precision": "float32-outside-autocast"}
132
+ return cls(backbone, head, metadata), tokenizer
133
+
134
+ @classmethod
135
+ def from_decision_checkpoint(cls, path, dtype=torch.bfloat16, attention="sdpa", head_dim=256):
136
+ """Warm-start the backbone of a trained v1 model; initialize a new head."""
137
+ path = Path(path)
138
+ metadata = json.loads((path / "decision_config.json").read_text())
139
+ backbone = Qwen3_5TextModel.from_pretrained(path / "backbone", dtype=dtype, local_files_only=True, attn_implementation=attention)
140
+ backbone.config.use_cache = False
141
+ head = CandidateHead(backbone.config.hidden_size, head_dim)
142
+ metadata.update({"prompt_version": PROMPT_VERSION, "head_dim": head_dim, "max_options": MAX_OPTIONS, "architecture": "contextual-candidate-endpoint-plus-global-query-shared-bilinear-mlp", "head_initialization": "random-shared-content-scorer", "warm_start": "trained-v1-text-backbone", "head_precision": "float32-outside-autocast"})
143
+ return cls(backbone, head, metadata), AutoTokenizer.from_pretrained(path, local_files_only=True)
144
+
145
+ @classmethod
146
+ def from_checkpoint(cls, path, dtype=torch.bfloat16, attention="sdpa"):
147
+ path = Path(path)
148
+ metadata = json.loads((path / "decision_config.json").read_text())
149
+ if metadata["prompt_version"] != PROMPT_VERSION:
150
+ raise ValueError("Not a pointer-v2 checkpoint; use from_decision_checkpoint for warm start")
151
+ backbone = Qwen3_5TextModel.from_pretrained(path / "backbone", dtype=dtype, local_files_only=True, attn_implementation=attention)
152
+ head = CandidateHead(backbone.config.hidden_size, metadata["head_dim"])
153
+ head.load_state_dict(load_file(path / "decision_head.safetensors"))
154
+ return cls(backbone, head, metadata), AutoTokenizer.from_pretrained(path, local_files_only=True)
155
+
156
+ def forward(self, input_ids, attention_mask, candidate_positions, candidate_mask, query_positions, **unused):
157
+ hidden = self.backbone(input_ids=input_ids, attention_mask=attention_mask, use_cache=False).last_hidden_state
158
+ batches = torch.arange(hidden.shape[0], device=hidden.device)
159
+ candidates = hidden[batches[:, None], candidate_positions]
160
+ query = hidden[batches, query_positions]
161
+ scores = self.head(candidates, query).float()
162
+ return scores.masked_fill(~candidate_mask, -float("inf"))
163
+
164
+ def save(self, path, tokenizer):
165
+ path = Path(path); path.mkdir(parents=True, exist_ok=True)
166
+ self.backbone.save_pretrained(path / "backbone", safe_serialization=True, max_shard_size="4GB")
167
+ save_file({n: v.detach().cpu().contiguous() for n, v in self.head.state_dict().items()}, str(path / "decision_head.safetensors"))
168
+ tokenizer.save_pretrained(path)
169
+ (path / "decision_config.json").write_text(json.dumps(self.metadata, indent=2) + "\n")
170
+
171
+
172
+ def classification_loss(logits, labels):
173
+ return F.cross_entropy(logits, labels)
code/profile_guard.py ADDED
@@ -0,0 +1,82 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Fail-closed official FLA strict-config setup for one isolated process.
2
+
3
+ No model/prompt/head changes. The guard prevents FLA STRICT's ordinary missing
4
+ key fallback and records actual configured calls. This is a diagnostic module,
5
+ not an installed change to the published wrapper or dependency environment.
6
+ """
7
+ import hashlib,importlib,json,os,sys
8
+ from pathlib import Path
9
+
10
+ def sha(p):return hashlib.sha256(Path(p).read_bytes()).hexdigest()
11
+ def serialized(key):return json.dumps(key,separators=(',',':'),sort_keys=True)
12
+ def config_fields(config):
13
+ if isinstance(config,dict):return {k:config.get(k) for k in ['kwargs','num_warps','num_stages','num_ctas','maxnreg','ir_override']}
14
+ return {k:getattr(config,k,None) for k in ['kwargs','num_warps','num_stages','num_ctas','maxnreg','ir_override']}
15
+ def validate_profile(path,expected_sha):
16
+ path=Path(path).resolve()
17
+ if sha(path)!=expected_sha:raise ValueError('Profile hash changed')
18
+ profile=json.loads(path.read_text())
19
+ if profile['format']!='decision-fla-l2norm-profile-v1' or profile.get('model_family')!='Qwen/Qwen3.5-9B' or profile['cache_mode']!='strict':raise ValueError('Profile format/mode unsupported')
20
+ if len(profile['files'])!=1 or profile['files'][0]['file']!='l2norm_fwd_kernel.json':raise ValueError('Unexpected profile file set')
21
+ f=path.parent/'l2norm_fwd_kernel.json'
22
+ if sha(f)!=profile['files'][0]['sha256']:raise ValueError('Explicit kernel config changed')
23
+ data=json.loads(f.read_text());entries={}
24
+ if data.get('default_config') is not None:raise ValueError('Implicit fallback defaults forbidden')
25
+ for h,item in data['autotune_entries'].items():
26
+ key=item['autotune_key'];encoded=serialized(key)
27
+ if hashlib.md5(encoded.encode()).hexdigest()!=h or encoded in entries:raise ValueError('Invalid/duplicate key')
28
+ if len(key)!=5 or key[0]!=128 or type(key[1]) is not int or not 1<=key[1]<=64 or key[2:]!=['torch.bfloat16','torch.bfloat16','torch.float32']:raise ValueError('Unsupported numerical key')
29
+ c=item['config']
30
+ if c['kwargs'].keys()!={'BT'} or c['kwargs']['BT'] not in [8,16,32,64] or c['num_warps'] not in [1,2,4,8,16] or c['num_stages']!=3 or c['num_ctas']!=1 or any(c.get(x) is not None for x in ['maxnreg','pre_hook','ir_override']):raise ValueError('Unexpected launch configuration')
31
+ entries[encoded]=c
32
+ if {json.loads(k)[1] for k in entries}!=set(range(1,65)):raise ValueError('Incomplete legal NB coverage')
33
+ return path,profile,entries
34
+
35
+ def attach_guard(kernel,cache_module,entries,telemetry):
36
+ original=kernel.run
37
+ def guarded(*args,**kwargs):
38
+ if cache_module.FLA_CACHE_MODE is not cache_module.FlaCacheMode.STRICT:raise RuntimeError('FLA strict mode changed')
39
+ key=cache_module.AutotuneKey.build(kernel.arg_names,kernel.keys,args,kwargs);encoded=serialized(list(key.autotune_key))
40
+ if encoded not in entries:raise RuntimeError('Uncontracted FLA l2norm key: '+encoded)
41
+ expected=entries[encoded];loaded=cache_module.load_cached_config(kernel.kernel_name,key)
42
+ if config_fields(loaded)!=config_fields(expected):raise RuntimeError('FLA exact config lookup mismatch')
43
+ if key.autotune_key in kernel.cache and config_fields(kernel.cache[key.autotune_key])!=config_fields(expected):raise RuntimeError('A conflicting in-process kernel cache exists')
44
+ # Explicit official configuration load guarantees that the following original
45
+ # run finds this exact cache entry and cannot perform timing-based autotune.
46
+ kernel.maybe_load_cached_config(key)
47
+ if key.autotune_key not in kernel.cache or config_fields(kernel.cache[key.autotune_key])!=config_fields(expected):raise RuntimeError('Official strict config did not load')
48
+ result=original(*args,**kwargs)
49
+ if config_fields(kernel.cache[key.autotune_key])!=config_fields(expected):raise RuntimeError('Kernel config changed during call')
50
+ telemetry['calls']+=1;telemetry['keys'][encoded]=telemetry['keys'].get(encoded,0)+1
51
+ return result
52
+ kernel.run=guarded
53
+ return original
54
+
55
+ def install(profile_path,expected_sha):
56
+ path,profile,entries=validate_profile(profile_path,expected_sha)
57
+ if any(n=='fla' or n.startswith('fla.') for n in sys.modules):raise RuntimeError('Install profile before importing FLA; use a fresh isolated process')
58
+ for name,wanted in {'FLA_CACHE_MODE':'strict','FLA_CONFIG_DIR':str(path.parent)}.items():
59
+ actual=os.environ.get(name)
60
+ if actual is not None and actual!=wanted:raise RuntimeError('Conflicting '+name)
61
+ os.environ[name]=wanted
62
+ import torch,triton,fla
63
+ actual={'torch':str(torch.__version__),'hip':torch.version.hip,'triton':triton.__version__,'fla':fla.__version__}
64
+ for name,value in actual.items():
65
+ if value!=profile['runtime'][name]:raise RuntimeError('Runtime mismatch: '+name)
66
+ if not torch.cuda.is_available():raise RuntimeError('Profile is only qualified for the specified ROCm GPU')
67
+ arch=torch.cuda.get_device_properties(0).gcnArchName.split(':')[0]
68
+ if arch!=profile['runtime']['gpu_arch']:raise RuntimeError('Unsupported GPU architecture '+arch)
69
+ module=importlib.import_module('fla.modules.l2norm');cache_module=importlib.import_module('fla.ops.utils.cache');root=Path(fla.__file__).parent
70
+ for name,value in profile['fla_source_sha256'].items():
71
+ if sha(root/name)!=value:raise RuntimeError('Pinned FLA source changed: '+name)
72
+ kernel=module.l2norm_fwd_kernel
73
+ if kernel.kernel_name!='l2norm_fwd_kernel' or kernel.keys!=['D','NB'] or kernel.cache:raise RuntimeError('Kernel identity or fresh-cache precondition failed')
74
+ telemetry={'profile_sha256':expected_sha,'status':'installed','calls':0,'keys':{},'strict_guard':True,'unknown_keys':'raise','autotune_fallback_permitted':False,'runtime':actual,'gpu_arch':arch,'process_scope':'one explicitly profiled Lux model; other model loading in this process is not supported'}
75
+ attach_guard(kernel,cache_module,entries,telemetry)
76
+ # The validated inference path only uses the vectorized D128 forward kernel.
77
+ # Other dimensions/backward must not silently enter a different autotuner.
78
+ for name in ['l2norm_fwd_kernel1','l2norm_bwd_kernel','l2norm_bwd_kernel1']:
79
+ other=getattr(module,name)
80
+ def reject(*args,_name=name,**kwargs):raise RuntimeError('Uncontracted normalization kernel: '+_name)
81
+ other.run=reject
82
+ return telemetry
code/runtime_profile.py ADDED
@@ -0,0 +1,36 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Bundle-local automatic launch-profile binding, before importing FLA.
2
+
3
+ One profile is active per Python process. Repeated loading of the same verified
4
+ profile is allowed; mixing with an unprofiled/different-profile model is not.
5
+ """
6
+ import hashlib,importlib.util,json,os,sys,types
7
+ from pathlib import Path
8
+ STATE='_decision_process_normalization_profile_v1'
9
+ def sha(p):return hashlib.sha256(Path(p).read_bytes()).hexdigest()
10
+ def safe(root,relative):
11
+ p=Path(relative)
12
+ if p.is_absolute() or '..' in p.parts:raise ValueError('Unsafe runtime-profile path')
13
+ return root/p
14
+
15
+ def ensure_profile(bundle):
16
+ bundle=Path(bundle).resolve();runtime=json.loads((bundle/'runtime.json').read_text());spec=runtime.get('normalization_profile');active=sys.modules.get(STATE)
17
+ if spec is None:
18
+ if active is not None:raise RuntimeError('Load unprofiled and profiled Decision models in separate processes')
19
+ return None
20
+ if spec.get('kind')!='decision-fla-l2norm-profile-v1' or spec.get('validated_arch')!='gfx942':raise ValueError('Unknown normalization profile contract')
21
+ profile=safe(bundle,spec['profile_file']);guard=safe(bundle,spec['guard_file'])
22
+ if sha(profile)!=spec['profile_sha256'] or sha(guard)!=spec['guard_sha256']:raise ValueError('Bound runtime profile bytes changed')
23
+ if json.loads((bundle/'decision_config.json').read_text()).get('base_model')!='Qwen/Qwen3.5-9B':raise ValueError('This bundle profile is bound to the validated Lux family only')
24
+ if active is not None:
25
+ if active.profile_sha256!=spec['profile_sha256'] or active.guard_sha256!=spec['guard_sha256']:raise RuntimeError('Different Decision normalization profile is already active; use a separate process')
26
+ active.guard.validate_profile(profile,spec['profile_sha256'])
27
+ if os.environ.get('FLA_CACHE_MODE')!='strict' or os.environ.get('FLA_CONFIG_DIR')!=active.profile_dir:raise RuntimeError('Active FLA profile environment changed')
28
+ return {'profile_sha256':active.profile_sha256,'guard_sha256':active.guard_sha256,'validated_arch':'gfx942','automatic_bundle_binding':True,'scope':'single profile per process'}
29
+ module_spec=importlib.util.spec_from_file_location('decision_profile_guard_'+spec['guard_sha256'][:16],guard);module=importlib.util.module_from_spec(module_spec);module_spec.loader.exec_module(module)
30
+ telemetry=module.install(profile,spec['profile_sha256'])
31
+ state=types.ModuleType(STATE);state.profile_sha256=spec['profile_sha256'];state.guard_sha256=spec['guard_sha256'];state.profile_dir=str(profile.parent);state.guard=module;state.telemetry=telemetry;sys.modules[STATE]=state
32
+ return {'profile_sha256':state.profile_sha256,'guard_sha256':state.guard_sha256,'validated_arch':'gfx942','automatic_bundle_binding':True,'scope':'single profile per process'}
33
+
34
+ def active_telemetry():
35
+ state=sys.modules.get(STATE)
36
+ return None if state is None else dict(state.telemetry)
decision_config.json ADDED
@@ -0,0 +1,17 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "base_revision": "c202236235762e1c871ad0ccb60c8ee5ba337b9a",
3
+ "text_parameter_count": 7936684544,
4
+ "prompt_version": "structured-segmented-candidate-endpoints-global-query-v2",
5
+ "attention": "sdpa",
6
+ "head_dim": 256,
7
+ "max_options": 255,
8
+ "architecture": "contextual-candidate-endpoint-plus-global-query-shared-bilinear-mlp",
9
+ "head_precision": "float32-outside-autocast",
10
+ "base_model": "Qwen/Qwen3.5-9B",
11
+ "model_name": "Decision-1.0-Lux",
12
+ "backbone_parameter_dtype": "bfloat16",
13
+ "head_parameter_dtype": "float32",
14
+ "training_master_parameter_dtype": "float32",
15
+ "calibration_file": "temperature.json",
16
+ "runtime_file": "runtime.json"
17
+ }
decision_head.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:66540f0bea02a117f1c0b5d7b5191735d4cde2a5b6daf1330ecefa6e10497dac
3
+ size 16845656
metrics/benchmark.json ADDED
The diff for this file is too large to render. See raw diff
 
metrics/evaluation-provenance.json ADDED
@@ -0,0 +1,1019 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "protocol": {
3
+ "version": "decision-priority-benchmark-v4",
4
+ "utc": "2026-09-22T05:05:12.118255+00:00",
5
+ "status": "User-requested product weighting, observed regression benchmark; release remains evidence-gated",
6
+ "authorization": "Latest user explicitly requested transfer20%\u219215% and assigned released5% to the stronger original-decision/composition area. Original decisions chosen from already observed same-size gaps. This is outcome-informed product weighting, not neutral/blind prospective benchmark design.",
7
+ "weights": {
8
+ "old_core": "3/10",
9
+ "v3_core": "1/4",
10
+ "v4": "3/20",
11
+ "v5": "3/20",
12
+ "transfer_v9_test": "3/20"
13
+ },
14
+ "accuracy_mean": "Exact weighted categorical accuracy; originalwithin-panelweights and requested denominators unchanged.",
15
+ "original_panel_registry_sha256": "1aefdf58d56e832e765194d0fa53a8647acea7d2ad2fc25ac95380bf85949fd2",
16
+ "transfer_protocol_sha256": "f0c0f8982bfbb9edfc5e486f682906612fae579f28c58c2fa656ad9d10874d16",
17
+ "transfer_payload_sha256": "97d82909bcbd8756904d6ed59b31e16d16086d62990ccdc981dc95b8c274e0c5",
18
+ "transfer_labels_sha256": "0b64882d58e2ac1e0f5e2e102f916907554c94c07f2774b8c092683cfc81973c",
19
+ "transfer_accuracy_questions": 1046,
20
+ "original_core_questions": 2720,
21
+ "diagnostics": "Preserve internal-v2 source-generalization/order/unknown/proper-score/efficiency axes separately; no millisecond or ECE averaging into accuracy.",
22
+ "roster": {
23
+ "Nox": {
24
+ "role": "Decision family",
25
+ "size": "4B",
26
+ "repo": "llm-semantic-router/Decision-1.0-Nox"
27
+ },
28
+ "Sol": {
29
+ "role": "Decision family",
30
+ "size": "2B",
31
+ "repo": "llm-semantic-router/Decision-1.0-Sol"
32
+ },
33
+ "Kai": {
34
+ "role": "Decision family",
35
+ "size": "0.572B encoder",
36
+ "repo": "llm-semantic-router/Decision-1.0-Kai",
37
+ "revision": "2079354070d02f4c2b1e60bce6b99f60dfad9622",
38
+ "contract": "shipped1024complete-token admission; FP32; defaultB8; no limit override"
39
+ },
40
+ "Lux": {
41
+ "role": "Decision family",
42
+ "size": "9B",
43
+ "repo": "llm-semantic-router/Decision-1.0-Lux",
44
+ "pending_unreleased": true
45
+ },
46
+ "kev-0.8b": {
47
+ "role": "open reference",
48
+ "size": "0.8B",
49
+ "repo": "jaredpalmer/kev-0.8b",
50
+ "revision": "54f4f8777356cd5bbbb6c6919c657f26e6f2f6d8"
51
+ },
52
+ "kev-4b": {
53
+ "role": "open reference",
54
+ "size": "4B",
55
+ "repo": "jaredpalmer/kev-4b",
56
+ "revision": "485ace8703592fcf405488b262449990824cfed1"
57
+ },
58
+ "kev-9b": {
59
+ "role": "open reference",
60
+ "size": "9B",
61
+ "repo": "jaredpalmer/kev-9b",
62
+ "revision": "2629c06a5aeb0feb3b9783bafed17ed8f39ecf5c"
63
+ },
64
+ "Decider": {
65
+ "role": "open reference",
66
+ "size": "2B",
67
+ "repo": "Mapika/decider-2b"
68
+ },
69
+ "Laya-base": {
70
+ "role": "open reference",
71
+ "size": "0.421B encoder",
72
+ "repo": "convaiinnovations/laya",
73
+ "mode": "english"
74
+ },
75
+ "Laya-multilingual": {
76
+ "role": "open reference",
77
+ "size": "0.322B encoder",
78
+ "repo": "convaiinnovations/laya",
79
+ "mode": "multilingual"
80
+ },
81
+ "Jev": {
82
+ "role": "closed frontier",
83
+ "size": "undisclosed",
84
+ "served_identity": "jev-1.13.0"
85
+ },
86
+ "Qwen3.5-2B": {
87
+ "role": "untuned reference",
88
+ "size": "2B",
89
+ "repo": "Qwen/Qwen3.5-2B",
90
+ "adapter": "original locked chat/LM-head letter readout; no training"
91
+ },
92
+ "Qwen3.5-4B": {
93
+ "role": "untuned reference",
94
+ "size": "4B",
95
+ "repo": "Qwen/Qwen3.5-4B",
96
+ "adapter": "original locked chat/LM-head letter readout; no training"
97
+ },
98
+ "Qwen3.5-9B": {
99
+ "role": "untuned reference",
100
+ "size": "9B",
101
+ "repo": "Qwen/Qwen3.5-9B",
102
+ "adapter": "original locked chat/LM-head letter readout; no training"
103
+ }
104
+ },
105
+ "internal_extra_references": [
106
+ "LLM2Jev2B",
107
+ "LLM2Jev4B",
108
+ "Nimble9B"
109
+ ],
110
+ "cohort_gates": {
111
+ "Sol": {
112
+ "required_same_size_open": [
113
+ "Decider"
114
+ ],
115
+ "additional_measured_internal_peer": [
116
+ "LLM2Jev2B"
117
+ ],
118
+ "required_same_size_untuned": [
119
+ "Qwen3.5-2B"
120
+ ]
121
+ },
122
+ "Nox": {
123
+ "required_same_size_open": [
124
+ "kev-4b"
125
+ ],
126
+ "additional_measured_internal_peer": [
127
+ "LLM2Jev4B"
128
+ ],
129
+ "required_same_size_untuned": [
130
+ "Qwen3.5-4B"
131
+ ]
132
+ },
133
+ "Lux": {
134
+ "must_exceed_current_family": [
135
+ "Nox",
136
+ "Sol"
137
+ ],
138
+ "required_same_size_open": [
139
+ "kev-9b"
140
+ ],
141
+ "additional_measured_internal_peer": [
142
+ "Nimble9B"
143
+ ],
144
+ "required_same_size_untuned": [
145
+ "Qwen3.5-9B"
146
+ ]
147
+ }
148
+ },
149
+ "publication": {
150
+ "first_expanded_cards": "Nox and Sol independently eligible only when their own required same-size comparisons are complete and strictly beaten; table includes complete available public roster. Unreleased Lux need not delay either card.",
151
+ "later_updates": "Require newaggregate strictly higher than the current published same model plus maintained same-size lead; show regressions and full diagnostic metrics honestly.",
152
+ "Lux_first_release": "Newweightedaggregate strictly exceeds then-currentNox/Sol and known same-sizeopen references; old4panelgate superseded before heldout/publication.",
153
+ "missing_scores": "Never fill from different benchmark, revision, successful-only denominator or inferred model size.",
154
+ "accuracy_uncertainty": "Report paired component bootstrap intervals; no new positive-CI gate invented.",
155
+ "versions": "No trainingversion labels in visiblecard/table/figures; retain immutable revision/weight/adapter identities in machine-readable provenance.",
156
+ "rank": "All public roster descending actual score. Decision family copper; distinct restrained colors for Kev, Decider, Laya, frontierJev and untunedQwen. White background, no logos or trainingversion labels.",
157
+ "matrix": "Minimal readable typography, full task coverage, no logo; labels omit trainingversions.",
158
+ "original_comparison": "Retain original four-panel and prior five-panel25/25/15/15/20 comparisons in methods/provenance. Explicit outcome-informed product-priority amendment; no universal or prospective performance claim."
159
+ },
160
+ "training_integrity": "Frozen ongoing training/SELECT/CAL unchanged. Testresults never select anothercheckpoint; future targetedtraining uses independent TRAIN/SELECT/CAL. Observed tests labeledregression.",
161
+ "supersedes": {
162
+ "protocol": {
163
+ "path": "eval/decision-benchmark-v3/PROTOCOL.json",
164
+ "sha256": "d0b7c667fd4de0dfa5268d2d83aebf40e2b0c43d615ad77f4c2264e35a137534"
165
+ },
166
+ "prior_full_statistics": {
167
+ "path": "analysis/decision-benchmark-v3-results/full01/scored/STATISTICS.json",
168
+ "sha256": "aaed8ca33a874300545b81566cb0ba37bf7cc8fcaac954beb9094aeef8ad82f7"
169
+ }
170
+ },
171
+ "sensitivity": "Retain full25/25/15/15/20 results alongside30/25/15/15/15 in methods/provenance. Reweighting gains are never described as training improvements. No model selection, calibration, predictions or denominators changed."
172
+ },
173
+ "statistics_sha256": "1db772aa92754ce458965f8b449553dde4717d1400dd8fd509ce32d9c76f1d5c",
174
+ "manifest_sha256": "3199a40b2719cad9b1e815629cf5ebad923a9a399188c1199a9d27e918bda9e3",
175
+ "qualified_runtime": {
176
+ "python": "3.12.13",
177
+ "numpy": "2.3.5"
178
+ },
179
+ "source_sha256": "b00c60f21ace9bd258a1a79673b8cdbce3101ada2fed5547811f70bf030fcd91",
180
+ "model_manifest": {
181
+ "Nox": {
182
+ "panels": {
183
+ "old_core": {
184
+ "path": "results/expanded/nox-retention-v1/candidate/job-0/worker/old_core/normalized.jsonl",
185
+ "sha256": "9b6b7db907e8e458cd33982c0f7c45e5052c3248285d927990f160d5570c8976"
186
+ },
187
+ "v3_core": {
188
+ "path": "results/expanded/nox-retention-v1/candidate/job-0/worker/v3_core/normalized.jsonl",
189
+ "sha256": "49575aef2261c191e289d781857664c57fd6181701c03f02515292f2fa31b40c"
190
+ },
191
+ "v4": {
192
+ "path": "results/expanded/nox-retention-v1/candidate/job-0/worker/v4/normalized.jsonl",
193
+ "sha256": "9d85618b1a05c2ed68f6a188dadf19a64d6ede6d5a1b0ef58afa27a13906238b"
194
+ },
195
+ "v5": {
196
+ "path": "results/expanded/nox-retention-v1/candidate/job-0/worker/v5/normalized.jsonl",
197
+ "sha256": "5eb7faaa45a22dad414af1363d6087b8c379c86599df3c14483a3a648ca6bfd5"
198
+ }
199
+ },
200
+ "transfer_rows": {
201
+ "path": "analysis/kev-reciprocal-v1/reports/Decision-1.0-Nox/v9-transfer-v9-test-published-temperature-rows.json",
202
+ "sha256": "77b92a18433b260ac0e6f4e0b6c9ec3a94e16aaed870a4078b1d7f32ef18b169"
203
+ },
204
+ "evidence": [
205
+ {
206
+ "path": "results/expanded/nox-retention-v1/QUALITY-COLLECTION.json",
207
+ "sha256": "7608258282ec1e9731e4d2ea4dee7206fda8900055dc64f9193d2566172ca86a"
208
+ },
209
+ {
210
+ "path": "analysis/kev-reciprocal-v1/reports/Decision-1.0-Nox/REPORT.json",
211
+ "sha256": "8e78910746c9e78862b3aadd94490054a681f3ad013b43ec19bfbc2434111ed7"
212
+ }
213
+ ]
214
+ },
215
+ "Sol": {
216
+ "panels": {
217
+ "old_core": {
218
+ "path": "results/expanded/sol-composition-v2/candidate/job-0/worker/old_core/normalized.jsonl",
219
+ "sha256": "5fc0b3e49a5de576f7621839130855c178b12d942dc357aaf6b4dd9735e0a6a8"
220
+ },
221
+ "v3_core": {
222
+ "path": "results/expanded/sol-composition-v2/candidate/job-0/worker/v3_core/normalized.jsonl",
223
+ "sha256": "b2c99fbd739b849a66a55929a658de4517a7ffaa4496255566e6972ba28d4f15"
224
+ },
225
+ "v4": {
226
+ "path": "results/expanded/sol-composition-v2/candidate/job-0/worker/v4/normalized.jsonl",
227
+ "sha256": "3c28302750e02d4cead0aacd66a323734e773f726c4c41db658dd6461251b28b"
228
+ },
229
+ "v5": {
230
+ "path": "results/expanded/sol-composition-v2/candidate/job-0/worker/v5/normalized.jsonl",
231
+ "sha256": "6ddb0496012b596ea122afa4db4c4158e57dfb934e9ae29eb006756b652e31e1"
232
+ }
233
+ },
234
+ "transfer_rows": {
235
+ "path": "analysis/kev-reciprocal-v1/reports/Decision-1.0-Sol-attempt02/v9-transfer-v9-test-published-temperature-rows.json",
236
+ "sha256": "b05799417cb3715de633135c58f713960b75bdb9ed30eb8912a197d33ee3d85b"
237
+ },
238
+ "evidence": [
239
+ {
240
+ "path": "results/expanded/sol-composition-v2/QUALITY-COLLECTION.json",
241
+ "sha256": "9b77aac87b85cbf3a74e86af4a2b9a03fd9e1ed854356a92ff54ccd0d47a16f4"
242
+ },
243
+ {
244
+ "path": "analysis/kev-reciprocal-v1/reports/Decision-1.0-Sol-attempt02/REPORT.json",
245
+ "sha256": "5d8ebb339293687cec9089f54fb96aabf3330f16111487a49489f5e52bbb6f29"
246
+ }
247
+ ]
248
+ },
249
+ "kev-0.8b": {
250
+ "panels": {
251
+ "old_core": {
252
+ "path": "analysis/kev-reciprocal-v1/our-panels/kev-0.8b-old_core-normalized.jsonl",
253
+ "sha256": "b492b4512ac164f83a319f31bbc9e086cae12f87ffe2c14f307481ffb240807c"
254
+ },
255
+ "v3_core": {
256
+ "path": "analysis/kev-reciprocal-v1/our-panels/kev-0.8b-v3_core-normalized.jsonl",
257
+ "sha256": "70cbfc69dfc72ac0c5ae01305ec63a9f31129c115b7068c6824cd22ead013d68"
258
+ },
259
+ "v4": {
260
+ "path": "analysis/kev-reciprocal-v1/our-panels/kev-0.8b-v4-normalized.jsonl",
261
+ "sha256": "761dfd69acbface101f2c21fd27d2db684c5c8b8cc425890a075729ae1f7db24"
262
+ },
263
+ "v5": {
264
+ "path": "analysis/kev-reciprocal-v1/our-panels/kev-0.8b-v5-normalized.jsonl",
265
+ "sha256": "51b6fc5f826ca8db44df43885f34d857ddd5b5a66453bbdc1a83fcf488928bdb"
266
+ }
267
+ },
268
+ "transfer_rows": {
269
+ "path": "analysis/kev-reciprocal-v1/reports/kev-0.8b/v9-transfer-v9-test-published-temperature-rows.json",
270
+ "sha256": "d6abed88001aeac630219466f02c1aad89fcbd354b3ae533507f3c2f7fe42898"
271
+ },
272
+ "evidence": [
273
+ {
274
+ "path": "analysis/kev-reciprocal-v1/reports/kev-0.8b/REPORT.json",
275
+ "sha256": "6760dcbaa7d85e6c2e893e4c80ab8e66e0b84d7cfa26ae9e25909f48a05f8de4"
276
+ },
277
+ {
278
+ "path": "analysis/kev-reciprocal-v1/our-panels/POINTS.json",
279
+ "sha256": "a3aa6d501a14df1cf5e7c98ecc6161c88483801310c56913465059d28c4d88ea"
280
+ }
281
+ ]
282
+ },
283
+ "kev-4b": {
284
+ "panels": {
285
+ "old_core": {
286
+ "path": "analysis/kev-reciprocal-v1/our-panels/kev-4b-old_core-normalized.jsonl",
287
+ "sha256": "41dfbb7c9c4381effc219f9123b51bd4d688b3e746cec974ef930730b9dfd6b1"
288
+ },
289
+ "v3_core": {
290
+ "path": "analysis/kev-reciprocal-v1/our-panels/kev-4b-v3_core-normalized.jsonl",
291
+ "sha256": "2ec54da92c4e72412e89b3c07013b54729f5899ba8542eeadf962b04a1f3a4b8"
292
+ },
293
+ "v4": {
294
+ "path": "analysis/kev-reciprocal-v1/our-panels/kev-4b-v4-normalized.jsonl",
295
+ "sha256": "88c576c3d62ec11686dd5ebf360971a9087779e45a346897e2badf874a754a12"
296
+ },
297
+ "v5": {
298
+ "path": "analysis/kev-reciprocal-v1/our-panels/kev-4b-v5-normalized.jsonl",
299
+ "sha256": "d33bbb60134314dcb550060176ed0a3ae91ec0c6b033cceb40e3255b33666898"
300
+ }
301
+ },
302
+ "transfer_rows": {
303
+ "path": "analysis/kev-reciprocal-v1/reports/kev-4b/v9-transfer-v9-test-published-temperature-rows.json",
304
+ "sha256": "c2cb16a936be4c10ca6b1870c69426ddd6b51629c401f71ae1da652b3bd2ac7e"
305
+ },
306
+ "evidence": [
307
+ {
308
+ "path": "analysis/kev-reciprocal-v1/reports/kev-4b/REPORT.json",
309
+ "sha256": "366fb0d27a4d732e5e270c264521d9a93ccabc4f0ddf5b7d3e9f1ecc5f460406"
310
+ },
311
+ {
312
+ "path": "analysis/kev-reciprocal-v1/our-panels/POINTS.json",
313
+ "sha256": "a3aa6d501a14df1cf5e7c98ecc6161c88483801310c56913465059d28c4d88ea"
314
+ }
315
+ ]
316
+ },
317
+ "kev-9b": {
318
+ "panels": {
319
+ "old_core": {
320
+ "path": "analysis/kev-reciprocal-v1/our-panels/kev-9b-old_core-normalized.jsonl",
321
+ "sha256": "a3b11fee4bc5e207cc42e7f4d8dde05df6cfc73d9446ac33cd6015e1ab9e587d"
322
+ },
323
+ "v3_core": {
324
+ "path": "analysis/kev-reciprocal-v1/our-panels/kev-9b-v3_core-normalized.jsonl",
325
+ "sha256": "bb90e5701a151272f97caa4107888ba2379969b4e3f22abb4670a528b50d7729"
326
+ },
327
+ "v4": {
328
+ "path": "analysis/kev-reciprocal-v1/our-panels/kev-9b-v4-normalized.jsonl",
329
+ "sha256": "b9c50612c0be1b7b5647779d8489c185ea7f02673da176b1d965673da604d644"
330
+ },
331
+ "v5": {
332
+ "path": "analysis/kev-reciprocal-v1/our-panels/kev-9b-v5-normalized.jsonl",
333
+ "sha256": "2893714030357718e662633fddf31a328e56a02fc3ae71ed34612f480e50d4ee"
334
+ }
335
+ },
336
+ "transfer_rows": {
337
+ "path": "analysis/kev-reciprocal-v1/reports/kev-9b/v9-transfer-v9-test-published-temperature-rows.json",
338
+ "sha256": "e712bf23acea6a9431001f5424639d1e3738789907ff2e367db7c9dad4e0d239"
339
+ },
340
+ "evidence": [
341
+ {
342
+ "path": "analysis/kev-reciprocal-v1/reports/kev-9b/REPORT.json",
343
+ "sha256": "96dcf11961f510d05c011aa3fd8dfa6043348471b8e800d6bac2d691d0815d54"
344
+ },
345
+ {
346
+ "path": "analysis/kev-reciprocal-v1/our-panels/POINTS.json",
347
+ "sha256": "a3aa6d501a14df1cf5e7c98ecc6161c88483801310c56913465059d28c4d88ea"
348
+ }
349
+ ]
350
+ },
351
+ "Jev": {
352
+ "panels": {
353
+ "old_core": {
354
+ "path": "eval/heldout/official-predictions.jsonl",
355
+ "sha256": "069e48e1fdf1b5f3d95cc0097e4179ac3faab97bb357b0ce897cfb463b87cbaf",
356
+ "bytes": 185263
357
+ },
358
+ "v3_core": {
359
+ "path": "eval/heldout/v3/official/normalized/core.jsonl",
360
+ "sha256": "54707354bfe5d96b5de56e6dc1b38cba0a610eea516b28fbff013b94f422b397",
361
+ "bytes": 368716
362
+ },
363
+ "v4": {
364
+ "path": "eval/heldout/v4/official/normalized/core.jsonl",
365
+ "sha256": "d6583f903e974d70747fae2944014828954dc4d7344ceddcbb2cd1b06fbfc799",
366
+ "bytes": 183671
367
+ },
368
+ "v5": {
369
+ "path": "eval/heldout/v5/official/normalized/core.jsonl",
370
+ "sha256": "f883164c4c83db8e42ebe03f7e165afbd2f81f2b541bcf8008a2cc06a60e1ee7",
371
+ "bytes": 200316
372
+ }
373
+ },
374
+ "transfer_rows": {
375
+ "path": "analysis/decision-benchmark-v3-official/scored/ROWS.json",
376
+ "sha256": "be29a7f3804ad84993ffb06f05ada7e2d8ed25317504c4fc17fe5dc306fd788c"
377
+ },
378
+ "evidence": [
379
+ {
380
+ "path": "eval/expanded-public-extension-v1/COMPARATOR-REGISTRY.json",
381
+ "sha256": "b80be4270721cbeeac9e28beb10c6f19f75c1cbf2553e38d3d59e3a42d56a91c"
382
+ },
383
+ {
384
+ "path": "analysis/decision-benchmark-v3-official/scored/REPORT.json",
385
+ "sha256": "2fdbaf7db1e4b877762e33a87b23a236a065eaa3078b14c2d7d579ad44f9ea24"
386
+ },
387
+ {
388
+ "path": "analysis/decision-benchmark-v3-official/INDEPENDENT-PROJECTION-REVIEW.json",
389
+ "sha256": "b9506539e2c4a863aef16f7dab0517c175e85863416bc622c60df05489a9eb52"
390
+ }
391
+ ]
392
+ },
393
+ "Kai": {
394
+ "panels": {
395
+ "old_core": {
396
+ "path": "results/public-roster-baselines-v1/Kai/worker/old_core/normalized.jsonl",
397
+ "sha256": "4758a99a75392b197c8902600a88edca65d6acea500e046e1ea8e4edd398a8a3"
398
+ },
399
+ "v3_core": {
400
+ "path": "results/public-roster-baselines-v1/Kai/worker/v3_core/normalized.jsonl",
401
+ "sha256": "7cdb285b88882a4c0d73c3da890056ef3b1d09e09bc01a893e73a12ff3bdbab8"
402
+ },
403
+ "v4": {
404
+ "path": "results/public-roster-baselines-v1/Kai/worker/v4/normalized.jsonl",
405
+ "sha256": "fa744ee2181fc4c379b748fdd4ee1080ffd877f43ebc9441d27a47013a2a9cec"
406
+ },
407
+ "v5": {
408
+ "path": "results/public-roster-baselines-v1/Kai/worker/v5/normalized.jsonl",
409
+ "sha256": "9d1e50d62733dbf06b824855938a9ee1311c030916cf173cc5b18fe066af48cd"
410
+ }
411
+ },
412
+ "transfer_rows": {
413
+ "path": "analysis/decision-benchmark-v3-baselines/Kai/ROWS.json",
414
+ "sha256": "854950cb3888a49a801982341576f4a2ebdacbdd7e7a7e32d16f5777e0cbda46"
415
+ },
416
+ "evidence": [
417
+ {
418
+ "path": "results/public-roster-baselines-v1/Kai/worker/COMPLETE.json",
419
+ "sha256": "39cc8bae7b59e02720123cd8efa9ba1939e02ef15a4bb3be89972b14eb7877cf"
420
+ },
421
+ {
422
+ "path": "analysis/public-roster-discovery-v1/ACTUAL-EVALUATION-COMPLETE.json",
423
+ "sha256": "686834a8289d3771908530a95c2ff85df71e61dbb1ce7ea95b7c9550a23ba82f"
424
+ },
425
+ {
426
+ "path": "analysis/decision-benchmark-v3-baselines/Kai/REPORT.json",
427
+ "sha256": "5ed1fd6a69d68fbd64bd45c4da9431a73fd6053574786ddd00323e0537fcefbc"
428
+ }
429
+ ]
430
+ },
431
+ "Laya-base": {
432
+ "panels": {
433
+ "v4": {
434
+ "path": "results/public-roster-baselines-v1/laya-english/worker/v4/normalized.jsonl",
435
+ "sha256": "31da1d4140da69d2677172707eb301843cf8fdfb43200cba93be910321f922ba"
436
+ },
437
+ "v5": {
438
+ "path": "results/public-roster-baselines-v1/laya-english/worker/v5/normalized.jsonl",
439
+ "sha256": "159614ceede814b578e1e025836a595c8c71a9bec4896caf6454d7bf54c4c44d"
440
+ },
441
+ "old_core": {
442
+ "path": "eval/heldout/primary-v2/laya-english/core/normalized.jsonl",
443
+ "sha256": "850643d2daeb47a6909985003736cab0f1f0e5a87631a35e27192f0733efd756"
444
+ },
445
+ "v3_core": {
446
+ "path": "eval/heldout/v3/baselines/laya-english/core/normalized.jsonl",
447
+ "sha256": "dec19caeb8074e25289e43a367392172da21f0c2ccdb403a059e2e4a2364a64b"
448
+ }
449
+ },
450
+ "transfer_rows": {
451
+ "path": "analysis/decision-benchmark-v3-baselines/laya-english/ROWS.json",
452
+ "sha256": "19138909259400facb5d3d5402d0d736f870695fe3ee46019f2e438088c57275"
453
+ },
454
+ "evidence": [
455
+ {
456
+ "path": "results/public-roster-baselines-v1/laya-english/worker/COMPLETE.json",
457
+ "sha256": "cef0da797ae14cebaf1e4b9edf9b16c3cd9a412a30435d7d2517b8ba6dbe8b98"
458
+ },
459
+ {
460
+ "path": "analysis/public-roster-discovery-v1/ACTUAL-EVALUATION-COMPLETE.json",
461
+ "sha256": "686834a8289d3771908530a95c2ff85df71e61dbb1ce7ea95b7c9550a23ba82f"
462
+ },
463
+ {
464
+ "path": "analysis/decision-benchmark-v3-baselines/laya-english/REPORT.json",
465
+ "sha256": "cb5935bb399dcf23ffe24f855452b2e074bcbb3d023cfc7dd439342e54f02c6b"
466
+ },
467
+ {
468
+ "path": "analysis/public-roster-discovery-v1/LAYA-CORE-REUSE-VERIFIED.json",
469
+ "sha256": "01d2be9f23540f3fbe5d06e6707bb3d908dd76d2c509ee4af86c9e0c6129862f"
470
+ }
471
+ ]
472
+ },
473
+ "Laya-multilingual": {
474
+ "panels": {
475
+ "v4": {
476
+ "path": "results/public-roster-baselines-v1/laya-multilingual/worker/v4/normalized.jsonl",
477
+ "sha256": "814c1aa1f4d400942984ecb3786d3fd5fd15286d9b767a519c91993577795290"
478
+ },
479
+ "v5": {
480
+ "path": "results/public-roster-baselines-v1/laya-multilingual/worker/v5/normalized.jsonl",
481
+ "sha256": "6d4dbc16c7b977b15895e0e313dabc6d9db43df75ec56e89373d91f4f13ec676"
482
+ },
483
+ "old_core": {
484
+ "path": "eval/heldout/primary-v2/laya-multilingual/core/normalized.jsonl",
485
+ "sha256": "9ea39841c4d5901cd9ab860c980d2f3b55358be289c7e01b9b47982927bf849d"
486
+ },
487
+ "v3_core": {
488
+ "path": "eval/heldout/v3/baselines/laya-multilingual/core/normalized.jsonl",
489
+ "sha256": "a0b7ce9daf4be2c7831bcab6b5e96c8dcb6bef133d39475f72241ab18faccdd2"
490
+ }
491
+ },
492
+ "transfer_rows": {
493
+ "path": "analysis/decision-benchmark-v3-baselines/laya-multilingual/ROWS.json",
494
+ "sha256": "d32db6cf6576f287210d31336456497f9d6926179366db29304659baf9783b43"
495
+ },
496
+ "evidence": [
497
+ {
498
+ "path": "results/public-roster-baselines-v1/laya-multilingual/worker/COMPLETE.json",
499
+ "sha256": "b190d683fc4c4762c2805a03d52628a53e125d228a3105c8ac2ddda00a5364e4"
500
+ },
501
+ {
502
+ "path": "analysis/public-roster-discovery-v1/ACTUAL-EVALUATION-COMPLETE.json",
503
+ "sha256": "686834a8289d3771908530a95c2ff85df71e61dbb1ce7ea95b7c9550a23ba82f"
504
+ },
505
+ {
506
+ "path": "analysis/decision-benchmark-v3-baselines/laya-multilingual/REPORT.json",
507
+ "sha256": "d20d3d58f76ca2758545f83d358abd747f86d581c897d1fb521371d1000d6154"
508
+ },
509
+ {
510
+ "path": "analysis/public-roster-discovery-v1/LAYA-CORE-REUSE-VERIFIED.json",
511
+ "sha256": "01d2be9f23540f3fbe5d06e6707bb3d908dd76d2c509ee4af86c9e0c6129862f"
512
+ }
513
+ ]
514
+ },
515
+ "Qwen3.5-9B": {
516
+ "panels": {
517
+ "old_core": {
518
+ "path": "results/lux-new-host-v1/quality/base9b/old_core/normalized.jsonl",
519
+ "sha256": "16ccd1c30bbf29560a7ce1a288143692c5d2ed69cc7afebdc146d5bcc4b77611"
520
+ },
521
+ "v3_core": {
522
+ "path": "results/lux-new-host-v1/quality/base9b/v3_core/normalized.jsonl",
523
+ "sha256": "d081a8158863e85d5e49cb853fee87192119767124f53f9e6c782392f8edb018"
524
+ },
525
+ "v4": {
526
+ "path": "results/lux-new-host-v1/quality/base9b/v4/normalized.jsonl",
527
+ "sha256": "cc2775884d84c49e5155afbebc7f4562ec86f6f6d32765ffcbdafe653ca7d132"
528
+ },
529
+ "v5": {
530
+ "path": "results/lux-new-host-v1/quality/base9b/v5/normalized.jsonl",
531
+ "sha256": "ce6bf5c286e5920055d7de36fbd910311f7d3185380c63681c7baea6d0e3494a"
532
+ }
533
+ },
534
+ "transfer_rows": {
535
+ "path": "analysis/decision-benchmark-v3-baselines/Qwen3.5-9B/ROWS.json",
536
+ "sha256": "5acdd41f616a3ba63a5407af9c8371268012792fd324b70230d899bbb37101b4"
537
+ },
538
+ "evidence": [
539
+ {
540
+ "path": "results/lux-new-host-v1/quality/base9b/COMPLETE.json",
541
+ "sha256": "702c65f5c290f5cdda0495479f361101dbe976330233c5359455ba736245c3aa"
542
+ },
543
+ {
544
+ "path": "results/lux-new-host-v1/quality/base9b/metadata.json",
545
+ "sha256": "eaa29f71aba60e86068e6c5d1b5782ea73b0214b81a297b57419c4f0a2fc6c04"
546
+ },
547
+ {
548
+ "path": "results/lux-new-host-v1/quality/base9b/old_core/complete.json",
549
+ "sha256": "d28f2a646a908fc775e3727569a608f30ae5820f88a6173750a1031604e84f47"
550
+ },
551
+ {
552
+ "path": "results/lux-new-host-v1/quality/base9b/v3_core/complete.json",
553
+ "sha256": "3a10bd5e2121a81f8ec6f76b2526723a17689ec34d13f062c5259983203992cc"
554
+ },
555
+ {
556
+ "path": "results/lux-new-host-v1/quality/base9b/v4/complete.json",
557
+ "sha256": "073239ab7dd21768bccdfa4aaf3217c99dc7ff9b00fc308ffbe889c72fe002b4"
558
+ },
559
+ {
560
+ "path": "results/lux-new-host-v1/quality/base9b/v5/complete.json",
561
+ "sha256": "9a8d1e8e8db066f1f6594f137a7a1571dd5184aae363a9f919e44b634b782bba"
562
+ },
563
+ {
564
+ "path": "analysis/decision-benchmark-v3-baselines/Qwen3.5-9B/REPORT.json",
565
+ "sha256": "31b0a92bbbe7a0d259cb8e296a10b04b2bdd685b394eabdc5fbf6da1556ff12d"
566
+ },
567
+ {
568
+ "path": "analysis/decision-benchmark-v3-baselines/Qwen3.5-9B-projected/COMPLETE.json",
569
+ "sha256": "eaff79adb3a6a6bad61253a01b4c4e22cfa3d9969cd963cd82a3685755b77ec1"
570
+ }
571
+ ]
572
+ },
573
+ "Lux": {
574
+ "panels": {
575
+ "old_core": {
576
+ "path": "results/lux-new-host-v1/quality/lux/old_core/normalized.jsonl",
577
+ "sha256": "a659a7c7fef2871bbf806e842ea8ca81f645dcf610f0f2f6622d8d32e76a0836"
578
+ },
579
+ "v3_core": {
580
+ "path": "results/lux-new-host-v1/quality/lux/v3_core/normalized.jsonl",
581
+ "sha256": "da8903a07e36a9687f6ee1d0fd8dbc684e61c49ef5fea30c346e46bff5a90e94"
582
+ },
583
+ "v4": {
584
+ "path": "results/lux-new-host-v1/quality/lux/v4/normalized.jsonl",
585
+ "sha256": "26cd7482693ec1ffa1a93a156abfc88378e4872e320f3b9b775763635086d488"
586
+ },
587
+ "v5": {
588
+ "path": "results/lux-new-host-v1/quality/lux/v5/normalized.jsonl",
589
+ "sha256": "759c773e629a16317271de850780a8734533c795046122f850914c4482d1e72b"
590
+ }
591
+ },
592
+ "transfer_rows": {
593
+ "path": "analysis/decision-benchmark-v3-baselines/Lux/ROWS.json",
594
+ "sha256": "bafb55fc86762061e537382360eda14a8a8f8c5e4e111c2fda1f82cd6c06793f"
595
+ },
596
+ "evidence": [
597
+ {
598
+ "path": "results/lux-new-host-v1/quality/lux/COMPLETE.json",
599
+ "sha256": "9600ff792cbda13af5505c804cf63af94d2ee895f2912f0645a7d493ae1b7a9c"
600
+ },
601
+ {
602
+ "path": "results/lux-new-host-v1/quality/lux/metadata.json",
603
+ "sha256": "0d4509fae7eb3630118eaebc24e75a7b02f717d9a0ad74811fa27358c33c9cc5"
604
+ },
605
+ {
606
+ "path": "analysis/decision-benchmark-v3-baselines/Lux-projected/COMPLETE.json",
607
+ "sha256": "f58d02df78f297d4feebe39facfac072951111d8d6a19e2c1069f1b90250c622"
608
+ },
609
+ {
610
+ "path": "analysis/decision-benchmark-v3-baselines/Lux/REPORT.json",
611
+ "sha256": "f31703c342efe043a1e1534da189808bb7e0de0641fbaf3f91faa4a1f928a299"
612
+ },
613
+ {
614
+ "path": "analysis/decoder4b/lux9b-training-v1/CHECKPOINT-HELDOUT-COLLECTED.json",
615
+ "sha256": "d1c42515ab225b27ea011a6cf2262fe12dd7001e3b4d2cb64a23e7642c7a0d91"
616
+ }
617
+ ]
618
+ },
619
+ "Decider": {
620
+ "panels": {
621
+ "old_core": {
622
+ "path": "eval/heldout/matched-fla-v1/runs/core-fla-v1-shard1/decider/quality/normalized.jsonl",
623
+ "sha256": "e1b14f0f9ec521bffbde5945684bc2410a413b6b6418d65ac4fc702a7b5e7937"
624
+ },
625
+ "v3_core": {
626
+ "path": "eval/heldout/v3/baselines/decider/core/normalized.jsonl",
627
+ "sha256": "12a1f554decf1aff608743e7b4a44681289ea084ae25f383e519b1f34917d46f"
628
+ },
629
+ "v4": {
630
+ "path": "eval/heldout/v4/baselines/decider/core/normalized.jsonl",
631
+ "sha256": "626ef366ae471ec10dfb89ef2ff2f7b29a6c879976d3ff14fcbb2bafd7d9b045"
632
+ },
633
+ "v5": {
634
+ "path": "eval/heldout/v5/open-baselines-v1/decider/core/normalized.jsonl",
635
+ "sha256": "ae16b17ac8e4208940c0c3e9e04f25e94686b5d16bf46f64f53669baf41aeb71"
636
+ }
637
+ },
638
+ "transfer_rows": {
639
+ "path": "analysis/decision-benchmark-v3-baselines/decider/ROWS.json",
640
+ "sha256": "6da5b5692a8a0cea3dc8f6d02d911fafd9bbf62e286a9be8db27c89f77a897e4"
641
+ },
642
+ "evidence": [
643
+ {
644
+ "path": "results/accelerated-transfer-baselines-v1/decider/worker/COMPLETE.json",
645
+ "sha256": "e1ee0df8df59ec3eb2b49ba947a8508994a24bce226b7f2cccbe25f6835b99af"
646
+ },
647
+ {
648
+ "path": "results/accelerated-transfer-baselines-v1/decider/worker/METADATA.json",
649
+ "sha256": "f00061c1b0bb642bf1353aae867d19df988c01b0428b58206b31249521bcbff8"
650
+ },
651
+ {
652
+ "path": "results/accelerated-transfer-baselines-v1/decider/worker/predictions.jsonl",
653
+ "sha256": "f69e87fa65b09289e6861c1d789270aea5a92d23ac47817342647ed08f406df7"
654
+ },
655
+ {
656
+ "path": "analysis/decision-benchmark-v3-baselines/decider/REPORT.json",
657
+ "sha256": "5aae029f34524dfebc8e79128da8c39229c5747461dfce07ca59956da4416951"
658
+ },
659
+ {
660
+ "path": "eval/expanded-public-extension-v1/COMPARATOR-REGISTRY.json",
661
+ "sha256": "b80be4270721cbeeac9e28beb10c6f19f75c1cbf2553e38d3d59e3a42d56a91c"
662
+ },
663
+ {
664
+ "path": "analysis/decision-benchmark-v3-baselines/decider/COMPOSABLE-MANIFEST.json",
665
+ "sha256": "a49848e86feafaa4f36a33287910035e7605e344791aae139305b78cb304a77f"
666
+ }
667
+ ]
668
+ },
669
+ "Qwen3.5-2B": {
670
+ "panels": {
671
+ "old_core": {
672
+ "path": "eval/heldout/primary-v2/base-2b/core/normalized.jsonl",
673
+ "sha256": "c11a94d8f833d83427ea984a4f069aefaa3fa10a2b1085fb3a03542727f819cf"
674
+ },
675
+ "v3_core": {
676
+ "path": "eval/heldout/v3/baselines/base-2b/core/normalized.jsonl",
677
+ "sha256": "cc51471716342ecd466b50d5dd9b4ba805f5a6873cd96075b56d7c2e8336175e"
678
+ },
679
+ "v4": {
680
+ "path": "eval/heldout/v4/baselines/base-2b/core/normalized.jsonl",
681
+ "sha256": "93d84cdbe607691e2cef35ece3aa2dd5900f6410b0d72796b86d74c9156f3cdd"
682
+ },
683
+ "v5": {
684
+ "path": "eval/heldout/v5/open-baselines-v1/base-2b/core/normalized.jsonl",
685
+ "sha256": "75740c4279eb8dcc9c90b55f155916b3748c887e8c9d9e1d3a47edf7f4dbc34b"
686
+ }
687
+ },
688
+ "transfer_rows": {
689
+ "path": "analysis/decision-benchmark-v3-baselines/base-2b/ROWS.json",
690
+ "sha256": "62bf4f977d03eabd4f051aca5c8f7204f350eb995545aedeefa7bfa00bcbd115"
691
+ },
692
+ "evidence": [
693
+ {
694
+ "path": "results/accelerated-transfer-baselines-v1/base-2b/worker/COMPLETE.json",
695
+ "sha256": "0353d786b760d491c2a5ff3f8363ecd27a537f25a42dbeb2cea4db5d5eb77910"
696
+ },
697
+ {
698
+ "path": "results/accelerated-transfer-baselines-v1/base-2b/worker/METADATA.json",
699
+ "sha256": "745151846ade07431bc8ac964894bbd6ed234fb6a4665fcce223d91e09acaeeb"
700
+ },
701
+ {
702
+ "path": "results/accelerated-transfer-baselines-v1/base-2b/worker/predictions.jsonl",
703
+ "sha256": "347fc785ee59d5ca7694357b68a6c2e7bc7f863f1e26d94fc9b989391ed7d5fe"
704
+ },
705
+ {
706
+ "path": "analysis/decision-benchmark-v3-baselines/base-2b/REPORT.json",
707
+ "sha256": "9eb1db12e826a75dcda322960478bb153c4c1c24e5f5853d29a761b82bea36e7"
708
+ },
709
+ {
710
+ "path": "eval/expanded-public-extension-v1/COMPARATOR-REGISTRY.json",
711
+ "sha256": "b80be4270721cbeeac9e28beb10c6f19f75c1cbf2553e38d3d59e3a42d56a91c"
712
+ },
713
+ {
714
+ "path": "analysis/decision-benchmark-v3-baselines/base-2b/COMPOSABLE-MANIFEST.json",
715
+ "sha256": "c3e6f28b600776686afc56c1386cabf0869e69d5ca79f3ae6a0b17189eee968e"
716
+ }
717
+ ]
718
+ },
719
+ "Qwen3.5-4B": {
720
+ "panels": {
721
+ "old_core": {
722
+ "path": "eval/heldout/primary-v2/base-4b/core/normalized.jsonl",
723
+ "sha256": "d20ad845066ff81caaf60ebee44df941872bdc642da78b5f3b3da4a33dd9a08d"
724
+ },
725
+ "v3_core": {
726
+ "path": "eval/heldout/v3/baselines/base-4b/core/normalized.jsonl",
727
+ "sha256": "7541893d2b2ec0ec2ffc513f50c568eaef860b65bcef9a1747996c116fff9661"
728
+ },
729
+ "v4": {
730
+ "path": "eval/heldout/v4/baselines/base-4b/core/normalized.jsonl",
731
+ "sha256": "7a6d252ccaa72bf54b94445190bbad023fb465a3eaf4cf657532e2cdd4f750f5"
732
+ },
733
+ "v5": {
734
+ "path": "eval/heldout/v5/open-baselines-v1/base-4b/core/normalized.jsonl",
735
+ "sha256": "c01ad5daed65267e9aec40ed9e97ece1b6511e4cab17ebff09339b41524fb7f1"
736
+ }
737
+ },
738
+ "transfer_rows": {
739
+ "path": "analysis/decision-benchmark-v3-baselines/base-4b/ROWS.json",
740
+ "sha256": "37ab17d94d640ee46283e1f23723b9ea05133c22a8a9b62a97c756a89c5de785"
741
+ },
742
+ "evidence": [
743
+ {
744
+ "path": "results/remaining-transfer-baselines-v1/base-4b/worker/COMPLETE.json",
745
+ "sha256": "72abc4f87a9234f95eb3b0a72a9e7d66798779d5982c757faffb610003b1b242"
746
+ },
747
+ {
748
+ "path": "results/remaining-transfer-baselines-v1/base-4b/worker/METADATA.json",
749
+ "sha256": "871fed92cf762187015e59174a042808dbf3fc88f7791aa0ddd530ae923196ab"
750
+ },
751
+ {
752
+ "path": "results/remaining-transfer-baselines-v1/base-4b/worker/predictions.jsonl",
753
+ "sha256": "8c47edda469f8914eccc2ecafa91c245752812e830cb9f29445653102522854a"
754
+ },
755
+ {
756
+ "path": "analysis/decision-benchmark-v3-baselines/base-4b/REPORT.json",
757
+ "sha256": "c2a0aafbe2b7c85b3c07a4b9843182a3361654f57883311155aa3b05cfc7002d"
758
+ },
759
+ {
760
+ "path": "eval/expanded-public-extension-v1/COMPARATOR-REGISTRY.json",
761
+ "sha256": "b80be4270721cbeeac9e28beb10c6f19f75c1cbf2553e38d3d59e3a42d56a91c"
762
+ },
763
+ {
764
+ "path": "analysis/decision-benchmark-v3-baselines/base-4b/COMPOSABLE-MANIFEST.json",
765
+ "sha256": "c5966aab01f3d82ee26ed56dc26ed1277f4d3cbc976ecf28d2e8d3d190fd8fc0"
766
+ }
767
+ ]
768
+ },
769
+ "llm2jev-2b": {
770
+ "panels": {
771
+ "old_core": {
772
+ "path": "eval/heldout/matched-fla-v1/runs/core-fla-v1-shard0/llm2jev-2b/quality/normalized.jsonl",
773
+ "sha256": "5f0aaf2be47d11c42ecb228aeb6f317bf2074c80d398a66969f06490d5847760"
774
+ },
775
+ "v3_core": {
776
+ "path": "eval/heldout/v3/baselines/llm2jev-2b/core/normalized.jsonl",
777
+ "sha256": "339cd97706404b3d1e5af9d07ae6d3efc79cbb40310cfcf1406c64477e24e460"
778
+ },
779
+ "v4": {
780
+ "path": "results/open-reference-gap-v1/llm2jev-2b/v4/normalized.jsonl",
781
+ "sha256": "98e6b3bc2113e18e00442a86b085e06981560f43cf6a8807c3d1818b3fd27caf"
782
+ },
783
+ "v5": {
784
+ "path": "results/open-reference-gap-v1/llm2jev-2b/v5/normalized.jsonl",
785
+ "sha256": "36b9d9e96bae101eb3db6600c412585f48551ff5908072963a5dd435632e4d65"
786
+ }
787
+ },
788
+ "transfer_rows": {
789
+ "path": "analysis/decision-benchmark-v3-baselines/llm2jev-2b/ROWS.json",
790
+ "sha256": "a37a6b19549896854fd0fb9bb32f594bb8cb60b840aa49ead0c004dcea00778e"
791
+ },
792
+ "evidence": [
793
+ {
794
+ "path": "results/remaining-transfer-baselines-v1/llm2jev-2b/worker/COMPLETE.json",
795
+ "sha256": "ccd7e1f0d6a4193c629046961eb7171b1278a8dbae015ff216396ceac1f008ef"
796
+ },
797
+ {
798
+ "path": "results/remaining-transfer-baselines-v1/llm2jev-2b/worker/METADATA.json",
799
+ "sha256": "b0d0aa9840976e0330e676fdd67ce90230dbfec7502452956260188924583b2b"
800
+ },
801
+ {
802
+ "path": "results/remaining-transfer-baselines-v1/llm2jev-2b/worker/predictions.jsonl",
803
+ "sha256": "67a217eea9af720846f653d780973a64fff93111d8dcb9fa09b8b6f9594328a4"
804
+ },
805
+ {
806
+ "path": "analysis/decision-benchmark-v3-baselines/llm2jev-2b/REPORT.json",
807
+ "sha256": "bb2bf635114fc805e6fef66dce70ac6ceca126008b3e8365871a6d12d9cd6e09"
808
+ },
809
+ {
810
+ "path": "eval/open-reference-comparison-v1/REGISTRY.json",
811
+ "sha256": "af1a488c3afed2e4b3e1cfca544af2b31518025ffe2c2dd469889711ee7f47f9"
812
+ },
813
+ {
814
+ "path": "analysis/decision-benchmark-v3-baselines/llm2jev-2b/COMPOSABLE-MANIFEST.json",
815
+ "sha256": "071d3b47bd6a81e256e11be2ed2965c7416244bd512bd35be8baef5b8ebcb2be"
816
+ }
817
+ ]
818
+ },
819
+ "llm2jev-4b": {
820
+ "panels": {
821
+ "old_core": {
822
+ "path": "eval/heldout/matched-fla-v1/runs/core-fla-v1-shard1/llm2jev-4b/quality/normalized.jsonl",
823
+ "sha256": "8545cdc4ab9e9f870fdce1c1859632b8bac7505f6e961b4cc07e4f3b816328ce"
824
+ },
825
+ "v3_core": {
826
+ "path": "eval/heldout/v3/baselines/llm2jev-4b/core/normalized.jsonl",
827
+ "sha256": "b3bd38c90daf0bc3d77fea9a3ed0612049b07865bde5102356f82f7456002c97"
828
+ },
829
+ "v4": {
830
+ "path": "results/open-reference-gap-v1/llm2jev-4b/v4/normalized.jsonl",
831
+ "sha256": "409f8d2c57a1f8a4ef7dd18f3adb6cb8ee2ed544566206ccc909d54236ca95b7"
832
+ },
833
+ "v5": {
834
+ "path": "results/open-reference-gap-v1/llm2jev-4b/v5/normalized.jsonl",
835
+ "sha256": "d458e7e11cd2b1d66c7db6044408042630a48a474c9d3997141fc85dfb427f24"
836
+ }
837
+ },
838
+ "transfer_rows": {
839
+ "path": "analysis/decision-benchmark-v3-baselines/llm2jev-4b/ROWS.json",
840
+ "sha256": "abc8f7c04170544f5c4bef7a9af70ce0d7a25bfc4b4b8b1ba3e552413273bb7d"
841
+ },
842
+ "evidence": [
843
+ {
844
+ "path": "results/remaining-transfer-baselines-v1/llm2jev-4b/worker/COMPLETE.json",
845
+ "sha256": "d7745d5ec1e08100629c55e2c7518ff408462e6adea4ebfb01b7b401b212c4b6"
846
+ },
847
+ {
848
+ "path": "results/remaining-transfer-baselines-v1/llm2jev-4b/worker/METADATA.json",
849
+ "sha256": "6a522387522b92a0c01cbd8f32cdaa154289318a98564297d2f5723bc4c3dee0"
850
+ },
851
+ {
852
+ "path": "results/remaining-transfer-baselines-v1/llm2jev-4b/worker/predictions.jsonl",
853
+ "sha256": "61bea4eb25014f6277cf81010872c1724e6d8f07a89803063649a21160b02edc"
854
+ },
855
+ {
856
+ "path": "analysis/decision-benchmark-v3-baselines/llm2jev-4b/REPORT.json",
857
+ "sha256": "1097d4eeb4abc792cac569e4793c469a332dc24bbdf5568a242ce4cee4c03759"
858
+ },
859
+ {
860
+ "path": "eval/open-reference-comparison-v1/REGISTRY.json",
861
+ "sha256": "af1a488c3afed2e4b3e1cfca544af2b31518025ffe2c2dd469889711ee7f47f9"
862
+ },
863
+ {
864
+ "path": "analysis/decision-benchmark-v3-baselines/llm2jev-4b/COMPOSABLE-MANIFEST.json",
865
+ "sha256": "82a7589e866610251b5b52f5adb20ab5c7ed9fa6fb2a184c9d6a3367576171e7"
866
+ }
867
+ ]
868
+ },
869
+ "nimble-9b": {
870
+ "panels": {
871
+ "old_core": {
872
+ "path": "eval/heldout/matched-fla-v1/runs/core-fla-v1-shard1/nimble-9b/quality/normalized.jsonl",
873
+ "sha256": "6fefc71c6fbea1106894c1081057c2c7a253fa4ebc2e6aecfe23abc347be08bb"
874
+ },
875
+ "v3_core": {
876
+ "path": "eval/heldout/v3/baselines/nimble-9b/core/normalized.jsonl",
877
+ "sha256": "ab5323762fc5d34961c0fbce78497627528b696c60e18b880d671b468db97bf9"
878
+ },
879
+ "v4": {
880
+ "path": "results/open-reference-gap-v1/nimble-9b/v4/normalized.jsonl",
881
+ "sha256": "6ae263cef9cafdbcd35cd1f1dcc5b1bd920d3e6023d4fc487970286c04ec8d05"
882
+ },
883
+ "v5": {
884
+ "path": "eval/heldout/v5/open-baselines-v1/nimble-9b/core/normalized.jsonl",
885
+ "sha256": "0fc6dae372265ef62ea87ce9dc07b63a77ccd33e3aa48e15754f5ab98c32580a"
886
+ }
887
+ },
888
+ "transfer_rows": {
889
+ "path": "analysis/decision-benchmark-v3-baselines/nimble-9b/ROWS.json",
890
+ "sha256": "f55f29f973b0974adc7fc3fd582e2e5c0414f440e4a094a7ec8e152685e62a28"
891
+ },
892
+ "evidence": [
893
+ {
894
+ "path": "results/accelerated-transfer-baselines-v1/nimble-9b/worker/COMPLETE.json",
895
+ "sha256": "efc5a1be536227db99ff9937b84e9a5bc688fc497c0e4a84d8261144cea8a51d"
896
+ },
897
+ {
898
+ "path": "results/accelerated-transfer-baselines-v1/nimble-9b/worker/METADATA.json",
899
+ "sha256": "77eaa373ec6a74b483ac9b0cda739ecd8cce95a046266cd7a28488a768b82940"
900
+ },
901
+ {
902
+ "path": "results/accelerated-transfer-baselines-v1/nimble-9b/worker/predictions.jsonl",
903
+ "sha256": "ad2212fa691c662c7917b33315439844339252778fdaefcc3080cdf46e00a434"
904
+ },
905
+ {
906
+ "path": "analysis/decision-benchmark-v3-baselines/nimble-9b/REPORT.json",
907
+ "sha256": "b6065189b7522e5d0d2b7b088b673ba26212bee5fda2f0d7c6cd67fc9cc93136"
908
+ },
909
+ {
910
+ "path": "eval/open-reference-comparison-v1/REGISTRY.json",
911
+ "sha256": "af1a488c3afed2e4b3e1cfca544af2b31518025ffe2c2dd469889711ee7f47f9"
912
+ },
913
+ {
914
+ "path": "analysis/decision-benchmark-v3-baselines/nimble-9b/COMPOSABLE-MANIFEST.json",
915
+ "sha256": "7423a2430984cf8e7eed2c5933ceeb0e2894d252e30e447d1294b6b6b08192bd"
916
+ }
917
+ ]
918
+ }
919
+ },
920
+ "task_rows": 54,
921
+ "rank_public_models": [
922
+ "Jev",
923
+ "Lux",
924
+ "Nox",
925
+ "kev-9b",
926
+ "kev-4b",
927
+ "Qwen3.5-9B",
928
+ "Decider",
929
+ "Qwen3.5-4B",
930
+ "Sol",
931
+ "kev-0.8b",
932
+ "Qwen3.5-2B",
933
+ "Laya-base",
934
+ "Laya-multilingual",
935
+ "Kai"
936
+ ],
937
+ "internal_extra_models_not_public_rank": [
938
+ "llm2jev-2b",
939
+ "llm2jev-4b",
940
+ "nimble-9b"
941
+ ],
942
+ "laya_parameter_audit": {
943
+ "utc": "2026-09-22T05:01:58.295720+00:00",
944
+ "torch": "2.12.0+git6bbd260",
945
+ "transformers": "5.17.0",
946
+ "device": "meta, CPU-only process; no GPU devices exposed",
947
+ "models": {
948
+ "laya-english": {
949
+ "serialized_scalar_count": 421293830,
950
+ "unique_parameter_count": 421293827,
951
+ "parameter_entries": 205,
952
+ "named_parameter_entries_with_duplicates": 205,
953
+ "tied_parameter_aliases": [],
954
+ "persistent_buffers": {
955
+ "temperature": [
956
+ 3
957
+ ]
958
+ },
959
+ "persistent_buffer_scalar_count": 3,
960
+ "nonpersistent_buffers": {
961
+ "encoder.rotary_emb.full_attention_inv_freq": [
962
+ 32
963
+ ],
964
+ "encoder.rotary_emb.full_attention_original_inv_freq": [
965
+ 32
966
+ ],
967
+ "encoder.rotary_emb.sliding_attention_inv_freq": [
968
+ 32
969
+ ],
970
+ "encoder.rotary_emb.sliding_attention_original_inv_freq": [
971
+ 32
972
+ ]
973
+ },
974
+ "exact_state_dict_shapes_match_header": true,
975
+ "header_sha256": "3a42d8e7a96d5aa22c5bc9475d7ff092e32a832ae7c6973688f66f7a84048c04",
976
+ "source_sha256": "8d83611d480c971d640a7b7d3aa2f2219c5e8455e9cc2329fd073681bd8be23e",
977
+ "encoder_config_sha256": "bf3ab80598fdccf414855a2ce80f22859e4492d06ca8a62ddd1cfb63972f8979",
978
+ "agent_config_sha256": "ae287b56bbcf5f8c4f4541ae9dfd00c914c4c48b940b8398c3058af37ba92bbd",
979
+ "architecture": "ModernBertModel",
980
+ "interpretation": "Complete decision model, including encoder and heads; temperature calibration buffer excluded. No tied parameter aliases in the instantiated published model."
981
+ },
982
+ "laya-multilingual": {
983
+ "serialized_scalar_count": 321908998,
984
+ "unique_parameter_count": 321908995,
985
+ "parameter_entries": 169,
986
+ "named_parameter_entries_with_duplicates": 169,
987
+ "tied_parameter_aliases": [],
988
+ "persistent_buffers": {
989
+ "temperature": [
990
+ 3
991
+ ]
992
+ },
993
+ "persistent_buffer_scalar_count": 3,
994
+ "nonpersistent_buffers": {
995
+ "encoder.rotary_emb.full_attention_inv_freq": [
996
+ 32
997
+ ],
998
+ "encoder.rotary_emb.full_attention_original_inv_freq": [
999
+ 32
1000
+ ],
1001
+ "encoder.rotary_emb.sliding_attention_inv_freq": [
1002
+ 32
1003
+ ],
1004
+ "encoder.rotary_emb.sliding_attention_original_inv_freq": [
1005
+ 32
1006
+ ]
1007
+ },
1008
+ "exact_state_dict_shapes_match_header": true,
1009
+ "header_sha256": "22ab94329063133fdd2944997b906bb6076d2cfd3e43eb05d7ea92187e9c3984",
1010
+ "source_sha256": "8d83611d480c971d640a7b7d3aa2f2219c5e8455e9cc2329fd073681bd8be23e",
1011
+ "encoder_config_sha256": "83f6916d13ef0f556ac461f28308dc2bffa7ebeadee8ec9e2db5812020ea5bb4",
1012
+ "agent_config_sha256": "25061739243b617ad88d1219ba6f8a9c86c5881ca28df024fa2d9b3b2fcc30c6",
1013
+ "architecture": "ModernBertModel",
1014
+ "interpretation": "Complete decision model, including encoder and heads; temperature calibration buffer excluded. No tied parameter aliases in the instantiated published model."
1015
+ }
1016
+ },
1017
+ "source_sha256": "14fecede5c917ea0b1b4a8f7cfb84938f984fb4a2c721acaf7b77e2cb198ec64"
1018
+ }
1019
+ }
metrics/question-scaling.json ADDED
@@ -0,0 +1,342 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "status": "complete",
3
+ "model": "Lux",
4
+ "protocol_sha256": "ee6c8bfa8b43087ae88e8c85b27d444f0713f3cb2de3aed7e65e5f40ecb67b0a",
5
+ "source_lock_sha256": "3bc7bc8c828911548efc12518f1a67ae61ebe69a53d51627fe39cd6559e76f20",
6
+ "actual_sequence_complete_sha256": "eaecc2cffd804f54af7ab530e17076fb372127d1d920f98cf29c337be3e4c11e",
7
+ "hardware": {
8
+ "arch": "gfx942",
9
+ "SKU": "M3250101",
10
+ "PCI_device": "0x74b9",
11
+ "VRAM_bytes": 274542362624
12
+ },
13
+ "timing_scope": "Native local public decide, tokenization+inference+response+final synchronize; excludes model load/network/warmup",
14
+ "whole_host_quiet_observed": true,
15
+ "cases": {
16
+ "published_repeated-q1": {
17
+ "count": 30,
18
+ "p50_ms": 32.272068000000004,
19
+ "p95_ms": 32.778058449999996,
20
+ "mean_ms": 32.4287599,
21
+ "std_ms": 0.824897883231798,
22
+ "minimum_ms": 31.825991,
23
+ "maximum_ms": 36.620389,
24
+ "q": 1,
25
+ "family": "published_repeated",
26
+ "tokens_per_question": 309,
27
+ "p50_process_block_bootstrap_ci95_ms": [
28
+ 32.185123,
29
+ 32.489545
30
+ ],
31
+ "peak_allocated_gib": 14.918961524963379,
32
+ "peak_reserved_gib": 15.490234375,
33
+ "per_process_medians_ms": [
34
+ 32.333903,
35
+ 32.657675,
36
+ 32.215193,
37
+ 32.087982,
38
+ 32.301293,
39
+ 32.039621
40
+ ],
41
+ "distinct_output_hashes_across_blocks": 1
42
+ },
43
+ "distinct_primary-q16": {
44
+ "count": 30,
45
+ "p50_ms": 301.057595,
46
+ "p95_ms": 302.2201631,
47
+ "mean_ms": 300.5559644333333,
48
+ "std_ms": 1.6441091883766683,
49
+ "minimum_ms": 297.06975,
50
+ "maximum_ms": 302.726055,
51
+ "q": 16,
52
+ "family": "distinct_primary",
53
+ "tokens_per_question": 499,
54
+ "p50_process_block_bootstrap_ci95_ms": [
55
+ 298.694705,
56
+ 301.263677
57
+ ],
58
+ "peak_allocated_gib": 15.456972122192383,
59
+ "peak_reserved_gib": 15.490234375,
60
+ "per_process_medians_ms": [
61
+ 297.09632,
62
+ 300.987813,
63
+ 301.051835,
64
+ 301.095694,
65
+ 301.099985,
66
+ 301.585418
67
+ ],
68
+ "distinct_output_hashes_across_blocks": 1
69
+ },
70
+ "published_repeated-q4": {
71
+ "count": 30,
72
+ "p50_ms": 56.301522000000006,
73
+ "p95_ms": 56.5526525,
74
+ "mean_ms": 56.25305936666667,
75
+ "std_ms": 0.2542873912960911,
76
+ "minimum_ms": 55.804055,
77
+ "maximum_ms": 56.644,
78
+ "q": 4,
79
+ "family": "published_repeated",
80
+ "tokens_per_question": 309,
81
+ "p50_process_block_bootstrap_ci95_ms": [
82
+ 56.0213505,
83
+ 56.4477135
84
+ ],
85
+ "peak_allocated_gib": 15.055374145507812,
86
+ "peak_reserved_gib": 15.490234375,
87
+ "per_process_medians_ms": [
88
+ 55.858455,
89
+ 56.442498,
90
+ 56.392988,
91
+ 56.479279,
92
+ 55.936955,
93
+ 56.290938
94
+ ],
95
+ "distinct_output_hashes_across_blocks": 1
96
+ },
97
+ "distinct_primary-q2": {
98
+ "count": 30,
99
+ "p50_ms": 49.326505499999996,
100
+ "p95_ms": 49.7152335,
101
+ "mean_ms": 49.35776336666667,
102
+ "std_ms": 0.23140268932667862,
103
+ "minimum_ms": 48.959943,
104
+ "maximum_ms": 49.822548,
105
+ "q": 2,
106
+ "family": "distinct_primary",
107
+ "tokens_per_question": 499,
108
+ "p50_process_block_bootstrap_ci95_ms": [
109
+ 49.205204,
110
+ 49.554927
111
+ ],
112
+ "peak_allocated_gib": 15.019116878509521,
113
+ "peak_reserved_gib": 15.490234375,
114
+ "per_process_medians_ms": [
115
+ 49.097525,
116
+ 49.301715,
117
+ 49.578607,
118
+ 49.326126,
119
+ 49.359936,
120
+ 49.161194
121
+ ],
122
+ "distinct_output_hashes_across_blocks": 1
123
+ },
124
+ "distinct_primary-q8": {
125
+ "count": 30,
126
+ "p50_ms": 150.4095365,
127
+ "p95_ms": 150.9155465,
128
+ "mean_ms": 150.15292296666667,
129
+ "std_ms": 0.7668097199527638,
130
+ "minimum_ms": 148.43428,
131
+ "maximum_ms": 151.120616,
132
+ "q": 8,
133
+ "family": "distinct_primary",
134
+ "tokens_per_question": 499,
135
+ "p50_process_block_bootstrap_ci95_ms": [
136
+ 149.360541,
137
+ 150.798794
138
+ ],
139
+ "peak_allocated_gib": 15.456967830657959,
140
+ "peak_reserved_gib": 15.490234375,
141
+ "per_process_medians_ms": [
142
+ 148.596671,
143
+ 150.32225,
144
+ 150.880795,
145
+ 150.413842,
146
+ 150.667853,
147
+ 149.911599
148
+ ],
149
+ "distinct_output_hashes_across_blocks": 1
150
+ },
151
+ "published_repeated-q16": {
152
+ "count": 30,
153
+ "p50_ms": 195.256766,
154
+ "p95_ms": 196.56293449999998,
155
+ "mean_ms": 194.85459923333335,
156
+ "std_ms": 1.5003195937600042,
157
+ "minimum_ms": 192.072561,
158
+ "maximum_ms": 196.869201,
159
+ "q": 16,
160
+ "family": "published_repeated",
161
+ "tokens_per_question": 309,
162
+ "p50_process_block_bootstrap_ci95_ms": [
163
+ 193.0342415,
164
+ 196.03670449999998
165
+ ],
166
+ "peak_allocated_gib": 15.237587451934814,
167
+ "peak_reserved_gib": 15.490234375,
168
+ "per_process_medians_ms": [
169
+ 192.165662,
170
+ 196.462468,
171
+ 195.309722,
172
+ 194.139094,
173
+ 195.970924,
174
+ 195.538012
175
+ ],
176
+ "distinct_output_hashes_across_blocks": 1
177
+ },
178
+ "distinct_primary-q32": {
179
+ "count": 30,
180
+ "p50_ms": 600.736787,
181
+ "p95_ms": 605.47260555,
182
+ "mean_ms": 600.7545997666666,
183
+ "std_ms": 2.769019121212867,
184
+ "minimum_ms": 594.776044,
185
+ "maximum_ms": 605.874823,
186
+ "q": 32,
187
+ "family": "distinct_primary",
188
+ "tokens_per_question": 499,
189
+ "p50_process_block_bootstrap_ci95_ms": [
190
+ 598.681272,
191
+ 603.4955924999999
192
+ ],
193
+ "peak_allocated_gib": 15.456972122192383,
194
+ "peak_reserved_gib": 15.490234375,
195
+ "per_process_medians_ms": [
196
+ 595.632779,
197
+ 599.420382,
198
+ 602.509201,
199
+ 605.287299,
200
+ 600.955803,
201
+ 600.294559
202
+ ],
203
+ "distinct_output_hashes_across_blocks": 1
204
+ },
205
+ "published_repeated-q8": {
206
+ "count": 30,
207
+ "p50_ms": 98.165803,
208
+ "p95_ms": 98.678712,
209
+ "mean_ms": 98.03377306666667,
210
+ "std_ms": 0.5543181468288665,
211
+ "minimum_ms": 96.92902,
212
+ "maximum_ms": 99.674767,
213
+ "q": 8,
214
+ "family": "published_repeated",
215
+ "tokens_per_question": 309,
216
+ "p50_process_block_bootstrap_ci95_ms": [
217
+ 97.741066,
218
+ 98.279879
219
+ ],
220
+ "peak_allocated_gib": 15.23758316040039,
221
+ "peak_reserved_gib": 15.490234375,
222
+ "per_process_medians_ms": [
223
+ 97.555833,
224
+ 98.357609,
225
+ 98.279879,
226
+ 97.446595,
227
+ 98.149458,
228
+ 98.228989
229
+ ],
230
+ "distinct_output_hashes_across_blocks": 1
231
+ },
232
+ "published_repeated-q2": {
233
+ "count": 30,
234
+ "p50_ms": 36.246312,
235
+ "p95_ms": 36.75535655,
236
+ "mean_ms": 36.354916966666664,
237
+ "std_ms": 0.4110907613540518,
238
+ "minimum_ms": 35.992496,
239
+ "maximum_ms": 38.32468,
240
+ "q": 2,
241
+ "family": "published_repeated",
242
+ "tokens_per_question": 309,
243
+ "p50_process_block_bootstrap_ci95_ms": [
244
+ 36.208107,
245
+ 36.391368
246
+ ],
247
+ "peak_allocated_gib": 14.964269638061523,
248
+ "peak_reserved_gib": 15.490234375,
249
+ "per_process_medians_ms": [
250
+ 36.147877,
251
+ 36.205907,
252
+ 36.565229,
253
+ 36.230487,
254
+ 36.414958,
255
+ 36.253187
256
+ ],
257
+ "distinct_output_hashes_across_blocks": 1
258
+ },
259
+ "distinct_primary-q1": {
260
+ "count": 30,
261
+ "p50_ms": 33.173294,
262
+ "p95_ms": 33.51943595,
263
+ "mean_ms": 33.18662556666666,
264
+ "std_ms": 0.1599538249331253,
265
+ "minimum_ms": 32.947508,
266
+ "maximum_ms": 33.555061,
267
+ "q": 1,
268
+ "family": "distinct_primary",
269
+ "tokens_per_question": 499,
270
+ "p50_process_block_bootstrap_ci95_ms": [
271
+ 33.068598,
272
+ 33.235259
273
+ ],
274
+ "peak_allocated_gib": 14.946141719818115,
275
+ "peak_reserved_gib": 15.490234375,
276
+ "per_process_medians_ms": [
277
+ 33.008438,
278
+ 33.29688,
279
+ 33.171449,
280
+ 33.148899,
281
+ 33.107828,
282
+ 33.285728
283
+ ],
284
+ "distinct_output_hashes_across_blocks": 1
285
+ },
286
+ "distinct_primary-q4": {
287
+ "count": 30,
288
+ "p50_ms": 82.0101965,
289
+ "p95_ms": 82.69186805,
290
+ "mean_ms": 82.03453213333333,
291
+ "std_ms": 0.383264555025658,
292
+ "minimum_ms": 81.289207,
293
+ "maximum_ms": 82.758045,
294
+ "q": 4,
295
+ "family": "distinct_primary",
296
+ "tokens_per_question": 499,
297
+ "p50_process_block_bootstrap_ci95_ms": [
298
+ 81.6675695,
299
+ 82.402964
300
+ ],
301
+ "peak_allocated_gib": 15.165067195892334,
302
+ "peak_reserved_gib": 15.490234375,
303
+ "per_process_medians_ms": [
304
+ 81.432948,
305
+ 81.92596,
306
+ 82.525494,
307
+ 82.009601,
308
+ 81.874702,
309
+ 82.161412
310
+ ],
311
+ "distinct_output_hashes_across_blocks": 1
312
+ },
313
+ "published_repeated-q32": {
314
+ "count": 30,
315
+ "p50_ms": 390.274604,
316
+ "p95_ms": 393.28575205000004,
317
+ "mean_ms": 390.39152386666666,
318
+ "std_ms": 1.9873352241345295,
319
+ "minimum_ms": 387.001278,
320
+ "maximum_ms": 393.502758,
321
+ "q": 32,
322
+ "family": "published_repeated",
323
+ "tokens_per_question": 309,
324
+ "p50_process_block_bootstrap_ci95_ms": [
325
+ 388.994192,
326
+ 392.883485
327
+ ],
328
+ "peak_allocated_gib": 15.237587451934814,
329
+ "peak_reserved_gib": 15.490234375,
330
+ "per_process_medians_ms": [
331
+ 388.177386,
332
+ 389.472613,
333
+ 390.929603,
334
+ 393.144967,
335
+ 392.304752,
336
+ 388.60779
337
+ ],
338
+ "distinct_output_hashes_across_blocks": 1
339
+ }
340
+ },
341
+ "cross_SKU_speedup_claim": false
342
+ }
model-card-example.json ADDED
@@ -0,0 +1,134 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "requests": {
3
+ "model_card": {
4
+ "state": "Customer reports a duplicate charge and asks for a refund.",
5
+ "questions": {
6
+ "route": {
7
+ "type": "choice",
8
+ "instructions": "Which team should handle this request?",
9
+ "criteria": {
10
+ "billing": "Payments and refunds",
11
+ "technical": "Product faults"
12
+ }
13
+ },
14
+ "refund_requested": {
15
+ "type": "noul",
16
+ "instructions": "Did the customer request a refund?"
17
+ }
18
+ }
19
+ },
20
+ "usage": {
21
+ "state": {
22
+ "owner": "Lee",
23
+ "status": "active",
24
+ "severity": "medium"
25
+ },
26
+ "questions": {
27
+ "owner": {
28
+ "type": "choice",
29
+ "instructions": "Choose the owner.",
30
+ "criteria": {
31
+ "lee": "Lee",
32
+ "sam": "Sam"
33
+ }
34
+ },
35
+ "active": {
36
+ "type": "noul",
37
+ "instructions": "Is the status active?"
38
+ },
39
+ "severity": {
40
+ "type": "score",
41
+ "instructions": "Apply the severity scale.",
42
+ "criteria": [
43
+ "low",
44
+ "medium",
45
+ "high"
46
+ ]
47
+ }
48
+ }
49
+ }
50
+ },
51
+ "actual_responses": {
52
+ "model_card": {
53
+ "model": "Decision-1.0-Lux",
54
+ "answers": {
55
+ "route": {
56
+ "type": "choice",
57
+ "probabilities": {
58
+ "billing": 0.9996062011015111,
59
+ "technical": 0.00039379889848886117
60
+ },
61
+ "confidence": 0.9992124022030222,
62
+ "choice": "billing"
63
+ },
64
+ "refund_requested": {
65
+ "type": "noul",
66
+ "noul": 0.9959241870863982
67
+ }
68
+ },
69
+ "usage": {
70
+ "input_tokens": 182,
71
+ "scored_questions": 2
72
+ }
73
+ },
74
+ "usage": {
75
+ "model": "Decision-1.0-Lux",
76
+ "answers": {
77
+ "owner": {
78
+ "type": "choice",
79
+ "probabilities": {
80
+ "lee": 0.9992592420970957,
81
+ "sam": 0.0007407579029042402
82
+ },
83
+ "confidence": 0.9985184841941914,
84
+ "choice": "lee"
85
+ },
86
+ "active": {
87
+ "type": "noul",
88
+ "noul": 0.9975156124461363
89
+ },
90
+ "severity": {
91
+ "type": "score",
92
+ "probabilities": {
93
+ "0": 0.006993044617138876,
94
+ "1": 0.9872190394607501,
95
+ "2": 0.005787915922111086
96
+ },
97
+ "confidence": 0.9808285591911252,
98
+ "score": 0.9987948713049722,
99
+ "legend": {
100
+ "0": "low",
101
+ "1": "medium",
102
+ "2": "high"
103
+ }
104
+ }
105
+ },
106
+ "usage": {
107
+ "input_tokens": 278,
108
+ "scored_questions": 3
109
+ }
110
+ }
111
+ },
112
+ "runtime": {
113
+ "actual": {
114
+ "torch": "2.12.0+git6bbd260",
115
+ "hip": "7.2.53211",
116
+ "transformers": "5.17.0",
117
+ "fla": "0.5.2",
118
+ "tokenizers": "0.23.2",
119
+ "safetensors": "0.8.0",
120
+ "triton": "3.7.1",
121
+ "gated_delta": "fla.ops.gated_delta_rule.chunk"
122
+ },
123
+ "differences": {},
124
+ "matches_validated_runtime": true,
125
+ "normalization_profile": {
126
+ "profile_sha256": "cd73a4df3146d950e900acce4e5cff7ab61f79d866cc86c77035e098ecedcb68",
127
+ "guard_sha256": "d3ceca1abe8f92dd378ecccff076f7a2c5438691d2fd734e9c026388d2188205",
128
+ "validated_arch": "gfx942",
129
+ "automatic_bundle_binding": true,
130
+ "scope": "single profile per process"
131
+ }
132
+ },
133
+ "mode": "isolated exact readme and usage Python examples"
134
+ }
pyproject.toml ADDED
@@ -0,0 +1,19 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ [build-system]
2
+ requires = ["setuptools>=68"]
3
+ build-backend = "setuptools.build_meta"
4
+
5
+ [project]
6
+ name = "decision-local"
7
+ version = "1.1.0"
8
+ description = "Local typed inference for exported Decision decoder models"
9
+ requires-python = ">=3.10"
10
+ dependencies = []
11
+
12
+ [project.optional-dependencies]
13
+ hub = ["huggingface-hub==1.31.0"]
14
+
15
+ [project.scripts]
16
+ decision-example = "decision.example:main"
17
+
18
+ [tool.setuptools.packages.find]
19
+ where = ["src"]
release-manifest.json ADDED
@@ -0,0 +1,48 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "model": "Decision-1.0-Lux",
3
+ "repository": "llm-semantic-router/Decision-1.0-Lux",
4
+ "base_model": "Qwen/Qwen3.5-9B",
5
+ "base_revision": "c202236235762e1c871ad0ccb60c8ee5ba337b9a",
6
+ "deployed_parameters": 7940895744,
7
+ "bundle_manifest_sha256": "cebb32a9d277c26a0682be3b14fbd242136818130b3150b54f4f39e09fad2299",
8
+ "public_proof_sha256": "506e9f9cfcfaf8475839f712b1ebad37e45444de4f81076ba68527e1b9cd82e7",
9
+ "benchmark_protocol_sha256": "ab97f652ab890eaefdd5373dd2deeb1e26588815f1f339eee4bcaa610846b0af",
10
+ "statistics_sha256": "1db772aa92754ce458965f8b449553dde4717d1400dd8fd509ce32d9c76f1d5c",
11
+ "displayed_materials": {
12
+ "ATTRIBUTIONS.md": "f13d3d134767e99988e721a35500a904ed0fdb72bb48aaf8519b06da3629473d",
13
+ "DIAGNOSTICS.md": "5492ed49869e45e45650c2501d02a78f9a2fbc17267c54f90cf034a01ebda2b3",
14
+ "Dockerfile.runtime": "6c781fdb5ad2abfb8f61c906751ac449d01644c62704cbe28312fd7b7fb57521",
15
+ "EVALUATION.md": "7ea30f8dee24e3c05a52e28c07612fc0cfe158e122cf8a0553078558da28d87b",
16
+ "LICENSE": "bbedc3fda3305820b977265f01b8619d87570a6739de3a5582c3464840f1e57a",
17
+ "METHODS.md": "8cc5a8edb9cad84736e19b712a9b9d141883ba7ead5aa3cdd45567180c1db7a4",
18
+ "QUESTION-SCALING.md": "2a04e4f06132108e022fdc694d9d66c4e30bbd5206149844649138f322e5ca18",
19
+ "QWEN-LICENSE": "bbedc3fda3305820b977265f01b8619d87570a6739de3a5582c3464840f1e57a",
20
+ "README.md": "128d35a759d269cfe88147c59b6335739155b3f86fb7182e5be6438bfc019ee4",
21
+ "RUNTIME.md": "91284e3281944e60330258fcbd2617905671da7330f02cac1b0217c987a30095",
22
+ "SENSITIVITY.md": "e03ec18140d5c18e606da23cc5b26ef41176ab38019912d85050677e4637f0a0",
23
+ "TASKS.md": "e87c5271049f09ce23fcb7fdcbaa9f03e5c50039b6d7a4005a6b4c7fbc1d30ca",
24
+ "USAGE.md": "49888ddf4ac8b87d60556a1f44083b4789c9c19d95de460514a26af1f264bc17",
25
+ "WEIGHTING.md": "e03ec18140d5c18e606da23cc5b26ef41176ab38019912d85050677e4637f0a0",
26
+ "assets/architecture.pdf": "482fac3bd270e5155e7d7fe61898c113279214222e4c2124b3e495b04b0fe1ea",
27
+ "assets/architecture.png": "1e8fbe7655e7c1aadb4b2323290cdb20e726c0a42b7914c0a1d5954a6a029931",
28
+ "assets/architecture.svg": "76a1f3165173ff07a4126947b9b6ae49da46f64ee2b1d740a3d9a35652fb6bae",
29
+ "assets/decision-matrix.pdf": "d1651d91db893ffbdf3261ec476241eb03c18c7a769d19978c668d60a1ac844f",
30
+ "assets/decision-matrix.png": "b6e20d8a2dd7a9bd6cf923e1b99cb1ac8f837532012d5f3e6f0cb245b12f0acc",
31
+ "assets/decision-matrix.svg": "dc92d1e3f3cf0518ee44f1bc512aae74732e1db82b919774a52c0da7c744b0b4",
32
+ "assets/decision-question-scaling.pdf": "d89a843d5b59b1bdd01124ec5e0733a41fa0cf8715ad9adf77cf8e525d749bc2",
33
+ "assets/decision-question-scaling.png": "fcab4ce4d6ae4876f9110d0231796926faa139a1fbfcd6c3e8ab4f44917be7e7",
34
+ "assets/decision-question-scaling.svg": "a88c94ace9d3acd6c6171943024d73cdee9e0301d5f4e7287bec488ddcf0b5db",
35
+ "assets/decision-ranking.pdf": "f3e037ada6c30871d9ebe0087fa3d122346f7ab1c927515878d524c492b7799a",
36
+ "assets/decision-ranking.png": "1bef771a47e7a770c44bd174aaf8c33db3ec8c5dfa20087d5eff15f5a18848e0",
37
+ "assets/decision-ranking.svg": "c8de982ba51fb9d2b729191018220964426a0f524a00a796553b8879db271d1a",
38
+ "assets/readout.png": "fb6b3d1cf149a1a3ccb584eb1c213df73c07bcd328732b4238928e2a2a4af085",
39
+ "assets/readout.svg": "701084bf7b0858a3adb247ae446df737b77d4b75004341dfd5f956b933f3d253",
40
+ "metrics/benchmark.json": "50dac7884849cbdb8be58c1c82b811f51543a08a548253cb8295f4e6444c4718",
41
+ "metrics/evaluation-provenance.json": "1f9c8b1e03d521160a8426798394418748daa72014dfc7272578d1e83fe237b4",
42
+ "metrics/question-scaling.json": "859c474c4c1fb571c6041574698f6f47d4982fb56f1f2e68820d8a99b3a41337",
43
+ "model-card-example.json": "23915dd0232955a2345f53a4e716f43c3816e536a3e5418030db24687644de88",
44
+ "runtime-fla-requirements.lock": "35e1fedff9ca49092a7ada277bbd794abe7e8474e14a678a7e4bcec6406d0eb5"
45
+ },
46
+ "weight_identity": "all weights and runtime bound by bundle-manifest.json",
47
+ "shared_diagnostics_materials_sha256": "47d511c7918112304b194e2797a7b0d9847312dacc705a0aca204b7bc53f4b0e"
48
+ }
runtime-fla-requirements.lock ADDED
@@ -0,0 +1,4 @@
 
 
 
 
 
1
+ # Only the FLA overlay; all other dependencies are fixed by the public base digest.
2
+ # Install with --no-deps --require-hashes to preserve the ROCm Torch/Triton builds.
3
+ fla-core @ https://files.pythonhosted.org/packages/2d/ed/dfe19c4da779957eb6a42a26812f9b4e2280bf757a17a71933ff59ffcb98/fla_core-0.5.2-py3-none-any.whl --hash=sha256:5e830c85bad3d0d34677f98ac7074d08687a3756f0f0499d95ceb96eb6920761
4
+ flash-linear-attention @ https://files.pythonhosted.org/packages/90/d2/2070e3cf2148c5cce99ca4876633c4b7b89ec085323600bd3d47aeacd306/flash_linear_attention-0.5.2-py3-none-any.whl --hash=sha256:dcf405d81f5426393b59037097aa700d0f4a841465d5028d5aa543f4502f2400