tnh0527 commited on
Commit
bf518ed
·
verified ·
1 Parent(s): 2e1d6d0

Publish ettin-150m-memory-reranker-ft-v1 (qualified ONNX export and complete attribution)

Browse files
MODIFICATIONS.md CHANGED
@@ -16,21 +16,24 @@ What changed:
16
  not human annotations.
17
  - **Export.** The merged model was exported to ONNX graphs taking
18
  `input_ids` and `attention_mask` and emitting one score per query-passage
19
- pair. CUDA uses FP16. DirectML widens the same stored weights and floating
20
- calculations to FP32, and uses compatible inferred reshapes. Both graphs
 
 
21
  share one external parameter file; the CUDA graph and weights are unchanged.
22
- - **Serving contracts.** `serving.json` and `serving.directml.json` use the
23
  same descriptor format for each graph: 1,153 total pair tokens and batch
24
- ceilings of 16 on CUDA and 8 on DirectML. CPU handles permitted control
25
  operations, not a fallback execution of the whole reranker.
26
 
27
  Files modified or generated by Daecore:
28
 
29
  - `model.onnx`: the generated ONNX graph of the fine-tuned reranker;
30
- - `model.directml.onnx`: the FP32 calculation graph for DirectML;
31
  - `model.onnx.data`: the fine-tuned parameters shared by both graphs;
32
  - `config.json`: the export configuration of the modified reranker;
33
- - `serving.json` and `serving.directml.json`: the runtime and sequence contracts.
 
34
 
35
  Unchanged: `tokenizer.json` and `tokenizer_config.json` preserve the upstream
36
  tokenizer contract.
 
16
  not human annotations.
17
  - **Export.** The merged model was exported to ONNX graphs taking
18
  `input_ids` and `attention_mask` and emitting one score per query-passage
19
+ pair. CUDA uses FP16. Vulkan widens the same stored weights and floating
20
+ calculations to FP32, and uses compatible inferred reshapes. Bounded
21
+ integer/boolean attention-mask calculations are expressed exactly in FP32
22
+ before restoring their original output types. Both graphs
23
  share one external parameter file; the CUDA graph and weights are unchanged.
24
+ - **Serving contracts.** `serving.json` and `serving.vulkan.json` use the
25
  same descriptor format for each graph: 1,153 total pair tokens and batch
26
+ ceilings of 16 on CUDA and 8 on Vulkan. CPU handles permitted control
27
  operations, not a fallback execution of the whole reranker.
28
 
29
  Files modified or generated by Daecore:
30
 
31
  - `model.onnx`: the generated ONNX graph of the fine-tuned reranker;
32
+ - `model.vulkan.onnx`: the derived FP32 calculation graph for Vulkan;
33
  - `model.onnx.data`: the fine-tuned parameters shared by both graphs;
34
  - `config.json`: the export configuration of the modified reranker;
35
+ - `serving.json` and `serving.vulkan.json`: the runtime and sequence contracts;
36
+ - `vulkan-derivation.json`: exact source and derived graph/weight identities.
37
 
38
  Unchanged: `tokenizer.json` and `tokenizer_config.json` preserve the upstream
39
  tokenizer contract.
README.md CHANGED
@@ -1,219 +1,122 @@
1
- ---
2
- license: apache-2.0
3
- base_model: cross-encoder/ettin-reranker-150m-v1
4
- pipeline_tag: text-ranking
5
- tags:
6
- - ettin
7
- - modernbert
8
- - cross-encoder
9
- - reranker
10
- - fine-tuned
11
- - onnx
12
- ---
13
-
14
- # Ettin-150M memory reranker v1
15
-
16
- ## Summary
17
-
18
- This model reorders the fixed 50-passage pool that a hybrid first stage
19
- produces for a memory query. It is a fine-tune of
20
- [`cross-encoder/ettin-reranker-150m-v1`](https://huggingface.co/cross-encoder/ettin-reranker-150m-v1)
21
- trained with a graded listwise objective on 144,100 unique query-passage pairs
22
- across 2,882 queries.
23
-
24
- On the 441 development queries whose fixed pool holds at least one useful
25
- passage, it raised nDCG@10 from 0.562 for the fused first stage to 0.762 and
26
- Hit@1 from 0.644 to 0.807. In the protected promotion read, on the same 865
27
- answerable pools, it scored nDCG@10 0.728 against 0.595 for the previous
28
- reranker, +0.133 with a one-sided 95% lower bound of +0.119.
29
-
30
- All relevance grades are frontier-model judgments under a frozen protocol,
31
- the corpus is mostly generated documents, and the per-query evidence behind
32
- the development and gate reads was lost after the aggregates were recorded.
33
- The limitations section spells out each.
34
-
35
- ## Model details
36
-
37
- | Item | Value |
38
- |---|---|
39
- | Architecture | Ettin-150M (ModernBERT-family) cross-encoder with a trained scoring head; rank-16 LoRA+ adapter merged into the weights |
40
- | Input | one query and one passage, 1,153 tokens total |
41
- | Pool | a fixed 50-passage candidate set; the model reorders it and never adds to it |
42
- | Output | one relevance score per pair, used for ordering |
43
- | Serving precision | FP16 on CUDA; FP32 calculations on DirectML, sharing the same stored weights |
44
- | Deployment state | CUDA release plus an approved DirectML graph; installation is a separate update |
45
-
46
- ## Intended use
47
-
48
- Final ordering inside a local memory-retrieval pipeline: the dense retriever
49
- and BM25 each retrieve, reciprocal-rank fusion merges them, and this model
50
- scores the top 50. It cannot recover a passage the first stage missed; on the
51
- full 553-query development panel the 112 queries with no useful passage in
52
- their pool stay at zero before and after reranking.
53
-
54
- ## Training data
55
-
56
- | Item | Value |
57
- |---|---|
58
- | Queries | 2,882 answerable, plus 30 explicit no-positive rows kept out of the ranking loss |
59
- | Unique pairs | 144,100, every one seen; 184,448 list presentations |
60
- | Unique passages | 35,006 |
61
- | Grade counts 0 / 1 / 2 / 3 | 32,331 / 37,258 / 43,731 / 30,780 |
62
- | Query mix | 1,318 anchor-conditioned, 1,270 unanchored, 294 lexical-style |
63
- | Protected queries or families in gradients | 0 |
64
-
65
- Grades come from final consensus ledgers, so every grade-1 label is usable
66
- supervision rather than a masked unknown. The lexical-style queries (terse,
67
- exact, pasted, or lightly misspelled) are there so the model learns the
68
- lexical lane's cases as well as rich semantic questions.
69
-
70
- ## Recipe
71
-
72
- Grade-gain ListNet over lists of 16 with an adjacent-grade RankNet auxiliary
73
- at weight 0.2; LoRA+ at rank 16 and alpha 32 on all linear layers, learning
74
- rate 5e-5 for the A matrices and the score head and 8e-4 for the B matrices;
75
- 2,882 optimizer updates; FP16. Merge and reload changed scores by 0.
76
-
77
- ## Evaluation
78
-
79
- **Development panel.** 553 queries with every one of the 27,650 pool
80
- passages graded; 441 have a useful passage in their pool.
81
-
82
- | Stage | nDCG@10 | Hit@1 | Hit@5 | Hit@10 | Hit@20 | Hit@50 |
83
- |---|---:|---:|---:|---:|---:|---:|
84
- | fused first stage, 441 answerable | 0.5617 | 0.6440 | 0.8662 | 0.9206 | 0.9569 | 1.0000 |
85
- | this model, 441 answerable | 0.7621 | 0.8073 | 0.9410 | 0.9615 | 0.9841 | 1.0000 |
86
- | fused first stage, all 553 | 0.4479 | 0.5136 | 0.6908 | 0.7342 | 0.7631 | 0.7975 |
87
- | this model, all 553 | 0.6078 | 0.6438 | 0.7505 | 0.7667 | 0.7848 | 0.7975 |
88
-
89
- Pairwise accuracy on the answerable pools is 0.820. The all-query rows are the
90
- honest end-to-end view; Hit@50 is unchanged by design.
91
-
92
- **How it was chosen.** Against the consensus hard-label baseline (0.680
93
- nDCG@10 on this panel), a Lambda-style objective lost, ListNet at width 8 tied,
94
- and width 16 without a higher learning rate lost. The adjacent-grade term at
95
- 5e-5 reached 0.715. Restarting from the base over every paid pair gave 0.728
96
- with pure ListNet, 0.731 with the auxiliary and standard LoRA, and 0.732 with
97
- DoRA. LoRA+ at rank 16 reached 0.762 and was the only large optimization gain;
98
- rank 32 tied it (+0.0006) and a further top-10 continuation lost 0.005.
99
-
100
- **Protected gate, same-pool comparison.** On the 865 sealed queries whose
101
- candidate pool holds a useful passage, this model scored nDCG@10 0.7280
102
- against 0.5952 for the previous reranker, +0.1328 with a one-sided 95% lower
103
- bound of +0.1191 against a registered 0.02 margin. The judge-noise band for
104
- this comparison was 0.0036. The full pipeline reached Hit@1 0.7609 and nDCG@10
105
- 0.6692 on the 895 union-answerable queries, against 0.6078 and 0.4988 before.
106
-
107
- ## Runtime
108
-
109
- The FP16 ONNX export removes CPU-hosted sequence operations from all 22 MLP
110
- layers so the whole graph runs on the GPU. It is bounded-quality equivalent
111
- rather than bit-exact to the Torch model: as recorded at qualification,
112
- +0.0002 nDCG@10 and +0.0023 Hit@1 on the development pools, 4 of 441 top
113
- choices and 15 of 441 top-10 memberships changed, and a 95th-percentile score
114
- drift of 0.0098. A top-50 pool reranks in 0.664 s median and 1.032 s at the
115
- 95th percentile on an RTX 3060 Ti; the one-second target is advisory.
116
-
117
- The export regenerated on 2026-09-04 reproduced the qualified graph byte for
118
- byte. A parity recheck the same day had to use the training partition (2,882
119
- queries, 144,100 pairs) because the qualification pools were lost, and it
120
- fails the frozen policy on two checks: Hit@1 moved by −0.0003, one query,
121
- and one passage moved 16 places against a limit of 4. All 10 top-1 changes
122
- and all 73 top-10 membership changes on that panel occur between passages
123
- whose FP16 scores are exactly tied in one run and one unit apart in the
124
- other, so the tie-break decides the order; the largest score change, 0.08,
125
- sits at rank 47. On that panel the export reranked a pool in 0.584 s median
126
- and 0.912 s at the 95th percentile, and the Torch FP16 model in 0.532 s and
127
- 0.777 s.
128
-
129
- ### DirectML comparison
130
-
131
- The DirectML package adds `model.directml.onnx` alongside the unchanged CUDA
132
- `model.onnx`. Both read the same `model.onnx.data`; no second parameter file
133
- or training framework is needed at runtime. DirectML uses full-precision
134
- calculations because the half-precision graph had larger score differences.
135
- CUDA keeps its faster half-precision graph. Each provider plans batches from
136
- available memory and pair length, with measured ceilings of 16 rows for CUDA
137
- and 8 for DirectML.
138
-
139
- On a current 970-query replay (48,500 fixed-pool pairs), DirectML completed
140
- without a failure: 2.231 seconds median, 3.247 seconds at the 95th percentile,
141
- and 4.286 seconds maximum on an RTX 3060 Ti. No call exceeded ten seconds.
142
- This panel retains its per-query evidence; 2,498 pairs are unjudged, and their
143
- grades remain unknown. It is separate from the lost historical gate evidence.
144
-
145
- Against the published CUDA graph, nDCG@10 changed by +0.000017. Eight first
146
- choices and 36 top-ten sets changed. Both versions returned known-useful
147
- evidence for 816 of 970 queries; DirectML returned 5,528 known-useful passages
148
- versus 5,524. The unchanged calibration was applied for comparison only; its
149
- DirectML registration is not implied by these results.
150
-
151
- Two strict cross-precision checks fail and were explicitly accepted:
152
- Hit@1 and Hit@5 each lose one query, and one passage moves five positions
153
- against a limit of four. The affected useful passages remain in the returned
154
- sets; the five-position move is rank 33 to 28, outside the returned three.
155
- On the 482 fully judged pools, no Hit metric decreases.
156
-
157
- Running the same FP32 graph on CUDA for all 165 queries with changed
158
- deliveries or relevant ranking differences isolates the provider error:
159
- 8,250 pairs have a 95th-percentile score difference of 0.00000954 and a
160
- maximum of 0.000149. One pair exceeds the diagnostic 0.0001 bound; one full
161
- ordering changes and no first choice changes. The two Hit losses and the
162
- five-position move also occur on CUDA FP32, identifying precision rather
163
- than DirectML as their cause.
164
-
165
- On the supported sequential build policy, concurrent embedding and recall
166
- kept reranking below 3.2 seconds in three rounds. Eight queued recalls across
167
- four threads completed within 9.8 seconds. Forcing all three GPU models to
168
- remain resident on this 8 GB card exceeded the memory budget and failed the
169
- ten-second limit; those forced-overload timings are not a passing result.
170
- Worker-kill recovery completed for reranking, embedding, and classification.
171
-
172
- ## Limitations, ranked
173
-
174
- 1. **Model-judged relevance.** No human grades; session agreement 0.858 and
175
- kappa 0.682 on a 419-query audit. Scores measure agreement with that
176
- instrument.
177
- 2. **Generated corpus.** Mostly fictional document worlds chosen for breadth;
178
- behavior on other real corpora is unmeasured.
179
- 3. **Lost per-query evidence.** The development ledger, gate judgments, pool
180
- scores, and gate report were destroyed in a local storage incident after
181
- these aggregates were recorded. The numbers stand as recorded and cannot be
182
- re-audited; the recipe-search arms survive only as recorded results.
183
- 4. **One training run per arm.** Several selection margins (rank 16 versus 32,
184
- width 8 versus 16) are inside a plausible seed effect; compute set that limit.
185
- 5. **First-stage ceiling.** 112 of 553 development queries and 95 of 990 gate
186
- queries have no useful passage in any top-50 pool; ranks 51 to 200 were
187
- never judged.
188
- 6. **Provider limits.** CPU reranking is not qualified. DirectML has been
189
- measured on an NVIDIA RTX 3060 Ti, not AMD or Intel hardware, and its
190
- cross-precision policy exceptions above were accepted with disclosure.
191
- Hosts without a qualified reranker serve fusion order with a typed notice.
192
- 7. **The reproducible parity read fails the frozen policy.** The recorded
193
- development-panel pass cannot be re-run, and the 2026-09-04 recheck on the
194
- training partition fails two checks at exact FP16 ties. The operator
195
- accepted the recorded qualification with this disclosure rather than
196
- exporting an FP32 graph or amending the policy. The current FP32 provider
197
- comparison is separate evidence and does not erase that historical result.
198
-
199
- ## Reproducibility
200
-
201
- - Model tree SHA-256 `bd8028b5dbc6ab2cd2ab9a6019346de7b25536f50022617194bcb8503d4467ad`;
202
- base revision `025501c4e0f9bbeb4c5b198318e0089ff061cc14`, tree
203
- `e0edfb38cc495f23507de50233d0965f4e72ad0872e70e6f4e8fc889d394ea01`.
204
- - Training data manifest `286318992c584ae790a72a5a86b0e7d03d8860f8ae93c35f6cead4b08b4d395a`,
205
- schedule `e98d3cec2533b6edcded7c0b5017cba432e4af0e73fd68fca84997f934fa8f63`.
206
- - Recipe id `final-listnet-adjacent-loraplus16-full-coverage-w16-lr5e5`;
207
- registry manifest `../model_registry/ettin-150m-memory-reranker-ft-v1.json`;
208
- experiment log `../history/README.md`.
209
- - CUDA graph SHA-256 `7df4a3b3cd89d216c016abd7d169f0fd215e9f79c06a2ffed57bac1abf556073`.
210
- - DirectML graph SHA-256 `05abdc03628b599712a993381919aef1e22fcc5ad6384608abc2e069c402e369`.
211
- - Shared weights SHA-256 `78ddab681c2f506d62bbcd475fec5c10c62d867ed60d97a0170828efaff5bdde`.
212
- - CUDA serving-contract SHA-256 `ce05a6a810007372c6de0f4c56dc94f652e1d28f2eeddf1878b64a170b4c48c2`.
213
- - Build stack: Transformers 5.2.0, Sentence Transformers 5.5.1, PyTorch
214
- 2.10.0; export opset 18.
215
-
216
- ## License
217
-
218
- Apache 2.0, following the upstream model; the distribution includes the
219
- license, a notice, and a modification disclosure.
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: cross-encoder/ettin-reranker-150m-v1
4
+ pipeline_tag: text-ranking
5
+ language:
6
+ - en
7
+ tags:
8
+ - ettin
9
+ - modernbert
10
+ - cross-encoder
11
+ - reranker
12
+ - fine-tuned
13
+ - onnx
14
+ ---
15
+
16
+ # Ettin-150M memory reranker
17
+
18
+ An English cross-encoder for ranking passages from notes, procedures, decisions and technical documentation. Give it a query and candidate passages; it returns one relevance score per pair. It is a fine-tune of [Ettin Reranker 150M](https://huggingface.co/cross-encoder/ettin-reranker-150m-v1) and runs locally through ONNX Runtime.
19
+
20
+ This revision adds a Vulkan FP32 graph, `model.vulkan.onnx`; the learned weights and the CUDA FP16 graph are unchanged. It omits the DirectML graph, which remains in the previous revision.
21
+
22
+ ## Quick start
23
+
24
+ Install `tokenizers`, `numpy` and `onnxruntime-gpu[cuda,cudnn]==1.24.4`. Keep `model.onnx`, `model.onnx.data` and the tokenizer from the same package together. This CUDA example assumes that package is in `downloaded-model` and scores two passages:
25
+
26
+ ```python
27
+ import numpy as np
28
+ import onnxruntime as ort
29
+ from pathlib import Path
30
+ from tokenizers import Tokenizer
31
+
32
+ root = Path("downloaded-model") # The exact downloaded serving package.
33
+ tokenizer = Tokenizer.from_file(f"{root}/tokenizer.json")
34
+ tokenizer.no_padding()
35
+ tokenizer.no_truncation()
36
+ ort.preload_dlls()
37
+ session = ort.InferenceSession(
38
+ f"{root}/model.onnx", providers=[("CUDAExecutionProvider", {"use_tf32": "0"})]
39
+ )
40
+ if "CUDAExecutionProvider" not in session.get_providers():
41
+ raise RuntimeError("CUDA did not initialize; check the ONNX Runtime CUDA dependencies.")
42
+ query = "What must happen before a database migration?"
43
+ passages = [
44
+ "Take a verified backup before applying the migration.",
45
+ "The dashboard theme uses a blue background.",
46
+ ]
47
+ query_tokens = tokenizer.encode(query, add_special_tokens=False)
48
+ if len(query_tokens.ids) > 128:
49
+ raise ValueError("Query exceeds the qualified 128-token content limit.")
50
+ scores = []
51
+ for passage in passages:
52
+ passage_tokens = tokenizer.encode(passage, add_special_tokens=False)
53
+ passage_tokens.truncate(1022)
54
+ pair = tokenizer.post_process(query_tokens, passage_tokens, add_special_tokens=True)
55
+ inputs = {
56
+ "input_ids": np.array([pair.ids], dtype=np.int64),
57
+ "attention_mask": np.array([pair.attention_mask], dtype=np.int64),
58
+ }
59
+ scores.append(float(session.run(None, inputs)[0].reshape(-1)[0]))
60
+ print(sorted(zip(passages, scores), key=lambda item: item[1], reverse=True))
61
+ ```
62
+
63
+ Higher scores rank first; raw scores are not probabilities. The input envelope is 128 query-content tokens plus 1,022 passage-content tokens and three special tokens, a maximum of **1,153 tokens**. The example rejects an oversized query and truncates the passage, matching the tested input policy.
64
+
65
+ ## Evaluation
66
+
67
+ Upstream and fine-tuned Ettin scored the same **970 pools of 50 candidates**, with all 48,500 pairs judged. The table uses the **873 answerable pools** that contain at least one useful passage, so it measures ordering after retrieval. “Upstream” is the exact starting checkpoint without Daecore fine-tuning.
68
+
69
+ | Metric, answerable pools | Upstream Ettin | Daecore fine-tune |
70
+ |---|---:|---:|
71
+ | Hit@3 | 83.16% | 92.55% |
72
+ | Precision@3 | 59.76% | 80.03% |
73
+ | nDCG@10 | 0.5519 | 0.7407 |
74
+
75
+ ![Ettin ordering quality across return depths](figures/ettin-comparison.svg)
76
+
77
+ Grades 2 and 3 count as useful. Hit asks whether at least one useful passage appears, precision measures the useful share of returned positions, and nDCG rewards higher grades near the top. Many pools contain several useful passages, so random ordering already reaches 97.30% Hit@20; early precision and graded ranking are more informative than deep Hit. The comparison runs both models in PyTorch FP16 with matched tokenization on reused development inputs. The [evaluation companion](evaluation/README.md) has every cutoff through 20, all-query results, model identities and metric code.
78
+
79
+ On public data, both models rerank the same 50 candidates that upstream Gemma retrieved for each query, scored with official relevance judgments:
80
+
81
+ | Dataset | Queries | Upstream Ettin | Daecore fine-tune |
82
+ |---|---:|---:|---:|
83
+ | FiQA | 648 | 0.4860 | 0.4527 |
84
+ | SciFact | 300 | 0.7487 | 0.7544 |
85
+
86
+ These are nDCG@10 scores for reranking fixed candidate pools, not dense Gemma scores. Fine-tuning lowered FiQA and slightly raised SciFact. Later continuation recipes did not recover FiQA while keeping Daecore ranking quality, so the original learned weights remain.
87
+
88
+ ## Training data and objective
89
+
90
+ Training passages come from generated organizational documents and one person's project documentation, much of it AI-written. Queries include full questions, keywords, identifiers and pasted fragments. Model judges assigned grades from 0 to 3, and several passages can be useful for one query. ListNet learns ordering within each candidate list, and a pairwise term emphasizes distinctions between neighboring grades. Lists rotate through the full labeled pools, so useful, partly useful and unhelpful passages stay in the comparison.
91
+
92
+ <details>
93
+ <summary>Training recipe and exposure</summary>
94
+
95
+ | Item | Value |
96
+ |---|---|
97
+ | Starting model | `cross-encoder/ettin-reranker-150m-v1` |
98
+ | Training queries | 2,882 with ranking supervision; 30 no-positive queries excluded from the ranking loss |
99
+ | Unique query–passage pairs | 144,100 |
100
+ | Unique passage texts | 35,006 |
101
+ | Supervision | Four relevance grades, 0–3; grades 2–3 count as useful |
102
+ | Objective | ListNet with an adjacent-grade pairwise term weighted 0.2 |
103
+ | Adapter | LoRA+, rank 16, alpha 32; merged for serving |
104
+ | Learning rates | 5e-5 for LoRA A and the scoring head; 8e-4 for LoRA B |
105
+ | Training steps | 2,882 scheduled; 2,881 applied updates, accumulating four lists of 16 per step |
106
+ | Exposure | 11,528 lists; 184,448 candidate presentations |
107
+
108
+ </details>
109
+
110
+ ## Runtime and limits
111
+
112
+ `model.onnx` is the CUDA FP16 graph and `model.vulkan.onnx` the Vulkan FP32 graph; both read `model.onnx.data`. The Vulkan derivation rewrites bounded mask operations and keeps the learned weights. The FP32 graph's results were also checked for correctness on the CPU execution provider; no CPU latency guidance is given.
113
+
114
+ Vulkan requires `onnxruntime==1.24.4` and the native WebGPU plugin [`onnxruntime-ep-webgpu==0.4.0`](https://pypi.org/project/onnxruntime-ep-webgpu/), registered explicitly. Add the plugin device to the session options with `dawnBackendType=Vulkan`, and create the session without a providers list; the plugin name alone does not select Vulkan. `serving.vulkan.json` records the tested provider options.
115
+
116
+ On 970 Daecore candidate pools, CUDA and Vulkan produced the same Hit@5, Hit@10 and Hit@20; Hit@3 differed by one query and nDCG@10 by 0.0004. The paths are not bit-exact, so close scores can swap order. Checks used Windows x64 with an NVIDIA RTX 3060 Ti (8 GB); other GPUs and Linux are untested, even where Vulkan is available. Inside Daecore, Ettin reorders up to 50 deduplicated BM25 and Gemma candidates; the [retrieval pipeline overview](https://huggingface.co/Daecore) describes that path and result selection.
117
+
118
+ The task-specific panel is mostly generated and model-labeled, with no human reference, and was reused during development; training-seed variation is unmeasured. A reranker cannot recover a passage missing from its candidates, and several relevant passages may repeat one fact. These results do not establish unique-evidence coverage or downstream agent success.
119
+
120
+ ## License
121
+
122
+ Apache-2.0. The package includes upstream attribution and a modification notice.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
evaluation/README.md ADDED
@@ -0,0 +1,83 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Model evaluation records
2
+
3
+ These files let readers recompute the comparisons in the Daecore model cards without model weights. Task-specific records contain anonymized row and group IDs, labels, and scores or ranked grades; none contains query or document text, private source identifiers or workspace paths. Model identities are recorded inside each file.
4
+
5
+ ## Recompute
6
+
7
+ Python 3.11 or later is enough; no packages or network access are needed. From the directory that holds the records:
8
+
9
+ ```sh
10
+ python metrics.py --directory .
11
+ ```
12
+
13
+ The command summarizes every record present and stops with an error if there is none. Each entry in its output equals the matching `*-summary.json`. The classifier and Ettin packages also include `figures.py`; `python figures.py --output <directory>` redraws their card chart. Gemma uses a table. Each model package contains only its own records and summaries; the source repository holds all four record types. Fractions are stored at full precision; cards round percentages to two decimals and ranking metrics to four.
14
+
15
+ The records support metric recomputation. Re-running inference would need the private text and corpus, which are not included. They come from task-specific development evaluations, not a new untouched test.
16
+
17
+ ## Records
18
+
19
+ | Record | Inputs and denominator | Comparison |
20
+ |---|---|---|
21
+ | `classifier.json` | 1,300 generated passages from three held-out source families; resolved labels only, per facet | Upstream GLiClass, a trained word-feature baseline and the fine-tune |
22
+ | `reranker.json` | 970 fixed pools of 50 passages, all 48,500 pairs judged; 873 pools have useful evidence | Upstream Ettin and the fine-tune, plus the expected result of random ordering |
23
+ | `promotion.json` | 970 hybrid-search queries over 82,719 passage texts, with reviewed labels and declared exclusions; five public dense-retrieval panels | Previous and updated Gemma with unchanged Ettin; the public panels also include upstream Gemma |
24
+ | `serving.json` | The same 970 queries with updated-Gemma candidate pools; reference extended by 32 grades | Ettin through CUDA FP16 and Vulkan FP32, with identical weights and score mapping |
25
+
26
+ `serving-qualification.json` summarizes provider and recovery checks by receipt hash and keeps the aggregate FiQA and SciFact results for upstream and fine-tuned Ettin.
27
+
28
+ “Upstream” means no Daecore fine-tuning, not an untrained network; upstream Ettin is already a trained reranker. Interim training checkpoints are not included.
29
+
30
+ ## How each comparison was run
31
+
32
+ **Classifier.** Evaluation families were excluded from training. Upstream uses the same five label definitions and 768-token input limit as the fine-tune. Its raw logits and the fine-tune's calibrated probabilities are used only to rank within each facet; their scales are not compared. All 1,300 saved predictions matched fresh CPU inference of the released graph within 5.62e-6. The baseline uses word unigram and bigram TF-IDF with one balanced logistic-regression model per facet, fitted on the same 59,886 training passages; it sees full passage text, while the neural models apply their 768-token limit.
33
+
34
+ **Ettin, fixed pools.** Candidates and grades are held fixed. Both models run in PyTorch FP16 with the same 1,153-token pair construction, Transformers 5.2.0 and Sentence Transformers 5.5.1; upstream's metadata names a newer library version, but both use the same reference runtime here. The fine-tune's saved scores were checked against fresh inference on three complete pools. This evaluates the trained models, not ONNX serving or latency. The public FiQA and SciFact results rerank fixed 50-candidate pools from upstream Gemma.
35
+
36
+ **Gemma update.** Only Gemma changes; BM25, fusion and Ettin's weights stay fixed. All 970 queries are retained, and seven reviewed grade corrections apply to both models. Within each original top-k or selected prefix, 65 unresolved query–passage abstentions are excluded without backfilling: `ranked_grades` keeps their positions as `null`, with the required `excluded` flag. Precision is pooled over retained positions; nDCG reindexes them and takes its ideal ordering at the retained depth from `reference_grade_counts`. Selected depths refer to the original score-selected prefixes.
37
+
38
+ | Metric, 970 queries | Previous Gemma | Updated Gemma |
39
+ |---|---:|---:|
40
+ | Hit@3 | 83.92% | 85.36% |
41
+ | Hit@5 | 85.88% | 88.45% |
42
+ | Hit@10 | 87.84% | 89.90% |
43
+ | Hit@20 | 90.10% | 91.03% |
44
+ | nDCG@10 | 0.6764 | 0.6611 |
45
+ | Selected-prefix precision | 80.41% | 80.34% |
46
+
47
+ The same record holds per-query dense nDCG@10 for the five public panels. The previous fine-tune was measured only on FiQA (0.4009) and SciFact (0.7679); missing baselines stay absent. These sets supplied no training examples but informed development, and recomputing their means is different from rerunning retrieval on the public corpora.
48
+
49
+ **Ettin providers.** Updated-Gemma candidate pools for the same 970 queries were scored through CUDA FP16 and Vulkan FP32 with unchanged weights and score mapping. The two paths' top-20 results included 32 query–passage pairs without a grade. Of these, 21 reuse grades from the hard-contrast Ettin evaluation. The other 11 received two independent GPT-6 Sol judgments plus a resolution step and were then reviewed against the full passage text by GPT-6 Astra; no person reviewed them. No existing grade changed, and the 65 abstentions remain excluded. The added grades slightly change nDCG's ideal ordering, so the promotion record keeps its original reference.
50
+
51
+ | Metric, 970 queries | CUDA FP16 | Vulkan FP32 |
52
+ |---|---:|---:|
53
+ | Hit@3 | 85.36% | 85.26% |
54
+ | Hit@5 | 88.45% | 88.45% |
55
+ | Hit@10 | 89.90% | 89.90% |
56
+ | Hit@20 | 91.03% | 91.03% |
57
+ | nDCG@10 | 0.6611 | 0.6615 |
58
+ | Selected-prefix precision | 80.34% | 80.28% |
59
+
60
+ Small numerical differences between the paths can reorder close scores. On 20 matched pools replayed twice, second-pass reranking took 1.56 s median and 2.11 s at the 95th percentile with Vulkan, against 1.92 s and 2.70 s with the previous package's DirectML graph. All provider measurements come from one Windows x64 machine with an NVIDIA RTX 3060 Ti (8 GB) and do not transfer to other GPUs or platforms.
61
+
62
+ ## Metric definitions
63
+
64
+ Relevance grades are 0–3, and grades **2 and 3** count as useful. An “answerable” query has at least one useful labeled passage in the specified pool; the absence of a useful judgment does not prove that no answer exists.
65
+
66
+ - **Hit@k:** fraction of queries with at least one useful passage in the first k positions.
67
+ - **Precision@k:** useful passages among the first k. Fixed-pool tables average it over queries; the promotion and serving records pool it over retained positions. Selected-prefix precision applies the same pooling to all returned passages.
68
+ - **Recall@k:** useful passages retrieved divided by the query's known useful passages, averaged over queries. It counts labeled passages, not every fact an answer needs.
69
+ - **nDCG@k:** gain `2**grade - 1`, discount `1/log2(rank + 1)`, divided by the ideal ordering at the same cutoff. Grade 1 adds a small gain although it is not useful for Hit, precision or recall. Each cutoff has its own ideal, so values need not change monotonically with k.
70
+ - **Average precision (AP):** area under the stepwise precision–recall curve, with tied scores grouped at one threshold; macro AP weights the five facets equally.
71
+ - **Precision at 90% or 95% recall:** the best measured precision at any threshold reaching that recall, taken from the evaluation curve rather than a threshold chosen in advance.
72
+
73
+ Ties keep the original candidate order. Classifier fields marked `null` are excluded identically for every model. Conditional reranker nDCG and recall use the 873 answerable pools; all-query Hit and precision use all 970. The random-order reference is exact within each pool: with N candidates and R useful passages, expected precision is R/N and Hit@k is `1 - C(N-R, k)/C(N, k)`. Random ordering already reaches 97.30% Hit@20 on the answerable reranker pools.
74
+
75
+ ## Limits
76
+
77
+ Task-specific labels are language-model judgments without a human-adjudicated reference. Much of the source material is generated. Project-document sources are narrow and often AI-written; they do not establish generalization across users. The panels were reused during development, and each released model was trained once. Several relevant passages can repeat one fact, so passage-level scores do not measure unique-fact coverage, answer completeness or downstream agent success.
78
+
79
+ ## Cards
80
+
81
+ - [Five-facet classifier](https://huggingface.co/Daecore/gliclass-std-base-v3-5facet-qint8-v2)
82
+ - [EmbeddingGemma retriever](https://huggingface.co/Daecore/embeddinggemma-300m-memory-ft-v1)
83
+ - [Ettin reranker](https://huggingface.co/Daecore/ettin-150m-memory-reranker-ft-v1)
evaluation/figures.py ADDED
@@ -0,0 +1,89 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Regenerate the classifier and reranker SVGs from the public labels and scores.
2
+
3
+ Run with Python 3.11+: python figures.py --output ../figures
4
+ The repository shares its renderer from eval/lib; HF packages include that
5
+ same renderer next to this script.
6
+ """
7
+
8
+ from __future__ import annotations
9
+
10
+ import argparse
11
+ import importlib.util
12
+ import json
13
+ import sys
14
+ from pathlib import Path
15
+
16
+ import metrics
17
+
18
+ HERE = Path(__file__).resolve().parent
19
+ renderer = HERE / 'svg_figures.py'
20
+ if not renderer.is_file():
21
+ renderer = HERE.parents[1] / 'lib/svg_figures.py'
22
+ spec = importlib.util.spec_from_file_location('model_card_svg', renderer)
23
+ svg = importlib.util.module_from_spec(spec)
24
+ sys.modules[spec.name] = svg
25
+ spec.loader.exec_module(svg)
26
+
27
+ UPSTREAM, FIT, REFERENCE = '#0072B2', '#D55E00', '#777777'
28
+
29
+
30
+ def classifier(data: dict) -> str:
31
+ summary = metrics.summarize_classifier(data)
32
+ facets = [*metrics.FACETS, 'macro']
33
+ def values(model):
34
+ result = summary['models'][model]
35
+ return [result['facets'][f]['average_precision'] for f in metrics.FACETS] + [result['macro']['average_precision']]
36
+ return svg.render_bars(
37
+ facets, [('Upstream', values('upstream'), UPSTREAM),
38
+ ('TF-IDF + logistic regression', values('linear'), REFERENCE),
39
+ ('Daecore fine-tune', values('finetuned'), FIT)],
40
+ title='Classifier: matched before and after fine-tuning',
41
+ subtitle='1,300 generated passages · 3 held-out families · resolved labels only',
42
+ ylabel='average precision', ylim=(0, 1.05), separator_before=5,
43
+ width=840, height=360,
44
+ notes=['Model-generated labels; these results do not establish accuracy on unrelated real documents.'],
45
+ )
46
+
47
+
48
+ def reranker(data: dict) -> str:
49
+ summary = metrics.summarize_reranker(data)
50
+ definitions = [('hit', 'At least one useful passage'), ('precision', 'Useful passages / returned passages'), ('ndcg', 'Graded ranking quality')]
51
+ panels = []
52
+ legend = [('Upstream', UPSTREAM, None), ('Daecore fine-tune', FIT, None), ('Random pool order', REFERENCE, '5 3')]
53
+ for field, title in definitions:
54
+ series = []
55
+ for model, (label, color, dash) in zip(('upstream', 'finetuned', 'random'), legend, strict=True):
56
+ values = summary['models'][model]['answerable']
57
+ series.append(svg.Series(label, list(range(5)), [values[str(k)][field] for k in metrics.CUTOFFS], color, dash=dash))
58
+ panels.append(svg.Panel(title, series, xlabel='rank cutoff',
59
+ ylabel={'hit': 'Hit@k', 'precision': 'Precision@k', 'ndcg': 'nDCG@k'}[field],
60
+ xlim=(0, 4), ylim=(0, 1.02), xticks=list(enumerate(map(str, metrics.CUTOFFS)))))
61
+ return svg.render_grid(panels, columns=3, title='Ettin: identical candidates, different ordering',
62
+ subtitle='873 answerable pools of 50 · 970 queries in total · useful = grade 2 or 3',
63
+ legend=legend, panel_width=300, panel_height=270)
64
+
65
+
66
+ def main() -> None:
67
+ parser = argparse.ArgumentParser(description=__doc__)
68
+ parser.add_argument('--directory', type=Path, default=HERE)
69
+ parser.add_argument('--output', type=Path, required=True)
70
+ parser.add_argument('--model', choices=('classifier', 'reranker'))
71
+ args = parser.parse_args()
72
+ args.output.mkdir(parents=True, exist_ok=True)
73
+ definitions = {'classifier': ('classifier', 'classifier-comparison.svg'),
74
+ 'reranker': ('reranker', 'ettin-comparison.svg')}
75
+ names = [args.model] if args.model else [
76
+ name for name, (record, _) in definitions.items()
77
+ if (args.directory / f'{record}.json').is_file()
78
+ ]
79
+ if not names:
80
+ parser.error('No model comparison records found in the selected directory')
81
+ for name in names:
82
+ record, filename = definitions[name]
83
+ data = json.loads((args.directory / f'{record}.json').read_text(encoding='utf-8'))
84
+ (args.output / filename).write_text(globals()[name](data), encoding='utf-8', newline='\n')
85
+
86
+
87
+
88
+ if __name__ == '__main__':
89
+ main()
evaluation/metrics.py ADDED
@@ -0,0 +1,219 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Recompute the published model comparisons from text-free evaluation records.
2
+
3
+ Python 3.11+, standard library only. Run: python metrics.py --directory .
4
+ """
5
+
6
+ from __future__ import annotations
7
+
8
+ import argparse
9
+ import json
10
+ import math
11
+ from pathlib import Path
12
+ from statistics import mean
13
+
14
+ CUTOFFS = (1, 3, 5, 10, 20)
15
+ FACETS = ("trap", "decision", "constraint", "mechanism", "procedure")
16
+
17
+
18
+ def classification_metrics(labels: list[int], scores: list[float]) -> dict[str, float]:
19
+ """Threshold-grouped AP and best precision at or above each recall target."""
20
+ if len(labels) != len(scores) or not labels or set(labels) - {0, 1}:
21
+ raise ValueError("Classification labels and scores must be aligned binary rows")
22
+ if not all(math.isfinite(s) for s in scores) or sum(labels) == 0:
23
+ raise ValueError("Classification scores must be finite with positive support")
24
+ order = sorted(range(len(labels)), key=lambda i: -scores[i])
25
+ positives = sum(labels)
26
+ true_positives = 0
27
+ previous_recall = 0.0
28
+ ap = 0.0
29
+ points = []
30
+ for rank, index in enumerate(order, 1):
31
+ true_positives += labels[index]
32
+ if rank < len(order) and scores[order[rank]] == scores[index]:
33
+ continue
34
+ precision = true_positives / rank
35
+ recall = true_positives / positives
36
+ ap += (recall - previous_recall) * precision
37
+ previous_recall = recall
38
+ points.append((precision, recall))
39
+ return {
40
+ "average_precision": ap,
41
+ "precision_at_recall_90": max(p for p, r in points if r >= 0.9),
42
+ "precision_at_recall_95": max(p for p, r in points if r >= 0.95),
43
+ "prevalence": positives / len(labels),
44
+ }
45
+
46
+
47
+ def dcg(grades: list[int], k: int) -> float:
48
+ return sum((2**g - 1) / math.log2(i + 2) for i, g in enumerate(grades[:k]))
49
+
50
+
51
+ def ranking_metrics(grades: list[int], scores: list[float], k: int) -> dict[str, float]:
52
+ if len(grades) != len(scores) or len(grades) < k or k < 1:
53
+ raise ValueError("Ranking inputs must be aligned and cover the cutoff")
54
+ if any(type(g) is not int or g not in range(4) for g in grades):
55
+ raise ValueError("Ranking requires fully judged integer grades 0 through 3")
56
+ if not all(math.isfinite(s) for s in scores):
57
+ raise ValueError("Ranking scores must be finite")
58
+ order = sorted(range(len(scores)), key=lambda i: (-scores[i], i))
59
+ ranked = [grades[i] for i in order]
60
+ useful = sum(g >= 2 for g in ranked[:k])
61
+ relevant = sum(g >= 2 for g in grades)
62
+ ideal = dcg(sorted(grades, reverse=True), k)
63
+ result = {"hit": float(useful > 0), "precision": useful / k, "useful": float(useful)}
64
+ if relevant:
65
+ result["recall"] = useful / relevant
66
+ result["ndcg"] = dcg(ranked, k) / ideal
67
+ return result
68
+
69
+
70
+ def random_ranking_metrics(grades: list[int], k: int) -> dict[str, float]:
71
+ """Exact expectation under a uniform permutation of one fixed judged pool."""
72
+ n = len(grades)
73
+ if k < 1 or k > n or any(type(g) is not int or g not in range(4) for g in grades):
74
+ raise ValueError("Random reference requires judged grades and a valid cutoff")
75
+ relevant = sum(g >= 2 for g in grades)
76
+ misses = math.comb(n - relevant, k) if n - relevant >= k else 0
77
+ result = {"hit": 1 - misses / math.comb(n, k), "precision": relevant / n, "useful": k * relevant / n}
78
+ if relevant:
79
+ expected_dcg = mean(2**g - 1 for g in grades) * sum(1 / math.log2(i + 2) for i in range(k))
80
+ result["recall"] = k / n
81
+ result["ndcg"] = expected_dcg / dcg(sorted(grades, reverse=True), k)
82
+ return result
83
+
84
+
85
+ def summarize_classifier(data: dict) -> dict:
86
+ rows = data["rows"]
87
+ result = {"rows": len(rows), "families": len({r["group"] for r in rows}), "models": {}}
88
+ for model in data["model_order"]:
89
+ by_facet = {}
90
+ for facet in FACETS:
91
+ resolved = [r for r in rows if r["labels"][facet] is not None]
92
+ if any(r["labels"][facet] not in (0, 1) for r in resolved):
93
+ raise ValueError("Invalid resolved classifier label")
94
+ by_facet[facet] = {"rows": len(resolved), **classification_metrics(
95
+ [r["labels"][facet] for r in resolved], [r["scores"][model][facet] for r in resolved]
96
+ )}
97
+ macro = {k: mean(v[k] for v in by_facet.values()) for k in ("average_precision", "precision_at_recall_90", "precision_at_recall_95")}
98
+ result["models"][model] = {"facets": by_facet, "macro": macro}
99
+ return result
100
+
101
+
102
+ def summarize_reranker(data: dict) -> dict:
103
+ rows = data["rows"]
104
+ answerable = [r for r in rows if any(g >= 2 for g in r["grades"])]
105
+ output = {"queries": len(rows), "answerable": len(answerable), "models": {}}
106
+ for model in ["random", *data["model_order"]]:
107
+ scopes = {}
108
+ for scope, subset in [("answerable", answerable), ("all", rows)]:
109
+ cuts = {}
110
+ for k in CUTOFFS:
111
+ values = [random_ranking_metrics(r["grades"], k) if model == "random" else ranking_metrics(r["grades"], r["scores"][model], k) for r in subset]
112
+ fields = ("hit", "precision", "useful", "recall", "ndcg") if scope == "answerable" else ("hit", "precision", "useful")
113
+ cuts[str(k)] = {field: mean(v[field] for v in values) for field in fields}
114
+ scopes[scope] = cuts
115
+ output["models"][model] = scopes
116
+ return output
117
+
118
+
119
+ def summarize_promotion(data: dict) -> dict:
120
+ """Hybrid search with reviewed exclusions inside each original prefix.
121
+
122
+ A null is permitted only for an explicitly excluded judging abstention.
123
+ Cut first, remove exclusions second, and never backfill from a deeper rank.
124
+ """
125
+ rows = data['rows']
126
+ if not rows or len({row['id'] for row in rows}) != len(rows):
127
+ raise ValueError('Promotion rows require unique nonempty query identities')
128
+ output = {'queries': len(rows), 'models': {}, 'public': {}}
129
+ for model in data['model_order']:
130
+ cutoffs = {}
131
+ for cutoff in (3, 5, 10, 20, 'selected'):
132
+ per_query = []
133
+ for row in rows:
134
+ counts = row['reference_grade_counts']
135
+ if set(counts) != {'0', '1', '2', '3'} or any(type(n) is not int or n < 0 for n in counts.values()):
136
+ raise ValueError('Reference grade counts must cover grades zero through three')
137
+ ranked = row['ranked_grades'][model]
138
+ excluded = row['excluded'][model]
139
+ depth = row['selected_depth'][model] if cutoff == 'selected' else cutoff
140
+ if (len(ranked) != 20 or len(excluded) != 20 or type(depth) is not int
141
+ or not 3 <= depth <= 20 or any(type(x) is not bool for x in excluded)):
142
+ raise ValueError('Promotion rows require a bounded original top twenty')
143
+ if any((grade is not None if drop else type(grade) is not int or grade not in range(4))
144
+ for grade, drop in zip(ranked, excluded, strict=True)):
145
+ raise ValueError('Only declared abstentions may lack grades')
146
+ kept = [grade for grade, drop in zip(ranked[:depth], excluded[:depth], strict=True) if not drop]
147
+ if not kept:
148
+ raise ValueError('Every scored prefix must retain judged passages')
149
+ ideal_grades = [g for g in (3, 2, 1, 0) for _ in range(min(counts[str(g)], len(kept)))][:len(kept)]
150
+ useful = sum(g >= 2 for g in kept)
151
+ positives = counts['2'] + counts['3']
152
+ ideal = dcg(ideal_grades, len(kept))
153
+ per_query.append({
154
+ 'hit': float(useful > 0), 'precision': useful / len(kept),
155
+ 'useful': useful, 'retained': len(kept), 'excluded': depth - len(kept),
156
+ 'ndcg': dcg(kept, len(kept)) / ideal if ideal else 0.0,
157
+ 'known_recall': useful / positives if positives else None,
158
+ })
159
+ useful = sum(row['useful'] for row in per_query)
160
+ retained = sum(row['retained'] for row in per_query)
161
+ recalls = [row['known_recall'] for row in per_query if row['known_recall'] is not None]
162
+ cutoffs[str(cutoff)] = {
163
+ 'hit': mean(row['hit'] for row in per_query), 'precision': useful / retained,
164
+ 'macro_precision': mean(row['precision'] for row in per_query),
165
+ 'ndcg': mean(row['ndcg'] for row in per_query),
166
+ 'known_positive_recall': mean(recalls) if recalls else None,
167
+ 'recall_queries': len(recalls), 'useful': useful, 'retained': retained,
168
+ 'excluded_positions': sum(row['excluded'] for row in per_query),
169
+ 'mean_useful': useful / len(rows), 'mean_retained': retained / len(rows),
170
+ }
171
+ output['models'][model] = cutoffs
172
+ for dataset, panel in data['public'].items():
173
+ if not panel['rows'] or len({row['id'] for row in panel['rows']}) != len(panel['rows']):
174
+ raise ValueError('Public panel requires distinct query identities')
175
+ measured = panel['model_order']
176
+ for row in panel['rows']:
177
+ if set(row['ndcg@10']) != set(measured) or any(
178
+ isinstance(value, bool) or not isinstance(value, int | float)
179
+ or not math.isfinite(value) or not 0 <= value <= 1
180
+ for value in row['ndcg@10'].values()
181
+ ):
182
+ raise ValueError('Public nDCG values must be finite measured scores')
183
+ output['public'][dataset] = {
184
+ 'queries': len(panel['rows']),
185
+ 'ndcg@10': {model: mean(row['ndcg@10'][model] for row in panel['rows']) for model in measured},
186
+ }
187
+ return output
188
+
189
+
190
+ def main() -> None:
191
+ parser = argparse.ArgumentParser(description=__doc__)
192
+ parser.add_argument("--directory", type=Path, default=Path(__file__).resolve().parent)
193
+ parser.add_argument("--output", type=Path)
194
+ parser.add_argument("--model", choices=("classifier", "reranker", "promotion", "serving"), help="Recompute one comparison; default: every record in this package")
195
+ args = parser.parse_args()
196
+ summarizers = {"classifier": summarize_classifier, "reranker": summarize_reranker,
197
+ "promotion": summarize_promotion, "serving": summarize_promotion}
198
+ names = [args.model] if args.model else [
199
+ name for name in summarizers if (args.directory / f"{name}.json").is_file()
200
+ ]
201
+ if not names:
202
+ parser.error("No evaluation records found in the selected directory")
203
+ result = {}
204
+ for name in names:
205
+ summarize = summarizers[name]
206
+ data = json.loads((args.directory / f"{name}.json").read_text(encoding="utf-8"))
207
+ ids = [r["id"] for r in data["rows"]]
208
+ if len(set(ids)) != len(ids):
209
+ raise ValueError(f"Repeated query/passage identity in {name}")
210
+ result[name] = summarize(data)
211
+ text = json.dumps(result, indent=2, allow_nan=False) + "\n"
212
+ if args.output:
213
+ args.output.write_text(text, encoding="utf-8")
214
+ else:
215
+ print(text, end="")
216
+
217
+
218
+ if __name__ == "__main__":
219
+ main()
evaluation/reranker-summary.json ADDED
@@ -0,0 +1,204 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "queries": 970,
3
+ "answerable": 873,
4
+ "models": {
5
+ "random": {
6
+ "answerable": {
7
+ "1": {
8
+ "hit": 0.4268041237113402,
9
+ "precision": 0.4268041237113402,
10
+ "useful": 0.4268041237113402,
11
+ "recall": 0.02,
12
+ "ndcg": 0.34859651993672613
13
+ },
14
+ "3": {
15
+ "hit": 0.726812305678285,
16
+ "precision": 0.4268041237113402,
17
+ "useful": 1.2804123711340205,
18
+ "recall": 0.06,
19
+ "ndcg": 0.3568668683201474
20
+ },
21
+ "5": {
22
+ "hit": 0.83263944209344,
23
+ "precision": 0.4268041237113402,
24
+ "useful": 2.134020618556701,
25
+ "recall": 0.1,
26
+ "ndcg": 0.3673965034684685
27
+ },
28
+ "10": {
29
+ "hit": 0.9245071219176823,
30
+ "precision": 0.4268041237113402,
31
+ "useful": 4.268041237113402,
32
+ "recall": 0.2,
33
+ "ndcg": 0.3973178040013556
34
+ },
35
+ "20": {
36
+ "hit": 0.9730113306184918,
37
+ "precision": 0.4268041237113402,
38
+ "useful": 8.536082474226804,
39
+ "recall": 0.4,
40
+ "ndcg": 0.46181890469238246
41
+ }
42
+ },
43
+ "all": {
44
+ "1": {
45
+ "hit": 0.3841237113402062,
46
+ "precision": 0.3841237113402062,
47
+ "useful": 0.3841237113402062
48
+ },
49
+ "3": {
50
+ "hit": 0.6541310751104565,
51
+ "precision": 0.3841237113402062,
52
+ "useful": 1.1523711340206186
53
+ },
54
+ "5": {
55
+ "hit": 0.749375497884096,
56
+ "precision": 0.3841237113402062,
57
+ "useful": 1.920618556701031
58
+ },
59
+ "10": {
60
+ "hit": 0.832056409725914,
61
+ "precision": 0.3841237113402062,
62
+ "useful": 3.841237113402062
63
+ },
64
+ "20": {
65
+ "hit": 0.8757101975566426,
66
+ "precision": 0.3841237113402062,
67
+ "useful": 7.682474226804124
68
+ }
69
+ }
70
+ },
71
+ "upstream": {
72
+ "answerable": {
73
+ "1": {
74
+ "hit": 0.6174112256586484,
75
+ "precision": 0.6174112256586484,
76
+ "useful": 0.6174112256586484,
77
+ "recall": 0.03681805740540042,
78
+ "ndcg": 0.5277368679430535
79
+ },
80
+ "3": {
81
+ "hit": 0.8316151202749141,
82
+ "precision": 0.5975563192058038,
83
+ "useful": 1.7926689576174113,
84
+ "recall": 0.10028426529744264,
85
+ "ndcg": 0.5180026543309457
86
+ },
87
+ "5": {
88
+ "hit": 0.8946162657502864,
89
+ "precision": 0.5857961053837343,
90
+ "useful": 2.928980526918671,
91
+ "recall": 0.15797988036502988,
92
+ "ndcg": 0.5264221881850474
93
+ },
94
+ "10": {
95
+ "hit": 0.9381443298969072,
96
+ "precision": 0.5673539518900343,
97
+ "useful": 5.673539518900344,
98
+ "recall": 0.2906998695460454,
99
+ "ndcg": 0.5518800541282721
100
+ },
101
+ "20": {
102
+ "hit": 0.981672394043528,
103
+ "precision": 0.5446162657502863,
104
+ "useful": 10.892325315005728,
105
+ "recall": 0.5378196846495221,
106
+ "ndcg": 0.6153644012290921
107
+ }
108
+ },
109
+ "all": {
110
+ "1": {
111
+ "hit": 0.5556701030927835,
112
+ "precision": 0.5556701030927835,
113
+ "useful": 0.5556701030927835
114
+ },
115
+ "3": {
116
+ "hit": 0.7484536082474227,
117
+ "precision": 0.5378006872852233,
118
+ "useful": 1.6134020618556701
119
+ },
120
+ "5": {
121
+ "hit": 0.8051546391752578,
122
+ "precision": 0.5272164948453608,
123
+ "useful": 2.636082474226804
124
+ },
125
+ "10": {
126
+ "hit": 0.8443298969072165,
127
+ "precision": 0.510618556701031,
128
+ "useful": 5.106185567010309
129
+ },
130
+ "20": {
131
+ "hit": 0.8835051546391752,
132
+ "precision": 0.49015463917525776,
133
+ "useful": 9.803092783505155
134
+ }
135
+ }
136
+ },
137
+ "finetuned": {
138
+ "answerable": {
139
+ "1": {
140
+ "hit": 0.8144329896907216,
141
+ "precision": 0.8144329896907216,
142
+ "useful": 0.8144329896907216,
143
+ "recall": 0.0523348139467294,
144
+ "ndcg": 0.718158511972945
145
+ },
146
+ "3": {
147
+ "hit": 0.9255441008018328,
148
+ "precision": 0.8003054600992745,
149
+ "useful": 2.4009163802978235,
150
+ "recall": 0.14284244957257694,
151
+ "ndcg": 0.7197739775683897
152
+ },
153
+ "5": {
154
+ "hit": 0.9404352806414662,
155
+ "precision": 0.7793814432989691,
156
+ "useful": 3.8969072164948453,
157
+ "recall": 0.22343247007861236,
158
+ "ndcg": 0.7212210726901884
159
+ },
160
+ "10": {
161
+ "hit": 0.9690721649484536,
162
+ "precision": 0.7418098510882016,
163
+ "useful": 7.418098510882016,
164
+ "recall": 0.39954728808322054,
165
+ "ndcg": 0.7406757160512017
166
+ },
167
+ "20": {
168
+ "hit": 0.9896907216494846,
169
+ "precision": 0.6734249713631157,
170
+ "useful": 13.468499427262314,
171
+ "recall": 0.6787919700186864,
172
+ "ndcg": 0.7868944732410001
173
+ }
174
+ },
175
+ "all": {
176
+ "1": {
177
+ "hit": 0.7329896907216494,
178
+ "precision": 0.7329896907216494,
179
+ "useful": 0.7329896907216494
180
+ },
181
+ "3": {
182
+ "hit": 0.8329896907216495,
183
+ "precision": 0.7202749140893471,
184
+ "useful": 2.160824742268041
185
+ },
186
+ "5": {
187
+ "hit": 0.8463917525773196,
188
+ "precision": 0.7014432989690722,
189
+ "useful": 3.507216494845361
190
+ },
191
+ "10": {
192
+ "hit": 0.8721649484536083,
193
+ "precision": 0.6676288659793814,
194
+ "useful": 6.676288659793815
195
+ },
196
+ "20": {
197
+ "hit": 0.8907216494845361,
198
+ "precision": 0.6060824742268042,
199
+ "useful": 12.121649484536082
200
+ }
201
+ }
202
+ }
203
+ }
204
+ }
evaluation/reranker.json ADDED
The diff for this file is too large to render. See raw diff
 
evaluation/serving-qualification.json ADDED
@@ -0,0 +1,134 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "schema": "daecore.serving-preparation-summary.v1",
3
+ "date": "2026-09-28",
4
+ "status": "local-model-serving-checks-complete-release-transition-pending",
5
+ "hardware": "Windows x64, NVIDIA RTX 3060 Ti 8 GiB",
6
+ "unmeasured_targets": [
7
+ "AMD",
8
+ "Intel",
9
+ "Linux x64"
10
+ ],
11
+ "runtimes": {
12
+ "onnxruntime": "1.24.4",
13
+ "vulkan_plugin": "0.4.0",
14
+ "backend": "Vulkan",
15
+ "storage_buffer_cache": "lazyRelease"
16
+ },
17
+ "models": {
18
+ "gemma": {
19
+ "vectors": 74,
20
+ "providers": {
21
+ "cpu": {
22
+ "receipt_sha256": "31851791104fa95d6b7b25d35f950b25e239a3d6f8fa458addb2a44917c25cb4",
23
+ "maximum_vector_delta": 3.0174851417541504e-07,
24
+ "concurrent_delta": 0.0
25
+ },
26
+ "cuda": {
27
+ "receipt_sha256": "6977d0aa0ebf9c39ea0fd24078c47367d1d50e81184653225fcd10758961903e",
28
+ "maximum_vector_delta": 2.644956111907959e-07,
29
+ "concurrent_delta": 2.2351741790771484e-08
30
+ },
31
+ "vulkan": {
32
+ "receipt_sha256": "567fb89ea195e8b188203d2334f3eda82994ae8c7fc97f2126fac3e6f4766079",
33
+ "maximum_vector_delta": 5.438923835754395e-07,
34
+ "concurrent_delta": 0.0
35
+ }
36
+ },
37
+ "admission": {
38
+ "cases": 3841,
39
+ "semantic_fallbacks": 679,
40
+ "new_useful_exclusions": 0,
41
+ "additional_noise_retained": 2,
42
+ "receipt_sha256": "71550f606c95c1e1192dd5939488b280d5f2f2f091f99413f6b63f4408a5bc2c"
43
+ }
44
+ },
45
+ "classifier": {
46
+ "cpu": {
47
+ "rows": 1300,
48
+ "maximum_posterior_delta": 5.612167303103988e-06,
49
+ "changed_threshold_labels": {
50
+ "contract": 0,
51
+ "recall_leaning": 0
52
+ },
53
+ "receipt_sha256": "ea545fb0a30ff25bf3da178b60593e934e6025e789c2db67290ff5360e8669a4"
54
+ },
55
+ "cuda": {
56
+ "rows": 1300,
57
+ "maximum_posterior_delta": 5.612167303103988e-06,
58
+ "changed_threshold_labels": {
59
+ "contract": 0,
60
+ "recall_leaning": 0
61
+ },
62
+ "receipt_sha256": "5616891e9dfdc3b6615951901cb24aa31189d6a81ed84b0080f461b85c045065"
63
+ },
64
+ "vulkan": {
65
+ "rows": 1300,
66
+ "maximum_posterior_delta": 5.612167303103988e-06,
67
+ "changed_threshold_labels": {
68
+ "contract": 0,
69
+ "recall_leaning": 0
70
+ },
71
+ "receipt_sha256": "c7ec971e020d9771adf0756316f468629cc05e300d7edf9abb52612901adf6b0"
72
+ }
73
+ },
74
+ "ettin": {
75
+ "queries": 970,
76
+ "unjudged_top20": 0,
77
+ "quality_evidence": "serving.json",
78
+ "quality_receipt_sha256": "805a1b0acd7baaf248e182908bc23cd52171e40aa8f1281799a20d15b58c6905",
79
+ "full_replay_seconds_p50_p95": [
80
+ 1.7189999999827705,
81
+ 2.610000000044238
82
+ ],
83
+ "matched_20_pool_seconds_p50_p95": {
84
+ "vulkan": [
85
+ 1.5565476999909151,
86
+ 2.1063475799834124
87
+ ],
88
+ "predecessor_directml": [
89
+ 1.9197023500164505,
90
+ 2.701977299965802
91
+ ]
92
+ },
93
+ "pressure": "One preflight refusal after 322 queries with lazy release; resumed all remaining rows with zero further retries. An earlier default-cache run stopped after 424. No native OOM observed.",
94
+ "precision": "FP32 Vulkan versus FP16 CUDA; rankings not bit-exact",
95
+ "selector": "Unchanged CUDA mapping transferred for measurement; new package/provider binding pending"
96
+ }
97
+ },
98
+ "operator_runtime_changed": false,
99
+ "published": false,
100
+ "remaining": [
101
+ "immutable-publication-revisions",
102
+ "profile-and-calibration-release-bindings",
103
+ "isolated-package-update-and-index-rebuild-recovery",
104
+ "DirectML-current-route-retirement"
105
+ ],
106
+ "worker_recovery": {
107
+ "receipt_sha256": "46205d39521bde40d1228c837e57f88e1a941d42f6846f4a244c48b513347cc6",
108
+ "all_three_consumer_outputs_identical_after_owned_worker_crash": true,
109
+ "ettin_maximum_envelope": {
110
+ "tokens": 1153,
111
+ "max_abs_logit_delta_vs_cpu": 5.91278076171875e-05
112
+ }
113
+ },
114
+ "retained_reranker_public": {
115
+ "scope": "Retained matched public pools retrieved by upstream Gemma; upstream and production Ettin PyTorch FP16; not a new Vulkan public replay",
116
+ "source_sha256": "0814db560081af8376e4b1f3e3d2bcd5c4d9b8b208f7fb8b7172f113e8123ef3",
117
+ "panels": {
118
+ "fiqa": {
119
+ "queries": 648,
120
+ "ndcg@10": {
121
+ "upstream": 0.48603492061219766,
122
+ "finetuned": 0.452680922415071
123
+ }
124
+ },
125
+ "scifact": {
126
+ "queries": 300,
127
+ "ndcg@10": {
128
+ "upstream": 0.7487436478294831,
129
+ "finetuned": 0.7543833016238701
130
+ }
131
+ }
132
+ }
133
+ }
134
+ }
evaluation/serving-summary.json ADDED
@@ -0,0 +1 @@
 
 
1
+ {"queries":970,"models":{"cuda":{"3":{"hit":0.8536082474226804,"precision":0.7157640565712314,"macro_precision":0.7166666666666667,"ndcg":0.6576832474683237,"known_positive_recall":0.09521315062968948,"recall_queries":895,"useful":2075,"retained":2899,"excluded_positions":11,"mean_useful":2.1391752577319587,"mean_retained":2.988659793814433},"5":{"hit":0.8845360824742268,"precision":0.7028311634635255,"macro_precision":0.7031786941580757,"ndcg":0.6554718901879515,"known_positive_recall":0.15383335671735354,"recall_queries":895,"useful":3401,"retained":4839,"excluded_positions":11,"mean_useful":3.5061855670103093,"mean_retained":4.988659793814433},"10":{"hit":0.8989690721649485,"precision":0.67180070291503,"macro_precision":0.6722631320569465,"ndcg":0.6610872123112499,"known_positive_recall":0.2802986435290336,"recall_queries":895,"useful":6499,"retained":9674,"excluded_positions":26,"mean_useful":6.7,"mean_retained":9.97319587628866},"20":{"hit":0.9103092783505154,"precision":0.61794500723589,"macro_precision":0.6184956360634603,"ndcg":0.6840908385828889,"known_positive_recall":0.4884616340614924,"recall_queries":895,"useful":11956,"retained":19348,"excluded_positions":52,"mean_useful":12.32577319587629,"mean_retained":19.94639175257732},"selected":{"hit":0.8618556701030928,"precision":0.8033770583310076,"macro_precision":0.7040325511660611,"ndcg":0.6598955189009045,"known_positive_recall":0.2079138769097556,"recall_queries":895,"useful":5757,"retained":7166,"excluded_positions":17,"mean_useful":5.935051546391753,"mean_retained":7.387628865979382}},"vulkan":{"3":{"hit":0.8525773195876288,"precision":0.7157640565712314,"macro_precision":0.7166666666666667,"ndcg":0.6577263968717306,"known_positive_recall":0.09519545071125639,"recall_queries":895,"useful":2075,"retained":2899,"excluded_positions":11,"mean_useful":2.1391752577319587,"mean_retained":2.988659793814433},"5":{"hit":0.8845360824742268,"precision":0.7023563455973543,"macro_precision":0.702766323024055,"ndcg":0.6550340124265787,"known_positive_recall":0.15374094654114448,"recall_queries":895,"useful":3398,"retained":4838,"excluded_positions":12,"mean_useful":3.5030927835051546,"mean_retained":4.9876288659793815},"10":{"hit":0.8989690721649485,"precision":0.6723514211886304,"macro_precision":0.6727589592538046,"ndcg":0.6615310405020679,"known_positive_recall":0.2805978824757964,"recall_queries":895,"useful":6505,"retained":9675,"excluded_positions":25,"mean_useful":6.706185567010309,"mean_retained":9.974226804123711},"20":{"hit":0.9103092783505154,"precision":0.6178416373785404,"macro_precision":0.6183925432799551,"ndcg":0.6840389893720067,"known_positive_recall":0.4883516349357555,"recall_queries":895,"useful":11954,"retained":19348,"excluded_positions":52,"mean_useful":12.323711340206186,"mean_retained":19.94639175257732},"selected":{"hit":0.8608247422680413,"precision":0.8027855153203343,"macro_precision":0.7038533836490732,"ndcg":0.6599601279630594,"known_positive_recall":0.2082001919384574,"recall_queries":895,"useful":5764,"retained":7180,"excluded_positions":17,"mean_useful":5.942268041237114,"mean_retained":7.402061855670103}}},"public":{}}
evaluation/serving.json ADDED
The diff for this file is too large to render. See raw diff
 
evaluation/svg_figures.py ADDED
@@ -0,0 +1,576 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Dependency-free SVG charts for tracked evaluation figures.
2
+
3
+ The evaluation figures are generated from private per-row evidence into tracked SVG files.
4
+ Keeping the renderer inside the repository, with no plotting dependency, means a figure can be
5
+ regenerated by any contributor with the private inputs and compared byte for byte.
6
+
7
+ Everything is emitted as presentation attributes rather than CSS. Markdown hosts sanitize
8
+ embedded stylesheets and scripts out of SVG, so a figure that carries its styling in attributes
9
+ renders the same in the repository, on a model-card host, and in a local viewer.
10
+
11
+ Layout is measured rather than assumed: legend entries, panel titles, and reference-line labels
12
+ are placed from an estimated text width, so a longer label reflows instead of overlapping its
13
+ neighbour. ``text_width`` approximates a sans-serif advance table, which is enough to keep
14
+ elements apart but is not a substitute for a real font metric.
15
+ """
16
+
17
+ from __future__ import annotations
18
+
19
+ import math
20
+ from dataclasses import dataclass, field
21
+ from xml.sax.saxutils import escape
22
+
23
+ # Okabe-Ito, chosen because it stays distinguishable under the common colour-vision
24
+ # deficiencies and prints legibly in greyscale.
25
+ PALETTE = ("#0072B2", "#D55E00", "#009E73", "#CC79A7", "#E69F00", "#56B4E9", "#000000")
26
+ FONT_STACK = "system-ui, Segoe UI, Roboto, Helvetica, Arial, sans-serif"
27
+ FONT = f'font-family="{FONT_STACK}"'
28
+
29
+ INK = "#1a1a1a"
30
+ MUTED = "#5c5c5c"
31
+ AXIS = "#8a8a8a"
32
+ GRID = "#e8e8e8"
33
+ GRID_STRONG = "#d0d0d0"
34
+
35
+ # Per-character advance as a fraction of font size, for a humanist sans at normal weight.
36
+ _NARROW = set("iljft.,;:|!()[]{}I '")
37
+ _WIDE = set("mwMW@%")
38
+ _DIGIT = set("0123456789")
39
+
40
+
41
+ def text_width(text: str, size: float, *, weight: str = "normal") -> float:
42
+ """Estimate rendered width in user units."""
43
+
44
+ total = 0.0
45
+ for character in text:
46
+ if character in _NARROW:
47
+ total += 0.30
48
+ elif character in _WIDE:
49
+ total += 0.86
50
+ elif character in _DIGIT:
51
+ total += 0.56
52
+ elif character.isupper():
53
+ total += 0.66
54
+ else:
55
+ total += 0.52
56
+ if weight in {"600", "700", "bold"}:
57
+ total *= 1.05
58
+ return total * size
59
+
60
+
61
+ def _fit_lines(text: str, size: float, limit: float, *, weight: str = "normal") -> list[str]:
62
+ """Wrap to at most two lines, breaking on whitespace."""
63
+
64
+ if text_width(text, size, weight=weight) <= limit:
65
+ return [text]
66
+ words = text.split(" ")
67
+ line: list[str] = []
68
+ for index, word in enumerate(words):
69
+ candidate = " ".join([*line, word])
70
+ if line and text_width(candidate, size, weight=weight) > limit:
71
+ return [" ".join(line), " ".join(words[index:])]
72
+ line.append(word)
73
+ return [" ".join(line)]
74
+
75
+
76
+ @dataclass
77
+ class Series:
78
+ label: str
79
+ x: list[float]
80
+ y: list[float]
81
+ color: str = PALETTE[0]
82
+ dash: str | None = None
83
+ marker: bool = True
84
+ width: float = 1.9
85
+ marker_size: float = 2.6
86
+
87
+
88
+ @dataclass
89
+ class Band:
90
+ """A shaded interval drawn behind its series."""
91
+
92
+ x: list[float]
93
+ low: list[float]
94
+ high: list[float]
95
+ color: str = PALETTE[0]
96
+ opacity: float = 0.16
97
+
98
+
99
+ @dataclass
100
+ class Counts:
101
+ """A support strip under the plot: how many rows sit behind each x position."""
102
+
103
+ x: list[float]
104
+ values: list[float]
105
+ color: str = MUTED
106
+ label: str = "rows per bin"
107
+
108
+
109
+ @dataclass
110
+ class Panel:
111
+ title: str
112
+ series: list[Series] = field(default_factory=list)
113
+ bands: list[Band] = field(default_factory=list)
114
+ xlabel: str = ""
115
+ ylabel: str = ""
116
+ xlim: tuple[float, float] | None = None
117
+ ylim: tuple[float, float] | None = None
118
+ xticks: list[tuple[float, str]] | None = None
119
+ yticks: list[tuple[float, str]] | None = None
120
+ xscale: str = "linear"
121
+ hlines: list[tuple[float, str, str]] = field(default_factory=list)
122
+ diagonal: bool = False
123
+ notes: list[str] = field(default_factory=list)
124
+ counts: Counts | None = None
125
+
126
+
127
+ def _fmt(value: float) -> str:
128
+ if abs(value) >= 1e6:
129
+ return f"{value:.3g}"
130
+ text = f"{value:.3f}".rstrip("0").rstrip(".")
131
+ return text or "0"
132
+
133
+
134
+ def _auto_ticks(low: float, high: float, count: int = 5) -> list[tuple[float, str]]:
135
+ if high <= low:
136
+ high = low + 1.0
137
+ step = (high - low) / count
138
+ magnitude = 10 ** math.floor(math.log10(step)) if step > 0 else 1.0
139
+ for factor in (1, 2, 2.5, 5, 10):
140
+ if step <= factor * magnitude:
141
+ step = factor * magnitude
142
+ break
143
+ start = math.ceil(low / step) * step
144
+ ticks = []
145
+ value = start
146
+ while value <= high + 1e-9:
147
+ ticks.append((value, _fmt(0.0 if abs(value) < step * 1e-6 else value)))
148
+ value += step
149
+ return ticks
150
+
151
+
152
+ def _text(
153
+ x: float,
154
+ y: float,
155
+ body: str,
156
+ *,
157
+ size: float,
158
+ fill: str = INK,
159
+ anchor: str = "start",
160
+ weight: str | None = None,
161
+ ) -> str:
162
+ weight_attr = f' font-weight="{weight}"' if weight else ""
163
+ return (
164
+ f'<text x="{x:.1f}" y="{y:.1f}" text-anchor="{anchor}" font-size="{size}" '
165
+ f'fill="{fill}"{weight_attr} {FONT}>{escape(body)}</text>'
166
+ )
167
+
168
+
169
+ _TITLE_SIZE = 12.5
170
+
171
+ # Everything below the plot box is stacked in fixed bands rather than placed at absolute
172
+ # offsets, so an x-axis label, a support strip, and a note can coexist without overlapping.
173
+ _TICK_BAND = 18.0
174
+ _XLABEL_BAND = 16.0
175
+ _COUNTS_BAND = 30.0
176
+ _NOTE_BAND = 12.0
177
+ _NOTE_LEAD = 6.0
178
+ _FLOOR_SLACK = 8.0
179
+
180
+
181
+ def _panel_bottom(panel: Panel) -> float:
182
+ bottom = _TICK_BAND + _FLOOR_SLACK
183
+ if panel.xlabel:
184
+ bottom += _XLABEL_BAND
185
+ if panel.counts:
186
+ bottom += _COUNTS_BAND
187
+ if panel.notes:
188
+ bottom += _NOTE_LEAD + _NOTE_BAND * len(panel.notes)
189
+ return bottom
190
+
191
+
192
+ def _panel_svg(panel: Panel, width: float, height: float, *, title_rows: int | None = None) -> str:
193
+ left, right = 54.0, 14.0
194
+ plot_w = width - left - right
195
+ title_lines = _fit_lines(panel.title, _TITLE_SIZE, plot_w, weight="600")
196
+ top = 18.0 + 14.0 * (title_rows or len(title_lines))
197
+ plot_h = height - top - _panel_bottom(panel)
198
+
199
+ xs = [x for s in panel.series for x in s.x] or [0.0, 1.0]
200
+ ys = [y for s in panel.series for y in s.y] or [0.0, 1.0]
201
+ ys += [value for band in panel.bands for value in (*band.low, *band.high)]
202
+ ys += [level for level, _, _ in panel.hlines]
203
+ xlim = panel.xlim or (min(xs), max(xs))
204
+ ylim = panel.ylim or (min(ys), max(ys))
205
+ if ylim[0] == ylim[1]:
206
+ ylim = (ylim[0] - 0.5, ylim[1] + 0.5)
207
+
208
+ def tx(value: float) -> float:
209
+ if panel.xscale == "log":
210
+ lo, hi = math.log(xlim[0]), math.log(xlim[1])
211
+ return left + (math.log(value) - lo) / (hi - lo) * plot_w
212
+ return left + (value - xlim[0]) / (xlim[1] - xlim[0]) * plot_w
213
+
214
+ def ty(value: float) -> float:
215
+ return top + plot_h - (value - ylim[0]) / (ylim[1] - ylim[0]) * plot_h
216
+
217
+ parts: list[str] = []
218
+ for index, line in enumerate(title_lines):
219
+ parts.append(
220
+ _text(
221
+ left + plot_w / 2,
222
+ 16.0 + 14.0 * index,
223
+ line,
224
+ size=_TITLE_SIZE,
225
+ anchor="middle",
226
+ weight="600",
227
+ )
228
+ )
229
+
230
+ xticks = panel.xticks or _auto_ticks(*xlim)
231
+ yticks = panel.yticks or _auto_ticks(*ylim)
232
+ for value, label in yticks:
233
+ if not ylim[0] - 1e-9 <= value <= ylim[1] + 1e-9:
234
+ continue
235
+ y = ty(value)
236
+ parts.append(
237
+ f'<line x1="{left}" y1="{y:.1f}" x2="{left + plot_w:.1f}" y2="{y:.1f}" '
238
+ f'stroke="{GRID}" stroke-width="1"/>'
239
+ )
240
+ parts.append(_text(left - 6, y + 3.4, label, size=10, fill=MUTED, anchor="end"))
241
+ for value, label in xticks:
242
+ if not xlim[0] - 1e-9 <= value <= xlim[1] + 1e-9:
243
+ continue
244
+ x = tx(value)
245
+ parts.append(
246
+ f'<line x1="{x:.1f}" y1="{top}" x2="{x:.1f}" y2="{top + plot_h:.1f}" '
247
+ f'stroke="{GRID}" stroke-width="1"/>'
248
+ )
249
+ parts.append(_text(x, top + plot_h + 13, label, size=10, fill=MUTED, anchor="middle"))
250
+
251
+ if panel.diagonal:
252
+ parts.append(
253
+ f'<line x1="{tx(xlim[0]):.1f}" y1="{ty(ylim[0]):.1f}" '
254
+ f'x2="{tx(xlim[1]):.1f}" y2="{ty(ylim[1]):.1f}" stroke="{AXIS}" '
255
+ f'stroke-width="1.1" stroke-dasharray="4 3"/>'
256
+ )
257
+
258
+ for band in panel.bands:
259
+ forward = " ".join(
260
+ f"{tx(x):.1f},{ty(y):.1f}" for x, y in zip(band.x, band.high, strict=True)
261
+ )
262
+ backward = " ".join(
263
+ f"{tx(x):.1f},{ty(y):.1f}"
264
+ for x, y in zip(reversed(band.x), reversed(band.low), strict=True)
265
+ )
266
+ parts.append(
267
+ f'<polygon points="{forward} {backward}" fill="{band.color}" '
268
+ f'fill-opacity="{band.opacity}" stroke="none"/>'
269
+ )
270
+
271
+ # A reference line carries its label at the left margin over a solid backing box, so the
272
+ # label never lands on the data it is a reference for.
273
+ for value, label, color in panel.hlines:
274
+ y = ty(value)
275
+ parts.append(
276
+ f'<line x1="{left}" y1="{y:.1f}" x2="{left + plot_w:.1f}" y2="{y:.1f}" '
277
+ f'stroke="{color}" stroke-width="1.2" stroke-dasharray="5 3"/>'
278
+ )
279
+ if not label:
280
+ continue
281
+ label_w = text_width(label, 9.5) + 8.0
282
+ parts.append(
283
+ f'<rect x="{left + 3:.1f}" y="{y - 12:.1f}" width="{label_w:.1f}" height="11.5" '
284
+ f'fill="white" fill-opacity="0.9" stroke="none"/>'
285
+ )
286
+ parts.append(_text(left + 7, y - 3.5, label, size=9.5, fill=color))
287
+
288
+ for series in panel.series:
289
+ points = " ".join(
290
+ f"{tx(x):.1f},{ty(y):.1f}" for x, y in zip(series.x, series.y, strict=True)
291
+ )
292
+ dash = f' stroke-dasharray="{series.dash}"' if series.dash else ""
293
+ parts.append(
294
+ f'<polyline points="{points}" fill="none" stroke="{series.color}" '
295
+ f'stroke-width="{series.width}" stroke-linejoin="round" '
296
+ f'stroke-linecap="round"{dash}/>'
297
+ )
298
+ if series.marker:
299
+ for x, y in zip(series.x, series.y, strict=True):
300
+ parts.append(
301
+ f'<circle cx="{tx(x):.1f}" cy="{ty(y):.1f}" r="{series.marker_size}" '
302
+ f'fill="{series.color}"/>'
303
+ )
304
+
305
+ parts.append(
306
+ f'<rect x="{left}" y="{top}" width="{plot_w:.1f}" height="{plot_h:.1f}" fill="none" '
307
+ f'stroke="{AXIS}" stroke-width="1"/>'
308
+ )
309
+
310
+ cursor = top + plot_h + _TICK_BAND
311
+ if panel.xlabel:
312
+ parts.append(
313
+ _text(
314
+ left + plot_w / 2,
315
+ cursor + 11,
316
+ panel.xlabel,
317
+ size=10.5,
318
+ fill=MUTED,
319
+ anchor="middle",
320
+ )
321
+ )
322
+ cursor += _XLABEL_BAND
323
+ if panel.counts:
324
+ counts = panel.counts
325
+ if not counts.values:
326
+ raise ValueError("a support strip needs at least one count")
327
+ strip_top = cursor + 2.0
328
+ strip_h = 15.0
329
+ peak = max(counts.values) or 1.0
330
+ slot = plot_w / max(len(counts.x), 1) * 0.7
331
+ for x, value in zip(counts.x, counts.values, strict=True):
332
+ bar_h = (value / peak) * strip_h
333
+ parts.append(
334
+ f'<rect x="{tx(x) - slot / 2:.1f}" y="{strip_top + strip_h - bar_h:.1f}" '
335
+ f'width="{slot:.1f}" height="{bar_h:.1f}" fill="{counts.color}" '
336
+ f'fill-opacity="0.5"/>'
337
+ )
338
+ parts.append(
339
+ f'<line x1="{left}" y1="{strip_top + strip_h:.1f}" x2="{left + plot_w:.1f}" '
340
+ f'y2="{strip_top + strip_h:.1f}" stroke="{GRID_STRONG}" stroke-width="1"/>'
341
+ )
342
+ parts.append(_text(left, strip_top + strip_h + 9, counts.label, size=9, fill=MUTED))
343
+ parts.append(
344
+ _text(
345
+ left + plot_w,
346
+ strip_top + strip_h + 9,
347
+ f"tallest {int(peak):,}",
348
+ size=9,
349
+ fill=MUTED,
350
+ anchor="end",
351
+ )
352
+ )
353
+ cursor += _COUNTS_BAND
354
+ for index, note in enumerate(panel.notes):
355
+ parts.append(
356
+ _text(left, cursor + _NOTE_LEAD + 9 + _NOTE_BAND * index, note, size=9.5, fill=MUTED)
357
+ )
358
+ if panel.ylabel:
359
+ parts.append(
360
+ f'<text transform="translate(13,{top + plot_h / 2:.1f}) rotate(-90)" '
361
+ f'text-anchor="middle" font-size="10.5" fill="{MUTED}" {FONT}>'
362
+ f"{escape(panel.ylabel)}</text>"
363
+ )
364
+ return "\n".join(parts)
365
+
366
+
367
+ def _legend_svg(
368
+ entries: list[tuple[str, str, str | None]],
369
+ *,
370
+ x: float,
371
+ y: float,
372
+ max_width: float,
373
+ ) -> tuple[str, float]:
374
+ """Flow legend entries across as many rows as their measured widths need."""
375
+
376
+ swatch, gap, pad = 22.0, 7.0, 22.0
377
+ parts: list[str] = []
378
+ cursor_x, cursor_y, rows = x, y, 1
379
+ for label, color, dash in entries:
380
+ entry_w = swatch + gap + text_width(label, 11) + pad
381
+ if cursor_x > x and cursor_x + entry_w - pad > x + max_width:
382
+ cursor_x, cursor_y, rows = x, cursor_y + 16.0, rows + 1
383
+ dash_attr = f' stroke-dasharray="{dash}"' if dash else ""
384
+ parts.append(
385
+ f'<line x1="{cursor_x:.1f}" y1="{cursor_y:.1f}" x2="{cursor_x + swatch:.1f}" '
386
+ f'y2="{cursor_y:.1f}" stroke="{color}" stroke-width="2.4" '
387
+ f'stroke-linecap="round"{dash_attr}/>'
388
+ )
389
+ parts.append(_text(cursor_x + swatch + gap, cursor_y + 3.8, label, size=11))
390
+ cursor_x += entry_w
391
+ return "\n".join(parts), 16.0 * rows
392
+
393
+
394
+ def _dedupe(entries: list[tuple[str, str, str | None]]) -> list[tuple[str, str, str | None]]:
395
+ seen: set[tuple[str, str, str | None]] = set()
396
+ unique = []
397
+ for entry in entries:
398
+ if entry not in seen:
399
+ seen.add(entry)
400
+ unique.append(entry)
401
+ return unique
402
+
403
+
404
+ def _open_svg(width: float, height: float, title: str, description: str) -> list[str]:
405
+ return [
406
+ f'<svg xmlns="http://www.w3.org/2000/svg" width="{width:.0f}" height="{height:.0f}" '
407
+ f'viewBox="0 0 {width:.0f} {height:.0f}" role="img" aria-label="{escape(title)}">',
408
+ f"<title>{escape(title)}</title>",
409
+ f"<desc>{escape(description)}</desc>",
410
+ '<rect width="100%" height="100%" fill="white"/>',
411
+ ]
412
+
413
+
414
+ def render_grid(
415
+ panels: list[Panel],
416
+ *,
417
+ columns: int,
418
+ title: str,
419
+ subtitle: str = "",
420
+ panel_width: float = 352.0,
421
+ panel_height: float = 256.0,
422
+ legend: list[tuple[str, str, str | None]] | None = None,
423
+ description: str = "",
424
+ ) -> str:
425
+ """Render panels on a grid with one shared title and a flowed legend.
426
+
427
+ A final short row is centred, so a five-panel figure on three columns has no empty cell.
428
+ """
429
+
430
+ if not panels:
431
+ raise ValueError("render_grid needs at least one panel")
432
+ if columns < 1:
433
+ raise ValueError("render_grid needs at least one column")
434
+ legend = _dedupe(legend or [])
435
+ rows = math.ceil(len(panels) / columns)
436
+ width = columns * panel_width
437
+ header = 26.0 + (16.0 if subtitle else 0.0)
438
+ legend_svg, legend_h = "", 0.0
439
+ if legend:
440
+ legend_svg, legend_h = _legend_svg(legend, x=18.0, y=header + 10.0, max_width=width - 36.0)
441
+ legend_h += 8.0
442
+ height = header + legend_h + rows * panel_height
443
+ parts = _open_svg(width, height, title, description or title)
444
+ parts.append(_text(width / 2, 19, title, size=14.5, anchor="middle", weight="700"))
445
+ if subtitle:
446
+ parts.append(_text(width / 2, 34, subtitle, size=11, fill=MUTED, anchor="middle"))
447
+ if legend_svg:
448
+ parts.append(legend_svg)
449
+ # One title row count for the whole grid, so a panel whose title wraps does not push its
450
+ # plot box below its neighbours'.
451
+ title_rows = max(
452
+ len(_fit_lines(panel.title, _TITLE_SIZE, panel_width - 68.0, weight="600"))
453
+ for panel in panels
454
+ )
455
+ for index, panel in enumerate(panels):
456
+ row, column = divmod(index, columns)
457
+ in_row = min(columns, len(panels) - row * columns)
458
+ offset = (columns - in_row) * panel_width / 2.0
459
+ px = offset + column * panel_width
460
+ py = header + legend_h + row * panel_height
461
+ parts.append(f'<g transform="translate({px:.1f},{py:.1f})">')
462
+ parts.append(_panel_svg(panel, panel_width, panel_height, title_rows=title_rows))
463
+ parts.append("</g>")
464
+ parts.append("</svg>")
465
+ return "\n".join(parts) + "\n"
466
+
467
+
468
+ def render_bars(
469
+ groups: list[str],
470
+ series: list[tuple[str, list[float], str]],
471
+ *,
472
+ title: str,
473
+ ylabel: str,
474
+ subtitle: str = "",
475
+ ylim: tuple[float, float] = (0.0, 1.0),
476
+ reference: list[tuple[str, list[float], str]] | None = None,
477
+ notes: list[str] | None = None,
478
+ separator_before: int | None = None,
479
+ width: float = 780.0,
480
+ height: float = 350.0,
481
+ description: str = "",
482
+ ) -> str:
483
+ """Render grouped bars with per-group dashed reference levels.
484
+
485
+ Bars keep a zero baseline. Value labels are drawn only where a bar is wide enough to hold
486
+ one, because a crowded label is worse than none.
487
+ """
488
+
489
+ if not groups or not series:
490
+ raise ValueError("render_bars needs at least one group and one series")
491
+ if any(len(values) != len(groups) for _, values, _ in series):
492
+ raise ValueError("every bar series must carry one value per group")
493
+ notes = notes or []
494
+ left, right, top = 54.0, 16.0, 26.0 + (15.0 if subtitle else 0.0)
495
+ legend_entries = [(label, color, None) for label, _, color in series]
496
+ legend_entries += [(label, color, "5 3") for label, _, color in reference or []]
497
+ legend_svg, legend_h = _legend_svg(
498
+ _dedupe(legend_entries), x=18.0, y=top + 12.0, max_width=width - 36.0
499
+ )
500
+ top += legend_h + 10.0
501
+ bottom = 46.0 + 12.0 * len(notes)
502
+ plot_w, plot_h = width - left - right, height - top - bottom
503
+ group_w = plot_w / len(groups)
504
+ bar_w = group_w * 0.74 / len(series)
505
+
506
+ def ty(value: float) -> float:
507
+ return top + plot_h - (value - ylim[0]) / (ylim[1] - ylim[0]) * plot_h
508
+
509
+ parts = _open_svg(width, height, title, description or title)
510
+ parts.append(_text(width / 2, 19, title, size=14.5, anchor="middle", weight="700"))
511
+ if subtitle:
512
+ parts.append(_text(width / 2, 34, subtitle, size=11, fill=MUTED, anchor="middle"))
513
+ parts.append(legend_svg)
514
+ for value, label in _auto_ticks(*ylim):
515
+ y = ty(value)
516
+ parts.append(
517
+ f'<line x1="{left}" y1="{y:.1f}" x2="{left + plot_w:.1f}" y2="{y:.1f}" '
518
+ f'stroke="{GRID}" stroke-width="1"/>'
519
+ )
520
+ parts.append(_text(left - 6, y + 3.4, label, size=10, fill=MUTED, anchor="end"))
521
+
522
+ label_fits = bar_w - 2 >= text_width("0.000", 8.5) + 2
523
+ for g_index, group in enumerate(groups):
524
+ gx = left + g_index * group_w + group_w * 0.13
525
+ for s_index, (_, values, color) in enumerate(series):
526
+ value = values[g_index]
527
+ x = gx + s_index * bar_w
528
+ parts.append(
529
+ f'<rect x="{x:.1f}" y="{ty(value):.1f}" width="{bar_w - 2:.1f}" '
530
+ f'height="{ty(ylim[0]) - ty(value):.1f}" fill="{color}"/>'
531
+ )
532
+ if label_fits:
533
+ parts.append(
534
+ _text(
535
+ x + (bar_w - 2) / 2,
536
+ ty(value) - 4,
537
+ f"{value:.3f}",
538
+ size=8.5,
539
+ anchor="middle",
540
+ )
541
+ )
542
+ for _, values, color in reference or []:
543
+ y = ty(values[g_index])
544
+ parts.append(
545
+ f'<line x1="{gx - 3:.1f}" y1="{y:.1f}" '
546
+ f'x2="{gx + group_w * 0.74 + 3:.1f}" y2="{y:.1f}" stroke="{color}" '
547
+ f'stroke-width="1.6" stroke-dasharray="5 3"/>'
548
+ )
549
+ parts.append(
550
+ _text(
551
+ left + g_index * group_w + group_w / 2,
552
+ top + plot_h + 16,
553
+ group,
554
+ size=11,
555
+ anchor="middle",
556
+ )
557
+ )
558
+ if separator_before is not None and 0 < separator_before < len(groups):
559
+ x = left + separator_before * group_w
560
+ parts.append(
561
+ f'<line x1="{x:.1f}" y1="{top}" x2="{x:.1f}" y2="{top + plot_h + 6:.1f}" '
562
+ f'stroke="{GRID_STRONG}" stroke-width="1.4"/>'
563
+ )
564
+ parts.append(
565
+ f'<rect x="{left}" y="{top}" width="{plot_w:.1f}" height="{plot_h:.1f}" fill="none" '
566
+ f'stroke="{AXIS}" stroke-width="1"/>'
567
+ )
568
+ parts.append(
569
+ f'<text transform="translate(13,{top + plot_h / 2:.1f}) rotate(-90)" '
570
+ f'text-anchor="middle" font-size="10.5" fill="{MUTED}" {FONT}>'
571
+ f"{escape(ylabel)}</text>"
572
+ )
573
+ for index, note in enumerate(notes):
574
+ parts.append(_text(left, top + plot_h + 34 + 12 * index, note, size=9.5, fill=MUTED))
575
+ parts.append("</svg>")
576
+ return "\n".join(parts) + "\n"
figures/ettin-comparison.svg ADDED
model.directml.onnx → model.vulkan.onnx RENAMED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:05abdc03628b599712a993381919aef1e22fcc5ad6384608abc2e069c402e369
3
- size 2242802
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:cf27725a15982be3a0c329aec48486e5e048de5b4d3028f0cf88915eafa79b40
3
+ size 2243288
publication-manifest.json CHANGED
@@ -1,6 +1,6 @@
1
  {
2
  "schema": "daecore.retrieval-publication-manifest",
3
- "tool_sha256": "2369a55f4fa3f007f800834880b6e4a65624c91589b469f48133ebc058c3d2ef",
4
  "model_id": "ettin-150m-memory-reranker-ft-v1",
5
  "public_repo": "Daecore/ettin-150m-memory-reranker-ft-v1",
6
  "role": "cross-encoder-reranker",
@@ -9,21 +9,30 @@
9
  "revision": "025501c4e0f9bbeb4c5b198318e0089ff061cc14",
10
  "tree_sha256": "e0edfb38cc495f23507de50233d0965f4e72ad0872e70e6f4e8fc889d394ea01"
11
  },
12
- "export_receipt_sha256": "4faf07d568cc5da6400e9aa2645d73b2a2562cb720742c64106b2a70f43b12ba",
13
  "serving_sha256": "ce05a6a810007372c6de0f4c56dc94f652e1d28f2eeddf1878b64a170b4c48c2",
14
  "source_model_tree_sha256": "bd8028b5dbc6ab2cd2ab9a6019346de7b25536f50022617194bcb8503d4467ad",
15
- "staged_at": "2026-09-11T14:57:19+00:00",
 
 
 
 
 
 
 
 
 
 
 
 
 
 
16
  "files": {
17
  "config.json": {
18
  "sha256": "a897acffc99f97238bb5bc0e1bccd6b7fe8fdfebe1450d4db3a9ab026cce6795",
19
  "size": 2020,
20
  "binding": "export receipt (file hash)"
21
  },
22
- "model.directml.onnx": {
23
- "sha256": "05abdc03628b599712a993381919aef1e22fcc5ad6384608abc2e069c402e369",
24
- "size": 2242802,
25
- "binding": "export receipt (file hash)"
26
- },
27
  "model.onnx": {
28
  "sha256": "7df4a3b3cd89d216c016abd7d169f0fd215e9f79c06a2ffed57bac1abf556073",
29
  "size": 2231565,
@@ -34,9 +43,9 @@
34
  "size": 299270144,
35
  "binding": "export receipt (file hash)"
36
  },
37
- "serving.directml.json": {
38
- "sha256": "7a43e53663efc6e82db9b29a13fa3775a35d73bd6731e1163601f8f36748925f",
39
- "size": 711,
40
  "binding": "export receipt (file hash)"
41
  },
42
  "serving.json": {
@@ -44,6 +53,11 @@
44
  "size": 704,
45
  "binding": "export receipt (file hash)"
46
  },
 
 
 
 
 
47
  "tokenizer.json": {
48
  "sha256": "6c8aaa9a542084f2457eab775d4eeb51f92a70c0fd9de28d5edb0ddec3c08d30",
49
  "size": 3583228,
@@ -54,10 +68,10 @@
54
  "size": 488,
55
  "binding": "export receipt (file hash)"
56
  },
57
- "README.md": {
58
- "sha256": "19b97e33bddbcb587446bbbb164cd091fc2cb534911cf0aafb0efc8013bda68c",
59
- "size": 11806,
60
- "binding": "packaging record"
61
  },
62
  "LICENSE": {
63
  "sha256": "c71d239df91726fc519c6eb72d318ec65820627232b2f796219e87dcf35d0ab4",
@@ -70,9 +84,65 @@
70
  "binding": "packaging record"
71
  },
72
  "MODIFICATIONS.md": {
73
- "sha256": "1f8c677387fb72ef31b7a107e866868a13821b6282debebc6f8150eac9866b18",
74
- "size": 1958,
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
75
  "binding": "packaging record"
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
76
  }
77
  }
78
  }
 
1
  {
2
  "schema": "daecore.retrieval-publication-manifest",
3
+ "tool_sha256": "5ac2ef9465fe9a3454bfd2e61c7aa651b08fbdfef9473fb5b3455585ba560a7b",
4
  "model_id": "ettin-150m-memory-reranker-ft-v1",
5
  "public_repo": "Daecore/ettin-150m-memory-reranker-ft-v1",
6
  "role": "cross-encoder-reranker",
 
9
  "revision": "025501c4e0f9bbeb4c5b198318e0089ff061cc14",
10
  "tree_sha256": "e0edfb38cc495f23507de50233d0965f4e72ad0872e70e6f4e8fc889d394ea01"
11
  },
12
+ "export_receipt_sha256": "92cc733460256c481547efc3e3aaeee460bd35b2f417caf40cd5b312693f7d46",
13
  "serving_sha256": "ce05a6a810007372c6de0f4c56dc94f652e1d28f2eeddf1878b64a170b4c48c2",
14
  "source_model_tree_sha256": "bd8028b5dbc6ab2cd2ab9a6019346de7b25536f50022617194bcb8503d4467ad",
15
+ "retirement": {
16
+ "parent_revision": "2e1d6d0bef0e3a1a801496fac1fbb9c0d87effd4",
17
+ "predecessor_manifest_sha256": "dfb9d18a795c431ea75dd9fde14509cd4e12e8be962ebc5b591e7b069c099999",
18
+ "files": {
19
+ "model.directml.onnx": {
20
+ "sha256": "05abdc03628b599712a993381919aef1e22fcc5ad6384608abc2e069c402e369",
21
+ "size": 2242802
22
+ },
23
+ "serving.directml.json": {
24
+ "sha256": "7a43e53663efc6e82db9b29a13fa3775a35d73bd6731e1163601f8f36748925f",
25
+ "size": 711
26
+ }
27
+ }
28
+ },
29
+ "staged_at": "2026-09-29T03:04:18+00:00",
30
  "files": {
31
  "config.json": {
32
  "sha256": "a897acffc99f97238bb5bc0e1bccd6b7fe8fdfebe1450d4db3a9ab026cce6795",
33
  "size": 2020,
34
  "binding": "export receipt (file hash)"
35
  },
 
 
 
 
 
36
  "model.onnx": {
37
  "sha256": "7df4a3b3cd89d216c016abd7d169f0fd215e9f79c06a2ffed57bac1abf556073",
38
  "size": 2231565,
 
43
  "size": 299270144,
44
  "binding": "export receipt (file hash)"
45
  },
46
+ "model.vulkan.onnx": {
47
+ "sha256": "cf27725a15982be3a0c329aec48486e5e048de5b4d3028f0cf88915eafa79b40",
48
+ "size": 2243288,
49
  "binding": "export receipt (file hash)"
50
  },
51
  "serving.json": {
 
53
  "size": 704,
54
  "binding": "export receipt (file hash)"
55
  },
56
+ "serving.vulkan.json": {
57
+ "sha256": "908e092f9c2197fe1028ece2aa8aa142dc088528dacc7237b25c81bfdd44a837",
58
+ "size": 946,
59
+ "binding": "export receipt (file hash)"
60
+ },
61
  "tokenizer.json": {
62
  "sha256": "6c8aaa9a542084f2457eab775d4eeb51f92a70c0fd9de28d5edb0ddec3c08d30",
63
  "size": 3583228,
 
68
  "size": 488,
69
  "binding": "export receipt (file hash)"
70
  },
71
+ "vulkan-derivation.json": {
72
+ "sha256": "e3b08005807fb6b9974d327fa8db1d90df317109879762be55397d8a4c42baa3",
73
+ "size": 972,
74
+ "binding": "export receipt (file hash)"
75
  },
76
  "LICENSE": {
77
  "sha256": "c71d239df91726fc519c6eb72d318ec65820627232b2f796219e87dcf35d0ab4",
 
84
  "binding": "packaging record"
85
  },
86
  "MODIFICATIONS.md": {
87
+ "sha256": "18e06084d4a6e34b6c7a976da739cb8b28383f7132552a89c2d4d0f0f7189257",
88
+ "size": 2164,
89
+ "binding": "packaging record"
90
+ },
91
+ "evaluation/README.md": {
92
+ "sha256": "3215e6b0e302fdecd86ef1457e46e7d033fa18fb1ebcbd0dcb47c2b8abe7c174",
93
+ "size": 8825,
94
+ "binding": "packaging record"
95
+ },
96
+ "evaluation/metrics.py": {
97
+ "sha256": "7c25f29a03e0f61d5f0e59f7781269489f80512ae6862920f91a2d333c058bc4",
98
+ "size": 11135,
99
+ "binding": "packaging record"
100
+ },
101
+ "evaluation/reranker.json": {
102
+ "sha256": "20b5760a462c3c2d2281607f2c5045d3f60210123a0c11c8b2d92580c3cccd3e",
103
+ "size": 1145459,
104
  "binding": "packaging record"
105
+ },
106
+ "evaluation/reranker-summary.json": {
107
+ "sha256": "05face1266e65d5a97a3053eed857c190fb32119581e395339770f3697bd5287",
108
+ "size": 5939,
109
+ "binding": "packaging record"
110
+ },
111
+ "evaluation/serving.json": {
112
+ "sha256": "0c017dbd0e66734c465d617eef24b2affb2670a96f7fd5b2c001514c3cc1381a",
113
+ "size": 500118,
114
+ "binding": "packaging record"
115
+ },
116
+ "evaluation/serving-summary.json": {
117
+ "sha256": "3a7c6a2c96495917e65a7dabde9c02c7c457b5416df51cbb58bb7267d5c74fda",
118
+ "size": 3162,
119
+ "binding": "packaging record"
120
+ },
121
+ "evaluation/serving-qualification.json": {
122
+ "sha256": "bf7deda3f25306d15454798ad4c8e785c25fba970fc774efbe557225fc5d42e0",
123
+ "size": 4667,
124
+ "binding": "packaging record"
125
+ },
126
+ "evaluation/figures.py": {
127
+ "sha256": "9944765ee42148f2bb135f428ee1dacb822af0ca5c77be9f82b507eb8da236d4",
128
+ "size": 4097,
129
+ "binding": "packaging record"
130
+ },
131
+ "evaluation/svg_figures.py": {
132
+ "sha256": "c68b7e26e072e3875f111b592f8031322261272fcf285ae2a3edbe378e70740d",
133
+ "size": 21536,
134
+ "binding": "packaging record"
135
+ },
136
+ "figures/ettin-comparison.svg": {
137
+ "sha256": "5d34b26e75afd868632cd1bca16081cdd6d28a60f85bbdc8592dffd80cea6814",
138
+ "size": 15067,
139
+ "binding": "packaging record"
140
+ },
141
+ "README.md": {
142
+ "sha256": "d81407ff4ec675405036dcd2179f825cc2dade7595ba55791a7d886d4ff41a89",
143
+ "size": 7628,
144
+ "binding": "model card with upload-relative links",
145
+ "source_sha256": "d1fe7d1b266f383ce8248e0682f4d22bb3cccb64d7076175d88f3dfa5351bda0"
146
  }
147
  }
148
  }
serving.directml.json → serving.vulkan.json RENAMED
@@ -6,7 +6,7 @@
6
  "max_batch": 8,
7
  "max_sequence": 1153
8
  },
9
- "graph": "model.directml.onnx",
10
  "inputs": [
11
  "input_ids",
12
  "attention_mask"
@@ -14,12 +14,21 @@
14
  "output": "scores",
15
  "precision": "float32",
16
  "provider_chain": [
17
- "DmlExecutionProvider",
18
  "CPUExecutionProvider"
19
  ],
 
 
 
 
 
20
  "role": "cross-encoder-reranker",
 
 
 
 
21
  "schema": "daecore.retrieval-onnx-serving.v1",
22
- "sha256": "b83ffa3a44f19d145ecb87e36a56da290b25def03db267e33c1c57a7a46dda76",
23
  "source_model_tree_sha256": "bd8028b5dbc6ab2cd2ab9a6019346de7b25536f50022617194bcb8503d4467ad",
24
  "torch_required_at_runtime": false
25
  }
 
6
  "max_batch": 8,
7
  "max_sequence": 1153
8
  },
9
+ "graph": "model.vulkan.onnx",
10
  "inputs": [
11
  "input_ids",
12
  "attention_mask"
 
14
  "output": "scores",
15
  "precision": "float32",
16
  "provider_chain": [
17
+ "WebGpuExecutionProvider",
18
  "CPUExecutionProvider"
19
  ],
20
+ "provider_options": {
21
+ "dawnBackendType": "Vulkan",
22
+ "enableInt64": "1",
23
+ "storageBufferCacheMode": "lazyRelease"
24
+ },
25
  "role": "cross-encoder-reranker",
26
+ "runtime_requirements": [
27
+ "onnxruntime==1.24.4",
28
+ "onnxruntime-ep-webgpu==0.4.0"
29
+ ],
30
  "schema": "daecore.retrieval-onnx-serving.v1",
31
+ "sha256": "bdace46b267a6fed5c084fde8798ed60c199e38c59553090d7f44618a65e2ca3",
32
  "source_model_tree_sha256": "bd8028b5dbc6ab2cd2ab9a6019346de7b25536f50022617194bcb8503d4467ad",
33
  "torch_required_at_runtime": false
34
  }
vulkan-derivation.json ADDED
@@ -0,0 +1,33 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "artifacts": [
3
+ {
4
+ "bytes": 2243288,
5
+ "path": "model.vulkan.onnx",
6
+ "sha256": "cf27725a15982be3a0c329aec48486e5e048de5b4d3028f0cf88915eafa79b40"
7
+ },
8
+ {
9
+ "bytes": 299270144,
10
+ "path": "model.onnx.data",
11
+ "sha256": "78ddab681c2f506d62bbcd475fec5c10c62d867ed60d97a0170828efaff5bdde"
12
+ }
13
+ ],
14
+ "changes": [
15
+ {
16
+ "dtype": 9,
17
+ "name": "node_GatherND_32",
18
+ "operation": "GatherND"
19
+ },
20
+ {
21
+ "dtype": 7,
22
+ "name": "node_abs_1",
23
+ "operation": "Abs"
24
+ }
25
+ ],
26
+ "max_sequence": 1153,
27
+ "role": "ettin",
28
+ "schema": "daecore.vulkan-mask-derivation.v1",
29
+ "source_derivation_sha256": "d874e9b6d92ae885733ae23d33e5612596af1f11c2bee725bf33649b8884357a",
30
+ "source_graph_sha256": "05abdc03628b599712a993381919aef1e22fcc5ad6384608abc2e069c402e369",
31
+ "source_weights_sha256": "78ddab681c2f506d62bbcd475fec5c10c62d867ed60d97a0170828efaff5bdde",
32
+ "weights_unchanged": true
33
+ }