tnh0527 commited on
Commit
13060d2
·
verified ·
1 Parent(s): cc80fc0

Update embeddinggemma-300m-memory-ft-v2 documentation

Browse files
README.md CHANGED
@@ -22,7 +22,15 @@ tags:
22
 
23
  An English embedding model for finding the passages that answer a question in notes, procedures, decision records and technical documentation. It is a fine-tune of [Google's EmbeddingGemma-300M](https://huggingface.co/google/embeddinggemma-300m), packaged as one ONNX graph with pooling and normalization built in, so it runs locally without PyTorch.
24
 
25
- On Daecore's **970-query dense-retrieval panel**, v2 raises nDCG@10 from **0.4506 to 0.5748** and precision@10 from **47.40% to 61.18%** compared with upstream Gemma. Both models search the same 82,719 passages, with reviewed labels through rank 20. V2 also recovers more than half of the first Daecore fine-tune's measured loss on FiQA and SciFact.
 
 
 
 
 
 
 
 
26
 
27
  Embed queries and passages with the same revision, and re-embed stored passages when switching to v2.
28
 
@@ -80,6 +88,13 @@ chooses the returned prefix. Chunking and candidate coverage therefore affect
80
  the final results alongside embedding quality. The package also works as a
81
  standalone embedder with other retrieval systems.
82
 
 
 
 
 
 
 
 
83
  ## Evaluation
84
 
85
  ### Gemma alone: dense retrieval on Daecore data
@@ -89,22 +104,32 @@ exact dense retrieval, the same frozen query forms and matched 128/1,024-token
89
  input limits. BM25 and Ettin do not contribute to these results. Each cell
90
  reads upstream → **Daecore v2**.
91
 
 
 
 
 
92
  | Cutoff | Hit | Precision | Known-positive recall | nDCG |
93
  |---|---:|---:|---:|---:|
94
- | 1 | 53.40% → **64.74%** | 53.40% → **64.74%** | 1.94% → **2.46%** | 0.4695 → **0.5421** |
95
- | 3 | 72.99% → **79.38%** | 50.83% → **63.08%** | 5.43% → **6.95%** | 0.4516 → **0.5524** |
96
- | 5 | 79.48% → **84.74%** | 49.95% → **62.66%** | 8.53% → **11.09%** | 0.4457 → **0.5586** |
97
- | 10 | 84.85% → **88.14%** | 47.40% → **61.18%** | 15.77% → **20.82%** | 0.4506 → **0.5748** |
98
- | 20 | 88.76% → **90.52%** | 43.94% → **58.11%** | 28.08% → **38.65%** | 0.4736 → **0.6065** |
 
99
 
100
  ![Gemma v2 dense retrieval compared with upstream](figures/gemma-comparison.svg)
101
 
102
- Every top-20 position has a relevance grade or a declared judging abstention;
103
- there are no missing-label ranges. All 970 queries remain eligible at every
104
- cutoff. At rank 20, 56 upstream positions and 53 v2 positions are excluded,
105
- without pulling in deeper results. Precision pools retained positions;
106
- known-positive recall averages the 911 queries with judged useful evidence.
107
- The corpus is not exhaustively labeled.
 
 
 
 
 
108
 
109
  Grades 2–3 count as useful. nDCG uses gains 0, 1, 3 and 7 against the same
110
  reviewed reference for both models, averaging all 970 queries. Rankings from
@@ -131,61 +156,80 @@ V2 is below upstream on all five datasets, most on FiQA (−0.0273). The first D
131
 
132
  The model was trained at 768, 512, 256 and 128 dimensions. To use a smaller
133
  width, keep the first dimensions and normalize the shortened vector again.
134
- Use the same width for queries and passages. The table measures **v2 alone**,
135
- re-ranking the five public corpora from its cached FP32 vectors; the 768-wide
136
- scores exactly reproduce the original per-query nDCG@10 results.
137
-
138
- | Dimensions | Bytes per FP32 vector | SciFact | FiQA | NFCorpus | SciDocs | ArguAna |
139
- |---|---:|---:|---:|---:|---:|---:|
140
- | 768 (default) | 3,072 | 0.7783 | 0.4468 | 0.3890 | 0.1852 | 0.6259 |
141
- | 512 | 2,048 | 0.7841 | 0.4374 | 0.3876 | 0.1793 | 0.6137 |
142
- | 256 | 1,024 | 0.7757 | 0.4091 | 0.3660 | 0.1708 | 0.5987 |
143
- | 128 | 512 | 0.7488 | 0.3668 | 0.3311 | 0.1484 | 0.5533 |
144
-
145
- At 512 dimensions, each raw FP32 vector uses one-third less storage; at 256,
146
- two-thirds less. These byte counts exclude index overhead and compression.
147
- Shortening vectors does not reduce the encoder's forward-pass cost. Lower
148
- widths lose quality on most of these panels; Daecore uses 768, and its private
149
- and full-pipeline scores in this card apply only to that width.
 
 
 
150
 
151
  ### Full pipeline: Gemma v2 + BM25 + Ettin
152
 
153
- This separate replay measures the complete retrieval path: semantic and lexical
154
- search, reciprocal-rank fusion, then Ettin reranking. It covers **all 970
155
- development queries over 82,719 passages**, including queries with no known
156
- useful evidence. The table reports fixed return depths on CUDA FP16 and is
157
- shared by the Gemma and Ettin cards; it is not either model's standalone score.
 
 
158
 
159
  | Return depth | Hit | Precision | Known-positive recall | nDCG |
160
  |---|---:|---:|---:|---:|
161
- | 1 | 72.33% | 72.33% | 3.46% | 0.6646 |
162
- | 3 | 85.36% | 71.58% | 9.52% | 0.6577 |
163
- | 5 | 88.45% | 70.28% | 15.38% | 0.6555 |
164
- | 10 | 89.90% | 67.18% | 28.03% | 0.6611 |
165
- | 20 | 91.03% | 61.79% | 48.85% | 0.6841 |
166
-
167
- Grades 2–3 count as useful. Depth 1 uses the same 965 queries for both
168
- providers: five are omitted because at least one first result was ungradable.
169
- At depths 3–20, Hit and nDCG average all 970 queries. Precision pools retained
170
- positions. Recall averages known-positive queries (890 at depth 1; 895 at
171
- depths 3–20) and measures judged passages, not every useful passage in the corpus. A shared judging-exclusion set removes 52 positions
172
- from the top 20 without backfilling. Daecore's selector chooses a prefix of
173
- 3–20; these fixed-depth scores do not evaluate that choice. The
174
- [evaluation companion](evaluation/README.md) provides the
175
- records, exclusions and CUDA/Vulkan comparison; the
176
- [pipeline overview](https://huggingface.co/Daecore) explains the retrieval path.
 
 
 
 
 
 
 
 
 
177
 
178
  ## Training data and objective
179
 
180
- Private supervision teaches passage relevance for agent retrieval. A question can have several useful passages; decisive evidence is graded above useful but incomplete evidence, and similar but unhelpful passages are explicit negatives. The documents are generated organizational material plus one person's project documentation, much of it AI-written. Four models from three provider families wrote the queries without seeing a designated answer, and model judges graded relevance under a common rubric. Construction, chunking, query writing and judging follow shared processes, so the data does not stand in for other users' workspaces.
 
 
 
 
 
 
 
 
181
 
182
- Private training passages were prepared with frozen versions of Daecore's
183
- structural chunker: headings and section context, cuts near paragraph or sentence
184
- boundaries, and handling for tables and code blocks. The corpus includes retained
185
- chunks from earlier parser versions; it was not regenerated for this release.
186
- Gemma applies its own tokenizer and input limits after passage preparation.
187
- Changing boundaries can change whether evidence stays together and how much the
188
- model sees, so these results are tied to the evaluated passages.
189
 
190
  Public supervision uses existing evidence annotations from [HotpotQA](https://hotpotqa.github.io/), [MultiDoc2Dial](https://github.com/doc2dial/multidoc2dial) and [FinQA](https://github.com/czyssrs/FinQA): multi-document questions, document-grounded dialogue, and textual or table evidence for numerical questions. The model learns to retrieve that evidence, not to perform FinQA's calculations; FinQA is unrelated to the FiQA evaluation above.
191
 
@@ -229,7 +273,7 @@ Vulkan runs through the native ONNX Runtime WebGPU plugin, [`onnxruntime-ep-webg
229
 
230
  Checks used Windows x64 with an NVIDIA RTX 3060 Ti (8 GB); other GPUs and Linux are untested, even where Vulkan is available.
231
 
232
- The task-specific data is mostly generated and model-judged, with no human reference panel, so it describes this development distribution rather than other users' documents. The dense and full-pipeline panels use the same queries but different judgment pools; compare models within each table. Some questions omit the project or task context needed for an unambiguous judgment. The 970-query pipeline panel uses natural-language questions and reports no separate terse-query or paraphrase robustness scores. Passage relevance does not measure distinct-fact coverage or downstream agent success.
233
 
234
  ## License
235
 
 
22
 
23
  An English embedding model for finding the passages that answer a question in notes, procedures, decision records and technical documentation. It is a fine-tune of [Google's EmbeddingGemma-300M](https://huggingface.co/google/embeddinggemma-300m), packaged as one ONNX graph with pooling and normalization built in, so it runs locally without PyTorch.
24
 
25
+ | Stronger private retrieval | Less forgetting | Local deployment |
26
+ |---|---|---|
27
+ | nDCG@10 **0.4363 → 0.5574**; precision@10 **47.40% → 61.18%** | Recovers **more than half** of the first fine-tune's loss on FiQA and SciFact | **300M parameters**, ONNX, CPU / CUDA / Vulkan |
28
+
29
+ The private comparison covers **970 queries and 82,719 passages**, using a
30
+ shared reference of **127,665 graded query–passage pairs** and reviewed results
31
+ through rank 50. Daecore accepted some general-retrieval loss for stronger
32
+ evidence retrieval on its document workflow. The tables below show both the gains
33
+ and the remaining public-benchmark gaps.
34
 
35
  Embed queries and passages with the same revision, and re-embed stored passages when switching to v2.
36
 
 
88
  the final results alongside embedding quality. The package also works as a
89
  standalone embedder with other retrieval systems.
90
 
91
+ Coverage at **50 candidates** measures what is available to a reranker. Hit@50
92
+ asks whether that pool contains any useful evidence; known-positive recall@50
93
+ measures how much of the judged useful evidence it contains. A reranker can
94
+ improve the order but cannot recover a passage outside its pool. Gemma's dense
95
+ top 50 describes a dense-only pipeline. Daecore's actual ceiling depends on
96
+ the 50 candidates admitted after Gemma and BM25 are fused and deduplicated.
97
+
98
  ## Evaluation
99
 
100
  ### Gemma alone: dense retrieval on Daecore data
 
104
  input limits. BM25 and Ettin do not contribute to these results. Each cell
105
  reads upstream → **Daecore v2**.
106
 
107
+ **Reference:** 127,665 graded query–passage pairs · shared top-50 reference
108
+ (September 2026) · abstentions excluded. This pooled reference covers both
109
+ models and is not an exhaustive labeling of the corpus.
110
+
111
  | Cutoff | Hit | Precision | Known-positive recall | nDCG |
112
  |---|---:|---:|---:|---:|
113
+ | 1 | 53.40% → **64.74%** | 53.40% → **64.74%** | 1.44% → **1.79%** | 0.4592 → **0.5313** |
114
+ | 3 | 72.99% → **79.38%** | 50.83% → **63.08%** | 3.94% → **5.03%** | 0.4414 → **0.5407** |
115
+ | 5 | 79.48% → **84.74%** | 49.95% → **62.66%** | 6.17% → **8.06%** | 0.4346 → **0.5453** |
116
+ | 10 | 84.85% → **88.14%** | 47.40% → **61.18%** | 11.18% → **15.07%** | 0.4363 → **0.5574** |
117
+ | 20 | 88.76% → **90.52%** | 43.94% → **58.11%** | 19.61% → **27.56%** | 0.4526 → **0.5819** |
118
+ | 50 | 93.20% → **94.43%** | 38.60% → **50.45%** | 41.12% → **57.92%** | 0.5107 → **0.6450** |
119
 
120
  ![Gemma v2 dense retrieval compared with upstream](figures/gemma-comparison.svg)
121
 
122
+ Every top-50 position is graded or explicitly excluded as ungradable, with no
123
+ missing-label ranges. All 970 queries remain eligible at every cutoff. At rank
124
+ 50, 164 upstream positions and 167 v2 positions are excluded without backfill.
125
+ Precision pools retained positions; recall averages the 925 queries with judged
126
+ useful evidence. The corpus is not exhaustively labeled.
127
+
128
+ At 50, v2 finds useful evidence for **94.43%** of queries and retrieves **57.92%**
129
+ of known useful passages, versus 93.20% and 41.12% upstream. The actual hybrid
130
+ pool below has lower coverage on this panel, despite stronger early precision
131
+ after reranking. Adding BM25 and capping the fused list at 50 can displace dense
132
+ candidates; the two pools should not be treated as interchangeable.
133
 
134
  Grades 2–3 count as useful. nDCG uses gains 0, 1, 3 and 7 against the same
135
  reviewed reference for both models, averaging all 970 queries. Rankings from
 
156
 
157
  The model was trained at 768, 512, 256 and 128 dimensions. To use a smaller
158
  width, keep the first dimensions and normalize the shortened vector again.
159
+ Use the same width for queries and passages. Each quality column below is
160
+ **v2 nDCG@10**, computed from exact rankings of cached FP32 vectors. The Daecore
161
+ column covers the same 970-query dense panel with reviewed top-10 labels at
162
+ each width, against the same **127,665-pair reference**. The 768-wide results
163
+ reproduce the corresponding full-width scores.
164
+
165
+ | Dimensions | Bytes per FP32 vector | Daecore | SciFact | FiQA | NFCorpus | SciDocs | ArguAna |
166
+ |---|---:|---:|---:|---:|---:|---:|---:|
167
+ | 768 (default) | 3,072 | 0.5574 | 0.7783 | 0.4468 | 0.3890 | 0.1852 | 0.6259 |
168
+ | 512 | 2,048 | 0.5605 | 0.7841 | 0.4374 | 0.3876 | 0.1793 | 0.6137 |
169
+ | 256 | 1,024 | 0.5292 | 0.7757 | 0.4091 | 0.3660 | 0.1708 | 0.5987 |
170
+ | 128 | 512 | 0.5035 | 0.7488 | 0.3668 | 0.3311 | 0.1484 | 0.5533 |
171
+
172
+ At 512 dimensions, each raw FP32 vector uses one-third less storage, with
173
+ slightly higher Daecore nDCG@10 in this measurement and lower scores on four
174
+ of the five public panels. At 256, vector storage falls by two-thirds with a
175
+ larger quality cost. These byte counts exclude index overhead and compression;
176
+ shorter vectors do not reduce encoder inference cost. Daecore uses 768, and
177
+ smaller widths have not been qualified through the full hybrid pipeline.
178
 
179
  ### Full pipeline: Gemma v2 + BM25 + Ettin
180
 
181
+ This replay measures semantic and lexical search, reciprocal-rank fusion, then
182
+ Ettin reranking over **970 development queries and 82,719 passages**. It includes
183
+ queries with no known useful evidence. Both model cards share this CUDA FP16
184
+ table; it is not either model's standalone score.
185
+
186
+ **Reference:** 970 queries · 82,719 passages · **127,665 graded query–passage
187
+ pairs** in the shared top-50 reference (September 2026). Abstentions are excluded.
188
 
189
  | Return depth | Hit | Precision | Known-positive recall | nDCG |
190
  |---|---:|---:|---:|---:|
191
+ | 1 | 72.33% | 72.33% | 2.12% | 0.6458 |
192
+ | 3 | 85.36% | 71.58% | 5.89% | 0.6374 |
193
+ | 5 | 88.45% | 70.28% | 9.35% | 0.6326 |
194
+ | 10 | 89.90% | 67.18% | 17.04% | 0.6315 |
195
+ | 20 | 91.03% | 61.79% | 29.72% | 0.6417 |
196
+ | 50 | 93.09% | 40.43% | 45.98% | 0.6051 |
197
+
198
+ **The hybrid candidate ceiling is 93.09% Hit@50 and 45.98% known-positive
199
+ recall@50.** These are coverage measures for the actual 50-candidate pool;
200
+ reranking cannot change them. The row at 50 exposes that pool, while Daecore
201
+ returns a selected prefix of 3–20. Fixed-depth scores do not evaluate the
202
+ selector's choice.
203
+
204
+ Grades 2–3 count as useful. At depth 1, both providers use the same 965 queries;
205
+ five are omitted because a first result was ungradable. Depths 3–50 retain all
206
+ 970. Precision pools judged positions; recall averages the 925 queries with
207
+ known useful evidence (920 at depth 1). Abstentions remove 52 positions at
208
+ rank 20 and 93 at rank 50 per provider, without drawing in deeper passages.
209
+ The corpus is not exhaustively labeled.
210
+
211
+ Dense retrieval and this replay use the same expanded relevance reference.
212
+ Earlier Hit and precision are unchanged; recall and nDCG were recomputed as
213
+ more relevant passages became known. The [evaluation companion](evaluation/README.md)
214
+ provides the records, exclusions and CUDA/Vulkan comparison; the
215
+ [pipeline overview](https://huggingface.co/Daecore) explains retrieval and selection.
216
 
217
  ## Training data and objective
218
 
219
+ Private supervision adapts Gemma to **evidence retrieval from document chunks**.
220
+ Several passages can be useful for one question, including partial evidence.
221
+
222
+ | Training ingredient | Purpose |
223
+ |---|---|
224
+ | Mostly synthetic organizational documents, plus one person's largely AI-written project docs | Exercise project notes, procedures, decisions and technical material |
225
+ | Frozen structural chunks with section context and table/code handling | Train on the passage units the retrieval workflow consumes |
226
+ | Queries written by four models from three provider families, without a designated answer | Ask for evidence without reducing every question to one target passage |
227
+ | Model-judged relevance grades | Distinguish decisive, partial, related and irrelevant evidence |
228
 
229
+ The [chunking guide](evaluation/chunking.md#retrieval-model-data) explains
230
+ the source data and methods. The corpus retains earlier parser outputs;
231
+ Gemma applies its own tokenizer and limits after chunking. These shared
232
+ generation and judging processes limit generalization beyond the tested data.
 
 
 
233
 
234
  Public supervision uses existing evidence annotations from [HotpotQA](https://hotpotqa.github.io/), [MultiDoc2Dial](https://github.com/doc2dial/multidoc2dial) and [FinQA](https://github.com/czyssrs/FinQA): multi-document questions, document-grounded dialogue, and textual or table evidence for numerical questions. The model learns to retrieve that evidence, not to perform FinQA's calculations; FinQA is unrelated to the FiQA evaluation above.
235
 
 
273
 
274
  Checks used Windows x64 with an NVIDIA RTX 3060 Ti (8 GB); other GPUs and Linux are untested, even where Vulkan is available.
275
 
276
+ The task-specific data is mostly generated and model-judged, with no human reference panel, so it describes this development distribution rather than other users' documents. The dense and full-pipeline panels share queries and a pooled relevance reference, but evaluate different retrieval paths. Some questions omit the project or task context needed for an unambiguous judgment. The 970-query pipeline panel uses natural-language questions and reports no separate terse-query or paraphrase robustness scores. Passage relevance does not measure distinct-fact coverage or downstream agent success.
277
 
278
  ## License
279
 
evaluation/README.md CHANGED
@@ -2,10 +2,14 @@
2
 
3
  These records support the [model card's](../README.md) results: Gemma's dense retrieval on private and public data, its smaller embedding widths, and the full retrieval pipeline. They contain scores and anonymized relevance grades, without private queries or document text.
4
 
 
 
 
 
5
  | File | Contents |
6
  |---|---|
7
- | `retriever.json` | Upstream and v2 reviewed top-20 dense rankings for all 970 queries over 82,719 passages |
8
- | `dimensions.json` | V2 per-query public nDCG@10 at 768, 512, 256 and 128 dimensions |
9
  | `promotion.json` | Ranked grades for the predecessor-to-v2 pipeline comparison and public nDCG@10 for upstream and v2 |
10
  | `serving.json` | The 970-query Gemma v2 + BM25 + Ettin replay used by both model cards, including CUDA/Vulkan results |
11
  | `*-summary.json` | Summaries that `metrics.py` reproduces |
@@ -32,59 +36,81 @@ prefixes and 128-query/1,024-passage token limits, and 768 dimensions. Exact
32
  cosine rankings take the top 50 for each frozen query form, then merge them
33
  with reciprocal-rank fusion (constant 60). Identical text is deduplicated;
34
  there is no parent-document filter, BM25 or reranker. The record includes
35
- model identities, source hashes, top-20 grades and reference grade counts.
36
  This is a reused development panel, not a fresh holdout.
37
 
38
- Completing the comparison required 11,480 new pair decisions. Two independent
39
- GPT-6 Sol judges graded each pair, and a fresh blind judgment resolved 3,207
40
- ordinal disagreements. The primary assistant read 179 complete query–passage
41
- pairs: all 18 final quote flags, all 72 abstentions, and samples spanning every
42
- primary grade combination. The 18 quote repairs changed no grades. The result
43
- adds 11,408 grades to a shared reference containing 87,434 known pairs. No
44
- person reviewed these judgments, and the corpus is not exhaustively labeled.
45
-
46
- Every top-20 position is graded or explicitly excluded as ungradable; no
47
- missing-label bounds remain. Cutoffs 1, 3, 5, 10 and 20 retain all 970 queries.
48
- At rank 20, the upstream lists exclude 56 positions and v2 excludes 53.
49
- Cut first, drop abstentions second, and never backfill. A wholly abstained
50
- prefix would omit that query from both models at that cutoff; none occurs in
51
- this comparison. Precision pools retained positions. Recall averages the 911
52
- queries with known useful passages. nDCG uses gains 0, 1, 3 and 7 against the
53
- same reference for both models, reindexes retained positions, and averages
54
- all 970 queries; an ideal gain of zero contributes zero. The dense reference
55
- is larger than the pipeline reference below, so compare models within each
56
- table rather than treating their nDCG values as one common scale.
 
 
 
 
 
 
 
57
 
58
  **Public data.**
59
  The model card compares v2 with upstream EmbeddingGemma on SciFact, FiQA, NFCorpus, SciDocs and ArguAna. `promotion.json` includes per-query nDCG@10 for all five datasets, using full corpora, matched preprocessing and official relevance judgments. V2 remains below upstream on every panel. The previous fine-tune was measured only on FiQA (0.4009) and SciFact (0.7679); its other three results are absent, not estimated.
60
 
61
  These datasets supplied no training examples, but informed development. They are not untouched tests.
62
 
63
- **Matryoshka widths.** `dimensions.json` contains per-query scores from new exact
64
- cosine rankings computed from the selected checkpoint's cached FP32 public vectors. For each
65
- smaller width, keep the first dimensions and L2-normalize query and passage
66
- vectors again. Self-ID matches are excluded; ties preserve corpus order.
67
- Official judgments use linear-gain nDCG, as in the original public scorer.
68
- All 1,406 ArguAna queries remain, including the five whose positives are absent
69
- from its corpus. The 768-wide per-query scores reproduce the original results
70
- exactly. Smaller widths were not evaluated on the private or full-pipeline
71
- panels and do not shorten the encoder's forward pass.
 
 
 
 
 
72
 
73
  ## Effect in the retrieval pipeline
74
 
75
- The identical full-pipeline table on both model cards comes from `serving.json`:
76
- Gemma v2 + BM25 + fusion + unchanged Ettin on all 970 queries, using CUDA FP16
77
- for reranking. At depths 3–20, Hit and nDCG average all queries, and
78
- known-positive recall averages the 895 with judged useful evidence. Depth 1
79
- uses 965 shared judged queries, including 890 with known positives; five
80
- queries are omitted from both providers because a first result was ungradable.
81
- Precision pools retained positions. The serving reference adds 32 grades to the earlier promotion
82
- reference. Its shared judging-exclusion set removes 52 top-20 positions per
83
- provider without backfilling. Fixed cutoffs measure ranking independently of
84
- the selector's choice of a 3–20-item prefix. Vulkan matches CUDA at Hit@5,
85
- Hit@10 and Hit@20; Hit@3 differs by one query and nDCG@10 by 0.0004.
86
-
87
- The separate predecessor comparison below comes from `promotion.json`.
 
 
 
 
 
 
 
 
 
 
88
  Only the embedding model changed in this comparison. BM25, reciprocal-rank fusion and the [Ettin reranker](https://huggingface.co/Daecore/ettin-150m-memory-reranker-ft-v1) stayed fixed. The same score mapping selected returned prefixes in both arms. These results measure the whole pipeline on 970 reused queries over 82,719 passage texts, rather than Gemma alone.
89
 
90
  | Metric | Previous Gemma | Gemma v2 |
@@ -117,4 +143,4 @@ retrieval method, not the query format.
117
 
118
  `serving-qualification.json` records this model's CPU, CUDA and Vulkan checks. The 74-vector panel covers token limits, mixed lengths and concurrent query and passage calls. It is a bounded serving check, separate from retrieval quality.
119
 
120
- The private documents and labels are mostly model-generated. Project-document sources come from one person's workspaces and are often AI-written. The panel was reused during development, training was run once, and there is no human reference panel. These results do not establish generalization across users, unique-fact coverage or downstream agent success.
 
2
 
3
  These records support the [model card's](../README.md) results: Gemma's dense retrieval on private and public data, its smaller embedding widths, and the full retrieval pipeline. They contain scores and anonymized relevance grades, without private queries or document text.
4
 
5
+ The [chunking guide](chunking.md#retrieval-model-data) explains the synthetic
6
+ source documents, passage preparation and relevance supervision behind the
7
+ private evaluations.
8
+
9
  | File | Contents |
10
  |---|---|
11
+ | `retriever.json` | Upstream and v2 reviewed top-50 dense rankings for all 970 queries over 82,719 passages |
12
+ | `dimensions.json` | V2 public nDCG@10 and private top-10 graded rankings at four embedding widths |
13
  | `promotion.json` | Ranked grades for the predecessor-to-v2 pipeline comparison and public nDCG@10 for upstream and v2 |
14
  | `serving.json` | The 970-query Gemma v2 + BM25 + Ettin replay used by both model cards, including CUDA/Vulkan results |
15
  | `*-summary.json` | Summaries that `metrics.py` reproduces |
 
36
  cosine rankings take the top 50 for each frozen query form, then merge them
37
  with reciprocal-rank fusion (constant 60). Identical text is deduplicated;
38
  there is no parent-document filter, BM25 or reranker. The record includes
39
+ model identities, source hashes, top-50 grades and reference grade counts.
40
  This is a reused development panel, not a fresh holdout.
41
 
42
+ The shared reference contains **127,665 graded query–passage pairs**. The top-50
43
+ extension added 40,231 grades and 222 abstentions from 40,453 previously
44
+ unjudged pairs, covering both dense models, the hybrid pool and smaller-width
45
+ top-10 results. Two independent GPT-6.1 Sol judges graded each pair; a fresh
46
+ blind adjudicator resolved 7,765 ordinal disagreements. A semantic abstention
47
+ from either primary judge remained excluded. The primary assistant read all
48
+ 349 selected pairs: every abstention and final quotation flag, plus a
49
+ stratified sample. Sixteen quotation repairs changed no grades. No person
50
+ reviewed the labels.
51
+
52
+ Every top-50 position is graded or explicitly excluded. All 970 queries remain
53
+ eligible at 1, 3, 5, 10, 20 and 50. Upstream excludes 164 positions at rank 50;
54
+ v2 excludes 167. Cut the original prefix first, drop abstentions second, and
55
+ never backfill. A wholly abstained prefix would omit the query from both arms
56
+ at that depth. Precision pools retained positions. Recall averages the 925
57
+ queries with known useful passages; these are not exhaustive corpus labels.
58
+ nDCG uses gains 0, 1, 3 and 7, reindexes retained positions and averages all
59
+ 970 queries; an ideal gain of zero contributes zero.
60
+
61
+ **Reference expansion, not model change.** The existing top-20 grades and
62
+ rankings are unchanged, so Hit and precision are unchanged. New relevant
63
+ passages enlarge recall's denominator and can strengthen the ideal nDCG
64
+ ranking. Dense nDCG@10 is now 0.4363 upstream and 0.5574 for v2, versus
65
+ 0.4506 and 0.5748 on the earlier 87,434-pair reference. Both models are recomputed together.
66
+ Dense, smaller-width and current pipeline results now share this reference;
67
+ the historical predecessor comparison below keeps its original reference.
68
 
69
  **Public data.**
70
  The model card compares v2 with upstream EmbeddingGemma on SciFact, FiQA, NFCorpus, SciDocs and ArguAna. `promotion.json` includes per-query nDCG@10 for all five datasets, using full corpora, matched preprocessing and official relevance judgments. V2 remains below upstream on every panel. The previous fine-tune was measured only on FiQA (0.4009) and SciFact (0.7679); its other three results are absent, not estimated.
71
 
72
  These datasets supplied no training examples, but informed development. They are not untouched tests.
73
 
74
+ **Matryoshka widths.** `dimensions.json` contains public per-query nDCG@10
75
+ and a `private` record with graded top-10 rankings at 768, 512, 256 and 128
76
+ dimensions. Vectors come from the selected checkpoint's cached FP32 outputs;
77
+ smaller widths retain the first dimensions and are L2-normalized again.
78
+
79
+ Public scoring uses exact cosine rankings, self-ID exclusion, stable corpus
80
+ order for ties and official linear-gain nDCG. All 1,406 ArguAna queries remain,
81
+ including five whose positives are absent from the corpus. These public
82
+ scores are unchanged. Private scoring uses the same frozen query forms,
83
+ dense-only fusion and expanded relevance reference as the 768-wide comparison.
84
+ All 970 queries remain eligible at top 10; excluded positions are 22, 26, 16
85
+ and 37 respectively. The 768-wide results reproduce both references exactly.
86
+ Smaller widths have not been qualified through the full pipeline and do not
87
+ shorten the encoder's forward pass.
88
 
89
  ## Effect in the retrieval pipeline
90
 
91
+ The identical table on both cards comes from `serving.json`: Gemma v2 +
92
+ BM25 + fusion + unchanged Ettin, on all 970 queries. CUDA FP16 and Vulkan FP32
93
+ use the same 50 candidates. At 50, Hit is 93.09% and known-positive recall is
94
+ 45.98%; those candidate-coverage measures cannot change through reranking.
95
+ Dense v2 reaches 94.43% and 57.92% on this panel. Fusion with a fixed 50-slot
96
+ budget can displace dense candidates even while reranking improves early
97
+ precision.
98
+
99
+ At depths 3–50, Hit and nDCG average all 970 queries; recall averages the 925
100
+ with known useful evidence. Depth 1 uses 965 shared judged queries, including
101
+ 920 with known positives; five are omitted because a first result was
102
+ ungradable. Precision pools retained positions. Each provider excludes 52
103
+ positions at rank 20 and 93 at rank 50 without backfill. Fixed cutoffs are
104
+ separate from the selector's choice of a 3–20-item prefix. Vulkan matches
105
+ CUDA at Hit@5, Hit@10, Hit@20 and Hit@50; Hit@3 differs by one query and
106
+ nDCG@10 by 0.0003.
107
+
108
+ ### Historical predecessor comparison
109
+
110
+ This table comes from `promotion.json`, with its **original 75,488-pair relevance
111
+ reference**. Its nDCG is not directly comparable with the expanded-reference
112
+ table on the main card.
113
+
114
  Only the embedding model changed in this comparison. BM25, reciprocal-rank fusion and the [Ettin reranker](https://huggingface.co/Daecore/ettin-150m-memory-reranker-ft-v1) stayed fixed. The same score mapping selected returned prefixes in both arms. These results measure the whole pipeline on 970 reused queries over 82,719 passage texts, rather than Gemma alone.
115
 
116
  | Metric | Previous Gemma | Gemma v2 |
 
143
 
144
  `serving-qualification.json` records this model's CPU, CUDA and Vulkan checks. The 74-vector panel covers token limits, mixed lengths and concurrent query and passage calls. It is a bounded serving check, separate from retrieval quality.
145
 
146
+ The private documents and labels are mostly model-generated. Project-document sources come from one person's workspaces and are often AI-written. Some queries omit the project or subject needed for a clear relevance judgment. The panel was reused during development, training was run once, and there is no human reference panel. These results do not establish generalization across users, unique-fact coverage or downstream agent success.
evaluation/chunking.md ADDED
@@ -0,0 +1,58 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Chunking and retrieval data
2
+
3
+ ## Retrieval-model data
4
+
5
+ Daecore fine-tunes retrieval models on passages prepared for its document
6
+ workflow. A passage is the exact chunk the model reads: a section, excerpt,
7
+ table fragment or code block. It can supply useful partial evidence without
8
+ answering the whole question.
9
+
10
+ **Source documents → structural chunks → graded query–passage pairs → retrieval training**
11
+
12
+ | Ingredient | What it contributes |
13
+ |---|---|
14
+ | Mostly synthetic workplace documentation | Notes, procedures, decisions and technical material across generated project scenarios |
15
+ | One person's project documents, much of them AI-written | Additional examples, not a representative sample of other users' workspaces |
16
+ | Frozen structural chunks | The retrieval units that Gemma embeds and Ettin ranks; some retain earlier parser outputs |
17
+ | Model-written questions and graded evidence | Several useful passages per question, with decisive and partial evidence distinguished from merely related text |
18
+ | Public evidence annotations used by Gemma v2 | HotpotQA paragraphs, MultiDoc2Dial grounding passages and FinQA text/table evidence |
19
+
20
+ Public retrieval data also comes in passages; it is not uniformly made of full
21
+ documents. The distinction is their source and preparation. Each model applies
22
+ its tokenizer and input limits after chunking, so a chunk boundary and a model's
23
+ truncation limit are separate constraints.
24
+
25
+ This is deliberate specialization of already capable upstream models. The
26
+ private comparisons measure gains on this workflow, while public benchmarks
27
+ show the general-retrieval cost. Daecore accepted that tradeoff for its intended
28
+ use; the cards report both sides. Generated questions and shared preparation
29
+ and judging processes do not establish gains for every real user or agent.
30
+
31
+ Training and evaluation use frozen passage snapshots. Current chunking fixes
32
+ do not retroactively change those texts or transfer their relevance labels to
33
+ different cuts.
34
+
35
+ ## Chunking methods
36
+
37
+ - **Heading-based splitting:** use the document's sections; oversized sections
38
+ can split further at their subheadings. This follows the author's structure.
39
+ - **Title and heading context:** add the document title and relevant heading
40
+ path to section chunks so their subject stays clear. Smaller searchable pieces
41
+ retain that metadata but do not automatically repeat it inside every piece.
42
+ - **Boundary-aware cuts and overlap:** prefer a nearby sentence end, paragraph
43
+ break or line break. Ordinary prose windows repeat up to 200 characters across
44
+ cuts to carry nearby context; oversized structural continuations have their
45
+ own framing rules.
46
+ - **Code-block and table handling:** keep blocks together when they fit. Split
47
+ oversized tables with repeated column headers and oversized code blocks with
48
+ reopened and closed code fences, so continuations remain readable.
49
+ - **Record-aware splitting:** preserve row, field or entry context for CSV,
50
+ JSON, JSONL and dated entries. Oversized records are divided with identifying
51
+ labels rather than losing their remaining content.
52
+ - **Paragraph packing for imported memory files:** group whole paragraphs into
53
+ bounded pieces so excerpts usually begin at a paragraph. A single oversized
54
+ paragraph falls back to structural splitting.
55
+
56
+ These methods work together: make supported text and record chunks fit first,
57
+ then create bounded searchable pieces for sections that are still too large.
58
+ Both operations use the same structural cutter.
evaluation/dimensions-summary.json CHANGED
@@ -1,56 +1,310 @@
1
- {
2
- "queries": 3677,
3
- "vector_bytes_fp32": {
4
- "768": 3072,
5
- "512": 2048,
6
- "256": 1024,
7
- "128": 512
8
- },
9
- "public": {
10
- "arguana": {
11
- "queries": 1406,
12
- "ndcg@10": {
13
- "768": 0.6258634411077135,
14
- "512": 0.613697054466976,
15
- "256": 0.5986725670753965,
16
- "128": 0.5533073210679951
17
- }
18
- },
19
- "fiqa": {
20
- "queries": 648,
21
- "ndcg@10": {
22
- "768": 0.446838124778398,
23
- "512": 0.43740349056283273,
24
- "256": 0.40912280098290843,
25
- "128": 0.3668286877036632
26
- }
27
- },
28
- "nfcorpus": {
29
- "queries": 323,
30
- "ndcg@10": {
31
- "768": 0.3889646280478944,
32
- "512": 0.3875786151617704,
33
- "256": 0.3660320811396263,
34
- "128": 0.3310697612326329
35
- }
36
- },
37
- "scidocs": {
38
- "queries": 1000,
39
- "ndcg@10": {
40
- "768": 0.18521860884727676,
41
- "512": 0.179327587641947,
42
- "256": 0.17081239437849563,
43
- "128": 0.1484461830477022
44
- }
45
- },
46
- "scifact": {
47
- "queries": 300,
48
- "ndcg@10": {
49
- "768": 0.7782845713792622,
50
- "512": 0.7840710444235555,
51
- "256": 0.7756812272647337,
52
- "128": 0.7488095898261483
53
- }
54
- }
55
- }
56
- }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "queries": 3677,
3
+ "vector_bytes_fp32": {
4
+ "768": 3072,
5
+ "512": 2048,
6
+ "256": 1024,
7
+ "128": 512
8
+ },
9
+ "public": {
10
+ "arguana": {
11
+ "queries": 1406,
12
+ "ndcg@10": {
13
+ "768": 0.6258634411077135,
14
+ "512": 0.613697054466976,
15
+ "256": 0.5986725670753965,
16
+ "128": 0.5533073210679951
17
+ }
18
+ },
19
+ "fiqa": {
20
+ "queries": 648,
21
+ "ndcg@10": {
22
+ "768": 0.446838124778398,
23
+ "512": 0.43740349056283273,
24
+ "256": 0.40912280098290843,
25
+ "128": 0.3668286877036632
26
+ }
27
+ },
28
+ "nfcorpus": {
29
+ "queries": 323,
30
+ "ndcg@10": {
31
+ "768": 0.3889646280478944,
32
+ "512": 0.3875786151617704,
33
+ "256": 0.3660320811396263,
34
+ "128": 0.3310697612326329
35
+ }
36
+ },
37
+ "scidocs": {
38
+ "queries": 1000,
39
+ "ndcg@10": {
40
+ "768": 0.18521860884727676,
41
+ "512": 0.179327587641947,
42
+ "256": 0.17081239437849563,
43
+ "128": 0.1484461830477022
44
+ }
45
+ },
46
+ "scifact": {
47
+ "queries": 300,
48
+ "ndcg@10": {
49
+ "768": 0.7782845713792622,
50
+ "512": 0.7840710444235555,
51
+ "256": 0.7756812272647337,
52
+ "128": 0.7488095898261483
53
+ }
54
+ }
55
+ },
56
+ "private": {
57
+ "queries": 970,
58
+ "models": {
59
+ "768": {
60
+ "1": {
61
+ "scored_queries": 964,
62
+ "excluded_queries": 6,
63
+ "hit": 0.6452282157676349,
64
+ "precision": 0.6452282157676349,
65
+ "macro_precision": 0.6452282157676349,
66
+ "ndcg": 0.5301323849041691,
67
+ "known_positive_recall": 0.01788016629109334,
68
+ "recall_queries": 919,
69
+ "useful": 622,
70
+ "retained": 964,
71
+ "excluded_positions": 0,
72
+ "mean_useful": 0.6452282157676349,
73
+ "mean_retained": 1.0
74
+ },
75
+ "3": {
76
+ "scored_queries": 970,
77
+ "excluded_queries": 0,
78
+ "hit": 0.7938144329896907,
79
+ "precision": 0.6308009625300791,
80
+ "macro_precision": 0.6309278350515464,
81
+ "ndcg": 0.5406982958852574,
82
+ "known_positive_recall": 0.05034823759106709,
83
+ "recall_queries": 925,
84
+ "useful": 1835,
85
+ "retained": 2909,
86
+ "excluded_positions": 1,
87
+ "mean_useful": 1.8917525773195876,
88
+ "mean_retained": 2.9989690721649485
89
+ },
90
+ "5": {
91
+ "scored_queries": 970,
92
+ "excluded_queries": 0,
93
+ "hit": 0.8474226804123711,
94
+ "precision": 0.6266005782734407,
95
+ "macro_precision": 0.6269759450171821,
96
+ "ndcg": 0.5453498298950709,
97
+ "known_positive_recall": 0.08055790386567772,
98
+ "recall_queries": 925,
99
+ "useful": 3034,
100
+ "retained": 4842,
101
+ "excluded_positions": 8,
102
+ "mean_useful": 3.1278350515463917,
103
+ "mean_retained": 4.991752577319588
104
+ },
105
+ "10": {
106
+ "scored_queries": 970,
107
+ "excluded_queries": 0,
108
+ "hit": 0.8814432989690721,
109
+ "precision": 0.6117999586691465,
110
+ "macro_precision": 0.6123367697594502,
111
+ "ndcg": 0.5574064002423099,
112
+ "known_positive_recall": 0.15074270958466607,
113
+ "recall_queries": 925,
114
+ "useful": 5921,
115
+ "retained": 9678,
116
+ "excluded_positions": 22,
117
+ "mean_useful": 6.104123711340206,
118
+ "mean_retained": 9.977319587628866
119
+ }
120
+ },
121
+ "512": {
122
+ "1": {
123
+ "scored_queries": 964,
124
+ "excluded_queries": 6,
125
+ "hit": 0.6462655601659751,
126
+ "precision": 0.6462655601659751,
127
+ "macro_precision": 0.6462655601659751,
128
+ "ndcg": 0.5478166370282552,
129
+ "known_positive_recall": 0.01775688496570127,
130
+ "recall_queries": 919,
131
+ "useful": 623,
132
+ "retained": 964,
133
+ "excluded_positions": 0,
134
+ "mean_useful": 0.6462655601659751,
135
+ "mean_retained": 1.0
136
+ },
137
+ "3": {
138
+ "scored_queries": 970,
139
+ "excluded_queries": 0,
140
+ "hit": 0.7969072164948454,
141
+ "precision": 0.635237439779766,
142
+ "macro_precision": 0.6355670103092783,
143
+ "ndcg": 0.5429708859076539,
144
+ "known_positive_recall": 0.05057307077004017,
145
+ "recall_queries": 925,
146
+ "useful": 1846,
147
+ "retained": 2906,
148
+ "excluded_positions": 4,
149
+ "mean_useful": 1.9030927835051545,
150
+ "mean_retained": 2.995876288659794
151
+ },
152
+ "5": {
153
+ "scored_queries": 970,
154
+ "excluded_queries": 0,
155
+ "hit": 0.8422680412371134,
156
+ "precision": 0.6308613922743235,
157
+ "macro_precision": 0.63106529209622,
158
+ "ndcg": 0.5504066702414289,
159
+ "known_positive_recall": 0.08121602068063129,
160
+ "recall_queries": 925,
161
+ "useful": 3054,
162
+ "retained": 4841,
163
+ "excluded_positions": 9,
164
+ "mean_useful": 3.1484536082474226,
165
+ "mean_retained": 4.990721649484536
166
+ },
167
+ "10": {
168
+ "scored_queries": 970,
169
+ "excluded_queries": 0,
170
+ "hit": 0.8824742268041237,
171
+ "precision": 0.6132933636551582,
172
+ "macro_precision": 0.613993618065783,
173
+ "ndcg": 0.5604826042172213,
174
+ "known_positive_recall": 0.15004173661513098,
175
+ "recall_queries": 925,
176
+ "useful": 5933,
177
+ "retained": 9674,
178
+ "excluded_positions": 26,
179
+ "mean_useful": 6.116494845360824,
180
+ "mean_retained": 9.97319587628866
181
+ }
182
+ },
183
+ "256": {
184
+ "1": {
185
+ "scored_queries": 964,
186
+ "excluded_queries": 6,
187
+ "hit": 0.6037344398340249,
188
+ "precision": 0.6037344398340249,
189
+ "macro_precision": 0.6037344398340249,
190
+ "ndcg": 0.5127445168938944,
191
+ "known_positive_recall": 0.015976956236435802,
192
+ "recall_queries": 919,
193
+ "useful": 582,
194
+ "retained": 964,
195
+ "excluded_positions": 0,
196
+ "mean_useful": 0.6037344398340249,
197
+ "mean_retained": 1.0
198
+ },
199
+ "3": {
200
+ "scored_queries": 970,
201
+ "excluded_queries": 0,
202
+ "hit": 0.7608247422680412,
203
+ "precision": 0.5955249569707401,
204
+ "macro_precision": 0.5962199312714777,
205
+ "ndcg": 0.5111829833045054,
206
+ "known_positive_recall": 0.046233615834183193,
207
+ "recall_queries": 925,
208
+ "useful": 1730,
209
+ "retained": 2905,
210
+ "excluded_positions": 5,
211
+ "mean_useful": 1.7835051546391754,
212
+ "mean_retained": 2.9948453608247423
213
+ },
214
+ "5": {
215
+ "scored_queries": 970,
216
+ "excluded_queries": 0,
217
+ "hit": 0.8216494845360824,
218
+ "precision": 0.5925696594427244,
219
+ "macro_precision": 0.5929381443298969,
220
+ "ndcg": 0.5174515106874455,
221
+ "known_positive_recall": 0.0744111972114428,
222
+ "recall_queries": 925,
223
+ "useful": 2871,
224
+ "retained": 4845,
225
+ "excluded_positions": 5,
226
+ "mean_useful": 2.95979381443299,
227
+ "mean_retained": 4.994845360824742
228
+ },
229
+ "10": {
230
+ "scored_queries": 970,
231
+ "excluded_queries": 0,
232
+ "hit": 0.8731958762886598,
233
+ "precision": 0.5818876497315159,
234
+ "macro_precision": 0.5823911798396334,
235
+ "ndcg": 0.5291576970298347,
236
+ "known_positive_recall": 0.1419199356602385,
237
+ "recall_queries": 925,
238
+ "useful": 5635,
239
+ "retained": 9684,
240
+ "excluded_positions": 16,
241
+ "mean_useful": 5.809278350515464,
242
+ "mean_retained": 9.983505154639175
243
+ }
244
+ },
245
+ "128": {
246
+ "1": {
247
+ "scored_queries": 964,
248
+ "excluded_queries": 6,
249
+ "hit": 0.5726141078838174,
250
+ "precision": 0.5726141078838174,
251
+ "macro_precision": 0.5726141078838174,
252
+ "ndcg": 0.47288085358624776,
253
+ "known_positive_recall": 0.015019269226427342,
254
+ "recall_queries": 919,
255
+ "useful": 552,
256
+ "retained": 964,
257
+ "excluded_positions": 0,
258
+ "mean_useful": 0.5726141078838174,
259
+ "mean_retained": 1.0
260
+ },
261
+ "3": {
262
+ "scored_queries": 970,
263
+ "excluded_queries": 0,
264
+ "hit": 0.7402061855670103,
265
+ "precision": 0.5779310344827586,
266
+ "macro_precision": 0.5780068728522336,
267
+ "ndcg": 0.4818931326910878,
268
+ "known_positive_recall": 0.044531751738185646,
269
+ "recall_queries": 925,
270
+ "useful": 1676,
271
+ "retained": 2900,
272
+ "excluded_positions": 10,
273
+ "mean_useful": 1.7278350515463918,
274
+ "mean_retained": 2.9896907216494846
275
+ },
276
+ "5": {
277
+ "scored_queries": 970,
278
+ "excluded_queries": 0,
279
+ "hit": 0.8041237113402062,
280
+ "precision": 0.574591351127664,
281
+ "macro_precision": 0.5746048109965636,
282
+ "ndcg": 0.48751134015621145,
283
+ "known_positive_recall": 0.07159737689223461,
284
+ "recall_queries": 925,
285
+ "useful": 2777,
286
+ "retained": 4833,
287
+ "excluded_positions": 17,
288
+ "mean_useful": 2.8628865979381444,
289
+ "mean_retained": 4.982474226804124
290
+ },
291
+ "10": {
292
+ "scored_queries": 970,
293
+ "excluded_queries": 0,
294
+ "hit": 0.8649484536082475,
295
+ "precision": 0.5659733002173238,
296
+ "macro_precision": 0.5665320733104239,
297
+ "ndcg": 0.5035498696905354,
298
+ "known_positive_recall": 0.13623796718702827,
299
+ "recall_queries": 925,
300
+ "useful": 5469,
301
+ "retained": 9663,
302
+ "excluded_positions": 37,
303
+ "mean_useful": 5.638144329896908,
304
+ "mean_retained": 9.961855670103093
305
+ }
306
+ }
307
+ },
308
+ "corpus_passages": 82719
309
+ }
310
+ }
evaluation/dimensions.json CHANGED
The diff for this file is too large to render. See raw diff
 
evaluation/figures.py CHANGED
@@ -63,7 +63,7 @@ def reranker(summary: dict) -> str:
63
 
64
 
65
  def retriever(summary: dict) -> str:
66
- cutoffs = [1, 3, 5, 10, 20]
67
  x = list(range(len(cutoffs)))
68
  panels = []
69
  legend = [('Upstream Gemma', UPSTREAM, None), ('Daecore v2', FIT, None)]
 
63
 
64
 
65
  def retriever(summary: dict) -> str:
66
+ cutoffs = sorted(map(int, summary['models']['upstream']))
67
  x = list(range(len(cutoffs)))
68
  panels = []
69
  legend = [('Upstream Gemma', UPSTREAM, None), ('Daecore v2', FIT, None)]
evaluation/metrics.py CHANGED
@@ -143,7 +143,13 @@ def summarize_dimensions(data: dict) -> dict:
143
  panels[name] = {'queries': len(panel), 'ndcg@10': {
144
  str(w): mean(row['ndcg@10'][str(w)] for row in panel) for w in widths
145
  }}
146
- return {'queries': len(rows), 'vector_bytes_fp32': {str(w): w * 4 for w in widths}, 'public': panels}
 
 
 
 
 
 
147
 
148
 
149
  def _summarize_ranked_grades(data: dict, *, include_selected: bool) -> dict:
@@ -155,6 +161,12 @@ def _summarize_ranked_grades(data: dict, *, include_selected: bool) -> dict:
155
  rows = data['rows']
156
  if not rows or len({row['id'] for row in rows}) != len(rows):
157
  raise ValueError('Ranked rows require unique nonempty query identities')
 
 
 
 
 
 
158
  # Validate before excluding fully abstained prefixes. An omitted query
159
  # must never conceal a malformed grade, model map or selected depth.
160
  for row in rows:
@@ -163,9 +175,9 @@ def _summarize_ranked_grades(data: dict, *, include_selected: bool) -> dict:
163
  raise ValueError('Reference grade counts must cover grades zero through three')
164
  for model in data['model_order']:
165
  ranked, excluded = row['ranked_grades'][model], row['excluded'][model]
166
- if (len(ranked) != 20 or len(excluded) != 20
167
  or any(type(x) is not bool for x in excluded)):
168
- raise ValueError('Ranked rows require a bounded original top twenty')
169
  if include_selected:
170
  depth = row['selected_depth'][model]
171
  if type(depth) is not int or not 3 <= depth <= 20:
@@ -174,7 +186,9 @@ def _summarize_ranked_grades(data: dict, *, include_selected: bool) -> dict:
174
  for grade, drop in zip(ranked, excluded, strict=True)):
175
  raise ValueError('Only declared abstentions may lack grades')
176
  output = {'queries': len(rows), 'models': {}}
177
- requested_cutoffs = (1, 3, 5, 10, 20, 'selected') if include_selected else (1, 3, 5, 10, 20)
 
 
178
  for model in data['model_order']:
179
  cutoffs = {}
180
  for cutoff in requested_cutoffs:
 
143
  panels[name] = {'queries': len(panel), 'ndcg@10': {
144
  str(w): mean(row['ndcg@10'][str(w)] for row in panel) for w in widths
145
  }}
146
+ result = {'queries': len(rows), 'vector_bytes_fp32': {str(w): w * 4 for w in widths}, 'public': panels}
147
+ if 'private' in data:
148
+ private = data['private']
149
+ if private['model_order'] != [str(w) for w in widths]:
150
+ raise ValueError('Private width rankings must cover the same ordered widths')
151
+ result['private'] = summarize_record('retriever', private)
152
+ return result
153
 
154
 
155
  def _summarize_ranked_grades(data: dict, *, include_selected: bool) -> dict:
 
161
  rows = data['rows']
162
  if not rows or len({row['id'] for row in rows}) != len(rows):
163
  raise ValueError('Ranked rows require unique nonempty query identities')
164
+ # Published promotion and serving v1 records contain twenty positions.
165
+ # Extended dense/serving records declare their measured depth explicitly.
166
+ depth_limit = data.get('ranking_depth', 20)
167
+ allowed_depths = (20, 50) if include_selected else (10, 20, 50)
168
+ if type(depth_limit) is not int or depth_limit not in allowed_depths:
169
+ raise ValueError('Unsupported measured ranking depth')
170
  # Validate before excluding fully abstained prefixes. An omitted query
171
  # must never conceal a malformed grade, model map or selected depth.
172
  for row in rows:
 
175
  raise ValueError('Reference grade counts must cover grades zero through three')
176
  for model in data['model_order']:
177
  ranked, excluded = row['ranked_grades'][model], row['excluded'][model]
178
+ if (len(ranked) != depth_limit or len(excluded) != depth_limit
179
  or any(type(x) is not bool for x in excluded)):
180
+ raise ValueError('Ranked rows must cover exactly the declared ranking depth')
181
  if include_selected:
182
  depth = row['selected_depth'][model]
183
  if type(depth) is not int or not 3 <= depth <= 20:
 
186
  for grade, drop in zip(ranked, excluded, strict=True)):
187
  raise ValueError('Only declared abstentions may lack grades')
188
  output = {'queries': len(rows), 'models': {}}
189
+ requested_cutoffs = tuple(k for k in CUTOFFS if k <= depth_limit)
190
+ if include_selected:
191
+ requested_cutoffs += ('selected',)
192
  for model in data['model_order']:
193
  cutoffs = {}
194
  for cutoff in requested_cutoffs:
evaluation/retriever-summary.json CHANGED
@@ -8,9 +8,9 @@
8
  "hit": 0.534020618556701,
9
  "precision": 0.534020618556701,
10
  "macro_precision": 0.534020618556701,
11
- "ndcg": 0.4695139911634757,
12
- "known_positive_recall": 0.019407636059998304,
13
- "recall_queries": 911,
14
  "useful": 518,
15
  "retained": 970,
16
  "excluded_positions": 0,
@@ -23,9 +23,9 @@
23
  "hit": 0.7298969072164948,
24
  "precision": 0.5082530949105915,
25
  "macro_precision": 0.5082474226804123,
26
- "ndcg": 0.4515657513806275,
27
- "known_positive_recall": 0.05434321461039996,
28
- "recall_queries": 911,
29
  "useful": 1478,
30
  "retained": 2908,
31
  "excluded_positions": 2,
@@ -38,9 +38,9 @@
38
  "hit": 0.7948453608247422,
39
  "precision": 0.49948379103861246,
40
  "macro_precision": 0.4996219931271478,
41
- "ndcg": 0.4457312507652407,
42
- "known_positive_recall": 0.08530117331466223,
43
- "recall_queries": 911,
44
  "useful": 2419,
45
  "retained": 4843,
46
  "excluded_positions": 7,
@@ -53,9 +53,9 @@
53
  "hit": 0.8484536082474227,
54
  "precision": 0.4739561802397685,
55
  "macro_precision": 0.47407543773523153,
56
- "ndcg": 0.45062269194348514,
57
- "known_positive_recall": 0.15770613185252844,
58
- "recall_queries": 911,
59
  "useful": 4586,
60
  "retained": 9676,
61
  "excluded_positions": 24,
@@ -68,14 +68,29 @@
68
  "hit": 0.8876288659793814,
69
  "precision": 0.4393610421836228,
70
  "macro_precision": 0.44005992535723476,
71
- "ndcg": 0.4736202687789316,
72
- "known_positive_recall": 0.2807541608701895,
73
- "recall_queries": 911,
74
  "useful": 8499,
75
  "retained": 19344,
76
  "excluded_positions": 56,
77
  "mean_useful": 8.761855670103094,
78
  "mean_retained": 19.942268041237114
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
79
  }
80
  },
81
  "finetuned": {
@@ -85,9 +100,9 @@
85
  "hit": 0.6474226804123712,
86
  "precision": 0.6474226804123712,
87
  "macro_precision": 0.6474226804123712,
88
- "ndcg": 0.5420716740304369,
89
- "known_positive_recall": 0.02462514980632024,
90
- "recall_queries": 911,
91
  "useful": 628,
92
  "retained": 970,
93
  "excluded_positions": 0,
@@ -100,9 +115,9 @@
100
  "hit": 0.7938144329896907,
101
  "precision": 0.6308009625300791,
102
  "macro_precision": 0.6309278350515464,
103
- "ndcg": 0.5523629493718011,
104
- "known_positive_recall": 0.06947139492907671,
105
- "recall_queries": 911,
106
  "useful": 1835,
107
  "retained": 2909,
108
  "excluded_positions": 1,
@@ -115,9 +130,9 @@
115
  "hit": 0.8474226804123711,
116
  "precision": 0.6266005782734407,
117
  "macro_precision": 0.6269759450171821,
118
- "ndcg": 0.558603868737523,
119
- "known_positive_recall": 0.11088922136775149,
120
- "recall_queries": 911,
121
  "useful": 3034,
122
  "retained": 4842,
123
  "excluded_positions": 8,
@@ -130,9 +145,9 @@
130
  "hit": 0.8814432989690721,
131
  "precision": 0.6117999586691465,
132
  "macro_precision": 0.6123367697594502,
133
- "ndcg": 0.5748078775849167,
134
- "known_positive_recall": 0.2082023156474653,
135
- "recall_queries": 911,
136
  "useful": 5921,
137
  "retained": 9678,
138
  "excluded_positions": 22,
@@ -145,14 +160,29 @@
145
  "hit": 0.9051546391752577,
146
  "precision": 0.5810720008270016,
147
  "macro_precision": 0.5815832221977094,
148
- "ndcg": 0.6064961581206146,
149
- "known_positive_recall": 0.38649540366398033,
150
- "recall_queries": 911,
151
  "useful": 11242,
152
  "retained": 19347,
153
  "excluded_positions": 53,
154
  "mean_useful": 11.589690721649484,
155
  "mean_retained": 19.945360824742266
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
156
  }
157
  }
158
  },
 
8
  "hit": 0.534020618556701,
9
  "precision": 0.534020618556701,
10
  "macro_precision": 0.534020618556701,
11
+ "ndcg": 0.4592047128129602,
12
+ "known_positive_recall": 0.014444554022554491,
13
+ "recall_queries": 925,
14
  "useful": 518,
15
  "retained": 970,
16
  "excluded_positions": 0,
 
23
  "hit": 0.7298969072164948,
24
  "precision": 0.5082530949105915,
25
  "macro_precision": 0.5082474226804123,
26
+ "ndcg": 0.4413745306026501,
27
+ "known_positive_recall": 0.03938697535513605,
28
+ "recall_queries": 925,
29
  "useful": 1478,
30
  "retained": 2908,
31
  "excluded_positions": 2,
 
38
  "hit": 0.7948453608247422,
39
  "precision": 0.49948379103861246,
40
  "macro_precision": 0.4996219931271478,
41
+ "ndcg": 0.43462885647387695,
42
+ "known_positive_recall": 0.06172226483192968,
43
+ "recall_queries": 925,
44
  "useful": 2419,
45
  "retained": 4843,
46
  "excluded_positions": 7,
 
53
  "hit": 0.8484536082474227,
54
  "precision": 0.4739561802397685,
55
  "macro_precision": 0.47407543773523153,
56
+ "ndcg": 0.43629894440575245,
57
+ "known_positive_recall": 0.11180689389376004,
58
+ "recall_queries": 925,
59
  "useful": 4586,
60
  "retained": 9676,
61
  "excluded_positions": 24,
 
68
  "hit": 0.8876288659793814,
69
  "precision": 0.4393610421836228,
70
  "macro_precision": 0.44005992535723476,
71
+ "ndcg": 0.4525661889832219,
72
+ "known_positive_recall": 0.19611550928151497,
73
+ "recall_queries": 925,
74
  "useful": 8499,
75
  "retained": 19344,
76
  "excluded_positions": 56,
77
  "mean_useful": 8.761855670103094,
78
  "mean_retained": 19.942268041237114
79
+ },
80
+ "50": {
81
+ "scored_queries": 970,
82
+ "excluded_queries": 0,
83
+ "hit": 0.931958762886598,
84
+ "precision": 0.38598560079443894,
85
+ "macro_precision": 0.3867178488702632,
86
+ "ndcg": 0.5107049932805754,
87
+ "known_positive_recall": 0.41118693689165037,
88
+ "recall_queries": 925,
89
+ "useful": 18657,
90
+ "retained": 48336,
91
+ "excluded_positions": 164,
92
+ "mean_useful": 19.234020618556702,
93
+ "mean_retained": 49.83092783505155
94
  }
95
  },
96
  "finetuned": {
 
100
  "hit": 0.6474226804123712,
101
  "precision": 0.6474226804123712,
102
  "macro_precision": 0.6474226804123712,
103
+ "ndcg": 0.5312714776632302,
104
+ "known_positive_recall": 0.017869431761799028,
105
+ "recall_queries": 925,
106
  "useful": 628,
107
  "retained": 970,
108
  "excluded_positions": 0,
 
115
  "hit": 0.7938144329896907,
116
  "precision": 0.6308009625300791,
117
  "macro_precision": 0.6309278350515464,
118
+ "ndcg": 0.5406982958852574,
119
+ "known_positive_recall": 0.05034823759106709,
120
+ "recall_queries": 925,
121
  "useful": 1835,
122
  "retained": 2909,
123
  "excluded_positions": 1,
 
130
  "hit": 0.8474226804123711,
131
  "precision": 0.6266005782734407,
132
  "macro_precision": 0.6269759450171821,
133
+ "ndcg": 0.5453498298950709,
134
+ "known_positive_recall": 0.08055790386567772,
135
+ "recall_queries": 925,
136
  "useful": 3034,
137
  "retained": 4842,
138
  "excluded_positions": 8,
 
145
  "hit": 0.8814432989690721,
146
  "precision": 0.6117999586691465,
147
  "macro_precision": 0.6123367697594502,
148
+ "ndcg": 0.5574064002423099,
149
+ "known_positive_recall": 0.15074270958466607,
150
+ "recall_queries": 925,
151
  "useful": 5921,
152
  "retained": 9678,
153
  "excluded_positions": 22,
 
160
  "hit": 0.9051546391752577,
161
  "precision": 0.5810720008270016,
162
  "macro_precision": 0.5815832221977094,
163
+ "ndcg": 0.5818739813983459,
164
+ "known_positive_recall": 0.27556154002694083,
165
+ "recall_queries": 925,
166
  "useful": 11242,
167
  "retained": 19347,
168
  "excluded_positions": 53,
169
  "mean_useful": 11.589690721649484,
170
  "mean_retained": 19.945360824742266
171
+ },
172
+ "50": {
173
+ "scored_queries": 970,
174
+ "excluded_queries": 0,
175
+ "hit": 0.9443298969072165,
176
+ "precision": 0.5045207208325575,
177
+ "macro_precision": 0.5051647365119746,
178
+ "ndcg": 0.6449716439245425,
179
+ "known_positive_recall": 0.5791516627126307,
180
+ "recall_queries": 925,
181
+ "useful": 24385,
182
+ "retained": 48333,
183
+ "excluded_positions": 167,
184
+ "mean_useful": 25.13917525773196,
185
+ "mean_retained": 49.827835051546394
186
  }
187
  }
188
  },
evaluation/retriever.json CHANGED
The diff for this file is too large to render. See raw diff
 
evaluation/serving-summary.json CHANGED
@@ -8,9 +8,9 @@
8
  "hit": 0.7233160621761658,
9
  "precision": 0.7233160621761658,
10
  "macro_precision": 0.7233160621761658,
11
- "ndcg": 0.664594127806563,
12
- "known_positive_recall": 0.0346020393358114,
13
- "recall_queries": 890,
14
  "useful": 698,
15
  "retained": 965,
16
  "excluded_positions": 0,
@@ -23,9 +23,9 @@
23
  "hit": 0.8536082474226804,
24
  "precision": 0.7157640565712314,
25
  "macro_precision": 0.7166666666666667,
26
- "ndcg": 0.6576832474683237,
27
- "known_positive_recall": 0.09521315062968948,
28
- "recall_queries": 895,
29
  "useful": 2075,
30
  "retained": 2899,
31
  "excluded_positions": 11,
@@ -38,9 +38,9 @@
38
  "hit": 0.8845360824742268,
39
  "precision": 0.7028311634635255,
40
  "macro_precision": 0.7031786941580757,
41
- "ndcg": 0.6554718901879515,
42
- "known_positive_recall": 0.15383335671735354,
43
- "recall_queries": 895,
44
  "useful": 3401,
45
  "retained": 4839,
46
  "excluded_positions": 11,
@@ -53,9 +53,9 @@
53
  "hit": 0.8989690721649485,
54
  "precision": 0.67180070291503,
55
  "macro_precision": 0.6722631320569465,
56
- "ndcg": 0.6610872123112499,
57
- "known_positive_recall": 0.2802986435290336,
58
- "recall_queries": 895,
59
  "useful": 6499,
60
  "retained": 9674,
61
  "excluded_positions": 26,
@@ -68,24 +68,39 @@
68
  "hit": 0.9103092783505154,
69
  "precision": 0.61794500723589,
70
  "macro_precision": 0.6184956360634603,
71
- "ndcg": 0.6840908385828889,
72
- "known_positive_recall": 0.4884616340614924,
73
- "recall_queries": 895,
74
  "useful": 11956,
75
  "retained": 19348,
76
  "excluded_positions": 52,
77
  "mean_useful": 12.32577319587629,
78
  "mean_retained": 19.94639175257732
79
  },
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
80
  "selected": {
81
  "scored_queries": 970,
82
  "excluded_queries": 0,
83
  "hit": 0.8618556701030928,
84
  "precision": 0.8033770583310076,
85
  "macro_precision": 0.7040325511660611,
86
- "ndcg": 0.6598955189009045,
87
- "known_positive_recall": 0.2079138769097556,
88
- "recall_queries": 895,
89
  "useful": 5757,
90
  "retained": 7166,
91
  "excluded_positions": 17,
@@ -100,9 +115,9 @@
100
  "hit": 0.7243523316062176,
101
  "precision": 0.7243523316062176,
102
  "macro_precision": 0.7243523316062176,
103
- "ndcg": 0.6648902047865778,
104
- "known_positive_recall": 0.03463608768446649,
105
- "recall_queries": 890,
106
  "useful": 699,
107
  "retained": 965,
108
  "excluded_positions": 0,
@@ -115,9 +130,9 @@
115
  "hit": 0.8525773195876288,
116
  "precision": 0.7157640565712314,
117
  "macro_precision": 0.7166666666666667,
118
- "ndcg": 0.6577263968717306,
119
- "known_positive_recall": 0.09519545071125639,
120
- "recall_queries": 895,
121
  "useful": 2075,
122
  "retained": 2899,
123
  "excluded_positions": 11,
@@ -130,9 +145,9 @@
130
  "hit": 0.8845360824742268,
131
  "precision": 0.7023563455973543,
132
  "macro_precision": 0.702766323024055,
133
- "ndcg": 0.6550340124265787,
134
- "known_positive_recall": 0.15374094654114448,
135
- "recall_queries": 895,
136
  "useful": 3398,
137
  "retained": 4838,
138
  "excluded_positions": 12,
@@ -145,9 +160,9 @@
145
  "hit": 0.8989690721649485,
146
  "precision": 0.6723514211886304,
147
  "macro_precision": 0.6727589592538046,
148
- "ndcg": 0.6615310405020679,
149
- "known_positive_recall": 0.2805978824757964,
150
- "recall_queries": 895,
151
  "useful": 6505,
152
  "retained": 9675,
153
  "excluded_positions": 25,
@@ -160,24 +175,39 @@
160
  "hit": 0.9103092783505154,
161
  "precision": 0.6178416373785404,
162
  "macro_precision": 0.6183925432799551,
163
- "ndcg": 0.6840389893720067,
164
- "known_positive_recall": 0.4883516349357555,
165
- "recall_queries": 895,
166
  "useful": 11954,
167
  "retained": 19348,
168
  "excluded_positions": 52,
169
  "mean_useful": 12.323711340206186,
170
  "mean_retained": 19.94639175257732
171
  },
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
172
  "selected": {
173
  "scored_queries": 970,
174
  "excluded_queries": 0,
175
  "hit": 0.8608247422680413,
176
  "precision": 0.8027855153203343,
177
  "macro_precision": 0.7038533836490732,
178
- "ndcg": 0.6599601279630594,
179
- "known_positive_recall": 0.2082001919384574,
180
- "recall_queries": 895,
181
  "useful": 5764,
182
  "retained": 7180,
183
  "excluded_positions": 17,
 
8
  "hit": 0.7233160621761658,
9
  "precision": 0.7233160621761658,
10
  "macro_precision": 0.7233160621761658,
11
+ "ndcg": 0.6458425857389588,
12
+ "known_positive_recall": 0.021209009540392988,
13
+ "recall_queries": 920,
14
  "useful": 698,
15
  "retained": 965,
16
  "excluded_positions": 0,
 
23
  "hit": 0.8536082474226804,
24
  "precision": 0.7157640565712314,
25
  "macro_precision": 0.7166666666666667,
26
+ "ndcg": 0.6374467700901159,
27
+ "known_positive_recall": 0.0589083117861917,
28
+ "recall_queries": 925,
29
  "useful": 2075,
30
  "retained": 2899,
31
  "excluded_positions": 11,
 
38
  "hit": 0.8845360824742268,
39
  "precision": 0.7028311634635255,
40
  "macro_precision": 0.7031786941580757,
41
+ "ndcg": 0.632593817782834,
42
+ "known_positive_recall": 0.09351524011264825,
43
+ "recall_queries": 925,
44
  "useful": 3401,
45
  "retained": 4839,
46
  "excluded_positions": 11,
 
53
  "hit": 0.8989690721649485,
54
  "precision": 0.67180070291503,
55
  "macro_precision": 0.6722631320569465,
56
+ "ndcg": 0.6315462046841925,
57
+ "known_positive_recall": 0.17039667688880922,
58
+ "recall_queries": 925,
59
  "useful": 6499,
60
  "retained": 9674,
61
  "excluded_positions": 26,
 
68
  "hit": 0.9103092783505154,
69
  "precision": 0.61794500723589,
70
  "macro_precision": 0.6184956360634603,
71
+ "ndcg": 0.6416731640651591,
72
+ "known_positive_recall": 0.29717756138114526,
73
+ "recall_queries": 925,
74
  "useful": 11956,
75
  "retained": 19348,
76
  "excluded_positions": 52,
77
  "mean_useful": 12.32577319587629,
78
  "mean_retained": 19.94639175257732
79
  },
80
+ "50": {
81
+ "scored_queries": 970,
82
+ "excluded_queries": 0,
83
+ "hit": 0.9309278350515464,
84
+ "precision": 0.4043216890119198,
85
+ "macro_precision": 0.4044939351610387,
86
+ "ndcg": 0.6051209510262134,
87
+ "known_positive_recall": 0.4598077844640693,
88
+ "recall_queries": 925,
89
+ "useful": 19572,
90
+ "retained": 48407,
91
+ "excluded_positions": 93,
92
+ "mean_useful": 20.177319587628865,
93
+ "mean_retained": 49.904123711340205
94
+ },
95
  "selected": {
96
  "scored_queries": 970,
97
  "excluded_queries": 0,
98
  "hit": 0.8618556701030928,
99
  "precision": 0.8033770583310076,
100
  "macro_precision": 0.7040325511660611,
101
+ "ndcg": 0.6359990281117442,
102
+ "known_positive_recall": 0.128026594631235,
103
+ "recall_queries": 925,
104
  "useful": 5757,
105
  "retained": 7166,
106
  "excluded_positions": 17,
 
115
  "hit": 0.7243523316062176,
116
  "precision": 0.7243523316062176,
117
  "macro_precision": 0.7243523316062176,
118
+ "ndcg": 0.6461386627189736,
119
+ "known_positive_recall": 0.021228078953055077,
120
+ "recall_queries": 920,
121
  "useful": 699,
122
  "retained": 965,
123
  "excluded_positions": 0,
 
130
  "hit": 0.8525773195876288,
131
  "precision": 0.7157640565712314,
132
  "macro_precision": 0.7166666666666667,
133
+ "ndcg": 0.6373247400469291,
134
+ "known_positive_recall": 0.05893952330659241,
135
+ "recall_queries": 925,
136
  "useful": 2075,
137
  "retained": 2899,
138
  "excluded_positions": 11,
 
145
  "hit": 0.8845360824742268,
146
  "precision": 0.7023563455973543,
147
  "macro_precision": 0.702766323024055,
148
+ "ndcg": 0.6321994825191731,
149
+ "known_positive_recall": 0.09344587015679974,
150
+ "recall_queries": 925,
151
  "useful": 3398,
152
  "retained": 4838,
153
  "excluded_positions": 12,
 
160
  "hit": 0.8989690721649485,
161
  "precision": 0.6723514211886304,
162
  "macro_precision": 0.6727589592538046,
163
+ "ndcg": 0.631891212140629,
164
+ "known_positive_recall": 0.17053838380216288,
165
+ "recall_queries": 925,
166
  "useful": 6505,
167
  "retained": 9675,
168
  "excluded_positions": 25,
 
175
  "hit": 0.9103092783505154,
176
  "precision": 0.6178416373785404,
177
  "macro_precision": 0.6183925432799551,
178
+ "ndcg": 0.6416203428445276,
179
+ "known_positive_recall": 0.2971249330393808,
180
+ "recall_queries": 925,
181
  "useful": 11954,
182
  "retained": 19348,
183
  "excluded_positions": 52,
184
  "mean_useful": 12.323711340206186,
185
  "mean_retained": 19.94639175257732
186
  },
187
+ "50": {
188
+ "scored_queries": 970,
189
+ "excluded_queries": 0,
190
+ "hit": 0.9309278350515464,
191
+ "precision": 0.4043216890119198,
192
+ "macro_precision": 0.4044939351610387,
193
+ "ndcg": 0.6050905860403181,
194
+ "known_positive_recall": 0.4598077844640693,
195
+ "recall_queries": 925,
196
+ "useful": 19572,
197
+ "retained": 48407,
198
+ "excluded_positions": 93,
199
+ "mean_useful": 20.177319587628865,
200
+ "mean_retained": 49.904123711340205
201
+ },
202
  "selected": {
203
  "scored_queries": 970,
204
  "excluded_queries": 0,
205
  "hit": 0.8608247422680413,
206
  "precision": 0.8027855153203343,
207
  "macro_precision": 0.7038533836490732,
208
+ "ndcg": 0.6358997417721292,
209
+ "known_positive_recall": 0.12825382711378852,
210
+ "recall_queries": 925,
211
  "useful": 5764,
212
  "retained": 7180,
213
  "excluded_positions": 17,
evaluation/serving.json CHANGED
The diff for this file is too large to render. See raw diff
 
figures/gemma-comparison.svg CHANGED
publication-manifest.json CHANGED
@@ -1,6 +1,6 @@
1
  {
2
  "schema": "daecore.retrieval-publication-manifest",
3
- "tool_sha256": "3ea615c62866883f911899f1e131d73ca8cc9701eea3540a3a016345ac5650f8",
4
  "model_id": "embeddinggemma-300m-memory-ft-v2",
5
  "public_repo": "Daecore/embeddinggemma-300m-memory-ft-v2",
6
  "role": "dense-retriever",
@@ -12,7 +12,7 @@
12
  "export_receipt_sha256": "e24e1692f7e03c9af8040bb3f6e310a061d305f485861007f1f37d36808b20d8",
13
  "serving_sha256": "8c088b1635875d5406e148782b303cfc5a961138600783fb4842b413b98a98c3",
14
  "source_model_tree_sha256": "30a7b9f1c42b1fcee29a6729742685d5b744e8fa87d13cefab2fd2d49ea9150e",
15
- "staged_at": "2026-09-29T16:47:59+00:00",
16
  "files": {
17
  "added_tokens.json": {
18
  "sha256": "50b2f405ba56a26d4913fd772089992252d7f942123cc0a034d96424221ba946",
@@ -80,8 +80,8 @@
80
  "binding": "packaging record"
81
  },
82
  "evaluation/README.md": {
83
- "sha256": "f7059b76c4e9956e899b5f8fb855a58262db42ffef301404c874c8ee9b1c19dc",
84
- "size": 8475,
85
  "binding": "packaging record"
86
  },
87
  "evaluation/serving-qualification.json": {
@@ -90,18 +90,23 @@
90
  "binding": "packaging record"
91
  },
92
  "evaluation/metrics.py": {
93
- "sha256": "72e7452766b491c247514df8ad1c8f52c8023e1e305d779655a9e1f3159ce52d",
94
- "size": 17556,
 
 
 
 
 
95
  "binding": "packaging record"
96
  },
97
  "evaluation/retriever-summary.json": {
98
- "sha256": "6d87cfa07056cb0316e21ed70a940994d1ad91e9a297161878482814811f736b",
99
- "size": 5063,
100
  "binding": "packaging record"
101
  },
102
  "evaluation/dimensions-summary.json": {
103
- "sha256": "d42adc93268a44675bcda3321a431f8cd92e6c9178efdb44e0314372468a7c6e",
104
- "size": 1253,
105
  "binding": "packaging record"
106
  },
107
  "evaluation/promotion-summary.json": {
@@ -110,18 +115,18 @@
110
  "binding": "packaging record"
111
  },
112
  "evaluation/serving-summary.json": {
113
- "sha256": "fc60b4e94fe01e0470dcced382046c4c30c329a87b5caa2b631b369015548198",
114
- "size": 6029,
115
  "binding": "packaging record"
116
  },
117
  "evaluation/retriever.json": {
118
- "sha256": "fcfbe8155cd30319ee38745f92376a45635ecaed08b86962ee8604680ce910b7",
119
- "size": 573086,
120
  "binding": "packaging record"
121
  },
122
  "evaluation/dimensions.json": {
123
- "sha256": "a6f999513a4b1e044901af4cb4a9797cb6f3d0ae6d1124e68bf406983a02f3ac",
124
- "size": 762884,
125
  "binding": "packaging record"
126
  },
127
  "evaluation/promotion.json": {
@@ -130,13 +135,13 @@
130
  "binding": "packaging record"
131
  },
132
  "evaluation/serving.json": {
133
- "sha256": "0c017dbd0e66734c465d617eef24b2affb2670a96f7fd5b2c001514c3cc1381a",
134
- "size": 500118,
135
  "binding": "packaging record"
136
  },
137
  "evaluation/figures.py": {
138
- "sha256": "53819eb947c3c5c25f6331b0e384e6e27f7b5da750c84f8bdf6fd6e7790c9d94",
139
- "size": 5378,
140
  "binding": "packaging record"
141
  },
142
  "evaluation/svg_figures.py": {
@@ -145,15 +150,15 @@
145
  "binding": "packaging record"
146
  },
147
  "figures/gemma-comparison.svg": {
148
- "sha256": "f4bf74ea9e2e044e124893d5ca8bd028c6f2f79ef53babfb7ee718f8b6786be2",
149
- "size": 13377,
150
  "binding": "packaging record"
151
  },
152
  "README.md": {
153
- "sha256": "e7d88fad6c2c12613c8e8a8108a38dc7fbb181c005ffadac172f1120f84c95f0",
154
- "size": 14331,
155
  "binding": "model card with upload-relative links",
156
- "source_sha256": "5ea17ec12b2673539eba89fc6e3f85c9ed8ba7d879d9ce94dbeebcb1392a8afa"
157
  }
158
  }
159
  }
 
1
  {
2
  "schema": "daecore.retrieval-publication-manifest",
3
+ "tool_sha256": "46dcb4b17750d5a4ba3d44b1236f33e9e5657673b648ae6cad48fe4ac4f147f8",
4
  "model_id": "embeddinggemma-300m-memory-ft-v2",
5
  "public_repo": "Daecore/embeddinggemma-300m-memory-ft-v2",
6
  "role": "dense-retriever",
 
12
  "export_receipt_sha256": "e24e1692f7e03c9af8040bb3f6e310a061d305f485861007f1f37d36808b20d8",
13
  "serving_sha256": "8c088b1635875d5406e148782b303cfc5a961138600783fb4842b413b98a98c3",
14
  "source_model_tree_sha256": "30a7b9f1c42b1fcee29a6729742685d5b744e8fa87d13cefab2fd2d49ea9150e",
15
+ "staged_at": "2026-09-30T01:53:48+00:00",
16
  "files": {
17
  "added_tokens.json": {
18
  "sha256": "50b2f405ba56a26d4913fd772089992252d7f942123cc0a034d96424221ba946",
 
80
  "binding": "packaging record"
81
  },
82
  "evaluation/README.md": {
83
+ "sha256": "390b5e5d525ae6c0b2c3370050af1cc14b4574127bb258db911fb0afed670494",
84
+ "size": 9703,
85
  "binding": "packaging record"
86
  },
87
  "evaluation/serving-qualification.json": {
 
90
  "binding": "packaging record"
91
  },
92
  "evaluation/metrics.py": {
93
+ "sha256": "bbccf29f34db24aefc5630adcaf9630e5f214010902dbe1df0032618aa329a45",
94
+ "size": 18329,
95
+ "binding": "packaging record"
96
+ },
97
+ "evaluation/chunking.md": {
98
+ "sha256": "cda82fca3a467f02e06625a607fc8a5bfc9b16d48f4fe225d7f2282330d4955e",
99
+ "size": 3509,
100
  "binding": "packaging record"
101
  },
102
  "evaluation/retriever-summary.json": {
103
+ "sha256": "b320430acd9938d08e31eb00693dd9f775e9890b17427f86577c9bbd91dd1bba",
104
+ "size": 6071,
105
  "binding": "packaging record"
106
  },
107
  "evaluation/dimensions-summary.json": {
108
+ "sha256": "3b160deb874b98e835b92b15070afe36e548b1ce692a7ef817af77026e5064f8",
109
+ "size": 9753,
110
  "binding": "packaging record"
111
  },
112
  "evaluation/promotion-summary.json": {
 
115
  "binding": "packaging record"
116
  },
117
  "evaluation/serving-summary.json": {
118
+ "sha256": "cd87757b53482e2ec9450a3c64f31a32afad2710900fbcbd7885828929aa9535",
119
+ "size": 7035,
120
  "binding": "packaging record"
121
  },
122
  "evaluation/retriever.json": {
123
+ "sha256": "c209c96fd17144a6f7d2c7174bc27db4b346288910d085e2b614286b80ae4559",
124
+ "size": 1156315,
125
  "binding": "packaging record"
126
  },
127
  "evaluation/dimensions.json": {
128
+ "sha256": "3d9b49d6bda19fe9f9b7f8859ce550defa230757fef85ee2bc1ff175f6f4fe4b",
129
+ "size": 1063026,
130
  "binding": "packaging record"
131
  },
132
  "evaluation/promotion.json": {
 
135
  "binding": "packaging record"
136
  },
137
  "evaluation/serving.json": {
138
+ "sha256": "f4870a5118400152d4a78c7fd7c96df4df73901bfe0c590fdcd00c2af477a4be",
139
+ "size": 1184937,
140
  "binding": "packaging record"
141
  },
142
  "evaluation/figures.py": {
143
+ "sha256": "6bc89020363df7a0cfcdebaa4e8be5b37385ab97b2b7bbf3f60cb9b46b2ac63e",
144
+ "size": 5408,
145
  "binding": "packaging record"
146
  },
147
  "evaluation/svg_figures.py": {
 
150
  "binding": "packaging record"
151
  },
152
  "figures/gemma-comparison.svg": {
153
+ "sha256": "4393db8afc6de32877a66b863708e396a1503d753932680dc7e4e590c4ccd513",
154
+ "size": 14507,
155
  "binding": "packaging record"
156
  },
157
  "README.md": {
158
+ "sha256": "b5fe1ab8b6d3821cf5ee42e9ffbb2d88d8509233d27834382ce63a385f5d60a4",
159
+ "size": 16388,
160
  "binding": "model card with upload-relative links",
161
+ "source_sha256": "58c1418eb860b7e8d2db5b84085780bb665ffbd92d06224a59cff700cebe0f02"
162
  }
163
  }
164
  }