tnh0527 commited on
Commit
766fddf
·
verified ·
1 Parent(s): d02f909

Update embeddinggemma-300m-memory-ft-v2 documentation

Browse files
README.md CHANGED
@@ -129,9 +129,9 @@ passage, which does not prove none exists.
129
 
130
  At 50, v2 finds useful evidence for **99.03%** of these queries and retrieves
131
  **57.92%** of their known useful passages, versus 97.73% and 41.12% upstream.
132
- The hybrid pool below has lower coverage on the same subset, despite stronger
133
- early precision after reranking. Adding BM25 and capping the fused list at 50
134
- can displace dense candidates.
135
 
136
  Grades 2–3 count as useful. nDCG uses gains 0, 1, 3 and 7 against the same
137
  reviewed reference for both models. Rankings from query forms are combined
@@ -178,44 +178,47 @@ larger quality cost. These byte counts exclude index overhead and compression;
178
  shorter vectors do not reduce encoder inference cost. Daecore uses 768, and
179
  smaller widths have not been qualified through the full hybrid pipeline.
180
 
181
- ### Full pipeline: Gemma v2 + BM25 + Ettin
182
 
183
- This replay measures semantic and lexical search, reciprocal-rank fusion, then
184
- Ettin reranking over **82,719 passages**. The table uses the same **925 queries
185
- with known useful evidence** as Gemma's dense comparison, selected from the
186
- 970-query source panel. A query remains included when this pipeline misses its
187
- answer. Both cards share this CUDA FP16 table; it is not either model's
188
- standalone score.
189
 
190
- **Reference:** 122,231 graded query–passage pairs for these 925 queries, within
191
- the shared 127,665-pair top-50 reference (September 2026). Abstentions are excluded.
 
192
 
193
  | Return depth | Hit | Precision | Known-positive recall | nDCG |
194
  |---|---:|---:|---:|---:|
195
- | 1 | 75.87% | 75.87% | 2.12% | 0.6513 |
196
- | 3 | 89.51% | 75.07% | 5.89% | 0.6423 |
197
- | 5 | 92.76% | 73.71% | 9.35% | 0.6376 |
198
- | 10 | 94.27% | 70.46% | 17.04% | 0.6378 |
199
- | 20 | 95.46% | 64.81% | 29.72% | 0.6491 |
200
- | 50 | 97.62% | 42.40% | 45.98% | 0.6130 |
201
-
202
- **The hybrid candidate ceiling is 97.62% Hit@50 and 45.98% known-positive
203
- recall@50** on this shared answerable subset. These coverage measures describe
204
- the actual 50-candidate pool; reranking cannot change them. Daecore returns a
205
- selected prefix of 3–20, so fixed-depth scores do not evaluate the selector's
206
- choice.
207
-
208
- Grades 2–3 count as useful. At depth 1, both providers use the same **920
209
- queries**; five are omitted because a first result was ungradable. Depths
210
- 3–50 use all 925. Precision pools judged positions. Abstentions remove 52
211
- positions at rank 20 and 93 at rank 50 per provider, without drawing in deeper
212
- passages. Known-positive recall counts judged passages, not every useful
213
- passage in the corpus.
214
-
215
- The [evaluation companion](evaluation/README.md) retains
216
- coverage across all 970 source queries and explains the shared reference,
217
- exclusions and CUDA/Vulkan comparison. The [pipeline overview](https://huggingface.co/Daecore)
218
- explains retrieval and selection.
 
 
 
 
219
 
220
  ## Training data and objective
221
 
 
129
 
130
  At 50, v2 finds useful evidence for **99.03%** of these queries and retrieves
131
  **57.92%** of their known useful passages, versus 97.73% and 41.12% upstream.
132
+ The capped hybrid pool below covers fewer queries at 50. Adding BM25 and
133
+ capping the fused list can displace dense candidates, while reranking can
134
+ improve early precision.
135
 
136
  Grades 2–3 count as useful. nDCG uses gains 0, 1, 3 and 7 against the same
137
  reviewed reference for both models. Rankings from query forms are combined
 
178
  shorter vectors do not reduce encoder inference cost. Daecore uses 768, and
179
  smaller widths have not been qualified through the full hybrid pipeline.
180
 
181
+ ### Full pipeline: the effect of Gemma v2
182
 
183
+ Only the embedder changes: **upstream Gemma → Daecore Gemma v2**. BM25,
184
+ reciprocal-rank fusion and the fine-tuned Ettin reranker stay fixed. Both
185
+ pipelines search the same **82,719 passages**. Each cell reads before →
186
+ **after**, so the change measures Gemma's contribution inside hybrid retrieval.
 
 
187
 
188
+ **Reference:** the same **925 known-answerable queries**, frozen before this
189
+ comparison · **122,783 graded query–passage pairs** · October 2026. Retrieval
190
+ misses remain included. Both cards share the same fully fine-tuned endpoint.
191
 
192
  | Return depth | Hit | Precision | Known-positive recall | nDCG |
193
  |---|---:|---:|---:|---:|
194
+ | 1 | 73.99% → **75.95%** | 73.99% → **75.95%** | 2.01% → **2.12%** | 0.6295 → **0.6527** |
195
+ | 3 | 86.81% → **89.41%** | 71.50% → **75.11%** | 5.49% → **5.89%** | 0.6121 → **0.6424** |
196
+ | 5 | 89.95% → **92.76%** | 70.10% → **73.67%** | 8.53% → **9.32%** | 0.6054 → **0.6373** |
197
+ | 10 | 93.08% → **94.27%** | 65.20% → **70.49%** | 15.25% → **16.99%** | 0.5959 → **0.6377** |
198
+ | 20 | 95.68% → **95.46%** | 56.86% → **64.80%** | 25.18% → **29.63%** | 0.5873 → **0.6489** |
199
+ | 50 | 96.97% → **97.62%** | 35.45% → **42.40%** | 36.85% → **45.83%** | 0.5374 → **0.6125** |
200
+
201
+ At three results, useful-passage precision changes from **71.50% to
202
+ 75.11%**; nDCG@10 changes from **0.5959 to 0.6377**.
203
+ At 50, the hybrid pool's known-positive recall changes from **36.85%
204
+ to 45.83%**. This measures evidence available to Ettin after
205
+ fusion; reranking cannot recover passages outside those 50 slots.
206
+
207
+ The gain is not uniform: Hit@20 falls from 95.68% to 95.46%, a difference of
208
+ two queries, while precision, known-positive recall and nDCG improve at that depth.
209
+
210
+ All three variants use matched **PyTorch FP16** reranking. Depth 1 uses **919 shared judged
211
+ queries** across the three pipeline variants; depths 3–50 use all 925.
212
+ Abstentions are excluded inside the original cutoff without backfill.
213
+ Grades 2–3 count as useful. Precision pools retained positions; recall counts
214
+ known useful passages.
215
+ The expanded reference and matched runtime distinguish this table from the
216
+ earlier CUDA/Vulkan serving check. Daecore returns 3–20 results; these fixed
217
+ cutoffs do not evaluate its selector.
218
+
219
+ The [evaluation companion](evaluation/README.md) provides
220
+ the paired records, methods and separate serving checks. The
221
+ [pipeline overview](https://huggingface.co/Daecore) explains retrieval and selection.
222
 
223
  ## Training data and objective
224
 
evaluation/README.md CHANGED
@@ -11,7 +11,8 @@ private evaluations.
11
  | `retriever.json` | Upstream and v2 reviewed top-50 dense rankings for all 970 queries over 82,719 passages |
12
  | `dimensions.json` | V2 public nDCG@10 and private top-10 graded rankings at four embedding widths |
13
  | `promotion.json` | Ranked grades for the predecessor-to-v2 pipeline comparison and public nDCG@10 for upstream and v2 |
14
- | `serving.json` | The 970-query Gemma v2 + BM25 + Ettin replay used by both model cards, including CUDA/Vulkan results |
 
15
  | `*-summary.json` | Summaries that `metrics.py` reproduces |
16
  | `serving-qualification.json` | CPU, CUDA and Vulkan serving checks for this model |
17
  | `metrics.py` | Metric code |
@@ -26,7 +27,7 @@ python metrics.py --directory .
26
  python figures.py --output ../figures
27
  ```
28
 
29
- No packages or network access are needed. Each result matches its corresponding summary. The `answerable` section of the dense and serving summaries supplies the card tables; `private.answerable` supplies the private width column. The outer `models` sections retain all-query results. This recomputes metrics from saved results; running private retrieval again would require the private corpus.
30
 
31
  ## Dense retrieval
32
 
@@ -70,8 +71,9 @@ reference contained 87,434 graded pairs.
70
 
71
  The subsequent decision to headline the same 925 known-answerable queries
72
  changes the averaging population, not the labels or rankings. On that subset,
73
- nDCG@10 is **0.4418 upstream and 0.5621 for v2**. The shared reference supports
74
- dense, smaller-width and current pipeline results. The historical predecessor
 
75
  comparison below keeps its original reference and population.
76
 
77
  **Public data.**
@@ -96,25 +98,44 @@ shorten the encoder's forward pass.
96
 
97
  ## Effect in the retrieval pipeline
98
 
99
- The identical table on both cards comes from `serving.json`: Gemma v2 +
100
- BM25 + fusion + unchanged Ettin. It uses the **same 925 known-answerable
101
- queries and 122,231 graded pairs** as the dense table. CUDA FP16 and Vulkan
102
- FP32 use the same 50 candidates for each query. At 50, Hit is **97.62%** and
103
- known-positive recall is **45.98%**; reranking cannot change candidate
104
- coverage. Dense v2 reaches 99.03% and 57.92% on this subset. Fusion with a
105
- fixed 50-slot budget can displace dense candidates even while reranking
106
- improves early precision.
107
-
108
- Depths 3–50 use all 925 queries. Depth 1 uses 920 shared judged queries; five
109
- are omitted because a first result was ungradable. Each provider excludes 52
110
- positions at rank 20 and 93 at rank 50 without backfill. Fixed cutoffs are
111
- separate from the selector's choice of a 3–20-item prefix. Vulkan matches
112
- CUDA at Hit@5, Hit@10, Hit@20 and Hit@50; Hit@3 differs by one query and
113
- nDCG@10 by 0.0004.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
114
 
115
  ### All-query coverage
116
 
117
- The records preserve all 970 source queries. This audit view includes the 45
118
  without known useful evidence and is separate from the card's answerable-only
119
  quality table. Hybrid top 1 uses 965 queries after shared abstention
120
  exclusions; dense top 1 retains all 970. The counts below use the full
 
11
  | `retriever.json` | Upstream and v2 reviewed top-50 dense rankings for all 970 queries over 82,719 passages |
12
  | `dimensions.json` | V2 public nDCG@10 and private top-10 graded rankings at four embedding widths |
13
  | `promotion.json` | Ranked grades for the predecessor-to-v2 pipeline comparison and public nDCG@10 for upstream and v2 |
14
+ | `pipeline.json` | Matched upstream-to-fine-tuned component changes on the frozen 925-query hybrid panel |
15
+ | `serving.json` | Separate 970-query CUDA/Vulkan replay with the September reference |
16
  | `*-summary.json` | Summaries that `metrics.py` reproduces |
17
  | `serving-qualification.json` | CPU, CUDA and Vulkan serving checks for this model |
18
  | `metrics.py` | Metric code |
 
27
  python figures.py --output ../figures
28
  ```
29
 
30
+ No packages or network access are needed. Each result matches its corresponding summary. The dense card table uses `retriever-summary.json`'s `answerable` section; the pipeline table uses `pipeline-summary.json`, comparing `upstream_gemma` with `finetuned`. `private.answerable` supplies the private width column. The dense and serving records retain all-query results in their outer `models` sections. This recomputes metrics from saved results; running private retrieval again would require the private corpus.
31
 
32
  ## Dense retrieval
33
 
 
71
 
72
  The subsequent decision to headline the same 925 known-answerable queries
73
  changes the averaging population, not the labels or rankings. On that subset,
74
+ nDCG@10 is **0.4418 upstream and 0.5621 for v2**. The September reference supports
75
+ dense, smaller-width and provider results. The new component comparison adds
76
+ its own reviewed overlay, described below. The historical predecessor
77
  comparison below keeps its original reference and population.
78
 
79
  **Public data.**
 
98
 
99
  ## Effect in the retrieval pipeline
100
 
101
+ `pipeline.json` holds three matched arms. For this card, compare
102
+ `upstream_gemma` with `finetuned`: only Gemma changes, while BM25, fusion and
103
+ the fine-tuned Ettin remain fixed. The Ettin card uses the other baseline
104
+ against the same fully fine-tuned endpoint.
105
+
106
+ The comparison keeps the **925-query cohort frozen before scoring** and uses
107
+ **122,783 graded query–passage pairs** for those queries. It adds 552 grades
108
+ and 10 abstentions to their September reference of 122,231 pairs; old
109
+ grades are unchanged. Two independent GPT-6.1 Sol judges and a fresh blind
110
+ adjudicator resolved the new pairs, with 84 ordinal disagreements.
111
+ The primary assistant reviewed 84 complete pairs, including every final quotation
112
+ flag and abstention plus a stratified sample. These are model judgments,
113
+ without a human reference panel.
114
+
115
+ All three variants were scored anew in **PyTorch FP16**, using Transformers
116
+ 5.2.0, Sentence Transformers 5.5.1, 128 query tokens and 1,024 passage tokens.
117
+ Gemma uses cached FP32 embeddings at 768 dimensions with matched input limits.
118
+ The actual first stage uses the frozen query forms, dense retrieval, BM25,
119
+ reciprocal-rank fusion with constant 60, exact-text deduplication and 50 slots.
120
+ Before substituting upstream Gemma's vectors, replay reproduced all 970 saved
121
+ v2 candidate lists exactly. Ties preserve candidate order.
122
+
123
+ Depth 1 uses **919 common judged queries across all three arms**; depths
124
+ 3–50 retain all 925. Original prefixes are cut before abstentions are removed,
125
+ with no backfill. Retrieval misses remain included. The reference supplies
126
+ known-positive recall and ideal DCG; it is not exhaustive corpus labeling.
127
+ Fixed cutoffs do not test the calibrated selector's choice of 3–20 results.
128
+ The earlier ONNX provider replay retains its own runtime and September
129
+ reference, so its absolute scores should not be substituted into this comparison.
130
+
131
+ `serving.json` separately compares ONNX CUDA FP16 with Vulkan FP32. It keeps
132
+ the September reference of 122,231 grades for the same 925 queries, with 920
133
+ eligible at top 1. Both providers use identical candidate pools. This checks
134
+ serving paths, not the effect of fine-tuning.
135
 
136
  ### All-query coverage
137
 
138
+ The September dense and serving records preserve all 970 source queries. This audit view includes the 45
139
  without known useful evidence and is separate from the card's answerable-only
140
  quality table. Hybrid top 1 uses 965 queries after shared abstention
141
  exclusions; dense top 1 retains all 970. The counts below use the full
evaluation/metrics.py CHANGED
@@ -267,10 +267,22 @@ def summarize_serving(data: dict) -> dict:
267
  'answerable': _summarize_ranked_grades(data, include_selected=True, answerable_only=True)}
268
 
269
 
 
 
 
 
 
 
 
 
 
 
 
 
270
  SUMMARIZERS = {
271
  'classifier': summarize_classifier, 'reranker': summarize_reranker,
272
  'retriever': summarize_retriever, 'dimensions': summarize_dimensions,
273
- 'promotion': summarize_promotion, 'serving': summarize_serving,
274
  }
275
 
276
 
@@ -283,6 +295,7 @@ def summarize_record(name: str, data: dict) -> dict:
283
  'dimensions': {'id', 'dataset', 'ndcg@10'},
284
  'promotion': {'id', 'group', 'reference_grade_counts', 'ranked_grades', 'excluded', 'selected_depth'},
285
  'serving': {'id', 'group', 'reference_grade_counts', 'ranked_grades', 'excluded', 'selected_depth'},
 
286
  }[name]
287
  rows = data['rows']
288
  if not rows or any(set(row) != fields for row in rows):
@@ -336,7 +349,7 @@ def main() -> None:
336
  parser = argparse.ArgumentParser(description=__doc__)
337
  parser.add_argument("--directory", type=Path, default=Path(__file__).resolve().parent)
338
  parser.add_argument("--output", type=Path)
339
- parser.add_argument("--model", choices=("classifier", "reranker", "retriever", "dimensions", "promotion", "serving"), help="Recompute one comparison; default: every record in this package")
340
  args = parser.parse_args()
341
  names = [args.model] if args.model else [
342
  name for name in SUMMARIZERS if (args.directory / f"{name}.json").is_file()
 
267
  'answerable': _summarize_ranked_grades(data, include_selected=True, answerable_only=True)}
268
 
269
 
270
+ def summarize_pipeline(data: dict) -> dict:
271
+ """Matched component changes on a frozen, known-answerable query cohort."""
272
+ if type(data['corpus_passages']) is not int or data['corpus_passages'] <= 0:
273
+ raise ValueError('Pipeline comparison requires a positive corpus size')
274
+ output = _summarize_ranked_grades(data, include_selected=False)
275
+ if any(row['reference_grade_counts']['2'] + row['reference_grade_counts']['3'] == 0
276
+ for row in data['rows']):
277
+ raise ValueError('Pipeline cohort requires known useful evidence for every query')
278
+ return {**output, 'corpus_passages': data['corpus_passages'],
279
+ 'reference_pairs': sum(sum(row['reference_grade_counts'].values()) for row in data['rows'])}
280
+
281
+
282
  SUMMARIZERS = {
283
  'classifier': summarize_classifier, 'reranker': summarize_reranker,
284
  'retriever': summarize_retriever, 'dimensions': summarize_dimensions,
285
+ 'promotion': summarize_promotion, 'serving': summarize_serving, 'pipeline': summarize_pipeline,
286
  }
287
 
288
 
 
295
  'dimensions': {'id', 'dataset', 'ndcg@10'},
296
  'promotion': {'id', 'group', 'reference_grade_counts', 'ranked_grades', 'excluded', 'selected_depth'},
297
  'serving': {'id', 'group', 'reference_grade_counts', 'ranked_grades', 'excluded', 'selected_depth'},
298
+ 'pipeline': {'id', 'group', 'reference_grade_counts', 'ranked_grades', 'excluded'},
299
  }[name]
300
  rows = data['rows']
301
  if not rows or any(set(row) != fields for row in rows):
 
349
  parser = argparse.ArgumentParser(description=__doc__)
350
  parser.add_argument("--directory", type=Path, default=Path(__file__).resolve().parent)
351
  parser.add_argument("--output", type=Path)
352
+ parser.add_argument("--model", choices=tuple(SUMMARIZERS), help="Recompute one comparison; default: every record in this package")
353
  args = parser.parse_args()
354
  names = [args.model] if args.model else [
355
  name for name in SUMMARIZERS if (args.directory / f"{name}.json").is_file()
evaluation/pipeline-summary.json ADDED
@@ -0,0 +1,283 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "queries": 925,
3
+ "models": {
4
+ "upstream_gemma": {
5
+ "1": {
6
+ "scored_queries": 919,
7
+ "excluded_queries": 6,
8
+ "hit": 0.7399347116430903,
9
+ "precision": 0.7399347116430903,
10
+ "macro_precision": 0.7399347116430903,
11
+ "ndcg": 0.6295144826156795,
12
+ "known_positive_recall": 0.02009148065967933,
13
+ "recall_queries": 919,
14
+ "useful": 680,
15
+ "retained": 919,
16
+ "excluded_positions": 0,
17
+ "mean_useful": 0.7399347116430903,
18
+ "mean_retained": 1.0
19
+ },
20
+ "3": {
21
+ "scored_queries": 925,
22
+ "excluded_queries": 0,
23
+ "hit": 0.8681081081081081,
24
+ "precision": 0.7149566473988439,
25
+ "macro_precision": 0.7153153153153153,
26
+ "ndcg": 0.6120552491562434,
27
+ "known_positive_recall": 0.05489711276806684,
28
+ "recall_queries": 925,
29
+ "useful": 1979,
30
+ "retained": 2768,
31
+ "excluded_positions": 7,
32
+ "mean_useful": 2.1394594594594594,
33
+ "mean_retained": 2.9924324324324325
34
+ },
35
+ "5": {
36
+ "scored_queries": 925,
37
+ "excluded_queries": 0,
38
+ "hit": 0.8994594594594595,
39
+ "precision": 0.7009527934170636,
40
+ "macro_precision": 0.7011351351351351,
41
+ "ndcg": 0.60538850988017,
42
+ "known_positive_recall": 0.08526196518888282,
43
+ "recall_queries": 925,
44
+ "useful": 3237,
45
+ "retained": 4618,
46
+ "excluded_positions": 7,
47
+ "mean_useful": 3.4994594594594592,
48
+ "mean_retained": 4.992432432432432
49
+ },
50
+ "10": {
51
+ "scored_queries": 925,
52
+ "excluded_queries": 0,
53
+ "hit": 0.9308108108108109,
54
+ "precision": 0.6519761775852734,
55
+ "macro_precision": 0.6521381381381381,
56
+ "ndcg": 0.5958896253275177,
57
+ "known_positive_recall": 0.15250767729827347,
58
+ "recall_queries": 925,
59
+ "useful": 6021,
60
+ "retained": 9235,
61
+ "excluded_positions": 15,
62
+ "mean_useful": 6.509189189189189,
63
+ "mean_retained": 9.983783783783784
64
+ },
65
+ "20": {
66
+ "scored_queries": 925,
67
+ "excluded_queries": 0,
68
+ "hit": 0.9567567567567568,
69
+ "precision": 0.5685743678596568,
70
+ "macro_precision": 0.5687161650814901,
71
+ "ndcg": 0.5873176954516321,
72
+ "known_positive_recall": 0.2518432399920873,
73
+ "recall_queries": 925,
74
+ "useful": 10501,
75
+ "retained": 18469,
76
+ "excluded_positions": 31,
77
+ "mean_useful": 11.352432432432433,
78
+ "mean_retained": 19.966486486486488
79
+ },
80
+ "50": {
81
+ "scored_queries": 925,
82
+ "excluded_queries": 0,
83
+ "hit": 0.9697297297297297,
84
+ "precision": 0.35454801464376234,
85
+ "macro_precision": 0.35468757089595926,
86
+ "ndcg": 0.5373826502849983,
87
+ "known_positive_recall": 0.36848413838619143,
88
+ "recall_queries": 925,
89
+ "useful": 16367,
90
+ "retained": 46163,
91
+ "excluded_positions": 87,
92
+ "mean_useful": 17.694054054054053,
93
+ "mean_retained": 49.905945945945945
94
+ }
95
+ },
96
+ "upstream_ettin": {
97
+ "1": {
98
+ "scored_queries": 919,
99
+ "excluded_queries": 6,
100
+ "hit": 0.5669205658324266,
101
+ "precision": 0.5669205658324266,
102
+ "macro_precision": 0.5669205658324266,
103
+ "ndcg": 0.4731851391263796,
104
+ "known_positive_recall": 0.01443636631579192,
105
+ "recall_queries": 919,
106
+ "useful": 521,
107
+ "retained": 919,
108
+ "excluded_positions": 0,
109
+ "mean_useful": 0.5669205658324266,
110
+ "mean_retained": 1.0
111
+ },
112
+ "3": {
113
+ "scored_queries": 925,
114
+ "excluded_queries": 0,
115
+ "hit": 0.7902702702702703,
116
+ "precision": 0.5603603603603604,
117
+ "macro_precision": 0.5603603603603603,
118
+ "ndcg": 0.4727968405371722,
119
+ "known_positive_recall": 0.04228504528013372,
120
+ "recall_queries": 925,
121
+ "useful": 1555,
122
+ "retained": 2775,
123
+ "excluded_positions": 0,
124
+ "mean_useful": 1.681081081081081,
125
+ "mean_retained": 3.0
126
+ },
127
+ "5": {
128
+ "scored_queries": 925,
129
+ "excluded_queries": 0,
130
+ "hit": 0.8562162162162162,
131
+ "precision": 0.5582954791261086,
132
+ "macro_precision": 0.5582162162162162,
133
+ "ndcg": 0.47778481822398317,
134
+ "known_positive_recall": 0.06742273722612116,
135
+ "recall_queries": 925,
136
+ "useful": 2581,
137
+ "retained": 4623,
138
+ "excluded_positions": 2,
139
+ "mean_useful": 2.79027027027027,
140
+ "mean_retained": 4.997837837837838
141
+ },
142
+ "10": {
143
+ "scored_queries": 925,
144
+ "excluded_queries": 0,
145
+ "hit": 0.9113513513513514,
146
+ "precision": 0.5447462395844606,
147
+ "macro_precision": 0.5447507507507507,
148
+ "ndcg": 0.49027329580696355,
149
+ "known_positive_recall": 0.12832495703473729,
150
+ "recall_queries": 925,
151
+ "useful": 5034,
152
+ "retained": 9241,
153
+ "excluded_positions": 9,
154
+ "mean_useful": 5.4421621621621625,
155
+ "mean_retained": 9.99027027027027
156
+ },
157
+ "20": {
158
+ "scored_queries": 925,
159
+ "excluded_queries": 0,
160
+ "hit": 0.9502702702702702,
161
+ "precision": 0.5220568335588633,
162
+ "macro_precision": 0.5220833774951422,
163
+ "ndcg": 0.5201039924104931,
164
+ "known_positive_recall": 0.23627621886671868,
165
+ "recall_queries": 925,
166
+ "useful": 9645,
167
+ "retained": 18475,
168
+ "excluded_positions": 25,
169
+ "mean_useful": 10.427027027027027,
170
+ "mean_retained": 19.972972972972972
171
+ },
172
+ "50": {
173
+ "scored_queries": 925,
174
+ "excluded_queries": 0,
175
+ "hit": 0.9762162162162162,
176
+ "precision": 0.424031024546656,
177
+ "macro_precision": 0.42417201849319736,
178
+ "ndcg": 0.5598208131173884,
179
+ "known_positive_recall": 0.4583210768438434,
180
+ "recall_queries": 925,
181
+ "useful": 19572,
182
+ "retained": 46157,
183
+ "excluded_positions": 93,
184
+ "mean_useful": 21.158918918918918,
185
+ "mean_retained": 49.89945945945946
186
+ }
187
+ },
188
+ "finetuned": {
189
+ "1": {
190
+ "scored_queries": 919,
191
+ "excluded_queries": 6,
192
+ "hit": 0.7595212187159956,
193
+ "precision": 0.7595212187159956,
194
+ "macro_precision": 0.7595212187159956,
195
+ "ndcg": 0.6526763044717343,
196
+ "known_positive_recall": 0.021181970033131128,
197
+ "recall_queries": 919,
198
+ "useful": 698,
199
+ "retained": 919,
200
+ "excluded_positions": 0,
201
+ "mean_useful": 0.7595212187159956,
202
+ "mean_retained": 1.0
203
+ },
204
+ "3": {
205
+ "scored_queries": 925,
206
+ "excluded_queries": 0,
207
+ "hit": 0.894054054054054,
208
+ "precision": 0.7510853835021708,
209
+ "macro_precision": 0.7518918918918919,
210
+ "ndcg": 0.6424089559634334,
211
+ "known_positive_recall": 0.058858902936843926,
212
+ "recall_queries": 925,
213
+ "useful": 2076,
214
+ "retained": 2764,
215
+ "excluded_positions": 11,
216
+ "mean_useful": 2.2443243243243245,
217
+ "mean_retained": 2.9881081081081082
218
+ },
219
+ "5": {
220
+ "scored_queries": 925,
221
+ "excluded_queries": 0,
222
+ "hit": 0.9275675675675675,
223
+ "precision": 0.7366710013003901,
224
+ "macro_precision": 0.736954954954955,
225
+ "ndcg": 0.6373095878906122,
226
+ "known_positive_recall": 0.0932034345948595,
227
+ "recall_queries": 925,
228
+ "useful": 3399,
229
+ "retained": 4614,
230
+ "excluded_positions": 11,
231
+ "mean_useful": 3.674594594594595,
232
+ "mean_retained": 4.988108108108108
233
+ },
234
+ "10": {
235
+ "scored_queries": 925,
236
+ "excluded_queries": 0,
237
+ "hit": 0.9427027027027027,
238
+ "precision": 0.7049002601908065,
239
+ "macro_precision": 0.7052921492921493,
240
+ "ndcg": 0.6377369246501603,
241
+ "known_positive_recall": 0.16987439814587393,
242
+ "recall_queries": 925,
243
+ "useful": 6502,
244
+ "retained": 9224,
245
+ "excluded_positions": 26,
246
+ "mean_useful": 7.029189189189189,
247
+ "mean_retained": 9.971891891891891
248
+ },
249
+ "20": {
250
+ "scored_queries": 925,
251
+ "excluded_queries": 0,
252
+ "hit": 0.9545945945945946,
253
+ "precision": 0.6479835212489159,
254
+ "macro_precision": 0.6484765048449259,
255
+ "ndcg": 0.6488792776497584,
256
+ "known_positive_recall": 0.2962743880804353,
257
+ "recall_queries": 925,
258
+ "useful": 11954,
259
+ "retained": 18448,
260
+ "excluded_positions": 52,
261
+ "mean_useful": 12.923243243243244,
262
+ "mean_retained": 19.943783783783783
263
+ },
264
+ "50": {
265
+ "scored_queries": 925,
266
+ "excluded_queries": 0,
267
+ "hit": 0.9762162162162162,
268
+ "precision": 0.424031024546656,
269
+ "macro_precision": 0.42417201849319736,
270
+ "ndcg": 0.612510329000538,
271
+ "known_positive_recall": 0.4583210768438434,
272
+ "recall_queries": 925,
273
+ "useful": 19572,
274
+ "retained": 46157,
275
+ "excluded_positions": 93,
276
+ "mean_useful": 21.158918918918918,
277
+ "mean_retained": 49.89945945945946
278
+ }
279
+ }
280
+ },
281
+ "corpus_passages": 82719,
282
+ "reference_pairs": 122783
283
+ }
evaluation/pipeline.json ADDED
The diff for this file is too large to render. See raw diff
 
publication-manifest.json CHANGED
@@ -1,6 +1,6 @@
1
  {
2
  "schema": "daecore.retrieval-publication-manifest",
3
- "tool_sha256": "46dcb4b17750d5a4ba3d44b1236f33e9e5657673b648ae6cad48fe4ac4f147f8",
4
  "model_id": "embeddinggemma-300m-memory-ft-v2",
5
  "public_repo": "Daecore/embeddinggemma-300m-memory-ft-v2",
6
  "role": "dense-retriever",
@@ -12,7 +12,7 @@
12
  "export_receipt_sha256": "e24e1692f7e03c9af8040bb3f6e310a061d305f485861007f1f37d36808b20d8",
13
  "serving_sha256": "8c088b1635875d5406e148782b303cfc5a961138600783fb4842b413b98a98c3",
14
  "source_model_tree_sha256": "30a7b9f1c42b1fcee29a6729742685d5b744e8fa87d13cefab2fd2d49ea9150e",
15
- "staged_at": "2026-09-30T02:12:24+00:00",
16
  "files": {
17
  "added_tokens.json": {
18
  "sha256": "50b2f405ba56a26d4913fd772089992252d7f942123cc0a034d96424221ba946",
@@ -80,8 +80,8 @@
80
  "binding": "packaging record"
81
  },
82
  "evaluation/README.md": {
83
- "sha256": "bf96501863235b426b679d84e0bba96bc5b5d7fdba76571966af92993ce4150e",
84
- "size": 11100,
85
  "binding": "packaging record"
86
  },
87
  "evaluation/serving-qualification.json": {
@@ -90,8 +90,8 @@
90
  "binding": "packaging record"
91
  },
92
  "evaluation/metrics.py": {
93
- "sha256": "22536c4437be0b8c678df0dee4c45e31c5413e91f1fecf104750486f3e372d46",
94
- "size": 19120,
95
  "binding": "packaging record"
96
  },
97
  "evaluation/chunking.md": {
@@ -119,6 +119,11 @@
119
  "size": 14521,
120
  "binding": "packaging record"
121
  },
 
 
 
 
 
122
  "evaluation/retriever.json": {
123
  "sha256": "c209c96fd17144a6f7d2c7174bc27db4b346288910d085e2b614286b80ae4559",
124
  "size": 1156315,
@@ -139,6 +144,11 @@
139
  "size": 1184937,
140
  "binding": "packaging record"
141
  },
 
 
 
 
 
142
  "evaluation/figures.py": {
143
  "sha256": "f13945b350e007deb18a2eafbe1c5708a71b9df62ee12b6833d07f3e2a06094e",
144
  "size": 5490,
@@ -155,10 +165,10 @@
155
  "binding": "packaging record"
156
  },
157
  "README.md": {
158
- "sha256": "7c9add0392d4d28e8e048a2308ac4cf88b5ac0d9ee2e5e49160625d3731975cb",
159
- "size": 16352,
160
  "binding": "model card with upload-relative links",
161
- "source_sha256": "38f074c37bde704f74971f782fa291850fdd325a4030671a344ef0ba4152a93d"
162
  }
163
  }
164
  }
 
1
  {
2
  "schema": "daecore.retrieval-publication-manifest",
3
+ "tool_sha256": "60165b669f2cce17d8830fe9b712fb60a8e82280c54b224d71a1ce8fd82142ac",
4
  "model_id": "embeddinggemma-300m-memory-ft-v2",
5
  "public_repo": "Daecore/embeddinggemma-300m-memory-ft-v2",
6
  "role": "dense-retriever",
 
12
  "export_receipt_sha256": "e24e1692f7e03c9af8040bb3f6e310a061d305f485861007f1f37d36808b20d8",
13
  "serving_sha256": "8c088b1635875d5406e148782b303cfc5a961138600783fb4842b413b98a98c3",
14
  "source_model_tree_sha256": "30a7b9f1c42b1fcee29a6729742685d5b744e8fa87d13cefab2fd2d49ea9150e",
15
+ "staged_at": "2026-10-01T16:22:42+00:00",
16
  "files": {
17
  "added_tokens.json": {
18
  "sha256": "50b2f405ba56a26d4913fd772089992252d7f942123cc0a034d96424221ba946",
 
80
  "binding": "packaging record"
81
  },
82
  "evaluation/README.md": {
83
+ "sha256": "eeda3234f34b38ccd765f67f0f11e71e125e09a383173cafcee49018884bf4a2",
84
+ "size": 12608,
85
  "binding": "packaging record"
86
  },
87
  "evaluation/serving-qualification.json": {
 
90
  "binding": "packaging record"
91
  },
92
  "evaluation/metrics.py": {
93
+ "sha256": "8123f75da6edb158db74c62c1c65b6eefbc2713572409ac8a9bfb5424497aaf1",
94
+ "size": 19927,
95
  "binding": "packaging record"
96
  },
97
  "evaluation/chunking.md": {
 
119
  "size": 14521,
120
  "binding": "packaging record"
121
  },
122
+ "evaluation/pipeline-summary.json": {
123
+ "sha256": "5f28ae25b5db2f4c0a8d11891df7117f95f644063092d58ba45e27a59747e3a2",
124
+ "size": 9373,
125
+ "binding": "packaging record"
126
+ },
127
  "evaluation/retriever.json": {
128
  "sha256": "c209c96fd17144a6f7d2c7174bc27db4b346288910d085e2b614286b80ae4559",
129
  "size": 1156315,
 
144
  "size": 1184937,
145
  "binding": "packaging record"
146
  },
147
+ "evaluation/pipeline.json": {
148
+ "sha256": "8f140901a6b151d9db97fd109a3774bf1c046b849d9a01c874cb0efdea297300",
149
+ "size": 1614738,
150
+ "binding": "packaging record"
151
+ },
152
  "evaluation/figures.py": {
153
  "sha256": "f13945b350e007deb18a2eafbe1c5708a71b9df62ee12b6833d07f3e2a06094e",
154
  "size": 5490,
 
165
  "binding": "packaging record"
166
  },
167
  "README.md": {
168
+ "sha256": "46843f44d3ca441406180c2f410a106601449ad092b99269f8c10ea751c1cfa3",
169
+ "size": 16894,
170
  "binding": "model card with upload-relative links",
171
+ "source_sha256": "9e4e473e9a7479006bc9d944d70f81a27152819a708cb66d78c2c5d4915e9a5e"
172
  }
173
  }
174
  }