tnh0527 commited on
Commit
c316336
·
verified ·
1 Parent(s): 6c038cf

Update ettin-150m-memory-reranker-ft-v1 documentation

Browse files
README.md CHANGED
@@ -129,44 +129,44 @@ Both models rerank the same 50 candidates that upstream Gemma retrieved for each
129
 
130
  These are nDCG@10 scores for reranking fixed candidate pools, not dense Gemma scores. Fine-tuning lowered FiQA and slightly raised SciFact. Later continuation recipes did not recover FiQA while keeping Daecore ranking quality, so the original learned weights remain.
131
 
132
- ### Full pipeline: Gemma v2 + BM25 + Ettin
133
 
134
- This replay measures semantic and lexical search, reciprocal-rank fusion, then
135
- Ettin reranking over **82,719 passages**. The table uses the same **925 queries
136
- with known useful evidence** as Gemma's dense comparison, selected from the
137
- 970-query source panel. A query remains included when this pipeline misses its
138
- answer. Both cards share this CUDA FP16 table; it is not either model's
139
- standalone score.
140
 
141
- **Reference:** 122,231 graded query–passage pairs for these 925 queries, within
142
- the shared 127,665-pair top-50 reference (September 2026). Abstentions are excluded.
 
143
 
144
  | Return depth | Hit | Precision | Known-positive recall | nDCG |
145
  |---|---:|---:|---:|---:|
146
- | 1 | 75.87% | 75.87% | 2.12% | 0.6513 |
147
- | 3 | 89.51% | 75.07% | 5.89% | 0.6423 |
148
- | 5 | 92.76% | 73.71% | 9.35% | 0.6376 |
149
- | 10 | 94.27% | 70.46% | 17.04% | 0.6378 |
150
- | 20 | 95.46% | 64.81% | 29.72% | 0.6491 |
151
- | 50 | 97.62% | 42.40% | 45.98% | 0.6130 |
152
-
153
- **The hybrid candidate ceiling is 97.62% Hit@50 and 45.98% known-positive
154
- recall@50** on this shared answerable subset. These coverage measures describe
155
- the actual 50-candidate pool; reranking cannot change them. Daecore returns a
156
- selected prefix of 3–20, so fixed-depth scores do not evaluate the selector's
157
- choice.
158
-
159
- Grades 2–3 count as useful. At depth 1, both providers use the same **920
160
- queries**; five are omitted because a first result was ungradable. Depths
161
- 3–50 use all 925. Precision pools judged positions. Abstentions remove 52
162
- positions at rank 20 and 93 at rank 50 per provider, without drawing in deeper
163
- passages. Known-positive recall counts judged passages, not every useful
164
- passage in the corpus.
165
-
166
- The [evaluation companion](evaluation/README.md) retains
167
- coverage across all 970 source queries and explains the shared reference,
168
- exclusions and CUDA/Vulkan comparison. The [pipeline overview](https://huggingface.co/Daecore)
169
- explains retrieval and selection.
 
170
 
171
  ## Training data and objective
172
 
 
129
 
130
  These are nDCG@10 scores for reranking fixed candidate pools, not dense Gemma scores. Fine-tuning lowered FiQA and slightly raised SciFact. Later continuation recipes did not recover FiQA while keeping Daecore ranking quality, so the original learned weights remain.
131
 
132
+ ### Full pipeline: the effect of fine-tuning Ettin
133
 
134
+ Only the reranker changes: **upstream Ettin → Daecore Ettin**. Gemma v2,
135
+ BM25 and reciprocal-rank fusion supply the same 50 candidates from
136
+ **82,719 passages**. Each cell reads before → **after**, isolating what
137
+ fine-tuning adds to the ordering of the actual hybrid candidates.
 
 
138
 
139
+ **Reference:** the same **925 known-answerable queries**, frozen before this
140
+ comparison · **122,783 graded query–passage pairs** · October 2026. Retrieval
141
+ misses remain included. Both cards share the same fully fine-tuned endpoint.
142
 
143
  | Return depth | Hit | Precision | Known-positive recall | nDCG |
144
  |---|---:|---:|---:|---:|
145
+ | 1 | 56.69% → **75.95%** | 56.69% → **75.95%** | 1.44% → **2.12%** | 0.4732 → **0.6527** |
146
+ | 3 | 79.03% → **89.41%** | 56.04% → **75.11%** | 4.23% → **5.89%** | 0.4728 → **0.6424** |
147
+ | 5 | 85.62% → **92.76%** | 55.83% → **73.67%** | 6.74% → **9.32%** | 0.4778 → **0.6373** |
148
+ | 10 | 91.14% → **94.27%** | 54.47% → **70.49%** | 12.83% → **16.99%** | 0.4903 → **0.6377** |
149
+ | 20 | 95.03% → **95.46%** | 52.21% → **64.80%** | 23.63% → **29.63%** | 0.5201 → **0.6489** |
150
+ | 50 | 97.62% → **97.62%** | 42.40% → **42.40%** | 45.83% → **45.83%** | 0.5598 → **0.6125** |
151
+
152
+ At three results, useful-passage precision changes from **56.04% to
153
+ 75.11%**; nDCG@10 changes from **0.4903 to 0.6377**.
154
+ At 50, Hit, precision and recall are identical because every candidate is
155
+ returned. nDCG@50 still measures order: **0.5598 → 0.6125**.
156
+ The pool's **97.62% Hit@50** is a retrieval ceiling, not an Ettin accuracy score.
157
+
158
+ All three variants use matched **PyTorch FP16** reranking. Depth 1 uses **919 shared judged
159
+ queries** across the three pipeline variants; depths 3–50 use all 925.
160
+ Abstentions are excluded inside the original cutoff without backfill.
161
+ Grades 2–3 count as useful. Precision pools retained positions; recall counts
162
+ known useful passages.
163
+ The expanded reference and matched runtime distinguish this table from the
164
+ earlier CUDA/Vulkan serving check. Daecore returns 3–20 results; these fixed
165
+ cutoffs do not evaluate its selector.
166
+
167
+ The [evaluation companion](evaluation/README.md) provides
168
+ the paired records, methods and separate serving checks. The
169
+ [pipeline overview](https://huggingface.co/Daecore) explains retrieval and selection.
170
 
171
  ## Training data and objective
172
 
evaluation/README.md CHANGED
@@ -1,6 +1,6 @@
1
  # Ettin reranker evaluation
2
 
3
- These records support the [model card's](../README.md) ranking results and compare Ettin's CUDA and Vulkan serving paths. They contain anonymized relevance grades and scores, without private queries or passage text.
4
 
5
  The [chunking guide](chunking.md#retrieval-model-data) explains the synthetic
6
  source documents and prepared chunks that form the private candidate pools.
@@ -8,8 +8,9 @@ source documents and prepared chunks that form the private candidate pools.
8
  | File | Contents |
9
  |---|---|
10
  | `reranker.json` | Grades and upstream and fine-tuned scores for 970 fixed pools of 50 passages |
 
11
  | `serving.json` | Reviewed top-50 hybrid rankings from CUDA and Vulkan on the same 970 queries |
12
- | `reranker-summary.json`, `serving-summary.json` | The summaries that `metrics.py` reproduces |
13
  | `serving-qualification.json` | Serving checks, public reranking results and tested limits |
14
  | `metrics.py`, `figures.py`, `svg_figures.py` | Metric and figure code |
15
 
@@ -22,7 +23,7 @@ python metrics.py --directory .
22
  python figures.py --output ../figures
23
  ```
24
 
25
- No packages or network access are needed. The outputs match `reranker-summary.json` and `serving-summary.json`; the second command redraws the card's comparison figure. This recomputes saved results, rather than running inference again. The shared pipeline table uses `serving-summary.json`'s `answerable` section; its outer `models` section preserves all-query coverage.
26
 
27
  ## Ranking quality
28
 
@@ -48,24 +49,45 @@ queries from an earlier lexical supplement. Its separate held-out artifacts
48
  are no longer available, so those historical checks cannot be recomputed and
49
  are not part of the comparison reported here.
50
 
51
- ## CUDA and Vulkan
52
-
53
- Both model cards share the CUDA fixed-depth table from `serving.json`,
54
- using **925 queries with known grade-2/3 evidence** in a shared reference
55
- across dense and hybrid retrieval. These queries have **122,231 graded
56
- query–passage pairs**, within a full reference of 127,665 pairs across 970
57
- queries. Membership comes from the reference, so pipeline retrieval misses
58
- remain included. This is a different population and candidate set from
59
- Ettin's historical 873-answerable-pool table.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
60
 
61
- Depths 3–50 use all 925 queries. Depth 1 uses the same 920 for both providers;
62
- five are omitted because a first result was ungradable. Fixed cutoffs do not
63
- evaluate the selector's choice of a 3–20-item prefix.
64
 
65
- At 50, the hybrid pool reaches **97.62% Hit and 45.98% known-positive recall**
66
- on the shared answerable subset. This is candidate coverage against the wider
67
- reference. Ettin's standalone fixed-pool Hit and recall instead reach 100%
68
- by construction on its answerable pools.
69
 
70
  `serving.json` compares CUDA FP16 and Vulkan FP32 on 970 candidate pools retrieved with [Gemma v2](https://huggingface.co/Daecore/embeddinggemma-300m-memory-ft-v2) and BM25. Ettin's learned weights and the score mapping are unchanged. This is a retrieval-pipeline check on fixed inputs; it does not measure a change in Gemma.
71
 
@@ -89,7 +111,7 @@ and final quotation flag. Sixteen quotation repairs changed no grades. These
89
  are model judgments without human review.
90
 
91
  Existing grades and rankings are unchanged. The earlier reference expansion
92
- changed recall and ideal DCG; the card now also conditions quality on the
93
  shared 925 known-answerable queries. Each provider omits 52 abstained positions
94
  at rank 20 and 93 at rank 50, with no backfill. Precision pools retained
95
  positions and nDCG reindexes them. The standalone fixed-pool Ettin comparison
@@ -97,7 +119,7 @@ and public benchmark labels are unchanged.
97
 
98
  ### All-query coverage
99
 
100
- The saved records also preserve all 970 source queries, including 45 with no
101
  known useful evidence. This separate audit view uses the full 127,665-pair
102
  reference; it does not replace the answerable-only card table.
103
 
 
1
  # Ettin reranker evaluation
2
 
3
+ These records support the [model card's](../README.md) standalone and full-pipeline ranking results, plus separate CUDA/Vulkan serving checks. They contain anonymized relevance grades and scores, without private queries or passage text.
4
 
5
  The [chunking guide](chunking.md#retrieval-model-data) explains the synthetic
6
  source documents and prepared chunks that form the private candidate pools.
 
8
  | File | Contents |
9
  |---|---|
10
  | `reranker.json` | Grades and upstream and fine-tuned scores for 970 fixed pools of 50 passages |
11
+ | `pipeline.json` | Matched component changes, including upstream and fine-tuned Ettin on Gemma v2 + BM25 candidates |
12
  | `serving.json` | Reviewed top-50 hybrid rankings from CUDA and Vulkan on the same 970 queries |
13
+ | `*-summary.json` | The summaries that `metrics.py` reproduces |
14
  | `serving-qualification.json` | Serving checks, public reranking results and tested limits |
15
  | `metrics.py`, `figures.py`, `svg_figures.py` | Metric and figure code |
16
 
 
23
  python figures.py --output ../figures
24
  ```
25
 
26
+ No packages or network access are needed. Each output matches its corresponding `*-summary.json`; the second command redraws the card's comparison figure. This recomputes saved results, rather than running inference again. The pipeline card table compares `upstream_ettin` with `finetuned` in `pipeline-summary.json`. The separate serving record preserves both answerable-only and all-query provider results.
27
 
28
  ## Ranking quality
29
 
 
49
  are no longer available, so those historical checks cannot be recomputed and
50
  are not part of the comparison reported here.
51
 
52
+ ## Effect in the retrieval pipeline
53
+
54
+ `pipeline.json` compares upstream and fine-tuned Ettin on identical Gemma v2 +
55
+ BM25 candidate pools (`upstream_ettin` → `finetuned`). A third arm changes
56
+ Gemma alone for its companion card; both comparisons end at the same pipeline.
57
+ At 50, Hit, precision and recall must match when only Ettin changes; nDCG
58
+ still measures how well it orders that pool.
59
+
60
+ The comparison keeps the **925-query cohort frozen before scoring** and uses
61
+ **122,783 graded query–passage pairs** for those queries. It adds 552 grades
62
+ and 10 abstentions to their September reference of 122,231 pairs; old
63
+ grades are unchanged. Two independent GPT-6.1 Sol judges and a fresh blind
64
+ adjudicator resolved the new pairs, with 84 ordinal disagreements.
65
+ The primary assistant reviewed 84 complete pairs, including every final quotation
66
+ flag and abstention plus a stratified sample. These are model judgments,
67
+ without a human reference panel.
68
+
69
+ All three variants were scored anew in **PyTorch FP16**, using Transformers
70
+ 5.2.0, Sentence Transformers 5.5.1, 128 query tokens and 1,024 passage tokens.
71
+ Gemma uses cached FP32 embeddings at 768 dimensions with matched input limits.
72
+ The actual first stage uses the frozen query forms, dense retrieval, BM25,
73
+ reciprocal-rank fusion with constant 60, exact-text deduplication and 50 slots.
74
+ Before substituting upstream Gemma's vectors, replay reproduced all 970 saved
75
+ v2 candidate lists exactly. Ties preserve candidate order.
76
+
77
+ Depth 1 uses **919 common judged queries across all three arms**; depths
78
+ 3–50 retain all 925. Original prefixes are cut before abstentions are removed,
79
+ with no backfill. Retrieval misses remain included. The reference supplies
80
+ known-positive recall and ideal DCG; it is not exhaustive corpus labeling.
81
+ Fixed cutoffs do not test the calibrated selector's choice of 3–20 results.
82
+ The earlier ONNX provider replay retains its own runtime and September
83
+ reference, so its absolute scores should not be substituted into this comparison.
84
 
85
+ ## CUDA and Vulkan
 
 
86
 
87
+ The separate provider check retains the September reference: **122,231 graded
88
+ pairs for 925 known-answerable queries**, within 127,665 pairs across all 970.
89
+ Depth 1 uses 920 common judged queries; depths 3–50 use all 925. It does not
90
+ compare upstream with fine-tuned Ettin.
91
 
92
  `serving.json` compares CUDA FP16 and Vulkan FP32 on 970 candidate pools retrieved with [Gemma v2](https://huggingface.co/Daecore/embeddinggemma-300m-memory-ft-v2) and BM25. Ettin's learned weights and the score mapping are unchanged. This is a retrieval-pipeline check on fixed inputs; it does not measure a change in Gemma.
93
 
 
111
  are model judgments without human review.
112
 
113
  Existing grades and rankings are unchanged. The earlier reference expansion
114
+ changed recall and ideal DCG; this provider table conditions quality on the
115
  shared 925 known-answerable queries. Each provider omits 52 abstained positions
116
  at rank 20 and 93 at rank 50, with no backfill. Precision pools retained
117
  positions and nDCG reindexes them. The standalone fixed-pool Ettin comparison
 
119
 
120
  ### All-query coverage
121
 
122
+ The September serving record also preserves all 970 source queries, including 45 with no
123
  known useful evidence. This separate audit view uses the full 127,665-pair
124
  reference; it does not replace the answerable-only card table.
125
 
evaluation/metrics.py CHANGED
@@ -267,10 +267,22 @@ def summarize_serving(data: dict) -> dict:
267
  'answerable': _summarize_ranked_grades(data, include_selected=True, answerable_only=True)}
268
 
269
 
 
 
 
 
 
 
 
 
 
 
 
 
270
  SUMMARIZERS = {
271
  'classifier': summarize_classifier, 'reranker': summarize_reranker,
272
  'retriever': summarize_retriever, 'dimensions': summarize_dimensions,
273
- 'promotion': summarize_promotion, 'serving': summarize_serving,
274
  }
275
 
276
 
@@ -283,6 +295,7 @@ def summarize_record(name: str, data: dict) -> dict:
283
  'dimensions': {'id', 'dataset', 'ndcg@10'},
284
  'promotion': {'id', 'group', 'reference_grade_counts', 'ranked_grades', 'excluded', 'selected_depth'},
285
  'serving': {'id', 'group', 'reference_grade_counts', 'ranked_grades', 'excluded', 'selected_depth'},
 
286
  }[name]
287
  rows = data['rows']
288
  if not rows or any(set(row) != fields for row in rows):
@@ -336,7 +349,7 @@ def main() -> None:
336
  parser = argparse.ArgumentParser(description=__doc__)
337
  parser.add_argument("--directory", type=Path, default=Path(__file__).resolve().parent)
338
  parser.add_argument("--output", type=Path)
339
- parser.add_argument("--model", choices=("classifier", "reranker", "retriever", "dimensions", "promotion", "serving"), help="Recompute one comparison; default: every record in this package")
340
  args = parser.parse_args()
341
  names = [args.model] if args.model else [
342
  name for name in SUMMARIZERS if (args.directory / f"{name}.json").is_file()
 
267
  'answerable': _summarize_ranked_grades(data, include_selected=True, answerable_only=True)}
268
 
269
 
270
+ def summarize_pipeline(data: dict) -> dict:
271
+ """Matched component changes on a frozen, known-answerable query cohort."""
272
+ if type(data['corpus_passages']) is not int or data['corpus_passages'] <= 0:
273
+ raise ValueError('Pipeline comparison requires a positive corpus size')
274
+ output = _summarize_ranked_grades(data, include_selected=False)
275
+ if any(row['reference_grade_counts']['2'] + row['reference_grade_counts']['3'] == 0
276
+ for row in data['rows']):
277
+ raise ValueError('Pipeline cohort requires known useful evidence for every query')
278
+ return {**output, 'corpus_passages': data['corpus_passages'],
279
+ 'reference_pairs': sum(sum(row['reference_grade_counts'].values()) for row in data['rows'])}
280
+
281
+
282
  SUMMARIZERS = {
283
  'classifier': summarize_classifier, 'reranker': summarize_reranker,
284
  'retriever': summarize_retriever, 'dimensions': summarize_dimensions,
285
+ 'promotion': summarize_promotion, 'serving': summarize_serving, 'pipeline': summarize_pipeline,
286
  }
287
 
288
 
 
295
  'dimensions': {'id', 'dataset', 'ndcg@10'},
296
  'promotion': {'id', 'group', 'reference_grade_counts', 'ranked_grades', 'excluded', 'selected_depth'},
297
  'serving': {'id', 'group', 'reference_grade_counts', 'ranked_grades', 'excluded', 'selected_depth'},
298
+ 'pipeline': {'id', 'group', 'reference_grade_counts', 'ranked_grades', 'excluded'},
299
  }[name]
300
  rows = data['rows']
301
  if not rows or any(set(row) != fields for row in rows):
 
349
  parser = argparse.ArgumentParser(description=__doc__)
350
  parser.add_argument("--directory", type=Path, default=Path(__file__).resolve().parent)
351
  parser.add_argument("--output", type=Path)
352
+ parser.add_argument("--model", choices=tuple(SUMMARIZERS), help="Recompute one comparison; default: every record in this package")
353
  args = parser.parse_args()
354
  names = [args.model] if args.model else [
355
  name for name in SUMMARIZERS if (args.directory / f"{name}.json").is_file()
evaluation/pipeline-summary.json ADDED
@@ -0,0 +1,283 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "queries": 925,
3
+ "models": {
4
+ "upstream_gemma": {
5
+ "1": {
6
+ "scored_queries": 919,
7
+ "excluded_queries": 6,
8
+ "hit": 0.7399347116430903,
9
+ "precision": 0.7399347116430903,
10
+ "macro_precision": 0.7399347116430903,
11
+ "ndcg": 0.6295144826156795,
12
+ "known_positive_recall": 0.02009148065967933,
13
+ "recall_queries": 919,
14
+ "useful": 680,
15
+ "retained": 919,
16
+ "excluded_positions": 0,
17
+ "mean_useful": 0.7399347116430903,
18
+ "mean_retained": 1.0
19
+ },
20
+ "3": {
21
+ "scored_queries": 925,
22
+ "excluded_queries": 0,
23
+ "hit": 0.8681081081081081,
24
+ "precision": 0.7149566473988439,
25
+ "macro_precision": 0.7153153153153153,
26
+ "ndcg": 0.6120552491562434,
27
+ "known_positive_recall": 0.05489711276806684,
28
+ "recall_queries": 925,
29
+ "useful": 1979,
30
+ "retained": 2768,
31
+ "excluded_positions": 7,
32
+ "mean_useful": 2.1394594594594594,
33
+ "mean_retained": 2.9924324324324325
34
+ },
35
+ "5": {
36
+ "scored_queries": 925,
37
+ "excluded_queries": 0,
38
+ "hit": 0.8994594594594595,
39
+ "precision": 0.7009527934170636,
40
+ "macro_precision": 0.7011351351351351,
41
+ "ndcg": 0.60538850988017,
42
+ "known_positive_recall": 0.08526196518888282,
43
+ "recall_queries": 925,
44
+ "useful": 3237,
45
+ "retained": 4618,
46
+ "excluded_positions": 7,
47
+ "mean_useful": 3.4994594594594592,
48
+ "mean_retained": 4.992432432432432
49
+ },
50
+ "10": {
51
+ "scored_queries": 925,
52
+ "excluded_queries": 0,
53
+ "hit": 0.9308108108108109,
54
+ "precision": 0.6519761775852734,
55
+ "macro_precision": 0.6521381381381381,
56
+ "ndcg": 0.5958896253275177,
57
+ "known_positive_recall": 0.15250767729827347,
58
+ "recall_queries": 925,
59
+ "useful": 6021,
60
+ "retained": 9235,
61
+ "excluded_positions": 15,
62
+ "mean_useful": 6.509189189189189,
63
+ "mean_retained": 9.983783783783784
64
+ },
65
+ "20": {
66
+ "scored_queries": 925,
67
+ "excluded_queries": 0,
68
+ "hit": 0.9567567567567568,
69
+ "precision": 0.5685743678596568,
70
+ "macro_precision": 0.5687161650814901,
71
+ "ndcg": 0.5873176954516321,
72
+ "known_positive_recall": 0.2518432399920873,
73
+ "recall_queries": 925,
74
+ "useful": 10501,
75
+ "retained": 18469,
76
+ "excluded_positions": 31,
77
+ "mean_useful": 11.352432432432433,
78
+ "mean_retained": 19.966486486486488
79
+ },
80
+ "50": {
81
+ "scored_queries": 925,
82
+ "excluded_queries": 0,
83
+ "hit": 0.9697297297297297,
84
+ "precision": 0.35454801464376234,
85
+ "macro_precision": 0.35468757089595926,
86
+ "ndcg": 0.5373826502849983,
87
+ "known_positive_recall": 0.36848413838619143,
88
+ "recall_queries": 925,
89
+ "useful": 16367,
90
+ "retained": 46163,
91
+ "excluded_positions": 87,
92
+ "mean_useful": 17.694054054054053,
93
+ "mean_retained": 49.905945945945945
94
+ }
95
+ },
96
+ "upstream_ettin": {
97
+ "1": {
98
+ "scored_queries": 919,
99
+ "excluded_queries": 6,
100
+ "hit": 0.5669205658324266,
101
+ "precision": 0.5669205658324266,
102
+ "macro_precision": 0.5669205658324266,
103
+ "ndcg": 0.4731851391263796,
104
+ "known_positive_recall": 0.01443636631579192,
105
+ "recall_queries": 919,
106
+ "useful": 521,
107
+ "retained": 919,
108
+ "excluded_positions": 0,
109
+ "mean_useful": 0.5669205658324266,
110
+ "mean_retained": 1.0
111
+ },
112
+ "3": {
113
+ "scored_queries": 925,
114
+ "excluded_queries": 0,
115
+ "hit": 0.7902702702702703,
116
+ "precision": 0.5603603603603604,
117
+ "macro_precision": 0.5603603603603603,
118
+ "ndcg": 0.4727968405371722,
119
+ "known_positive_recall": 0.04228504528013372,
120
+ "recall_queries": 925,
121
+ "useful": 1555,
122
+ "retained": 2775,
123
+ "excluded_positions": 0,
124
+ "mean_useful": 1.681081081081081,
125
+ "mean_retained": 3.0
126
+ },
127
+ "5": {
128
+ "scored_queries": 925,
129
+ "excluded_queries": 0,
130
+ "hit": 0.8562162162162162,
131
+ "precision": 0.5582954791261086,
132
+ "macro_precision": 0.5582162162162162,
133
+ "ndcg": 0.47778481822398317,
134
+ "known_positive_recall": 0.06742273722612116,
135
+ "recall_queries": 925,
136
+ "useful": 2581,
137
+ "retained": 4623,
138
+ "excluded_positions": 2,
139
+ "mean_useful": 2.79027027027027,
140
+ "mean_retained": 4.997837837837838
141
+ },
142
+ "10": {
143
+ "scored_queries": 925,
144
+ "excluded_queries": 0,
145
+ "hit": 0.9113513513513514,
146
+ "precision": 0.5447462395844606,
147
+ "macro_precision": 0.5447507507507507,
148
+ "ndcg": 0.49027329580696355,
149
+ "known_positive_recall": 0.12832495703473729,
150
+ "recall_queries": 925,
151
+ "useful": 5034,
152
+ "retained": 9241,
153
+ "excluded_positions": 9,
154
+ "mean_useful": 5.4421621621621625,
155
+ "mean_retained": 9.99027027027027
156
+ },
157
+ "20": {
158
+ "scored_queries": 925,
159
+ "excluded_queries": 0,
160
+ "hit": 0.9502702702702702,
161
+ "precision": 0.5220568335588633,
162
+ "macro_precision": 0.5220833774951422,
163
+ "ndcg": 0.5201039924104931,
164
+ "known_positive_recall": 0.23627621886671868,
165
+ "recall_queries": 925,
166
+ "useful": 9645,
167
+ "retained": 18475,
168
+ "excluded_positions": 25,
169
+ "mean_useful": 10.427027027027027,
170
+ "mean_retained": 19.972972972972972
171
+ },
172
+ "50": {
173
+ "scored_queries": 925,
174
+ "excluded_queries": 0,
175
+ "hit": 0.9762162162162162,
176
+ "precision": 0.424031024546656,
177
+ "macro_precision": 0.42417201849319736,
178
+ "ndcg": 0.5598208131173884,
179
+ "known_positive_recall": 0.4583210768438434,
180
+ "recall_queries": 925,
181
+ "useful": 19572,
182
+ "retained": 46157,
183
+ "excluded_positions": 93,
184
+ "mean_useful": 21.158918918918918,
185
+ "mean_retained": 49.89945945945946
186
+ }
187
+ },
188
+ "finetuned": {
189
+ "1": {
190
+ "scored_queries": 919,
191
+ "excluded_queries": 6,
192
+ "hit": 0.7595212187159956,
193
+ "precision": 0.7595212187159956,
194
+ "macro_precision": 0.7595212187159956,
195
+ "ndcg": 0.6526763044717343,
196
+ "known_positive_recall": 0.021181970033131128,
197
+ "recall_queries": 919,
198
+ "useful": 698,
199
+ "retained": 919,
200
+ "excluded_positions": 0,
201
+ "mean_useful": 0.7595212187159956,
202
+ "mean_retained": 1.0
203
+ },
204
+ "3": {
205
+ "scored_queries": 925,
206
+ "excluded_queries": 0,
207
+ "hit": 0.894054054054054,
208
+ "precision": 0.7510853835021708,
209
+ "macro_precision": 0.7518918918918919,
210
+ "ndcg": 0.6424089559634334,
211
+ "known_positive_recall": 0.058858902936843926,
212
+ "recall_queries": 925,
213
+ "useful": 2076,
214
+ "retained": 2764,
215
+ "excluded_positions": 11,
216
+ "mean_useful": 2.2443243243243245,
217
+ "mean_retained": 2.9881081081081082
218
+ },
219
+ "5": {
220
+ "scored_queries": 925,
221
+ "excluded_queries": 0,
222
+ "hit": 0.9275675675675675,
223
+ "precision": 0.7366710013003901,
224
+ "macro_precision": 0.736954954954955,
225
+ "ndcg": 0.6373095878906122,
226
+ "known_positive_recall": 0.0932034345948595,
227
+ "recall_queries": 925,
228
+ "useful": 3399,
229
+ "retained": 4614,
230
+ "excluded_positions": 11,
231
+ "mean_useful": 3.674594594594595,
232
+ "mean_retained": 4.988108108108108
233
+ },
234
+ "10": {
235
+ "scored_queries": 925,
236
+ "excluded_queries": 0,
237
+ "hit": 0.9427027027027027,
238
+ "precision": 0.7049002601908065,
239
+ "macro_precision": 0.7052921492921493,
240
+ "ndcg": 0.6377369246501603,
241
+ "known_positive_recall": 0.16987439814587393,
242
+ "recall_queries": 925,
243
+ "useful": 6502,
244
+ "retained": 9224,
245
+ "excluded_positions": 26,
246
+ "mean_useful": 7.029189189189189,
247
+ "mean_retained": 9.971891891891891
248
+ },
249
+ "20": {
250
+ "scored_queries": 925,
251
+ "excluded_queries": 0,
252
+ "hit": 0.9545945945945946,
253
+ "precision": 0.6479835212489159,
254
+ "macro_precision": 0.6484765048449259,
255
+ "ndcg": 0.6488792776497584,
256
+ "known_positive_recall": 0.2962743880804353,
257
+ "recall_queries": 925,
258
+ "useful": 11954,
259
+ "retained": 18448,
260
+ "excluded_positions": 52,
261
+ "mean_useful": 12.923243243243244,
262
+ "mean_retained": 19.943783783783783
263
+ },
264
+ "50": {
265
+ "scored_queries": 925,
266
+ "excluded_queries": 0,
267
+ "hit": 0.9762162162162162,
268
+ "precision": 0.424031024546656,
269
+ "macro_precision": 0.42417201849319736,
270
+ "ndcg": 0.612510329000538,
271
+ "known_positive_recall": 0.4583210768438434,
272
+ "recall_queries": 925,
273
+ "useful": 19572,
274
+ "retained": 46157,
275
+ "excluded_positions": 93,
276
+ "mean_useful": 21.158918918918918,
277
+ "mean_retained": 49.89945945945946
278
+ }
279
+ }
280
+ },
281
+ "corpus_passages": 82719,
282
+ "reference_pairs": 122783
283
+ }
evaluation/pipeline.json ADDED
The diff for this file is too large to render. See raw diff
 
publication-manifest.json CHANGED
@@ -1,6 +1,6 @@
1
  {
2
  "schema": "daecore.retrieval-publication-manifest",
3
- "tool_sha256": "46dcb4b17750d5a4ba3d44b1236f33e9e5657673b648ae6cad48fe4ac4f147f8",
4
  "model_id": "ettin-150m-memory-reranker-ft-v1",
5
  "public_repo": "Daecore/ettin-150m-memory-reranker-ft-v1",
6
  "role": "cross-encoder-reranker",
@@ -26,7 +26,7 @@
26
  }
27
  }
28
  },
29
- "staged_at": "2026-09-30T02:12:28+00:00",
30
  "files": {
31
  "config.json": {
32
  "sha256": "a897acffc99f97238bb5bc0e1bccd6b7fe8fdfebe1450d4db3a9ab026cce6795",
@@ -89,8 +89,8 @@
89
  "binding": "packaging record"
90
  },
91
  "evaluation/README.md": {
92
- "sha256": "de60f0b772883cce4b15269b3476a2d30953ea63fa23252675fb755a3d939515",
93
- "size": 7903,
94
  "binding": "packaging record"
95
  },
96
  "evaluation/serving-qualification.json": {
@@ -99,8 +99,8 @@
99
  "binding": "packaging record"
100
  },
101
  "evaluation/metrics.py": {
102
- "sha256": "22536c4437be0b8c678df0dee4c45e31c5413e91f1fecf104750486f3e372d46",
103
- "size": 19120,
104
  "binding": "packaging record"
105
  },
106
  "evaluation/chunking.md": {
@@ -118,6 +118,11 @@
118
  "size": 14521,
119
  "binding": "packaging record"
120
  },
 
 
 
 
 
121
  "evaluation/reranker.json": {
122
  "sha256": "20b5760a462c3c2d2281607f2c5045d3f60210123a0c11c8b2d92580c3cccd3e",
123
  "size": 1145459,
@@ -128,6 +133,11 @@
128
  "size": 1184937,
129
  "binding": "packaging record"
130
  },
 
 
 
 
 
131
  "evaluation/figures.py": {
132
  "sha256": "f13945b350e007deb18a2eafbe1c5708a71b9df62ee12b6833d07f3e2a06094e",
133
  "size": 5490,
@@ -144,10 +154,10 @@
144
  "binding": "packaging record"
145
  },
146
  "README.md": {
147
- "sha256": "7f08c4025cd386a64b904f0bb038af20df6c2bcf56aef1401dae4293a3fb22cb",
148
- "size": 13036,
149
  "binding": "model card with upload-relative links",
150
- "source_sha256": "accc91b7d2afd4d6097dcc8f6038ddddfdf1ca7faf9fb09e09bdf0ec5d90e9cb"
151
  }
152
  }
153
  }
 
1
  {
2
  "schema": "daecore.retrieval-publication-manifest",
3
+ "tool_sha256": "60165b669f2cce17d8830fe9b712fb60a8e82280c54b224d71a1ce8fd82142ac",
4
  "model_id": "ettin-150m-memory-reranker-ft-v1",
5
  "public_repo": "Daecore/ettin-150m-memory-reranker-ft-v1",
6
  "role": "cross-encoder-reranker",
 
26
  }
27
  }
28
  },
29
+ "staged_at": "2026-10-01T16:23:47+00:00",
30
  "files": {
31
  "config.json": {
32
  "sha256": "a897acffc99f97238bb5bc0e1bccd6b7fe8fdfebe1450d4db3a9ab026cce6795",
 
89
  "binding": "packaging record"
90
  },
91
  "evaluation/README.md": {
92
+ "sha256": "5399753e1e6e28be3997c1ebf4ee017bfda96d61950bebd89f2b175b29512cad",
93
+ "size": 9377,
94
  "binding": "packaging record"
95
  },
96
  "evaluation/serving-qualification.json": {
 
99
  "binding": "packaging record"
100
  },
101
  "evaluation/metrics.py": {
102
+ "sha256": "8123f75da6edb158db74c62c1c65b6eefbc2713572409ac8a9bfb5424497aaf1",
103
+ "size": 19927,
104
  "binding": "packaging record"
105
  },
106
  "evaluation/chunking.md": {
 
118
  "size": 14521,
119
  "binding": "packaging record"
120
  },
121
+ "evaluation/pipeline-summary.json": {
122
+ "sha256": "5f28ae25b5db2f4c0a8d11891df7117f95f644063092d58ba45e27a59747e3a2",
123
+ "size": 9373,
124
+ "binding": "packaging record"
125
+ },
126
  "evaluation/reranker.json": {
127
  "sha256": "20b5760a462c3c2d2281607f2c5045d3f60210123a0c11c8b2d92580c3cccd3e",
128
  "size": 1145459,
 
133
  "size": 1184937,
134
  "binding": "packaging record"
135
  },
136
+ "evaluation/pipeline.json": {
137
+ "sha256": "8f140901a6b151d9db97fd109a3774bf1c046b849d9a01c874cb0efdea297300",
138
+ "size": 1614738,
139
+ "binding": "packaging record"
140
+ },
141
  "evaluation/figures.py": {
142
  "sha256": "f13945b350e007deb18a2eafbe1c5708a71b9df62ee12b6833d07f3e2a06094e",
143
  "size": 5490,
 
154
  "binding": "packaging record"
155
  },
156
  "README.md": {
157
+ "sha256": "c8a24bd6c07a34e68d33ba94e476dd75634a72660b21e575435865c25f0002f4",
158
+ "size": 13443,
159
  "binding": "model card with upload-relative links",
160
+ "source_sha256": "d3c45681e0ba1ef22b0bfe9eef372b298e1223082e790e25ace134d6a6c760e3"
161
  }
162
  }
163
  }