Update ettin-150m-memory-reranker-ft-v1 documentation
Browse files- README.md +33 -33
- evaluation/README.md +43 -21
- evaluation/metrics.py +15 -2
- evaluation/pipeline-summary.json +283 -0
- evaluation/pipeline.json +0 -0
- publication-manifest.json +19 -9
README.md
CHANGED
|
@@ -129,44 +129,44 @@ Both models rerank the same 50 candidates that upstream Gemma retrieved for each
|
|
| 129 |
|
| 130 |
These are nDCG@10 scores for reranking fixed candidate pools, not dense Gemma scores. Fine-tuning lowered FiQA and slightly raised SciFact. Later continuation recipes did not recover FiQA while keeping Daecore ranking quality, so the original learned weights remain.
|
| 131 |
|
| 132 |
-
### Full pipeline:
|
| 133 |
|
| 134 |
-
|
| 135 |
-
|
| 136 |
-
|
| 137 |
-
|
| 138 |
-
answer. Both cards share this CUDA FP16 table; it is not either model's
|
| 139 |
-
standalone score.
|
| 140 |
|
| 141 |
-
**Reference:**
|
| 142 |
-
|
|
|
|
| 143 |
|
| 144 |
| Return depth | Hit | Precision | Known-positive recall | nDCG |
|
| 145 |
|---|---:|---:|---:|---:|
|
| 146 |
-
| 1 | 75.
|
| 147 |
-
| 3 | 89.
|
| 148 |
-
| 5 | 92.76% | 73.
|
| 149 |
-
| 10 | 94.27% | 70.
|
| 150 |
-
| 20 | 95.46% | 64.
|
| 151 |
-
| 50 | 97.62% | 42.40% | 45.
|
| 152 |
-
|
| 153 |
-
|
| 154 |
-
|
| 155 |
-
|
| 156 |
-
|
| 157 |
-
|
| 158 |
-
|
| 159 |
-
|
| 160 |
-
queries**
|
| 161 |
-
|
| 162 |
-
|
| 163 |
-
|
| 164 |
-
|
| 165 |
-
|
| 166 |
-
|
| 167 |
-
|
| 168 |
-
|
| 169 |
-
|
|
|
|
| 170 |
|
| 171 |
## Training data and objective
|
| 172 |
|
|
|
|
| 129 |
|
| 130 |
These are nDCG@10 scores for reranking fixed candidate pools, not dense Gemma scores. Fine-tuning lowered FiQA and slightly raised SciFact. Later continuation recipes did not recover FiQA while keeping Daecore ranking quality, so the original learned weights remain.
|
| 131 |
|
| 132 |
+
### Full pipeline: the effect of fine-tuning Ettin
|
| 133 |
|
| 134 |
+
Only the reranker changes: **upstream Ettin → Daecore Ettin**. Gemma v2,
|
| 135 |
+
BM25 and reciprocal-rank fusion supply the same 50 candidates from
|
| 136 |
+
**82,719 passages**. Each cell reads before → **after**, isolating what
|
| 137 |
+
fine-tuning adds to the ordering of the actual hybrid candidates.
|
|
|
|
|
|
|
| 138 |
|
| 139 |
+
**Reference:** the same **925 known-answerable queries**, frozen before this
|
| 140 |
+
comparison · **122,783 graded query–passage pairs** · October 2026. Retrieval
|
| 141 |
+
misses remain included. Both cards share the same fully fine-tuned endpoint.
|
| 142 |
|
| 143 |
| Return depth | Hit | Precision | Known-positive recall | nDCG |
|
| 144 |
|---|---:|---:|---:|---:|
|
| 145 |
+
| 1 | 56.69% → **75.95%** | 56.69% → **75.95%** | 1.44% → **2.12%** | 0.4732 → **0.6527** |
|
| 146 |
+
| 3 | 79.03% → **89.41%** | 56.04% → **75.11%** | 4.23% → **5.89%** | 0.4728 → **0.6424** |
|
| 147 |
+
| 5 | 85.62% → **92.76%** | 55.83% → **73.67%** | 6.74% → **9.32%** | 0.4778 → **0.6373** |
|
| 148 |
+
| 10 | 91.14% → **94.27%** | 54.47% → **70.49%** | 12.83% → **16.99%** | 0.4903 → **0.6377** |
|
| 149 |
+
| 20 | 95.03% → **95.46%** | 52.21% → **64.80%** | 23.63% → **29.63%** | 0.5201 → **0.6489** |
|
| 150 |
+
| 50 | 97.62% → **97.62%** | 42.40% → **42.40%** | 45.83% → **45.83%** | 0.5598 → **0.6125** |
|
| 151 |
+
|
| 152 |
+
At three results, useful-passage precision changes from **56.04% to
|
| 153 |
+
75.11%**; nDCG@10 changes from **0.4903 to 0.6377**.
|
| 154 |
+
At 50, Hit, precision and recall are identical because every candidate is
|
| 155 |
+
returned. nDCG@50 still measures order: **0.5598 → 0.6125**.
|
| 156 |
+
The pool's **97.62% Hit@50** is a retrieval ceiling, not an Ettin accuracy score.
|
| 157 |
+
|
| 158 |
+
All three variants use matched **PyTorch FP16** reranking. Depth 1 uses **919 shared judged
|
| 159 |
+
queries** across the three pipeline variants; depths 3–50 use all 925.
|
| 160 |
+
Abstentions are excluded inside the original cutoff without backfill.
|
| 161 |
+
Grades 2–3 count as useful. Precision pools retained positions; recall counts
|
| 162 |
+
known useful passages.
|
| 163 |
+
The expanded reference and matched runtime distinguish this table from the
|
| 164 |
+
earlier CUDA/Vulkan serving check. Daecore returns 3–20 results; these fixed
|
| 165 |
+
cutoffs do not evaluate its selector.
|
| 166 |
+
|
| 167 |
+
The [evaluation companion](evaluation/README.md) provides
|
| 168 |
+
the paired records, methods and separate serving checks. The
|
| 169 |
+
[pipeline overview](https://huggingface.co/Daecore) explains retrieval and selection.
|
| 170 |
|
| 171 |
## Training data and objective
|
| 172 |
|
evaluation/README.md
CHANGED
|
@@ -1,6 +1,6 @@
|
|
| 1 |
# Ettin reranker evaluation
|
| 2 |
|
| 3 |
-
These records support the [model card's](../README.md)
|
| 4 |
|
| 5 |
The [chunking guide](chunking.md#retrieval-model-data) explains the synthetic
|
| 6 |
source documents and prepared chunks that form the private candidate pools.
|
|
@@ -8,8 +8,9 @@ source documents and prepared chunks that form the private candidate pools.
|
|
| 8 |
| File | Contents |
|
| 9 |
|---|---|
|
| 10 |
| `reranker.json` | Grades and upstream and fine-tuned scores for 970 fixed pools of 50 passages |
|
|
|
|
| 11 |
| `serving.json` | Reviewed top-50 hybrid rankings from CUDA and Vulkan on the same 970 queries |
|
| 12 |
-
| `
|
| 13 |
| `serving-qualification.json` | Serving checks, public reranking results and tested limits |
|
| 14 |
| `metrics.py`, `figures.py`, `svg_figures.py` | Metric and figure code |
|
| 15 |
|
|
@@ -22,7 +23,7 @@ python metrics.py --directory .
|
|
| 22 |
python figures.py --output ../figures
|
| 23 |
```
|
| 24 |
|
| 25 |
-
No packages or network access are needed.
|
| 26 |
|
| 27 |
## Ranking quality
|
| 28 |
|
|
@@ -48,24 +49,45 @@ queries from an earlier lexical supplement. Its separate held-out artifacts
|
|
| 48 |
are no longer available, so those historical checks cannot be recomputed and
|
| 49 |
are not part of the comparison reported here.
|
| 50 |
|
| 51 |
-
##
|
| 52 |
-
|
| 53 |
-
|
| 54 |
-
|
| 55 |
-
|
| 56 |
-
|
| 57 |
-
|
| 58 |
-
|
| 59 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 60 |
|
| 61 |
-
|
| 62 |
-
five are omitted because a first result was ungradable. Fixed cutoffs do not
|
| 63 |
-
evaluate the selector's choice of a 3–20-item prefix.
|
| 64 |
|
| 65 |
-
|
| 66 |
-
|
| 67 |
-
|
| 68 |
-
|
| 69 |
|
| 70 |
`serving.json` compares CUDA FP16 and Vulkan FP32 on 970 candidate pools retrieved with [Gemma v2](https://huggingface.co/Daecore/embeddinggemma-300m-memory-ft-v2) and BM25. Ettin's learned weights and the score mapping are unchanged. This is a retrieval-pipeline check on fixed inputs; it does not measure a change in Gemma.
|
| 71 |
|
|
@@ -89,7 +111,7 @@ and final quotation flag. Sixteen quotation repairs changed no grades. These
|
|
| 89 |
are model judgments without human review.
|
| 90 |
|
| 91 |
Existing grades and rankings are unchanged. The earlier reference expansion
|
| 92 |
-
changed recall and ideal DCG;
|
| 93 |
shared 925 known-answerable queries. Each provider omits 52 abstained positions
|
| 94 |
at rank 20 and 93 at rank 50, with no backfill. Precision pools retained
|
| 95 |
positions and nDCG reindexes them. The standalone fixed-pool Ettin comparison
|
|
@@ -97,7 +119,7 @@ and public benchmark labels are unchanged.
|
|
| 97 |
|
| 98 |
### All-query coverage
|
| 99 |
|
| 100 |
-
The
|
| 101 |
known useful evidence. This separate audit view uses the full 127,665-pair
|
| 102 |
reference; it does not replace the answerable-only card table.
|
| 103 |
|
|
|
|
| 1 |
# Ettin reranker evaluation
|
| 2 |
|
| 3 |
+
These records support the [model card's](../README.md) standalone and full-pipeline ranking results, plus separate CUDA/Vulkan serving checks. They contain anonymized relevance grades and scores, without private queries or passage text.
|
| 4 |
|
| 5 |
The [chunking guide](chunking.md#retrieval-model-data) explains the synthetic
|
| 6 |
source documents and prepared chunks that form the private candidate pools.
|
|
|
|
| 8 |
| File | Contents |
|
| 9 |
|---|---|
|
| 10 |
| `reranker.json` | Grades and upstream and fine-tuned scores for 970 fixed pools of 50 passages |
|
| 11 |
+
| `pipeline.json` | Matched component changes, including upstream and fine-tuned Ettin on Gemma v2 + BM25 candidates |
|
| 12 |
| `serving.json` | Reviewed top-50 hybrid rankings from CUDA and Vulkan on the same 970 queries |
|
| 13 |
+
| `*-summary.json` | The summaries that `metrics.py` reproduces |
|
| 14 |
| `serving-qualification.json` | Serving checks, public reranking results and tested limits |
|
| 15 |
| `metrics.py`, `figures.py`, `svg_figures.py` | Metric and figure code |
|
| 16 |
|
|
|
|
| 23 |
python figures.py --output ../figures
|
| 24 |
```
|
| 25 |
|
| 26 |
+
No packages or network access are needed. Each output matches its corresponding `*-summary.json`; the second command redraws the card's comparison figure. This recomputes saved results, rather than running inference again. The pipeline card table compares `upstream_ettin` with `finetuned` in `pipeline-summary.json`. The separate serving record preserves both answerable-only and all-query provider results.
|
| 27 |
|
| 28 |
## Ranking quality
|
| 29 |
|
|
|
|
| 49 |
are no longer available, so those historical checks cannot be recomputed and
|
| 50 |
are not part of the comparison reported here.
|
| 51 |
|
| 52 |
+
## Effect in the retrieval pipeline
|
| 53 |
+
|
| 54 |
+
`pipeline.json` compares upstream and fine-tuned Ettin on identical Gemma v2 +
|
| 55 |
+
BM25 candidate pools (`upstream_ettin` → `finetuned`). A third arm changes
|
| 56 |
+
Gemma alone for its companion card; both comparisons end at the same pipeline.
|
| 57 |
+
At 50, Hit, precision and recall must match when only Ettin changes; nDCG
|
| 58 |
+
still measures how well it orders that pool.
|
| 59 |
+
|
| 60 |
+
The comparison keeps the **925-query cohort frozen before scoring** and uses
|
| 61 |
+
**122,783 graded query–passage pairs** for those queries. It adds 552 grades
|
| 62 |
+
and 10 abstentions to their September reference of 122,231 pairs; old
|
| 63 |
+
grades are unchanged. Two independent GPT-6.1 Sol judges and a fresh blind
|
| 64 |
+
adjudicator resolved the new pairs, with 84 ordinal disagreements.
|
| 65 |
+
The primary assistant reviewed 84 complete pairs, including every final quotation
|
| 66 |
+
flag and abstention plus a stratified sample. These are model judgments,
|
| 67 |
+
without a human reference panel.
|
| 68 |
+
|
| 69 |
+
All three variants were scored anew in **PyTorch FP16**, using Transformers
|
| 70 |
+
5.2.0, Sentence Transformers 5.5.1, 128 query tokens and 1,024 passage tokens.
|
| 71 |
+
Gemma uses cached FP32 embeddings at 768 dimensions with matched input limits.
|
| 72 |
+
The actual first stage uses the frozen query forms, dense retrieval, BM25,
|
| 73 |
+
reciprocal-rank fusion with constant 60, exact-text deduplication and 50 slots.
|
| 74 |
+
Before substituting upstream Gemma's vectors, replay reproduced all 970 saved
|
| 75 |
+
v2 candidate lists exactly. Ties preserve candidate order.
|
| 76 |
+
|
| 77 |
+
Depth 1 uses **919 common judged queries across all three arms**; depths
|
| 78 |
+
3–50 retain all 925. Original prefixes are cut before abstentions are removed,
|
| 79 |
+
with no backfill. Retrieval misses remain included. The reference supplies
|
| 80 |
+
known-positive recall and ideal DCG; it is not exhaustive corpus labeling.
|
| 81 |
+
Fixed cutoffs do not test the calibrated selector's choice of 3–20 results.
|
| 82 |
+
The earlier ONNX provider replay retains its own runtime and September
|
| 83 |
+
reference, so its absolute scores should not be substituted into this comparison.
|
| 84 |
|
| 85 |
+
## CUDA and Vulkan
|
|
|
|
|
|
|
| 86 |
|
| 87 |
+
The separate provider check retains the September reference: **122,231 graded
|
| 88 |
+
pairs for 925 known-answerable queries**, within 127,665 pairs across all 970.
|
| 89 |
+
Depth 1 uses 920 common judged queries; depths 3–50 use all 925. It does not
|
| 90 |
+
compare upstream with fine-tuned Ettin.
|
| 91 |
|
| 92 |
`serving.json` compares CUDA FP16 and Vulkan FP32 on 970 candidate pools retrieved with [Gemma v2](https://huggingface.co/Daecore/embeddinggemma-300m-memory-ft-v2) and BM25. Ettin's learned weights and the score mapping are unchanged. This is a retrieval-pipeline check on fixed inputs; it does not measure a change in Gemma.
|
| 93 |
|
|
|
|
| 111 |
are model judgments without human review.
|
| 112 |
|
| 113 |
Existing grades and rankings are unchanged. The earlier reference expansion
|
| 114 |
+
changed recall and ideal DCG; this provider table conditions quality on the
|
| 115 |
shared 925 known-answerable queries. Each provider omits 52 abstained positions
|
| 116 |
at rank 20 and 93 at rank 50, with no backfill. Precision pools retained
|
| 117 |
positions and nDCG reindexes them. The standalone fixed-pool Ettin comparison
|
|
|
|
| 119 |
|
| 120 |
### All-query coverage
|
| 121 |
|
| 122 |
+
The September serving record also preserves all 970 source queries, including 45 with no
|
| 123 |
known useful evidence. This separate audit view uses the full 127,665-pair
|
| 124 |
reference; it does not replace the answerable-only card table.
|
| 125 |
|
evaluation/metrics.py
CHANGED
|
@@ -267,10 +267,22 @@ def summarize_serving(data: dict) -> dict:
|
|
| 267 |
'answerable': _summarize_ranked_grades(data, include_selected=True, answerable_only=True)}
|
| 268 |
|
| 269 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 270 |
SUMMARIZERS = {
|
| 271 |
'classifier': summarize_classifier, 'reranker': summarize_reranker,
|
| 272 |
'retriever': summarize_retriever, 'dimensions': summarize_dimensions,
|
| 273 |
-
'promotion': summarize_promotion, 'serving': summarize_serving,
|
| 274 |
}
|
| 275 |
|
| 276 |
|
|
@@ -283,6 +295,7 @@ def summarize_record(name: str, data: dict) -> dict:
|
|
| 283 |
'dimensions': {'id', 'dataset', 'ndcg@10'},
|
| 284 |
'promotion': {'id', 'group', 'reference_grade_counts', 'ranked_grades', 'excluded', 'selected_depth'},
|
| 285 |
'serving': {'id', 'group', 'reference_grade_counts', 'ranked_grades', 'excluded', 'selected_depth'},
|
|
|
|
| 286 |
}[name]
|
| 287 |
rows = data['rows']
|
| 288 |
if not rows or any(set(row) != fields for row in rows):
|
|
@@ -336,7 +349,7 @@ def main() -> None:
|
|
| 336 |
parser = argparse.ArgumentParser(description=__doc__)
|
| 337 |
parser.add_argument("--directory", type=Path, default=Path(__file__).resolve().parent)
|
| 338 |
parser.add_argument("--output", type=Path)
|
| 339 |
-
parser.add_argument("--model", choices=(
|
| 340 |
args = parser.parse_args()
|
| 341 |
names = [args.model] if args.model else [
|
| 342 |
name for name in SUMMARIZERS if (args.directory / f"{name}.json").is_file()
|
|
|
|
| 267 |
'answerable': _summarize_ranked_grades(data, include_selected=True, answerable_only=True)}
|
| 268 |
|
| 269 |
|
| 270 |
+
def summarize_pipeline(data: dict) -> dict:
|
| 271 |
+
"""Matched component changes on a frozen, known-answerable query cohort."""
|
| 272 |
+
if type(data['corpus_passages']) is not int or data['corpus_passages'] <= 0:
|
| 273 |
+
raise ValueError('Pipeline comparison requires a positive corpus size')
|
| 274 |
+
output = _summarize_ranked_grades(data, include_selected=False)
|
| 275 |
+
if any(row['reference_grade_counts']['2'] + row['reference_grade_counts']['3'] == 0
|
| 276 |
+
for row in data['rows']):
|
| 277 |
+
raise ValueError('Pipeline cohort requires known useful evidence for every query')
|
| 278 |
+
return {**output, 'corpus_passages': data['corpus_passages'],
|
| 279 |
+
'reference_pairs': sum(sum(row['reference_grade_counts'].values()) for row in data['rows'])}
|
| 280 |
+
|
| 281 |
+
|
| 282 |
SUMMARIZERS = {
|
| 283 |
'classifier': summarize_classifier, 'reranker': summarize_reranker,
|
| 284 |
'retriever': summarize_retriever, 'dimensions': summarize_dimensions,
|
| 285 |
+
'promotion': summarize_promotion, 'serving': summarize_serving, 'pipeline': summarize_pipeline,
|
| 286 |
}
|
| 287 |
|
| 288 |
|
|
|
|
| 295 |
'dimensions': {'id', 'dataset', 'ndcg@10'},
|
| 296 |
'promotion': {'id', 'group', 'reference_grade_counts', 'ranked_grades', 'excluded', 'selected_depth'},
|
| 297 |
'serving': {'id', 'group', 'reference_grade_counts', 'ranked_grades', 'excluded', 'selected_depth'},
|
| 298 |
+
'pipeline': {'id', 'group', 'reference_grade_counts', 'ranked_grades', 'excluded'},
|
| 299 |
}[name]
|
| 300 |
rows = data['rows']
|
| 301 |
if not rows or any(set(row) != fields for row in rows):
|
|
|
|
| 349 |
parser = argparse.ArgumentParser(description=__doc__)
|
| 350 |
parser.add_argument("--directory", type=Path, default=Path(__file__).resolve().parent)
|
| 351 |
parser.add_argument("--output", type=Path)
|
| 352 |
+
parser.add_argument("--model", choices=tuple(SUMMARIZERS), help="Recompute one comparison; default: every record in this package")
|
| 353 |
args = parser.parse_args()
|
| 354 |
names = [args.model] if args.model else [
|
| 355 |
name for name in SUMMARIZERS if (args.directory / f"{name}.json").is_file()
|
evaluation/pipeline-summary.json
ADDED
|
@@ -0,0 +1,283 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"queries": 925,
|
| 3 |
+
"models": {
|
| 4 |
+
"upstream_gemma": {
|
| 5 |
+
"1": {
|
| 6 |
+
"scored_queries": 919,
|
| 7 |
+
"excluded_queries": 6,
|
| 8 |
+
"hit": 0.7399347116430903,
|
| 9 |
+
"precision": 0.7399347116430903,
|
| 10 |
+
"macro_precision": 0.7399347116430903,
|
| 11 |
+
"ndcg": 0.6295144826156795,
|
| 12 |
+
"known_positive_recall": 0.02009148065967933,
|
| 13 |
+
"recall_queries": 919,
|
| 14 |
+
"useful": 680,
|
| 15 |
+
"retained": 919,
|
| 16 |
+
"excluded_positions": 0,
|
| 17 |
+
"mean_useful": 0.7399347116430903,
|
| 18 |
+
"mean_retained": 1.0
|
| 19 |
+
},
|
| 20 |
+
"3": {
|
| 21 |
+
"scored_queries": 925,
|
| 22 |
+
"excluded_queries": 0,
|
| 23 |
+
"hit": 0.8681081081081081,
|
| 24 |
+
"precision": 0.7149566473988439,
|
| 25 |
+
"macro_precision": 0.7153153153153153,
|
| 26 |
+
"ndcg": 0.6120552491562434,
|
| 27 |
+
"known_positive_recall": 0.05489711276806684,
|
| 28 |
+
"recall_queries": 925,
|
| 29 |
+
"useful": 1979,
|
| 30 |
+
"retained": 2768,
|
| 31 |
+
"excluded_positions": 7,
|
| 32 |
+
"mean_useful": 2.1394594594594594,
|
| 33 |
+
"mean_retained": 2.9924324324324325
|
| 34 |
+
},
|
| 35 |
+
"5": {
|
| 36 |
+
"scored_queries": 925,
|
| 37 |
+
"excluded_queries": 0,
|
| 38 |
+
"hit": 0.8994594594594595,
|
| 39 |
+
"precision": 0.7009527934170636,
|
| 40 |
+
"macro_precision": 0.7011351351351351,
|
| 41 |
+
"ndcg": 0.60538850988017,
|
| 42 |
+
"known_positive_recall": 0.08526196518888282,
|
| 43 |
+
"recall_queries": 925,
|
| 44 |
+
"useful": 3237,
|
| 45 |
+
"retained": 4618,
|
| 46 |
+
"excluded_positions": 7,
|
| 47 |
+
"mean_useful": 3.4994594594594592,
|
| 48 |
+
"mean_retained": 4.992432432432432
|
| 49 |
+
},
|
| 50 |
+
"10": {
|
| 51 |
+
"scored_queries": 925,
|
| 52 |
+
"excluded_queries": 0,
|
| 53 |
+
"hit": 0.9308108108108109,
|
| 54 |
+
"precision": 0.6519761775852734,
|
| 55 |
+
"macro_precision": 0.6521381381381381,
|
| 56 |
+
"ndcg": 0.5958896253275177,
|
| 57 |
+
"known_positive_recall": 0.15250767729827347,
|
| 58 |
+
"recall_queries": 925,
|
| 59 |
+
"useful": 6021,
|
| 60 |
+
"retained": 9235,
|
| 61 |
+
"excluded_positions": 15,
|
| 62 |
+
"mean_useful": 6.509189189189189,
|
| 63 |
+
"mean_retained": 9.983783783783784
|
| 64 |
+
},
|
| 65 |
+
"20": {
|
| 66 |
+
"scored_queries": 925,
|
| 67 |
+
"excluded_queries": 0,
|
| 68 |
+
"hit": 0.9567567567567568,
|
| 69 |
+
"precision": 0.5685743678596568,
|
| 70 |
+
"macro_precision": 0.5687161650814901,
|
| 71 |
+
"ndcg": 0.5873176954516321,
|
| 72 |
+
"known_positive_recall": 0.2518432399920873,
|
| 73 |
+
"recall_queries": 925,
|
| 74 |
+
"useful": 10501,
|
| 75 |
+
"retained": 18469,
|
| 76 |
+
"excluded_positions": 31,
|
| 77 |
+
"mean_useful": 11.352432432432433,
|
| 78 |
+
"mean_retained": 19.966486486486488
|
| 79 |
+
},
|
| 80 |
+
"50": {
|
| 81 |
+
"scored_queries": 925,
|
| 82 |
+
"excluded_queries": 0,
|
| 83 |
+
"hit": 0.9697297297297297,
|
| 84 |
+
"precision": 0.35454801464376234,
|
| 85 |
+
"macro_precision": 0.35468757089595926,
|
| 86 |
+
"ndcg": 0.5373826502849983,
|
| 87 |
+
"known_positive_recall": 0.36848413838619143,
|
| 88 |
+
"recall_queries": 925,
|
| 89 |
+
"useful": 16367,
|
| 90 |
+
"retained": 46163,
|
| 91 |
+
"excluded_positions": 87,
|
| 92 |
+
"mean_useful": 17.694054054054053,
|
| 93 |
+
"mean_retained": 49.905945945945945
|
| 94 |
+
}
|
| 95 |
+
},
|
| 96 |
+
"upstream_ettin": {
|
| 97 |
+
"1": {
|
| 98 |
+
"scored_queries": 919,
|
| 99 |
+
"excluded_queries": 6,
|
| 100 |
+
"hit": 0.5669205658324266,
|
| 101 |
+
"precision": 0.5669205658324266,
|
| 102 |
+
"macro_precision": 0.5669205658324266,
|
| 103 |
+
"ndcg": 0.4731851391263796,
|
| 104 |
+
"known_positive_recall": 0.01443636631579192,
|
| 105 |
+
"recall_queries": 919,
|
| 106 |
+
"useful": 521,
|
| 107 |
+
"retained": 919,
|
| 108 |
+
"excluded_positions": 0,
|
| 109 |
+
"mean_useful": 0.5669205658324266,
|
| 110 |
+
"mean_retained": 1.0
|
| 111 |
+
},
|
| 112 |
+
"3": {
|
| 113 |
+
"scored_queries": 925,
|
| 114 |
+
"excluded_queries": 0,
|
| 115 |
+
"hit": 0.7902702702702703,
|
| 116 |
+
"precision": 0.5603603603603604,
|
| 117 |
+
"macro_precision": 0.5603603603603603,
|
| 118 |
+
"ndcg": 0.4727968405371722,
|
| 119 |
+
"known_positive_recall": 0.04228504528013372,
|
| 120 |
+
"recall_queries": 925,
|
| 121 |
+
"useful": 1555,
|
| 122 |
+
"retained": 2775,
|
| 123 |
+
"excluded_positions": 0,
|
| 124 |
+
"mean_useful": 1.681081081081081,
|
| 125 |
+
"mean_retained": 3.0
|
| 126 |
+
},
|
| 127 |
+
"5": {
|
| 128 |
+
"scored_queries": 925,
|
| 129 |
+
"excluded_queries": 0,
|
| 130 |
+
"hit": 0.8562162162162162,
|
| 131 |
+
"precision": 0.5582954791261086,
|
| 132 |
+
"macro_precision": 0.5582162162162162,
|
| 133 |
+
"ndcg": 0.47778481822398317,
|
| 134 |
+
"known_positive_recall": 0.06742273722612116,
|
| 135 |
+
"recall_queries": 925,
|
| 136 |
+
"useful": 2581,
|
| 137 |
+
"retained": 4623,
|
| 138 |
+
"excluded_positions": 2,
|
| 139 |
+
"mean_useful": 2.79027027027027,
|
| 140 |
+
"mean_retained": 4.997837837837838
|
| 141 |
+
},
|
| 142 |
+
"10": {
|
| 143 |
+
"scored_queries": 925,
|
| 144 |
+
"excluded_queries": 0,
|
| 145 |
+
"hit": 0.9113513513513514,
|
| 146 |
+
"precision": 0.5447462395844606,
|
| 147 |
+
"macro_precision": 0.5447507507507507,
|
| 148 |
+
"ndcg": 0.49027329580696355,
|
| 149 |
+
"known_positive_recall": 0.12832495703473729,
|
| 150 |
+
"recall_queries": 925,
|
| 151 |
+
"useful": 5034,
|
| 152 |
+
"retained": 9241,
|
| 153 |
+
"excluded_positions": 9,
|
| 154 |
+
"mean_useful": 5.4421621621621625,
|
| 155 |
+
"mean_retained": 9.99027027027027
|
| 156 |
+
},
|
| 157 |
+
"20": {
|
| 158 |
+
"scored_queries": 925,
|
| 159 |
+
"excluded_queries": 0,
|
| 160 |
+
"hit": 0.9502702702702702,
|
| 161 |
+
"precision": 0.5220568335588633,
|
| 162 |
+
"macro_precision": 0.5220833774951422,
|
| 163 |
+
"ndcg": 0.5201039924104931,
|
| 164 |
+
"known_positive_recall": 0.23627621886671868,
|
| 165 |
+
"recall_queries": 925,
|
| 166 |
+
"useful": 9645,
|
| 167 |
+
"retained": 18475,
|
| 168 |
+
"excluded_positions": 25,
|
| 169 |
+
"mean_useful": 10.427027027027027,
|
| 170 |
+
"mean_retained": 19.972972972972972
|
| 171 |
+
},
|
| 172 |
+
"50": {
|
| 173 |
+
"scored_queries": 925,
|
| 174 |
+
"excluded_queries": 0,
|
| 175 |
+
"hit": 0.9762162162162162,
|
| 176 |
+
"precision": 0.424031024546656,
|
| 177 |
+
"macro_precision": 0.42417201849319736,
|
| 178 |
+
"ndcg": 0.5598208131173884,
|
| 179 |
+
"known_positive_recall": 0.4583210768438434,
|
| 180 |
+
"recall_queries": 925,
|
| 181 |
+
"useful": 19572,
|
| 182 |
+
"retained": 46157,
|
| 183 |
+
"excluded_positions": 93,
|
| 184 |
+
"mean_useful": 21.158918918918918,
|
| 185 |
+
"mean_retained": 49.89945945945946
|
| 186 |
+
}
|
| 187 |
+
},
|
| 188 |
+
"finetuned": {
|
| 189 |
+
"1": {
|
| 190 |
+
"scored_queries": 919,
|
| 191 |
+
"excluded_queries": 6,
|
| 192 |
+
"hit": 0.7595212187159956,
|
| 193 |
+
"precision": 0.7595212187159956,
|
| 194 |
+
"macro_precision": 0.7595212187159956,
|
| 195 |
+
"ndcg": 0.6526763044717343,
|
| 196 |
+
"known_positive_recall": 0.021181970033131128,
|
| 197 |
+
"recall_queries": 919,
|
| 198 |
+
"useful": 698,
|
| 199 |
+
"retained": 919,
|
| 200 |
+
"excluded_positions": 0,
|
| 201 |
+
"mean_useful": 0.7595212187159956,
|
| 202 |
+
"mean_retained": 1.0
|
| 203 |
+
},
|
| 204 |
+
"3": {
|
| 205 |
+
"scored_queries": 925,
|
| 206 |
+
"excluded_queries": 0,
|
| 207 |
+
"hit": 0.894054054054054,
|
| 208 |
+
"precision": 0.7510853835021708,
|
| 209 |
+
"macro_precision": 0.7518918918918919,
|
| 210 |
+
"ndcg": 0.6424089559634334,
|
| 211 |
+
"known_positive_recall": 0.058858902936843926,
|
| 212 |
+
"recall_queries": 925,
|
| 213 |
+
"useful": 2076,
|
| 214 |
+
"retained": 2764,
|
| 215 |
+
"excluded_positions": 11,
|
| 216 |
+
"mean_useful": 2.2443243243243245,
|
| 217 |
+
"mean_retained": 2.9881081081081082
|
| 218 |
+
},
|
| 219 |
+
"5": {
|
| 220 |
+
"scored_queries": 925,
|
| 221 |
+
"excluded_queries": 0,
|
| 222 |
+
"hit": 0.9275675675675675,
|
| 223 |
+
"precision": 0.7366710013003901,
|
| 224 |
+
"macro_precision": 0.736954954954955,
|
| 225 |
+
"ndcg": 0.6373095878906122,
|
| 226 |
+
"known_positive_recall": 0.0932034345948595,
|
| 227 |
+
"recall_queries": 925,
|
| 228 |
+
"useful": 3399,
|
| 229 |
+
"retained": 4614,
|
| 230 |
+
"excluded_positions": 11,
|
| 231 |
+
"mean_useful": 3.674594594594595,
|
| 232 |
+
"mean_retained": 4.988108108108108
|
| 233 |
+
},
|
| 234 |
+
"10": {
|
| 235 |
+
"scored_queries": 925,
|
| 236 |
+
"excluded_queries": 0,
|
| 237 |
+
"hit": 0.9427027027027027,
|
| 238 |
+
"precision": 0.7049002601908065,
|
| 239 |
+
"macro_precision": 0.7052921492921493,
|
| 240 |
+
"ndcg": 0.6377369246501603,
|
| 241 |
+
"known_positive_recall": 0.16987439814587393,
|
| 242 |
+
"recall_queries": 925,
|
| 243 |
+
"useful": 6502,
|
| 244 |
+
"retained": 9224,
|
| 245 |
+
"excluded_positions": 26,
|
| 246 |
+
"mean_useful": 7.029189189189189,
|
| 247 |
+
"mean_retained": 9.971891891891891
|
| 248 |
+
},
|
| 249 |
+
"20": {
|
| 250 |
+
"scored_queries": 925,
|
| 251 |
+
"excluded_queries": 0,
|
| 252 |
+
"hit": 0.9545945945945946,
|
| 253 |
+
"precision": 0.6479835212489159,
|
| 254 |
+
"macro_precision": 0.6484765048449259,
|
| 255 |
+
"ndcg": 0.6488792776497584,
|
| 256 |
+
"known_positive_recall": 0.2962743880804353,
|
| 257 |
+
"recall_queries": 925,
|
| 258 |
+
"useful": 11954,
|
| 259 |
+
"retained": 18448,
|
| 260 |
+
"excluded_positions": 52,
|
| 261 |
+
"mean_useful": 12.923243243243244,
|
| 262 |
+
"mean_retained": 19.943783783783783
|
| 263 |
+
},
|
| 264 |
+
"50": {
|
| 265 |
+
"scored_queries": 925,
|
| 266 |
+
"excluded_queries": 0,
|
| 267 |
+
"hit": 0.9762162162162162,
|
| 268 |
+
"precision": 0.424031024546656,
|
| 269 |
+
"macro_precision": 0.42417201849319736,
|
| 270 |
+
"ndcg": 0.612510329000538,
|
| 271 |
+
"known_positive_recall": 0.4583210768438434,
|
| 272 |
+
"recall_queries": 925,
|
| 273 |
+
"useful": 19572,
|
| 274 |
+
"retained": 46157,
|
| 275 |
+
"excluded_positions": 93,
|
| 276 |
+
"mean_useful": 21.158918918918918,
|
| 277 |
+
"mean_retained": 49.89945945945946
|
| 278 |
+
}
|
| 279 |
+
}
|
| 280 |
+
},
|
| 281 |
+
"corpus_passages": 82719,
|
| 282 |
+
"reference_pairs": 122783
|
| 283 |
+
}
|
evaluation/pipeline.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
publication-manifest.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
| 1 |
{
|
| 2 |
"schema": "daecore.retrieval-publication-manifest",
|
| 3 |
-
"tool_sha256": "
|
| 4 |
"model_id": "ettin-150m-memory-reranker-ft-v1",
|
| 5 |
"public_repo": "Daecore/ettin-150m-memory-reranker-ft-v1",
|
| 6 |
"role": "cross-encoder-reranker",
|
|
@@ -26,7 +26,7 @@
|
|
| 26 |
}
|
| 27 |
}
|
| 28 |
},
|
| 29 |
-
"staged_at": "2026-
|
| 30 |
"files": {
|
| 31 |
"config.json": {
|
| 32 |
"sha256": "a897acffc99f97238bb5bc0e1bccd6b7fe8fdfebe1450d4db3a9ab026cce6795",
|
|
@@ -89,8 +89,8 @@
|
|
| 89 |
"binding": "packaging record"
|
| 90 |
},
|
| 91 |
"evaluation/README.md": {
|
| 92 |
-
"sha256": "
|
| 93 |
-
"size":
|
| 94 |
"binding": "packaging record"
|
| 95 |
},
|
| 96 |
"evaluation/serving-qualification.json": {
|
|
@@ -99,8 +99,8 @@
|
|
| 99 |
"binding": "packaging record"
|
| 100 |
},
|
| 101 |
"evaluation/metrics.py": {
|
| 102 |
-
"sha256": "
|
| 103 |
-
"size":
|
| 104 |
"binding": "packaging record"
|
| 105 |
},
|
| 106 |
"evaluation/chunking.md": {
|
|
@@ -118,6 +118,11 @@
|
|
| 118 |
"size": 14521,
|
| 119 |
"binding": "packaging record"
|
| 120 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 121 |
"evaluation/reranker.json": {
|
| 122 |
"sha256": "20b5760a462c3c2d2281607f2c5045d3f60210123a0c11c8b2d92580c3cccd3e",
|
| 123 |
"size": 1145459,
|
|
@@ -128,6 +133,11 @@
|
|
| 128 |
"size": 1184937,
|
| 129 |
"binding": "packaging record"
|
| 130 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 131 |
"evaluation/figures.py": {
|
| 132 |
"sha256": "f13945b350e007deb18a2eafbe1c5708a71b9df62ee12b6833d07f3e2a06094e",
|
| 133 |
"size": 5490,
|
|
@@ -144,10 +154,10 @@
|
|
| 144 |
"binding": "packaging record"
|
| 145 |
},
|
| 146 |
"README.md": {
|
| 147 |
-
"sha256": "
|
| 148 |
-
"size":
|
| 149 |
"binding": "model card with upload-relative links",
|
| 150 |
-
"source_sha256": "
|
| 151 |
}
|
| 152 |
}
|
| 153 |
}
|
|
|
|
| 1 |
{
|
| 2 |
"schema": "daecore.retrieval-publication-manifest",
|
| 3 |
+
"tool_sha256": "60165b669f2cce17d8830fe9b712fb60a8e82280c54b224d71a1ce8fd82142ac",
|
| 4 |
"model_id": "ettin-150m-memory-reranker-ft-v1",
|
| 5 |
"public_repo": "Daecore/ettin-150m-memory-reranker-ft-v1",
|
| 6 |
"role": "cross-encoder-reranker",
|
|
|
|
| 26 |
}
|
| 27 |
}
|
| 28 |
},
|
| 29 |
+
"staged_at": "2026-10-01T16:23:47+00:00",
|
| 30 |
"files": {
|
| 31 |
"config.json": {
|
| 32 |
"sha256": "a897acffc99f97238bb5bc0e1bccd6b7fe8fdfebe1450d4db3a9ab026cce6795",
|
|
|
|
| 89 |
"binding": "packaging record"
|
| 90 |
},
|
| 91 |
"evaluation/README.md": {
|
| 92 |
+
"sha256": "5399753e1e6e28be3997c1ebf4ee017bfda96d61950bebd89f2b175b29512cad",
|
| 93 |
+
"size": 9377,
|
| 94 |
"binding": "packaging record"
|
| 95 |
},
|
| 96 |
"evaluation/serving-qualification.json": {
|
|
|
|
| 99 |
"binding": "packaging record"
|
| 100 |
},
|
| 101 |
"evaluation/metrics.py": {
|
| 102 |
+
"sha256": "8123f75da6edb158db74c62c1c65b6eefbc2713572409ac8a9bfb5424497aaf1",
|
| 103 |
+
"size": 19927,
|
| 104 |
"binding": "packaging record"
|
| 105 |
},
|
| 106 |
"evaluation/chunking.md": {
|
|
|
|
| 118 |
"size": 14521,
|
| 119 |
"binding": "packaging record"
|
| 120 |
},
|
| 121 |
+
"evaluation/pipeline-summary.json": {
|
| 122 |
+
"sha256": "5f28ae25b5db2f4c0a8d11891df7117f95f644063092d58ba45e27a59747e3a2",
|
| 123 |
+
"size": 9373,
|
| 124 |
+
"binding": "packaging record"
|
| 125 |
+
},
|
| 126 |
"evaluation/reranker.json": {
|
| 127 |
"sha256": "20b5760a462c3c2d2281607f2c5045d3f60210123a0c11c8b2d92580c3cccd3e",
|
| 128 |
"size": 1145459,
|
|
|
|
| 133 |
"size": 1184937,
|
| 134 |
"binding": "packaging record"
|
| 135 |
},
|
| 136 |
+
"evaluation/pipeline.json": {
|
| 137 |
+
"sha256": "8f140901a6b151d9db97fd109a3774bf1c046b849d9a01c874cb0efdea297300",
|
| 138 |
+
"size": 1614738,
|
| 139 |
+
"binding": "packaging record"
|
| 140 |
+
},
|
| 141 |
"evaluation/figures.py": {
|
| 142 |
"sha256": "f13945b350e007deb18a2eafbe1c5708a71b9df62ee12b6833d07f3e2a06094e",
|
| 143 |
"size": 5490,
|
|
|
|
| 154 |
"binding": "packaging record"
|
| 155 |
},
|
| 156 |
"README.md": {
|
| 157 |
+
"sha256": "c8a24bd6c07a34e68d33ba94e476dd75634a72660b21e575435865c25f0002f4",
|
| 158 |
+
"size": 13443,
|
| 159 |
"binding": "model card with upload-relative links",
|
| 160 |
+
"source_sha256": "d3c45681e0ba1ef22b0bfe9eef372b298e1223082e790e25ace134d6a6c760e3"
|
| 161 |
}
|
| 162 |
}
|
| 163 |
}
|