Update embeddinggemma-300m-memory-ft-v2 documentation
Browse files- README.md +39 -36
- evaluation/README.md +41 -20
- evaluation/metrics.py +15 -2
- evaluation/pipeline-summary.json +283 -0
- evaluation/pipeline.json +0 -0
- publication-manifest.json +19 -9
README.md
CHANGED
|
@@ -129,9 +129,9 @@ passage, which does not prove none exists.
|
|
| 129 |
|
| 130 |
At 50, v2 finds useful evidence for **99.03%** of these queries and retrieves
|
| 131 |
**57.92%** of their known useful passages, versus 97.73% and 41.12% upstream.
|
| 132 |
-
The hybrid pool below
|
| 133 |
-
|
| 134 |
-
|
| 135 |
|
| 136 |
Grades 2–3 count as useful. nDCG uses gains 0, 1, 3 and 7 against the same
|
| 137 |
reviewed reference for both models. Rankings from query forms are combined
|
|
@@ -178,44 +178,47 @@ larger quality cost. These byte counts exclude index overhead and compression;
|
|
| 178 |
shorter vectors do not reduce encoder inference cost. Daecore uses 768, and
|
| 179 |
smaller widths have not been qualified through the full hybrid pipeline.
|
| 180 |
|
| 181 |
-
### Full pipeline:
|
| 182 |
|
| 183 |
-
|
| 184 |
-
|
| 185 |
-
|
| 186 |
-
|
| 187 |
-
answer. Both cards share this CUDA FP16 table; it is not either model's
|
| 188 |
-
standalone score.
|
| 189 |
|
| 190 |
-
**Reference:**
|
| 191 |
-
|
|
|
|
| 192 |
|
| 193 |
| Return depth | Hit | Precision | Known-positive recall | nDCG |
|
| 194 |
|---|---:|---:|---:|---:|
|
| 195 |
-
| 1 | 75.
|
| 196 |
-
| 3 | 89.
|
| 197 |
-
| 5 | 92.76% | 73.
|
| 198 |
-
| 10 | 94.27% | 70.
|
| 199 |
-
| 20 | 95.46% | 64.
|
| 200 |
-
| 50 | 97.62% | 42.40% | 45.
|
| 201 |
-
|
| 202 |
-
|
| 203 |
-
|
| 204 |
-
|
| 205 |
-
|
| 206 |
-
|
| 207 |
-
|
| 208 |
-
|
| 209 |
-
queries
|
| 210 |
-
|
| 211 |
-
|
| 212 |
-
|
| 213 |
-
|
| 214 |
-
|
| 215 |
-
|
| 216 |
-
|
| 217 |
-
|
| 218 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 219 |
|
| 220 |
## Training data and objective
|
| 221 |
|
|
|
|
| 129 |
|
| 130 |
At 50, v2 finds useful evidence for **99.03%** of these queries and retrieves
|
| 131 |
**57.92%** of their known useful passages, versus 97.73% and 41.12% upstream.
|
| 132 |
+
The capped hybrid pool below covers fewer queries at 50. Adding BM25 and
|
| 133 |
+
capping the fused list can displace dense candidates, while reranking can
|
| 134 |
+
improve early precision.
|
| 135 |
|
| 136 |
Grades 2–3 count as useful. nDCG uses gains 0, 1, 3 and 7 against the same
|
| 137 |
reviewed reference for both models. Rankings from query forms are combined
|
|
|
|
| 178 |
shorter vectors do not reduce encoder inference cost. Daecore uses 768, and
|
| 179 |
smaller widths have not been qualified through the full hybrid pipeline.
|
| 180 |
|
| 181 |
+
### Full pipeline: the effect of Gemma v2
|
| 182 |
|
| 183 |
+
Only the embedder changes: **upstream Gemma → Daecore Gemma v2**. BM25,
|
| 184 |
+
reciprocal-rank fusion and the fine-tuned Ettin reranker stay fixed. Both
|
| 185 |
+
pipelines search the same **82,719 passages**. Each cell reads before →
|
| 186 |
+
**after**, so the change measures Gemma's contribution inside hybrid retrieval.
|
|
|
|
|
|
|
| 187 |
|
| 188 |
+
**Reference:** the same **925 known-answerable queries**, frozen before this
|
| 189 |
+
comparison · **122,783 graded query–passage pairs** · October 2026. Retrieval
|
| 190 |
+
misses remain included. Both cards share the same fully fine-tuned endpoint.
|
| 191 |
|
| 192 |
| Return depth | Hit | Precision | Known-positive recall | nDCG |
|
| 193 |
|---|---:|---:|---:|---:|
|
| 194 |
+
| 1 | 73.99% → **75.95%** | 73.99% → **75.95%** | 2.01% → **2.12%** | 0.6295 → **0.6527** |
|
| 195 |
+
| 3 | 86.81% → **89.41%** | 71.50% → **75.11%** | 5.49% → **5.89%** | 0.6121 → **0.6424** |
|
| 196 |
+
| 5 | 89.95% → **92.76%** | 70.10% → **73.67%** | 8.53% → **9.32%** | 0.6054 → **0.6373** |
|
| 197 |
+
| 10 | 93.08% → **94.27%** | 65.20% → **70.49%** | 15.25% → **16.99%** | 0.5959 → **0.6377** |
|
| 198 |
+
| 20 | 95.68% → **95.46%** | 56.86% → **64.80%** | 25.18% → **29.63%** | 0.5873 → **0.6489** |
|
| 199 |
+
| 50 | 96.97% → **97.62%** | 35.45% → **42.40%** | 36.85% → **45.83%** | 0.5374 → **0.6125** |
|
| 200 |
+
|
| 201 |
+
At three results, useful-passage precision changes from **71.50% to
|
| 202 |
+
75.11%**; nDCG@10 changes from **0.5959 to 0.6377**.
|
| 203 |
+
At 50, the hybrid pool's known-positive recall changes from **36.85%
|
| 204 |
+
to 45.83%**. This measures evidence available to Ettin after
|
| 205 |
+
fusion; reranking cannot recover passages outside those 50 slots.
|
| 206 |
+
|
| 207 |
+
The gain is not uniform: Hit@20 falls from 95.68% to 95.46%, a difference of
|
| 208 |
+
two queries, while precision, known-positive recall and nDCG improve at that depth.
|
| 209 |
+
|
| 210 |
+
All three variants use matched **PyTorch FP16** reranking. Depth 1 uses **919 shared judged
|
| 211 |
+
queries** across the three pipeline variants; depths 3–50 use all 925.
|
| 212 |
+
Abstentions are excluded inside the original cutoff without backfill.
|
| 213 |
+
Grades 2–3 count as useful. Precision pools retained positions; recall counts
|
| 214 |
+
known useful passages.
|
| 215 |
+
The expanded reference and matched runtime distinguish this table from the
|
| 216 |
+
earlier CUDA/Vulkan serving check. Daecore returns 3–20 results; these fixed
|
| 217 |
+
cutoffs do not evaluate its selector.
|
| 218 |
+
|
| 219 |
+
The [evaluation companion](evaluation/README.md) provides
|
| 220 |
+
the paired records, methods and separate serving checks. The
|
| 221 |
+
[pipeline overview](https://huggingface.co/Daecore) explains retrieval and selection.
|
| 222 |
|
| 223 |
## Training data and objective
|
| 224 |
|
evaluation/README.md
CHANGED
|
@@ -11,7 +11,8 @@ private evaluations.
|
|
| 11 |
| `retriever.json` | Upstream and v2 reviewed top-50 dense rankings for all 970 queries over 82,719 passages |
|
| 12 |
| `dimensions.json` | V2 public nDCG@10 and private top-10 graded rankings at four embedding widths |
|
| 13 |
| `promotion.json` | Ranked grades for the predecessor-to-v2 pipeline comparison and public nDCG@10 for upstream and v2 |
|
| 14 |
-
| `
|
|
|
|
| 15 |
| `*-summary.json` | Summaries that `metrics.py` reproduces |
|
| 16 |
| `serving-qualification.json` | CPU, CUDA and Vulkan serving checks for this model |
|
| 17 |
| `metrics.py` | Metric code |
|
|
@@ -26,7 +27,7 @@ python metrics.py --directory .
|
|
| 26 |
python figures.py --output ../figures
|
| 27 |
```
|
| 28 |
|
| 29 |
-
No packages or network access are needed. Each result matches its corresponding summary. The `answerable` section
|
| 30 |
|
| 31 |
## Dense retrieval
|
| 32 |
|
|
@@ -70,8 +71,9 @@ reference contained 87,434 graded pairs.
|
|
| 70 |
|
| 71 |
The subsequent decision to headline the same 925 known-answerable queries
|
| 72 |
changes the averaging population, not the labels or rankings. On that subset,
|
| 73 |
-
nDCG@10 is **0.4418 upstream and 0.5621 for v2**. The
|
| 74 |
-
dense, smaller-width and
|
|
|
|
| 75 |
comparison below keeps its original reference and population.
|
| 76 |
|
| 77 |
**Public data.**
|
|
@@ -96,25 +98,44 @@ shorten the encoder's forward pass.
|
|
| 96 |
|
| 97 |
## Effect in the retrieval pipeline
|
| 98 |
|
| 99 |
-
|
| 100 |
-
|
| 101 |
-
|
| 102 |
-
|
| 103 |
-
|
| 104 |
-
|
| 105 |
-
|
| 106 |
-
|
| 107 |
-
|
| 108 |
-
|
| 109 |
-
|
| 110 |
-
|
| 111 |
-
|
| 112 |
-
|
| 113 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 114 |
|
| 115 |
### All-query coverage
|
| 116 |
|
| 117 |
-
The records preserve all 970 source queries. This audit view includes the 45
|
| 118 |
without known useful evidence and is separate from the card's answerable-only
|
| 119 |
quality table. Hybrid top 1 uses 965 queries after shared abstention
|
| 120 |
exclusions; dense top 1 retains all 970. The counts below use the full
|
|
|
|
| 11 |
| `retriever.json` | Upstream and v2 reviewed top-50 dense rankings for all 970 queries over 82,719 passages |
|
| 12 |
| `dimensions.json` | V2 public nDCG@10 and private top-10 graded rankings at four embedding widths |
|
| 13 |
| `promotion.json` | Ranked grades for the predecessor-to-v2 pipeline comparison and public nDCG@10 for upstream and v2 |
|
| 14 |
+
| `pipeline.json` | Matched upstream-to-fine-tuned component changes on the frozen 925-query hybrid panel |
|
| 15 |
+
| `serving.json` | Separate 970-query CUDA/Vulkan replay with the September reference |
|
| 16 |
| `*-summary.json` | Summaries that `metrics.py` reproduces |
|
| 17 |
| `serving-qualification.json` | CPU, CUDA and Vulkan serving checks for this model |
|
| 18 |
| `metrics.py` | Metric code |
|
|
|
|
| 27 |
python figures.py --output ../figures
|
| 28 |
```
|
| 29 |
|
| 30 |
+
No packages or network access are needed. Each result matches its corresponding summary. The dense card table uses `retriever-summary.json`'s `answerable` section; the pipeline table uses `pipeline-summary.json`, comparing `upstream_gemma` with `finetuned`. `private.answerable` supplies the private width column. The dense and serving records retain all-query results in their outer `models` sections. This recomputes metrics from saved results; running private retrieval again would require the private corpus.
|
| 31 |
|
| 32 |
## Dense retrieval
|
| 33 |
|
|
|
|
| 71 |
|
| 72 |
The subsequent decision to headline the same 925 known-answerable queries
|
| 73 |
changes the averaging population, not the labels or rankings. On that subset,
|
| 74 |
+
nDCG@10 is **0.4418 upstream and 0.5621 for v2**. The September reference supports
|
| 75 |
+
dense, smaller-width and provider results. The new component comparison adds
|
| 76 |
+
its own reviewed overlay, described below. The historical predecessor
|
| 77 |
comparison below keeps its original reference and population.
|
| 78 |
|
| 79 |
**Public data.**
|
|
|
|
| 98 |
|
| 99 |
## Effect in the retrieval pipeline
|
| 100 |
|
| 101 |
+
`pipeline.json` holds three matched arms. For this card, compare
|
| 102 |
+
`upstream_gemma` with `finetuned`: only Gemma changes, while BM25, fusion and
|
| 103 |
+
the fine-tuned Ettin remain fixed. The Ettin card uses the other baseline
|
| 104 |
+
against the same fully fine-tuned endpoint.
|
| 105 |
+
|
| 106 |
+
The comparison keeps the **925-query cohort frozen before scoring** and uses
|
| 107 |
+
**122,783 graded query–passage pairs** for those queries. It adds 552 grades
|
| 108 |
+
and 10 abstentions to their September reference of 122,231 pairs; old
|
| 109 |
+
grades are unchanged. Two independent GPT-6.1 Sol judges and a fresh blind
|
| 110 |
+
adjudicator resolved the new pairs, with 84 ordinal disagreements.
|
| 111 |
+
The primary assistant reviewed 84 complete pairs, including every final quotation
|
| 112 |
+
flag and abstention plus a stratified sample. These are model judgments,
|
| 113 |
+
without a human reference panel.
|
| 114 |
+
|
| 115 |
+
All three variants were scored anew in **PyTorch FP16**, using Transformers
|
| 116 |
+
5.2.0, Sentence Transformers 5.5.1, 128 query tokens and 1,024 passage tokens.
|
| 117 |
+
Gemma uses cached FP32 embeddings at 768 dimensions with matched input limits.
|
| 118 |
+
The actual first stage uses the frozen query forms, dense retrieval, BM25,
|
| 119 |
+
reciprocal-rank fusion with constant 60, exact-text deduplication and 50 slots.
|
| 120 |
+
Before substituting upstream Gemma's vectors, replay reproduced all 970 saved
|
| 121 |
+
v2 candidate lists exactly. Ties preserve candidate order.
|
| 122 |
+
|
| 123 |
+
Depth 1 uses **919 common judged queries across all three arms**; depths
|
| 124 |
+
3–50 retain all 925. Original prefixes are cut before abstentions are removed,
|
| 125 |
+
with no backfill. Retrieval misses remain included. The reference supplies
|
| 126 |
+
known-positive recall and ideal DCG; it is not exhaustive corpus labeling.
|
| 127 |
+
Fixed cutoffs do not test the calibrated selector's choice of 3–20 results.
|
| 128 |
+
The earlier ONNX provider replay retains its own runtime and September
|
| 129 |
+
reference, so its absolute scores should not be substituted into this comparison.
|
| 130 |
+
|
| 131 |
+
`serving.json` separately compares ONNX CUDA FP16 with Vulkan FP32. It keeps
|
| 132 |
+
the September reference of 122,231 grades for the same 925 queries, with 920
|
| 133 |
+
eligible at top 1. Both providers use identical candidate pools. This checks
|
| 134 |
+
serving paths, not the effect of fine-tuning.
|
| 135 |
|
| 136 |
### All-query coverage
|
| 137 |
|
| 138 |
+
The September dense and serving records preserve all 970 source queries. This audit view includes the 45
|
| 139 |
without known useful evidence and is separate from the card's answerable-only
|
| 140 |
quality table. Hybrid top 1 uses 965 queries after shared abstention
|
| 141 |
exclusions; dense top 1 retains all 970. The counts below use the full
|
evaluation/metrics.py
CHANGED
|
@@ -267,10 +267,22 @@ def summarize_serving(data: dict) -> dict:
|
|
| 267 |
'answerable': _summarize_ranked_grades(data, include_selected=True, answerable_only=True)}
|
| 268 |
|
| 269 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 270 |
SUMMARIZERS = {
|
| 271 |
'classifier': summarize_classifier, 'reranker': summarize_reranker,
|
| 272 |
'retriever': summarize_retriever, 'dimensions': summarize_dimensions,
|
| 273 |
-
'promotion': summarize_promotion, 'serving': summarize_serving,
|
| 274 |
}
|
| 275 |
|
| 276 |
|
|
@@ -283,6 +295,7 @@ def summarize_record(name: str, data: dict) -> dict:
|
|
| 283 |
'dimensions': {'id', 'dataset', 'ndcg@10'},
|
| 284 |
'promotion': {'id', 'group', 'reference_grade_counts', 'ranked_grades', 'excluded', 'selected_depth'},
|
| 285 |
'serving': {'id', 'group', 'reference_grade_counts', 'ranked_grades', 'excluded', 'selected_depth'},
|
|
|
|
| 286 |
}[name]
|
| 287 |
rows = data['rows']
|
| 288 |
if not rows or any(set(row) != fields for row in rows):
|
|
@@ -336,7 +349,7 @@ def main() -> None:
|
|
| 336 |
parser = argparse.ArgumentParser(description=__doc__)
|
| 337 |
parser.add_argument("--directory", type=Path, default=Path(__file__).resolve().parent)
|
| 338 |
parser.add_argument("--output", type=Path)
|
| 339 |
-
parser.add_argument("--model", choices=(
|
| 340 |
args = parser.parse_args()
|
| 341 |
names = [args.model] if args.model else [
|
| 342 |
name for name in SUMMARIZERS if (args.directory / f"{name}.json").is_file()
|
|
|
|
| 267 |
'answerable': _summarize_ranked_grades(data, include_selected=True, answerable_only=True)}
|
| 268 |
|
| 269 |
|
| 270 |
+
def summarize_pipeline(data: dict) -> dict:
|
| 271 |
+
"""Matched component changes on a frozen, known-answerable query cohort."""
|
| 272 |
+
if type(data['corpus_passages']) is not int or data['corpus_passages'] <= 0:
|
| 273 |
+
raise ValueError('Pipeline comparison requires a positive corpus size')
|
| 274 |
+
output = _summarize_ranked_grades(data, include_selected=False)
|
| 275 |
+
if any(row['reference_grade_counts']['2'] + row['reference_grade_counts']['3'] == 0
|
| 276 |
+
for row in data['rows']):
|
| 277 |
+
raise ValueError('Pipeline cohort requires known useful evidence for every query')
|
| 278 |
+
return {**output, 'corpus_passages': data['corpus_passages'],
|
| 279 |
+
'reference_pairs': sum(sum(row['reference_grade_counts'].values()) for row in data['rows'])}
|
| 280 |
+
|
| 281 |
+
|
| 282 |
SUMMARIZERS = {
|
| 283 |
'classifier': summarize_classifier, 'reranker': summarize_reranker,
|
| 284 |
'retriever': summarize_retriever, 'dimensions': summarize_dimensions,
|
| 285 |
+
'promotion': summarize_promotion, 'serving': summarize_serving, 'pipeline': summarize_pipeline,
|
| 286 |
}
|
| 287 |
|
| 288 |
|
|
|
|
| 295 |
'dimensions': {'id', 'dataset', 'ndcg@10'},
|
| 296 |
'promotion': {'id', 'group', 'reference_grade_counts', 'ranked_grades', 'excluded', 'selected_depth'},
|
| 297 |
'serving': {'id', 'group', 'reference_grade_counts', 'ranked_grades', 'excluded', 'selected_depth'},
|
| 298 |
+
'pipeline': {'id', 'group', 'reference_grade_counts', 'ranked_grades', 'excluded'},
|
| 299 |
}[name]
|
| 300 |
rows = data['rows']
|
| 301 |
if not rows or any(set(row) != fields for row in rows):
|
|
|
|
| 349 |
parser = argparse.ArgumentParser(description=__doc__)
|
| 350 |
parser.add_argument("--directory", type=Path, default=Path(__file__).resolve().parent)
|
| 351 |
parser.add_argument("--output", type=Path)
|
| 352 |
+
parser.add_argument("--model", choices=tuple(SUMMARIZERS), help="Recompute one comparison; default: every record in this package")
|
| 353 |
args = parser.parse_args()
|
| 354 |
names = [args.model] if args.model else [
|
| 355 |
name for name in SUMMARIZERS if (args.directory / f"{name}.json").is_file()
|
evaluation/pipeline-summary.json
ADDED
|
@@ -0,0 +1,283 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"queries": 925,
|
| 3 |
+
"models": {
|
| 4 |
+
"upstream_gemma": {
|
| 5 |
+
"1": {
|
| 6 |
+
"scored_queries": 919,
|
| 7 |
+
"excluded_queries": 6,
|
| 8 |
+
"hit": 0.7399347116430903,
|
| 9 |
+
"precision": 0.7399347116430903,
|
| 10 |
+
"macro_precision": 0.7399347116430903,
|
| 11 |
+
"ndcg": 0.6295144826156795,
|
| 12 |
+
"known_positive_recall": 0.02009148065967933,
|
| 13 |
+
"recall_queries": 919,
|
| 14 |
+
"useful": 680,
|
| 15 |
+
"retained": 919,
|
| 16 |
+
"excluded_positions": 0,
|
| 17 |
+
"mean_useful": 0.7399347116430903,
|
| 18 |
+
"mean_retained": 1.0
|
| 19 |
+
},
|
| 20 |
+
"3": {
|
| 21 |
+
"scored_queries": 925,
|
| 22 |
+
"excluded_queries": 0,
|
| 23 |
+
"hit": 0.8681081081081081,
|
| 24 |
+
"precision": 0.7149566473988439,
|
| 25 |
+
"macro_precision": 0.7153153153153153,
|
| 26 |
+
"ndcg": 0.6120552491562434,
|
| 27 |
+
"known_positive_recall": 0.05489711276806684,
|
| 28 |
+
"recall_queries": 925,
|
| 29 |
+
"useful": 1979,
|
| 30 |
+
"retained": 2768,
|
| 31 |
+
"excluded_positions": 7,
|
| 32 |
+
"mean_useful": 2.1394594594594594,
|
| 33 |
+
"mean_retained": 2.9924324324324325
|
| 34 |
+
},
|
| 35 |
+
"5": {
|
| 36 |
+
"scored_queries": 925,
|
| 37 |
+
"excluded_queries": 0,
|
| 38 |
+
"hit": 0.8994594594594595,
|
| 39 |
+
"precision": 0.7009527934170636,
|
| 40 |
+
"macro_precision": 0.7011351351351351,
|
| 41 |
+
"ndcg": 0.60538850988017,
|
| 42 |
+
"known_positive_recall": 0.08526196518888282,
|
| 43 |
+
"recall_queries": 925,
|
| 44 |
+
"useful": 3237,
|
| 45 |
+
"retained": 4618,
|
| 46 |
+
"excluded_positions": 7,
|
| 47 |
+
"mean_useful": 3.4994594594594592,
|
| 48 |
+
"mean_retained": 4.992432432432432
|
| 49 |
+
},
|
| 50 |
+
"10": {
|
| 51 |
+
"scored_queries": 925,
|
| 52 |
+
"excluded_queries": 0,
|
| 53 |
+
"hit": 0.9308108108108109,
|
| 54 |
+
"precision": 0.6519761775852734,
|
| 55 |
+
"macro_precision": 0.6521381381381381,
|
| 56 |
+
"ndcg": 0.5958896253275177,
|
| 57 |
+
"known_positive_recall": 0.15250767729827347,
|
| 58 |
+
"recall_queries": 925,
|
| 59 |
+
"useful": 6021,
|
| 60 |
+
"retained": 9235,
|
| 61 |
+
"excluded_positions": 15,
|
| 62 |
+
"mean_useful": 6.509189189189189,
|
| 63 |
+
"mean_retained": 9.983783783783784
|
| 64 |
+
},
|
| 65 |
+
"20": {
|
| 66 |
+
"scored_queries": 925,
|
| 67 |
+
"excluded_queries": 0,
|
| 68 |
+
"hit": 0.9567567567567568,
|
| 69 |
+
"precision": 0.5685743678596568,
|
| 70 |
+
"macro_precision": 0.5687161650814901,
|
| 71 |
+
"ndcg": 0.5873176954516321,
|
| 72 |
+
"known_positive_recall": 0.2518432399920873,
|
| 73 |
+
"recall_queries": 925,
|
| 74 |
+
"useful": 10501,
|
| 75 |
+
"retained": 18469,
|
| 76 |
+
"excluded_positions": 31,
|
| 77 |
+
"mean_useful": 11.352432432432433,
|
| 78 |
+
"mean_retained": 19.966486486486488
|
| 79 |
+
},
|
| 80 |
+
"50": {
|
| 81 |
+
"scored_queries": 925,
|
| 82 |
+
"excluded_queries": 0,
|
| 83 |
+
"hit": 0.9697297297297297,
|
| 84 |
+
"precision": 0.35454801464376234,
|
| 85 |
+
"macro_precision": 0.35468757089595926,
|
| 86 |
+
"ndcg": 0.5373826502849983,
|
| 87 |
+
"known_positive_recall": 0.36848413838619143,
|
| 88 |
+
"recall_queries": 925,
|
| 89 |
+
"useful": 16367,
|
| 90 |
+
"retained": 46163,
|
| 91 |
+
"excluded_positions": 87,
|
| 92 |
+
"mean_useful": 17.694054054054053,
|
| 93 |
+
"mean_retained": 49.905945945945945
|
| 94 |
+
}
|
| 95 |
+
},
|
| 96 |
+
"upstream_ettin": {
|
| 97 |
+
"1": {
|
| 98 |
+
"scored_queries": 919,
|
| 99 |
+
"excluded_queries": 6,
|
| 100 |
+
"hit": 0.5669205658324266,
|
| 101 |
+
"precision": 0.5669205658324266,
|
| 102 |
+
"macro_precision": 0.5669205658324266,
|
| 103 |
+
"ndcg": 0.4731851391263796,
|
| 104 |
+
"known_positive_recall": 0.01443636631579192,
|
| 105 |
+
"recall_queries": 919,
|
| 106 |
+
"useful": 521,
|
| 107 |
+
"retained": 919,
|
| 108 |
+
"excluded_positions": 0,
|
| 109 |
+
"mean_useful": 0.5669205658324266,
|
| 110 |
+
"mean_retained": 1.0
|
| 111 |
+
},
|
| 112 |
+
"3": {
|
| 113 |
+
"scored_queries": 925,
|
| 114 |
+
"excluded_queries": 0,
|
| 115 |
+
"hit": 0.7902702702702703,
|
| 116 |
+
"precision": 0.5603603603603604,
|
| 117 |
+
"macro_precision": 0.5603603603603603,
|
| 118 |
+
"ndcg": 0.4727968405371722,
|
| 119 |
+
"known_positive_recall": 0.04228504528013372,
|
| 120 |
+
"recall_queries": 925,
|
| 121 |
+
"useful": 1555,
|
| 122 |
+
"retained": 2775,
|
| 123 |
+
"excluded_positions": 0,
|
| 124 |
+
"mean_useful": 1.681081081081081,
|
| 125 |
+
"mean_retained": 3.0
|
| 126 |
+
},
|
| 127 |
+
"5": {
|
| 128 |
+
"scored_queries": 925,
|
| 129 |
+
"excluded_queries": 0,
|
| 130 |
+
"hit": 0.8562162162162162,
|
| 131 |
+
"precision": 0.5582954791261086,
|
| 132 |
+
"macro_precision": 0.5582162162162162,
|
| 133 |
+
"ndcg": 0.47778481822398317,
|
| 134 |
+
"known_positive_recall": 0.06742273722612116,
|
| 135 |
+
"recall_queries": 925,
|
| 136 |
+
"useful": 2581,
|
| 137 |
+
"retained": 4623,
|
| 138 |
+
"excluded_positions": 2,
|
| 139 |
+
"mean_useful": 2.79027027027027,
|
| 140 |
+
"mean_retained": 4.997837837837838
|
| 141 |
+
},
|
| 142 |
+
"10": {
|
| 143 |
+
"scored_queries": 925,
|
| 144 |
+
"excluded_queries": 0,
|
| 145 |
+
"hit": 0.9113513513513514,
|
| 146 |
+
"precision": 0.5447462395844606,
|
| 147 |
+
"macro_precision": 0.5447507507507507,
|
| 148 |
+
"ndcg": 0.49027329580696355,
|
| 149 |
+
"known_positive_recall": 0.12832495703473729,
|
| 150 |
+
"recall_queries": 925,
|
| 151 |
+
"useful": 5034,
|
| 152 |
+
"retained": 9241,
|
| 153 |
+
"excluded_positions": 9,
|
| 154 |
+
"mean_useful": 5.4421621621621625,
|
| 155 |
+
"mean_retained": 9.99027027027027
|
| 156 |
+
},
|
| 157 |
+
"20": {
|
| 158 |
+
"scored_queries": 925,
|
| 159 |
+
"excluded_queries": 0,
|
| 160 |
+
"hit": 0.9502702702702702,
|
| 161 |
+
"precision": 0.5220568335588633,
|
| 162 |
+
"macro_precision": 0.5220833774951422,
|
| 163 |
+
"ndcg": 0.5201039924104931,
|
| 164 |
+
"known_positive_recall": 0.23627621886671868,
|
| 165 |
+
"recall_queries": 925,
|
| 166 |
+
"useful": 9645,
|
| 167 |
+
"retained": 18475,
|
| 168 |
+
"excluded_positions": 25,
|
| 169 |
+
"mean_useful": 10.427027027027027,
|
| 170 |
+
"mean_retained": 19.972972972972972
|
| 171 |
+
},
|
| 172 |
+
"50": {
|
| 173 |
+
"scored_queries": 925,
|
| 174 |
+
"excluded_queries": 0,
|
| 175 |
+
"hit": 0.9762162162162162,
|
| 176 |
+
"precision": 0.424031024546656,
|
| 177 |
+
"macro_precision": 0.42417201849319736,
|
| 178 |
+
"ndcg": 0.5598208131173884,
|
| 179 |
+
"known_positive_recall": 0.4583210768438434,
|
| 180 |
+
"recall_queries": 925,
|
| 181 |
+
"useful": 19572,
|
| 182 |
+
"retained": 46157,
|
| 183 |
+
"excluded_positions": 93,
|
| 184 |
+
"mean_useful": 21.158918918918918,
|
| 185 |
+
"mean_retained": 49.89945945945946
|
| 186 |
+
}
|
| 187 |
+
},
|
| 188 |
+
"finetuned": {
|
| 189 |
+
"1": {
|
| 190 |
+
"scored_queries": 919,
|
| 191 |
+
"excluded_queries": 6,
|
| 192 |
+
"hit": 0.7595212187159956,
|
| 193 |
+
"precision": 0.7595212187159956,
|
| 194 |
+
"macro_precision": 0.7595212187159956,
|
| 195 |
+
"ndcg": 0.6526763044717343,
|
| 196 |
+
"known_positive_recall": 0.021181970033131128,
|
| 197 |
+
"recall_queries": 919,
|
| 198 |
+
"useful": 698,
|
| 199 |
+
"retained": 919,
|
| 200 |
+
"excluded_positions": 0,
|
| 201 |
+
"mean_useful": 0.7595212187159956,
|
| 202 |
+
"mean_retained": 1.0
|
| 203 |
+
},
|
| 204 |
+
"3": {
|
| 205 |
+
"scored_queries": 925,
|
| 206 |
+
"excluded_queries": 0,
|
| 207 |
+
"hit": 0.894054054054054,
|
| 208 |
+
"precision": 0.7510853835021708,
|
| 209 |
+
"macro_precision": 0.7518918918918919,
|
| 210 |
+
"ndcg": 0.6424089559634334,
|
| 211 |
+
"known_positive_recall": 0.058858902936843926,
|
| 212 |
+
"recall_queries": 925,
|
| 213 |
+
"useful": 2076,
|
| 214 |
+
"retained": 2764,
|
| 215 |
+
"excluded_positions": 11,
|
| 216 |
+
"mean_useful": 2.2443243243243245,
|
| 217 |
+
"mean_retained": 2.9881081081081082
|
| 218 |
+
},
|
| 219 |
+
"5": {
|
| 220 |
+
"scored_queries": 925,
|
| 221 |
+
"excluded_queries": 0,
|
| 222 |
+
"hit": 0.9275675675675675,
|
| 223 |
+
"precision": 0.7366710013003901,
|
| 224 |
+
"macro_precision": 0.736954954954955,
|
| 225 |
+
"ndcg": 0.6373095878906122,
|
| 226 |
+
"known_positive_recall": 0.0932034345948595,
|
| 227 |
+
"recall_queries": 925,
|
| 228 |
+
"useful": 3399,
|
| 229 |
+
"retained": 4614,
|
| 230 |
+
"excluded_positions": 11,
|
| 231 |
+
"mean_useful": 3.674594594594595,
|
| 232 |
+
"mean_retained": 4.988108108108108
|
| 233 |
+
},
|
| 234 |
+
"10": {
|
| 235 |
+
"scored_queries": 925,
|
| 236 |
+
"excluded_queries": 0,
|
| 237 |
+
"hit": 0.9427027027027027,
|
| 238 |
+
"precision": 0.7049002601908065,
|
| 239 |
+
"macro_precision": 0.7052921492921493,
|
| 240 |
+
"ndcg": 0.6377369246501603,
|
| 241 |
+
"known_positive_recall": 0.16987439814587393,
|
| 242 |
+
"recall_queries": 925,
|
| 243 |
+
"useful": 6502,
|
| 244 |
+
"retained": 9224,
|
| 245 |
+
"excluded_positions": 26,
|
| 246 |
+
"mean_useful": 7.029189189189189,
|
| 247 |
+
"mean_retained": 9.971891891891891
|
| 248 |
+
},
|
| 249 |
+
"20": {
|
| 250 |
+
"scored_queries": 925,
|
| 251 |
+
"excluded_queries": 0,
|
| 252 |
+
"hit": 0.9545945945945946,
|
| 253 |
+
"precision": 0.6479835212489159,
|
| 254 |
+
"macro_precision": 0.6484765048449259,
|
| 255 |
+
"ndcg": 0.6488792776497584,
|
| 256 |
+
"known_positive_recall": 0.2962743880804353,
|
| 257 |
+
"recall_queries": 925,
|
| 258 |
+
"useful": 11954,
|
| 259 |
+
"retained": 18448,
|
| 260 |
+
"excluded_positions": 52,
|
| 261 |
+
"mean_useful": 12.923243243243244,
|
| 262 |
+
"mean_retained": 19.943783783783783
|
| 263 |
+
},
|
| 264 |
+
"50": {
|
| 265 |
+
"scored_queries": 925,
|
| 266 |
+
"excluded_queries": 0,
|
| 267 |
+
"hit": 0.9762162162162162,
|
| 268 |
+
"precision": 0.424031024546656,
|
| 269 |
+
"macro_precision": 0.42417201849319736,
|
| 270 |
+
"ndcg": 0.612510329000538,
|
| 271 |
+
"known_positive_recall": 0.4583210768438434,
|
| 272 |
+
"recall_queries": 925,
|
| 273 |
+
"useful": 19572,
|
| 274 |
+
"retained": 46157,
|
| 275 |
+
"excluded_positions": 93,
|
| 276 |
+
"mean_useful": 21.158918918918918,
|
| 277 |
+
"mean_retained": 49.89945945945946
|
| 278 |
+
}
|
| 279 |
+
}
|
| 280 |
+
},
|
| 281 |
+
"corpus_passages": 82719,
|
| 282 |
+
"reference_pairs": 122783
|
| 283 |
+
}
|
evaluation/pipeline.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
publication-manifest.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
| 1 |
{
|
| 2 |
"schema": "daecore.retrieval-publication-manifest",
|
| 3 |
-
"tool_sha256": "
|
| 4 |
"model_id": "embeddinggemma-300m-memory-ft-v2",
|
| 5 |
"public_repo": "Daecore/embeddinggemma-300m-memory-ft-v2",
|
| 6 |
"role": "dense-retriever",
|
|
@@ -12,7 +12,7 @@
|
|
| 12 |
"export_receipt_sha256": "e24e1692f7e03c9af8040bb3f6e310a061d305f485861007f1f37d36808b20d8",
|
| 13 |
"serving_sha256": "8c088b1635875d5406e148782b303cfc5a961138600783fb4842b413b98a98c3",
|
| 14 |
"source_model_tree_sha256": "30a7b9f1c42b1fcee29a6729742685d5b744e8fa87d13cefab2fd2d49ea9150e",
|
| 15 |
-
"staged_at": "2026-
|
| 16 |
"files": {
|
| 17 |
"added_tokens.json": {
|
| 18 |
"sha256": "50b2f405ba56a26d4913fd772089992252d7f942123cc0a034d96424221ba946",
|
|
@@ -80,8 +80,8 @@
|
|
| 80 |
"binding": "packaging record"
|
| 81 |
},
|
| 82 |
"evaluation/README.md": {
|
| 83 |
-
"sha256": "
|
| 84 |
-
"size":
|
| 85 |
"binding": "packaging record"
|
| 86 |
},
|
| 87 |
"evaluation/serving-qualification.json": {
|
|
@@ -90,8 +90,8 @@
|
|
| 90 |
"binding": "packaging record"
|
| 91 |
},
|
| 92 |
"evaluation/metrics.py": {
|
| 93 |
-
"sha256": "
|
| 94 |
-
"size":
|
| 95 |
"binding": "packaging record"
|
| 96 |
},
|
| 97 |
"evaluation/chunking.md": {
|
|
@@ -119,6 +119,11 @@
|
|
| 119 |
"size": 14521,
|
| 120 |
"binding": "packaging record"
|
| 121 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 122 |
"evaluation/retriever.json": {
|
| 123 |
"sha256": "c209c96fd17144a6f7d2c7174bc27db4b346288910d085e2b614286b80ae4559",
|
| 124 |
"size": 1156315,
|
|
@@ -139,6 +144,11 @@
|
|
| 139 |
"size": 1184937,
|
| 140 |
"binding": "packaging record"
|
| 141 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 142 |
"evaluation/figures.py": {
|
| 143 |
"sha256": "f13945b350e007deb18a2eafbe1c5708a71b9df62ee12b6833d07f3e2a06094e",
|
| 144 |
"size": 5490,
|
|
@@ -155,10 +165,10 @@
|
|
| 155 |
"binding": "packaging record"
|
| 156 |
},
|
| 157 |
"README.md": {
|
| 158 |
-
"sha256": "
|
| 159 |
-
"size":
|
| 160 |
"binding": "model card with upload-relative links",
|
| 161 |
-
"source_sha256": "
|
| 162 |
}
|
| 163 |
}
|
| 164 |
}
|
|
|
|
| 1 |
{
|
| 2 |
"schema": "daecore.retrieval-publication-manifest",
|
| 3 |
+
"tool_sha256": "60165b669f2cce17d8830fe9b712fb60a8e82280c54b224d71a1ce8fd82142ac",
|
| 4 |
"model_id": "embeddinggemma-300m-memory-ft-v2",
|
| 5 |
"public_repo": "Daecore/embeddinggemma-300m-memory-ft-v2",
|
| 6 |
"role": "dense-retriever",
|
|
|
|
| 12 |
"export_receipt_sha256": "e24e1692f7e03c9af8040bb3f6e310a061d305f485861007f1f37d36808b20d8",
|
| 13 |
"serving_sha256": "8c088b1635875d5406e148782b303cfc5a961138600783fb4842b413b98a98c3",
|
| 14 |
"source_model_tree_sha256": "30a7b9f1c42b1fcee29a6729742685d5b744e8fa87d13cefab2fd2d49ea9150e",
|
| 15 |
+
"staged_at": "2026-10-01T16:22:42+00:00",
|
| 16 |
"files": {
|
| 17 |
"added_tokens.json": {
|
| 18 |
"sha256": "50b2f405ba56a26d4913fd772089992252d7f942123cc0a034d96424221ba946",
|
|
|
|
| 80 |
"binding": "packaging record"
|
| 81 |
},
|
| 82 |
"evaluation/README.md": {
|
| 83 |
+
"sha256": "eeda3234f34b38ccd765f67f0f11e71e125e09a383173cafcee49018884bf4a2",
|
| 84 |
+
"size": 12608,
|
| 85 |
"binding": "packaging record"
|
| 86 |
},
|
| 87 |
"evaluation/serving-qualification.json": {
|
|
|
|
| 90 |
"binding": "packaging record"
|
| 91 |
},
|
| 92 |
"evaluation/metrics.py": {
|
| 93 |
+
"sha256": "8123f75da6edb158db74c62c1c65b6eefbc2713572409ac8a9bfb5424497aaf1",
|
| 94 |
+
"size": 19927,
|
| 95 |
"binding": "packaging record"
|
| 96 |
},
|
| 97 |
"evaluation/chunking.md": {
|
|
|
|
| 119 |
"size": 14521,
|
| 120 |
"binding": "packaging record"
|
| 121 |
},
|
| 122 |
+
"evaluation/pipeline-summary.json": {
|
| 123 |
+
"sha256": "5f28ae25b5db2f4c0a8d11891df7117f95f644063092d58ba45e27a59747e3a2",
|
| 124 |
+
"size": 9373,
|
| 125 |
+
"binding": "packaging record"
|
| 126 |
+
},
|
| 127 |
"evaluation/retriever.json": {
|
| 128 |
"sha256": "c209c96fd17144a6f7d2c7174bc27db4b346288910d085e2b614286b80ae4559",
|
| 129 |
"size": 1156315,
|
|
|
|
| 144 |
"size": 1184937,
|
| 145 |
"binding": "packaging record"
|
| 146 |
},
|
| 147 |
+
"evaluation/pipeline.json": {
|
| 148 |
+
"sha256": "8f140901a6b151d9db97fd109a3774bf1c046b849d9a01c874cb0efdea297300",
|
| 149 |
+
"size": 1614738,
|
| 150 |
+
"binding": "packaging record"
|
| 151 |
+
},
|
| 152 |
"evaluation/figures.py": {
|
| 153 |
"sha256": "f13945b350e007deb18a2eafbe1c5708a71b9df62ee12b6833d07f3e2a06094e",
|
| 154 |
"size": 5490,
|
|
|
|
| 165 |
"binding": "packaging record"
|
| 166 |
},
|
| 167 |
"README.md": {
|
| 168 |
+
"sha256": "46843f44d3ca441406180c2f410a106601449ad092b99269f8c10ea751c1cfa3",
|
| 169 |
+
"size": 16894,
|
| 170 |
"binding": "model card with upload-relative links",
|
| 171 |
+
"source_sha256": "9e4e473e9a7479006bc9d944d70f81a27152819a708cb66d78c2c5d4915e9a5e"
|
| 172 |
}
|
| 173 |
}
|
| 174 |
}
|