Update embeddinggemma-300m-memory-ft-v2 documentation
Browse files- README.md +102 -58
- evaluation/README.md +71 -45
- evaluation/chunking.md +58 -0
- evaluation/dimensions-summary.json +310 -56
- evaluation/dimensions.json +0 -0
- evaluation/figures.py +1 -1
- evaluation/metrics.py +18 -4
- evaluation/retriever-summary.json +60 -30
- evaluation/retriever.json +0 -0
- evaluation/serving-summary.json +66 -36
- evaluation/serving.json +0 -0
- figures/gemma-comparison.svg +65 -53
- publication-manifest.json +30 -25
README.md
CHANGED
|
@@ -22,7 +22,15 @@ tags:
|
|
| 22 |
|
| 23 |
An English embedding model for finding the passages that answer a question in notes, procedures, decision records and technical documentation. It is a fine-tune of [Google's EmbeddingGemma-300M](https://huggingface.co/google/embeddinggemma-300m), packaged as one ONNX graph with pooling and normalization built in, so it runs locally without PyTorch.
|
| 24 |
|
| 25 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 26 |
|
| 27 |
Embed queries and passages with the same revision, and re-embed stored passages when switching to v2.
|
| 28 |
|
|
@@ -80,6 +88,13 @@ chooses the returned prefix. Chunking and candidate coverage therefore affect
|
|
| 80 |
the final results alongside embedding quality. The package also works as a
|
| 81 |
standalone embedder with other retrieval systems.
|
| 82 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 83 |
## Evaluation
|
| 84 |
|
| 85 |
### Gemma alone: dense retrieval on Daecore data
|
|
@@ -89,22 +104,32 @@ exact dense retrieval, the same frozen query forms and matched 128/1,024-token
|
|
| 89 |
input limits. BM25 and Ettin do not contribute to these results. Each cell
|
| 90 |
reads upstream → **Daecore v2**.
|
| 91 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 92 |
| Cutoff | Hit | Precision | Known-positive recall | nDCG |
|
| 93 |
|---|---:|---:|---:|---:|
|
| 94 |
-
| 1 | 53.40% → **64.74%** | 53.40% → **64.74%** | 1.
|
| 95 |
-
| 3 | 72.99% → **79.38%** | 50.83% → **63.08%** |
|
| 96 |
-
| 5 | 79.48% → **84.74%** | 49.95% → **62.66%** |
|
| 97 |
-
| 10 | 84.85% → **88.14%** | 47.40% → **61.18%** |
|
| 98 |
-
| 20 | 88.76% → **90.52%** | 43.94% → **58.11%** |
|
|
|
|
| 99 |
|
| 100 |

|
| 101 |
|
| 102 |
-
Every top-
|
| 103 |
-
|
| 104 |
-
|
| 105 |
-
|
| 106 |
-
|
| 107 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 108 |
|
| 109 |
Grades 2–3 count as useful. nDCG uses gains 0, 1, 3 and 7 against the same
|
| 110 |
reviewed reference for both models, averaging all 970 queries. Rankings from
|
|
@@ -131,61 +156,80 @@ V2 is below upstream on all five datasets, most on FiQA (−0.0273). The first D
|
|
| 131 |
|
| 132 |
The model was trained at 768, 512, 256 and 128 dimensions. To use a smaller
|
| 133 |
width, keep the first dimensions and normalize the shortened vector again.
|
| 134 |
-
Use the same width for queries and passages.
|
| 135 |
-
|
| 136 |
-
|
| 137 |
-
|
| 138 |
-
|
| 139 |
-
|
| 140 |
-
|
|
| 141 |
-
|
|
| 142 |
-
|
|
| 143 |
-
|
|
| 144 |
-
|
| 145 |
-
|
| 146 |
-
|
| 147 |
-
|
| 148 |
-
|
| 149 |
-
|
|
|
|
|
|
|
|
|
|
| 150 |
|
| 151 |
### Full pipeline: Gemma v2 + BM25 + Ettin
|
| 152 |
|
| 153 |
-
This
|
| 154 |
-
|
| 155 |
-
|
| 156 |
-
|
| 157 |
-
|
|
|
|
|
|
|
| 158 |
|
| 159 |
| Return depth | Hit | Precision | Known-positive recall | nDCG |
|
| 160 |
|---|---:|---:|---:|---:|
|
| 161 |
-
| 1 | 72.33% | 72.33% |
|
| 162 |
-
| 3 | 85.36% | 71.58% |
|
| 163 |
-
| 5 | 88.45% | 70.28% |
|
| 164 |
-
| 10 | 89.90% | 67.18% |
|
| 165 |
-
| 20 | 91.03% | 61.79% |
|
| 166 |
-
|
| 167 |
-
|
| 168 |
-
|
| 169 |
-
|
| 170 |
-
|
| 171 |
-
|
| 172 |
-
|
| 173 |
-
|
| 174 |
-
|
| 175 |
-
|
| 176 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 177 |
|
| 178 |
## Training data and objective
|
| 179 |
|
| 180 |
-
Private supervision
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 181 |
|
| 182 |
-
|
| 183 |
-
|
| 184 |
-
|
| 185 |
-
|
| 186 |
-
Gemma applies its own tokenizer and input limits after passage preparation.
|
| 187 |
-
Changing boundaries can change whether evidence stays together and how much the
|
| 188 |
-
model sees, so these results are tied to the evaluated passages.
|
| 189 |
|
| 190 |
Public supervision uses existing evidence annotations from [HotpotQA](https://hotpotqa.github.io/), [MultiDoc2Dial](https://github.com/doc2dial/multidoc2dial) and [FinQA](https://github.com/czyssrs/FinQA): multi-document questions, document-grounded dialogue, and textual or table evidence for numerical questions. The model learns to retrieve that evidence, not to perform FinQA's calculations; FinQA is unrelated to the FiQA evaluation above.
|
| 191 |
|
|
@@ -229,7 +273,7 @@ Vulkan runs through the native ONNX Runtime WebGPU plugin, [`onnxruntime-ep-webg
|
|
| 229 |
|
| 230 |
Checks used Windows x64 with an NVIDIA RTX 3060 Ti (8 GB); other GPUs and Linux are untested, even where Vulkan is available.
|
| 231 |
|
| 232 |
-
The task-specific data is mostly generated and model-judged, with no human reference panel, so it describes this development distribution rather than other users' documents. The dense and full-pipeline panels
|
| 233 |
|
| 234 |
## License
|
| 235 |
|
|
|
|
| 22 |
|
| 23 |
An English embedding model for finding the passages that answer a question in notes, procedures, decision records and technical documentation. It is a fine-tune of [Google's EmbeddingGemma-300M](https://huggingface.co/google/embeddinggemma-300m), packaged as one ONNX graph with pooling and normalization built in, so it runs locally without PyTorch.
|
| 24 |
|
| 25 |
+
| Stronger private retrieval | Less forgetting | Local deployment |
|
| 26 |
+
|---|---|---|
|
| 27 |
+
| nDCG@10 **0.4363 → 0.5574**; precision@10 **47.40% → 61.18%** | Recovers **more than half** of the first fine-tune's loss on FiQA and SciFact | **300M parameters**, ONNX, CPU / CUDA / Vulkan |
|
| 28 |
+
|
| 29 |
+
The private comparison covers **970 queries and 82,719 passages**, using a
|
| 30 |
+
shared reference of **127,665 graded query–passage pairs** and reviewed results
|
| 31 |
+
through rank 50. Daecore accepted some general-retrieval loss for stronger
|
| 32 |
+
evidence retrieval on its document workflow. The tables below show both the gains
|
| 33 |
+
and the remaining public-benchmark gaps.
|
| 34 |
|
| 35 |
Embed queries and passages with the same revision, and re-embed stored passages when switching to v2.
|
| 36 |
|
|
|
|
| 88 |
the final results alongside embedding quality. The package also works as a
|
| 89 |
standalone embedder with other retrieval systems.
|
| 90 |
|
| 91 |
+
Coverage at **50 candidates** measures what is available to a reranker. Hit@50
|
| 92 |
+
asks whether that pool contains any useful evidence; known-positive recall@50
|
| 93 |
+
measures how much of the judged useful evidence it contains. A reranker can
|
| 94 |
+
improve the order but cannot recover a passage outside its pool. Gemma's dense
|
| 95 |
+
top 50 describes a dense-only pipeline. Daecore's actual ceiling depends on
|
| 96 |
+
the 50 candidates admitted after Gemma and BM25 are fused and deduplicated.
|
| 97 |
+
|
| 98 |
## Evaluation
|
| 99 |
|
| 100 |
### Gemma alone: dense retrieval on Daecore data
|
|
|
|
| 104 |
input limits. BM25 and Ettin do not contribute to these results. Each cell
|
| 105 |
reads upstream → **Daecore v2**.
|
| 106 |
|
| 107 |
+
**Reference:** 127,665 graded query–passage pairs · shared top-50 reference
|
| 108 |
+
(September 2026) · abstentions excluded. This pooled reference covers both
|
| 109 |
+
models and is not an exhaustive labeling of the corpus.
|
| 110 |
+
|
| 111 |
| Cutoff | Hit | Precision | Known-positive recall | nDCG |
|
| 112 |
|---|---:|---:|---:|---:|
|
| 113 |
+
| 1 | 53.40% → **64.74%** | 53.40% → **64.74%** | 1.44% → **1.79%** | 0.4592 → **0.5313** |
|
| 114 |
+
| 3 | 72.99% → **79.38%** | 50.83% → **63.08%** | 3.94% → **5.03%** | 0.4414 → **0.5407** |
|
| 115 |
+
| 5 | 79.48% → **84.74%** | 49.95% → **62.66%** | 6.17% → **8.06%** | 0.4346 → **0.5453** |
|
| 116 |
+
| 10 | 84.85% → **88.14%** | 47.40% → **61.18%** | 11.18% → **15.07%** | 0.4363 → **0.5574** |
|
| 117 |
+
| 20 | 88.76% → **90.52%** | 43.94% → **58.11%** | 19.61% → **27.56%** | 0.4526 → **0.5819** |
|
| 118 |
+
| 50 | 93.20% → **94.43%** | 38.60% → **50.45%** | 41.12% → **57.92%** | 0.5107 → **0.6450** |
|
| 119 |
|
| 120 |

|
| 121 |
|
| 122 |
+
Every top-50 position is graded or explicitly excluded as ungradable, with no
|
| 123 |
+
missing-label ranges. All 970 queries remain eligible at every cutoff. At rank
|
| 124 |
+
50, 164 upstream positions and 167 v2 positions are excluded without backfill.
|
| 125 |
+
Precision pools retained positions; recall averages the 925 queries with judged
|
| 126 |
+
useful evidence. The corpus is not exhaustively labeled.
|
| 127 |
+
|
| 128 |
+
At 50, v2 finds useful evidence for **94.43%** of queries and retrieves **57.92%**
|
| 129 |
+
of known useful passages, versus 93.20% and 41.12% upstream. The actual hybrid
|
| 130 |
+
pool below has lower coverage on this panel, despite stronger early precision
|
| 131 |
+
after reranking. Adding BM25 and capping the fused list at 50 can displace dense
|
| 132 |
+
candidates; the two pools should not be treated as interchangeable.
|
| 133 |
|
| 134 |
Grades 2–3 count as useful. nDCG uses gains 0, 1, 3 and 7 against the same
|
| 135 |
reviewed reference for both models, averaging all 970 queries. Rankings from
|
|
|
|
| 156 |
|
| 157 |
The model was trained at 768, 512, 256 and 128 dimensions. To use a smaller
|
| 158 |
width, keep the first dimensions and normalize the shortened vector again.
|
| 159 |
+
Use the same width for queries and passages. Each quality column below is
|
| 160 |
+
**v2 nDCG@10**, computed from exact rankings of cached FP32 vectors. The Daecore
|
| 161 |
+
column covers the same 970-query dense panel with reviewed top-10 labels at
|
| 162 |
+
each width, against the same **127,665-pair reference**. The 768-wide results
|
| 163 |
+
reproduce the corresponding full-width scores.
|
| 164 |
+
|
| 165 |
+
| Dimensions | Bytes per FP32 vector | Daecore | SciFact | FiQA | NFCorpus | SciDocs | ArguAna |
|
| 166 |
+
|---|---:|---:|---:|---:|---:|---:|---:|
|
| 167 |
+
| 768 (default) | 3,072 | 0.5574 | 0.7783 | 0.4468 | 0.3890 | 0.1852 | 0.6259 |
|
| 168 |
+
| 512 | 2,048 | 0.5605 | 0.7841 | 0.4374 | 0.3876 | 0.1793 | 0.6137 |
|
| 169 |
+
| 256 | 1,024 | 0.5292 | 0.7757 | 0.4091 | 0.3660 | 0.1708 | 0.5987 |
|
| 170 |
+
| 128 | 512 | 0.5035 | 0.7488 | 0.3668 | 0.3311 | 0.1484 | 0.5533 |
|
| 171 |
+
|
| 172 |
+
At 512 dimensions, each raw FP32 vector uses one-third less storage, with
|
| 173 |
+
slightly higher Daecore nDCG@10 in this measurement and lower scores on four
|
| 174 |
+
of the five public panels. At 256, vector storage falls by two-thirds with a
|
| 175 |
+
larger quality cost. These byte counts exclude index overhead and compression;
|
| 176 |
+
shorter vectors do not reduce encoder inference cost. Daecore uses 768, and
|
| 177 |
+
smaller widths have not been qualified through the full hybrid pipeline.
|
| 178 |
|
| 179 |
### Full pipeline: Gemma v2 + BM25 + Ettin
|
| 180 |
|
| 181 |
+
This replay measures semantic and lexical search, reciprocal-rank fusion, then
|
| 182 |
+
Ettin reranking over **970 development queries and 82,719 passages**. It includes
|
| 183 |
+
queries with no known useful evidence. Both model cards share this CUDA FP16
|
| 184 |
+
table; it is not either model's standalone score.
|
| 185 |
+
|
| 186 |
+
**Reference:** 970 queries · 82,719 passages · **127,665 graded query–passage
|
| 187 |
+
pairs** in the shared top-50 reference (September 2026). Abstentions are excluded.
|
| 188 |
|
| 189 |
| Return depth | Hit | Precision | Known-positive recall | nDCG |
|
| 190 |
|---|---:|---:|---:|---:|
|
| 191 |
+
| 1 | 72.33% | 72.33% | 2.12% | 0.6458 |
|
| 192 |
+
| 3 | 85.36% | 71.58% | 5.89% | 0.6374 |
|
| 193 |
+
| 5 | 88.45% | 70.28% | 9.35% | 0.6326 |
|
| 194 |
+
| 10 | 89.90% | 67.18% | 17.04% | 0.6315 |
|
| 195 |
+
| 20 | 91.03% | 61.79% | 29.72% | 0.6417 |
|
| 196 |
+
| 50 | 93.09% | 40.43% | 45.98% | 0.6051 |
|
| 197 |
+
|
| 198 |
+
**The hybrid candidate ceiling is 93.09% Hit@50 and 45.98% known-positive
|
| 199 |
+
recall@50.** These are coverage measures for the actual 50-candidate pool;
|
| 200 |
+
reranking cannot change them. The row at 50 exposes that pool, while Daecore
|
| 201 |
+
returns a selected prefix of 3–20. Fixed-depth scores do not evaluate the
|
| 202 |
+
selector's choice.
|
| 203 |
+
|
| 204 |
+
Grades 2–3 count as useful. At depth 1, both providers use the same 965 queries;
|
| 205 |
+
five are omitted because a first result was ungradable. Depths 3–50 retain all
|
| 206 |
+
970. Precision pools judged positions; recall averages the 925 queries with
|
| 207 |
+
known useful evidence (920 at depth 1). Abstentions remove 52 positions at
|
| 208 |
+
rank 20 and 93 at rank 50 per provider, without drawing in deeper passages.
|
| 209 |
+
The corpus is not exhaustively labeled.
|
| 210 |
+
|
| 211 |
+
Dense retrieval and this replay use the same expanded relevance reference.
|
| 212 |
+
Earlier Hit and precision are unchanged; recall and nDCG were recomputed as
|
| 213 |
+
more relevant passages became known. The [evaluation companion](evaluation/README.md)
|
| 214 |
+
provides the records, exclusions and CUDA/Vulkan comparison; the
|
| 215 |
+
[pipeline overview](https://huggingface.co/Daecore) explains retrieval and selection.
|
| 216 |
|
| 217 |
## Training data and objective
|
| 218 |
|
| 219 |
+
Private supervision adapts Gemma to **evidence retrieval from document chunks**.
|
| 220 |
+
Several passages can be useful for one question, including partial evidence.
|
| 221 |
+
|
| 222 |
+
| Training ingredient | Purpose |
|
| 223 |
+
|---|---|
|
| 224 |
+
| Mostly synthetic organizational documents, plus one person's largely AI-written project docs | Exercise project notes, procedures, decisions and technical material |
|
| 225 |
+
| Frozen structural chunks with section context and table/code handling | Train on the passage units the retrieval workflow consumes |
|
| 226 |
+
| Queries written by four models from three provider families, without a designated answer | Ask for evidence without reducing every question to one target passage |
|
| 227 |
+
| Model-judged relevance grades | Distinguish decisive, partial, related and irrelevant evidence |
|
| 228 |
|
| 229 |
+
The [chunking guide](evaluation/chunking.md#retrieval-model-data) explains
|
| 230 |
+
the source data and methods. The corpus retains earlier parser outputs;
|
| 231 |
+
Gemma applies its own tokenizer and limits after chunking. These shared
|
| 232 |
+
generation and judging processes limit generalization beyond the tested data.
|
|
|
|
|
|
|
|
|
|
| 233 |
|
| 234 |
Public supervision uses existing evidence annotations from [HotpotQA](https://hotpotqa.github.io/), [MultiDoc2Dial](https://github.com/doc2dial/multidoc2dial) and [FinQA](https://github.com/czyssrs/FinQA): multi-document questions, document-grounded dialogue, and textual or table evidence for numerical questions. The model learns to retrieve that evidence, not to perform FinQA's calculations; FinQA is unrelated to the FiQA evaluation above.
|
| 235 |
|
|
|
|
| 273 |
|
| 274 |
Checks used Windows x64 with an NVIDIA RTX 3060 Ti (8 GB); other GPUs and Linux are untested, even where Vulkan is available.
|
| 275 |
|
| 276 |
+
The task-specific data is mostly generated and model-judged, with no human reference panel, so it describes this development distribution rather than other users' documents. The dense and full-pipeline panels share queries and a pooled relevance reference, but evaluate different retrieval paths. Some questions omit the project or task context needed for an unambiguous judgment. The 970-query pipeline panel uses natural-language questions and reports no separate terse-query or paraphrase robustness scores. Passage relevance does not measure distinct-fact coverage or downstream agent success.
|
| 277 |
|
| 278 |
## License
|
| 279 |
|
evaluation/README.md
CHANGED
|
@@ -2,10 +2,14 @@
|
|
| 2 |
|
| 3 |
These records support the [model card's](../README.md) results: Gemma's dense retrieval on private and public data, its smaller embedding widths, and the full retrieval pipeline. They contain scores and anonymized relevance grades, without private queries or document text.
|
| 4 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 5 |
| File | Contents |
|
| 6 |
|---|---|
|
| 7 |
-
| `retriever.json` | Upstream and v2 reviewed top-
|
| 8 |
-
| `dimensions.json` | V2
|
| 9 |
| `promotion.json` | Ranked grades for the predecessor-to-v2 pipeline comparison and public nDCG@10 for upstream and v2 |
|
| 10 |
| `serving.json` | The 970-query Gemma v2 + BM25 + Ettin replay used by both model cards, including CUDA/Vulkan results |
|
| 11 |
| `*-summary.json` | Summaries that `metrics.py` reproduces |
|
|
@@ -32,59 +36,81 @@ prefixes and 128-query/1,024-passage token limits, and 768 dimensions. Exact
|
|
| 32 |
cosine rankings take the top 50 for each frozen query form, then merge them
|
| 33 |
with reciprocal-rank fusion (constant 60). Identical text is deduplicated;
|
| 34 |
there is no parent-document filter, BM25 or reranker. The record includes
|
| 35 |
-
model identities, source hashes, top-
|
| 36 |
This is a reused development panel, not a fresh holdout.
|
| 37 |
|
| 38 |
-
|
| 39 |
-
|
| 40 |
-
|
| 41 |
-
|
| 42 |
-
|
| 43 |
-
|
| 44 |
-
|
| 45 |
-
|
| 46 |
-
|
| 47 |
-
|
| 48 |
-
|
| 49 |
-
|
| 50 |
-
|
| 51 |
-
|
| 52 |
-
|
| 53 |
-
|
| 54 |
-
|
| 55 |
-
|
| 56 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 57 |
|
| 58 |
**Public data.**
|
| 59 |
The model card compares v2 with upstream EmbeddingGemma on SciFact, FiQA, NFCorpus, SciDocs and ArguAna. `promotion.json` includes per-query nDCG@10 for all five datasets, using full corpora, matched preprocessing and official relevance judgments. V2 remains below upstream on every panel. The previous fine-tune was measured only on FiQA (0.4009) and SciFact (0.7679); its other three results are absent, not estimated.
|
| 60 |
|
| 61 |
These datasets supplied no training examples, but informed development. They are not untouched tests.
|
| 62 |
|
| 63 |
-
**Matryoshka widths.** `dimensions.json` contains per-query
|
| 64 |
-
|
| 65 |
-
|
| 66 |
-
|
| 67 |
-
|
| 68 |
-
|
| 69 |
-
|
| 70 |
-
|
| 71 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 72 |
|
| 73 |
## Effect in the retrieval pipeline
|
| 74 |
|
| 75 |
-
The identical
|
| 76 |
-
|
| 77 |
-
|
| 78 |
-
|
| 79 |
-
|
| 80 |
-
|
| 81 |
-
|
| 82 |
-
|
| 83 |
-
|
| 84 |
-
|
| 85 |
-
|
| 86 |
-
|
| 87 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 88 |
Only the embedding model changed in this comparison. BM25, reciprocal-rank fusion and the [Ettin reranker](https://huggingface.co/Daecore/ettin-150m-memory-reranker-ft-v1) stayed fixed. The same score mapping selected returned prefixes in both arms. These results measure the whole pipeline on 970 reused queries over 82,719 passage texts, rather than Gemma alone.
|
| 89 |
|
| 90 |
| Metric | Previous Gemma | Gemma v2 |
|
|
@@ -117,4 +143,4 @@ retrieval method, not the query format.
|
|
| 117 |
|
| 118 |
`serving-qualification.json` records this model's CPU, CUDA and Vulkan checks. The 74-vector panel covers token limits, mixed lengths and concurrent query and passage calls. It is a bounded serving check, separate from retrieval quality.
|
| 119 |
|
| 120 |
-
The private documents and labels are mostly model-generated. Project-document sources come from one person's workspaces and are often AI-written. The panel was reused during development, training was run once, and there is no human reference panel. These results do not establish generalization across users, unique-fact coverage or downstream agent success.
|
|
|
|
| 2 |
|
| 3 |
These records support the [model card's](../README.md) results: Gemma's dense retrieval on private and public data, its smaller embedding widths, and the full retrieval pipeline. They contain scores and anonymized relevance grades, without private queries or document text.
|
| 4 |
|
| 5 |
+
The [chunking guide](chunking.md#retrieval-model-data) explains the synthetic
|
| 6 |
+
source documents, passage preparation and relevance supervision behind the
|
| 7 |
+
private evaluations.
|
| 8 |
+
|
| 9 |
| File | Contents |
|
| 10 |
|---|---|
|
| 11 |
+
| `retriever.json` | Upstream and v2 reviewed top-50 dense rankings for all 970 queries over 82,719 passages |
|
| 12 |
+
| `dimensions.json` | V2 public nDCG@10 and private top-10 graded rankings at four embedding widths |
|
| 13 |
| `promotion.json` | Ranked grades for the predecessor-to-v2 pipeline comparison and public nDCG@10 for upstream and v2 |
|
| 14 |
| `serving.json` | The 970-query Gemma v2 + BM25 + Ettin replay used by both model cards, including CUDA/Vulkan results |
|
| 15 |
| `*-summary.json` | Summaries that `metrics.py` reproduces |
|
|
|
|
| 36 |
cosine rankings take the top 50 for each frozen query form, then merge them
|
| 37 |
with reciprocal-rank fusion (constant 60). Identical text is deduplicated;
|
| 38 |
there is no parent-document filter, BM25 or reranker. The record includes
|
| 39 |
+
model identities, source hashes, top-50 grades and reference grade counts.
|
| 40 |
This is a reused development panel, not a fresh holdout.
|
| 41 |
|
| 42 |
+
The shared reference contains **127,665 graded query–passage pairs**. The top-50
|
| 43 |
+
extension added 40,231 grades and 222 abstentions from 40,453 previously
|
| 44 |
+
unjudged pairs, covering both dense models, the hybrid pool and smaller-width
|
| 45 |
+
top-10 results. Two independent GPT-6.1 Sol judges graded each pair; a fresh
|
| 46 |
+
blind adjudicator resolved 7,765 ordinal disagreements. A semantic abstention
|
| 47 |
+
from either primary judge remained excluded. The primary assistant read all
|
| 48 |
+
349 selected pairs: every abstention and final quotation flag, plus a
|
| 49 |
+
stratified sample. Sixteen quotation repairs changed no grades. No person
|
| 50 |
+
reviewed the labels.
|
| 51 |
+
|
| 52 |
+
Every top-50 position is graded or explicitly excluded. All 970 queries remain
|
| 53 |
+
eligible at 1, 3, 5, 10, 20 and 50. Upstream excludes 164 positions at rank 50;
|
| 54 |
+
v2 excludes 167. Cut the original prefix first, drop abstentions second, and
|
| 55 |
+
never backfill. A wholly abstained prefix would omit the query from both arms
|
| 56 |
+
at that depth. Precision pools retained positions. Recall averages the 925
|
| 57 |
+
queries with known useful passages; these are not exhaustive corpus labels.
|
| 58 |
+
nDCG uses gains 0, 1, 3 and 7, reindexes retained positions and averages all
|
| 59 |
+
970 queries; an ideal gain of zero contributes zero.
|
| 60 |
+
|
| 61 |
+
**Reference expansion, not model change.** The existing top-20 grades and
|
| 62 |
+
rankings are unchanged, so Hit and precision are unchanged. New relevant
|
| 63 |
+
passages enlarge recall's denominator and can strengthen the ideal nDCG
|
| 64 |
+
ranking. Dense nDCG@10 is now 0.4363 upstream and 0.5574 for v2, versus
|
| 65 |
+
0.4506 and 0.5748 on the earlier 87,434-pair reference. Both models are recomputed together.
|
| 66 |
+
Dense, smaller-width and current pipeline results now share this reference;
|
| 67 |
+
the historical predecessor comparison below keeps its original reference.
|
| 68 |
|
| 69 |
**Public data.**
|
| 70 |
The model card compares v2 with upstream EmbeddingGemma on SciFact, FiQA, NFCorpus, SciDocs and ArguAna. `promotion.json` includes per-query nDCG@10 for all five datasets, using full corpora, matched preprocessing and official relevance judgments. V2 remains below upstream on every panel. The previous fine-tune was measured only on FiQA (0.4009) and SciFact (0.7679); its other three results are absent, not estimated.
|
| 71 |
|
| 72 |
These datasets supplied no training examples, but informed development. They are not untouched tests.
|
| 73 |
|
| 74 |
+
**Matryoshka widths.** `dimensions.json` contains public per-query nDCG@10
|
| 75 |
+
and a `private` record with graded top-10 rankings at 768, 512, 256 and 128
|
| 76 |
+
dimensions. Vectors come from the selected checkpoint's cached FP32 outputs;
|
| 77 |
+
smaller widths retain the first dimensions and are L2-normalized again.
|
| 78 |
+
|
| 79 |
+
Public scoring uses exact cosine rankings, self-ID exclusion, stable corpus
|
| 80 |
+
order for ties and official linear-gain nDCG. All 1,406 ArguAna queries remain,
|
| 81 |
+
including five whose positives are absent from the corpus. These public
|
| 82 |
+
scores are unchanged. Private scoring uses the same frozen query forms,
|
| 83 |
+
dense-only fusion and expanded relevance reference as the 768-wide comparison.
|
| 84 |
+
All 970 queries remain eligible at top 10; excluded positions are 22, 26, 16
|
| 85 |
+
and 37 respectively. The 768-wide results reproduce both references exactly.
|
| 86 |
+
Smaller widths have not been qualified through the full pipeline and do not
|
| 87 |
+
shorten the encoder's forward pass.
|
| 88 |
|
| 89 |
## Effect in the retrieval pipeline
|
| 90 |
|
| 91 |
+
The identical table on both cards comes from `serving.json`: Gemma v2 +
|
| 92 |
+
BM25 + fusion + unchanged Ettin, on all 970 queries. CUDA FP16 and Vulkan FP32
|
| 93 |
+
use the same 50 candidates. At 50, Hit is 93.09% and known-positive recall is
|
| 94 |
+
45.98%; those candidate-coverage measures cannot change through reranking.
|
| 95 |
+
Dense v2 reaches 94.43% and 57.92% on this panel. Fusion with a fixed 50-slot
|
| 96 |
+
budget can displace dense candidates even while reranking improves early
|
| 97 |
+
precision.
|
| 98 |
+
|
| 99 |
+
At depths 3–50, Hit and nDCG average all 970 queries; recall averages the 925
|
| 100 |
+
with known useful evidence. Depth 1 uses 965 shared judged queries, including
|
| 101 |
+
920 with known positives; five are omitted because a first result was
|
| 102 |
+
ungradable. Precision pools retained positions. Each provider excludes 52
|
| 103 |
+
positions at rank 20 and 93 at rank 50 without backfill. Fixed cutoffs are
|
| 104 |
+
separate from the selector's choice of a 3–20-item prefix. Vulkan matches
|
| 105 |
+
CUDA at Hit@5, Hit@10, Hit@20 and Hit@50; Hit@3 differs by one query and
|
| 106 |
+
nDCG@10 by 0.0003.
|
| 107 |
+
|
| 108 |
+
### Historical predecessor comparison
|
| 109 |
+
|
| 110 |
+
This table comes from `promotion.json`, with its **original 75,488-pair relevance
|
| 111 |
+
reference**. Its nDCG is not directly comparable with the expanded-reference
|
| 112 |
+
table on the main card.
|
| 113 |
+
|
| 114 |
Only the embedding model changed in this comparison. BM25, reciprocal-rank fusion and the [Ettin reranker](https://huggingface.co/Daecore/ettin-150m-memory-reranker-ft-v1) stayed fixed. The same score mapping selected returned prefixes in both arms. These results measure the whole pipeline on 970 reused queries over 82,719 passage texts, rather than Gemma alone.
|
| 115 |
|
| 116 |
| Metric | Previous Gemma | Gemma v2 |
|
|
|
|
| 143 |
|
| 144 |
`serving-qualification.json` records this model's CPU, CUDA and Vulkan checks. The 74-vector panel covers token limits, mixed lengths and concurrent query and passage calls. It is a bounded serving check, separate from retrieval quality.
|
| 145 |
|
| 146 |
+
The private documents and labels are mostly model-generated. Project-document sources come from one person's workspaces and are often AI-written. Some queries omit the project or subject needed for a clear relevance judgment. The panel was reused during development, training was run once, and there is no human reference panel. These results do not establish generalization across users, unique-fact coverage or downstream agent success.
|
evaluation/chunking.md
ADDED
|
@@ -0,0 +1,58 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Chunking and retrieval data
|
| 2 |
+
|
| 3 |
+
## Retrieval-model data
|
| 4 |
+
|
| 5 |
+
Daecore fine-tunes retrieval models on passages prepared for its document
|
| 6 |
+
workflow. A passage is the exact chunk the model reads: a section, excerpt,
|
| 7 |
+
table fragment or code block. It can supply useful partial evidence without
|
| 8 |
+
answering the whole question.
|
| 9 |
+
|
| 10 |
+
**Source documents → structural chunks → graded query–passage pairs → retrieval training**
|
| 11 |
+
|
| 12 |
+
| Ingredient | What it contributes |
|
| 13 |
+
|---|---|
|
| 14 |
+
| Mostly synthetic workplace documentation | Notes, procedures, decisions and technical material across generated project scenarios |
|
| 15 |
+
| One person's project documents, much of them AI-written | Additional examples, not a representative sample of other users' workspaces |
|
| 16 |
+
| Frozen structural chunks | The retrieval units that Gemma embeds and Ettin ranks; some retain earlier parser outputs |
|
| 17 |
+
| Model-written questions and graded evidence | Several useful passages per question, with decisive and partial evidence distinguished from merely related text |
|
| 18 |
+
| Public evidence annotations used by Gemma v2 | HotpotQA paragraphs, MultiDoc2Dial grounding passages and FinQA text/table evidence |
|
| 19 |
+
|
| 20 |
+
Public retrieval data also comes in passages; it is not uniformly made of full
|
| 21 |
+
documents. The distinction is their source and preparation. Each model applies
|
| 22 |
+
its tokenizer and input limits after chunking, so a chunk boundary and a model's
|
| 23 |
+
truncation limit are separate constraints.
|
| 24 |
+
|
| 25 |
+
This is deliberate specialization of already capable upstream models. The
|
| 26 |
+
private comparisons measure gains on this workflow, while public benchmarks
|
| 27 |
+
show the general-retrieval cost. Daecore accepted that tradeoff for its intended
|
| 28 |
+
use; the cards report both sides. Generated questions and shared preparation
|
| 29 |
+
and judging processes do not establish gains for every real user or agent.
|
| 30 |
+
|
| 31 |
+
Training and evaluation use frozen passage snapshots. Current chunking fixes
|
| 32 |
+
do not retroactively change those texts or transfer their relevance labels to
|
| 33 |
+
different cuts.
|
| 34 |
+
|
| 35 |
+
## Chunking methods
|
| 36 |
+
|
| 37 |
+
- **Heading-based splitting:** use the document's sections; oversized sections
|
| 38 |
+
can split further at their subheadings. This follows the author's structure.
|
| 39 |
+
- **Title and heading context:** add the document title and relevant heading
|
| 40 |
+
path to section chunks so their subject stays clear. Smaller searchable pieces
|
| 41 |
+
retain that metadata but do not automatically repeat it inside every piece.
|
| 42 |
+
- **Boundary-aware cuts and overlap:** prefer a nearby sentence end, paragraph
|
| 43 |
+
break or line break. Ordinary prose windows repeat up to 200 characters across
|
| 44 |
+
cuts to carry nearby context; oversized structural continuations have their
|
| 45 |
+
own framing rules.
|
| 46 |
+
- **Code-block and table handling:** keep blocks together when they fit. Split
|
| 47 |
+
oversized tables with repeated column headers and oversized code blocks with
|
| 48 |
+
reopened and closed code fences, so continuations remain readable.
|
| 49 |
+
- **Record-aware splitting:** preserve row, field or entry context for CSV,
|
| 50 |
+
JSON, JSONL and dated entries. Oversized records are divided with identifying
|
| 51 |
+
labels rather than losing their remaining content.
|
| 52 |
+
- **Paragraph packing for imported memory files:** group whole paragraphs into
|
| 53 |
+
bounded pieces so excerpts usually begin at a paragraph. A single oversized
|
| 54 |
+
paragraph falls back to structural splitting.
|
| 55 |
+
|
| 56 |
+
These methods work together: make supported text and record chunks fit first,
|
| 57 |
+
then create bounded searchable pieces for sections that are still too large.
|
| 58 |
+
Both operations use the same structural cutter.
|
evaluation/dimensions-summary.json
CHANGED
|
@@ -1,56 +1,310 @@
|
|
| 1 |
-
{
|
| 2 |
-
"queries": 3677,
|
| 3 |
-
"vector_bytes_fp32": {
|
| 4 |
-
"768": 3072,
|
| 5 |
-
"512": 2048,
|
| 6 |
-
"256": 1024,
|
| 7 |
-
"128": 512
|
| 8 |
-
},
|
| 9 |
-
"public": {
|
| 10 |
-
"arguana": {
|
| 11 |
-
"queries": 1406,
|
| 12 |
-
"ndcg@10": {
|
| 13 |
-
"768": 0.6258634411077135,
|
| 14 |
-
"512": 0.613697054466976,
|
| 15 |
-
"256": 0.5986725670753965,
|
| 16 |
-
"128": 0.5533073210679951
|
| 17 |
-
}
|
| 18 |
-
},
|
| 19 |
-
"fiqa": {
|
| 20 |
-
"queries": 648,
|
| 21 |
-
"ndcg@10": {
|
| 22 |
-
"768": 0.446838124778398,
|
| 23 |
-
"512": 0.43740349056283273,
|
| 24 |
-
"256": 0.40912280098290843,
|
| 25 |
-
"128": 0.3668286877036632
|
| 26 |
-
}
|
| 27 |
-
},
|
| 28 |
-
"nfcorpus": {
|
| 29 |
-
"queries": 323,
|
| 30 |
-
"ndcg@10": {
|
| 31 |
-
"768": 0.3889646280478944,
|
| 32 |
-
"512": 0.3875786151617704,
|
| 33 |
-
"256": 0.3660320811396263,
|
| 34 |
-
"128": 0.3310697612326329
|
| 35 |
-
}
|
| 36 |
-
},
|
| 37 |
-
"scidocs": {
|
| 38 |
-
"queries": 1000,
|
| 39 |
-
"ndcg@10": {
|
| 40 |
-
"768": 0.18521860884727676,
|
| 41 |
-
"512": 0.179327587641947,
|
| 42 |
-
"256": 0.17081239437849563,
|
| 43 |
-
"128": 0.1484461830477022
|
| 44 |
-
}
|
| 45 |
-
},
|
| 46 |
-
"scifact": {
|
| 47 |
-
"queries": 300,
|
| 48 |
-
"ndcg@10": {
|
| 49 |
-
"768": 0.7782845713792622,
|
| 50 |
-
"512": 0.7840710444235555,
|
| 51 |
-
"256": 0.7756812272647337,
|
| 52 |
-
"128": 0.7488095898261483
|
| 53 |
-
}
|
| 54 |
-
}
|
| 55 |
-
}
|
| 56 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"queries": 3677,
|
| 3 |
+
"vector_bytes_fp32": {
|
| 4 |
+
"768": 3072,
|
| 5 |
+
"512": 2048,
|
| 6 |
+
"256": 1024,
|
| 7 |
+
"128": 512
|
| 8 |
+
},
|
| 9 |
+
"public": {
|
| 10 |
+
"arguana": {
|
| 11 |
+
"queries": 1406,
|
| 12 |
+
"ndcg@10": {
|
| 13 |
+
"768": 0.6258634411077135,
|
| 14 |
+
"512": 0.613697054466976,
|
| 15 |
+
"256": 0.5986725670753965,
|
| 16 |
+
"128": 0.5533073210679951
|
| 17 |
+
}
|
| 18 |
+
},
|
| 19 |
+
"fiqa": {
|
| 20 |
+
"queries": 648,
|
| 21 |
+
"ndcg@10": {
|
| 22 |
+
"768": 0.446838124778398,
|
| 23 |
+
"512": 0.43740349056283273,
|
| 24 |
+
"256": 0.40912280098290843,
|
| 25 |
+
"128": 0.3668286877036632
|
| 26 |
+
}
|
| 27 |
+
},
|
| 28 |
+
"nfcorpus": {
|
| 29 |
+
"queries": 323,
|
| 30 |
+
"ndcg@10": {
|
| 31 |
+
"768": 0.3889646280478944,
|
| 32 |
+
"512": 0.3875786151617704,
|
| 33 |
+
"256": 0.3660320811396263,
|
| 34 |
+
"128": 0.3310697612326329
|
| 35 |
+
}
|
| 36 |
+
},
|
| 37 |
+
"scidocs": {
|
| 38 |
+
"queries": 1000,
|
| 39 |
+
"ndcg@10": {
|
| 40 |
+
"768": 0.18521860884727676,
|
| 41 |
+
"512": 0.179327587641947,
|
| 42 |
+
"256": 0.17081239437849563,
|
| 43 |
+
"128": 0.1484461830477022
|
| 44 |
+
}
|
| 45 |
+
},
|
| 46 |
+
"scifact": {
|
| 47 |
+
"queries": 300,
|
| 48 |
+
"ndcg@10": {
|
| 49 |
+
"768": 0.7782845713792622,
|
| 50 |
+
"512": 0.7840710444235555,
|
| 51 |
+
"256": 0.7756812272647337,
|
| 52 |
+
"128": 0.7488095898261483
|
| 53 |
+
}
|
| 54 |
+
}
|
| 55 |
+
},
|
| 56 |
+
"private": {
|
| 57 |
+
"queries": 970,
|
| 58 |
+
"models": {
|
| 59 |
+
"768": {
|
| 60 |
+
"1": {
|
| 61 |
+
"scored_queries": 964,
|
| 62 |
+
"excluded_queries": 6,
|
| 63 |
+
"hit": 0.6452282157676349,
|
| 64 |
+
"precision": 0.6452282157676349,
|
| 65 |
+
"macro_precision": 0.6452282157676349,
|
| 66 |
+
"ndcg": 0.5301323849041691,
|
| 67 |
+
"known_positive_recall": 0.01788016629109334,
|
| 68 |
+
"recall_queries": 919,
|
| 69 |
+
"useful": 622,
|
| 70 |
+
"retained": 964,
|
| 71 |
+
"excluded_positions": 0,
|
| 72 |
+
"mean_useful": 0.6452282157676349,
|
| 73 |
+
"mean_retained": 1.0
|
| 74 |
+
},
|
| 75 |
+
"3": {
|
| 76 |
+
"scored_queries": 970,
|
| 77 |
+
"excluded_queries": 0,
|
| 78 |
+
"hit": 0.7938144329896907,
|
| 79 |
+
"precision": 0.6308009625300791,
|
| 80 |
+
"macro_precision": 0.6309278350515464,
|
| 81 |
+
"ndcg": 0.5406982958852574,
|
| 82 |
+
"known_positive_recall": 0.05034823759106709,
|
| 83 |
+
"recall_queries": 925,
|
| 84 |
+
"useful": 1835,
|
| 85 |
+
"retained": 2909,
|
| 86 |
+
"excluded_positions": 1,
|
| 87 |
+
"mean_useful": 1.8917525773195876,
|
| 88 |
+
"mean_retained": 2.9989690721649485
|
| 89 |
+
},
|
| 90 |
+
"5": {
|
| 91 |
+
"scored_queries": 970,
|
| 92 |
+
"excluded_queries": 0,
|
| 93 |
+
"hit": 0.8474226804123711,
|
| 94 |
+
"precision": 0.6266005782734407,
|
| 95 |
+
"macro_precision": 0.6269759450171821,
|
| 96 |
+
"ndcg": 0.5453498298950709,
|
| 97 |
+
"known_positive_recall": 0.08055790386567772,
|
| 98 |
+
"recall_queries": 925,
|
| 99 |
+
"useful": 3034,
|
| 100 |
+
"retained": 4842,
|
| 101 |
+
"excluded_positions": 8,
|
| 102 |
+
"mean_useful": 3.1278350515463917,
|
| 103 |
+
"mean_retained": 4.991752577319588
|
| 104 |
+
},
|
| 105 |
+
"10": {
|
| 106 |
+
"scored_queries": 970,
|
| 107 |
+
"excluded_queries": 0,
|
| 108 |
+
"hit": 0.8814432989690721,
|
| 109 |
+
"precision": 0.6117999586691465,
|
| 110 |
+
"macro_precision": 0.6123367697594502,
|
| 111 |
+
"ndcg": 0.5574064002423099,
|
| 112 |
+
"known_positive_recall": 0.15074270958466607,
|
| 113 |
+
"recall_queries": 925,
|
| 114 |
+
"useful": 5921,
|
| 115 |
+
"retained": 9678,
|
| 116 |
+
"excluded_positions": 22,
|
| 117 |
+
"mean_useful": 6.104123711340206,
|
| 118 |
+
"mean_retained": 9.977319587628866
|
| 119 |
+
}
|
| 120 |
+
},
|
| 121 |
+
"512": {
|
| 122 |
+
"1": {
|
| 123 |
+
"scored_queries": 964,
|
| 124 |
+
"excluded_queries": 6,
|
| 125 |
+
"hit": 0.6462655601659751,
|
| 126 |
+
"precision": 0.6462655601659751,
|
| 127 |
+
"macro_precision": 0.6462655601659751,
|
| 128 |
+
"ndcg": 0.5478166370282552,
|
| 129 |
+
"known_positive_recall": 0.01775688496570127,
|
| 130 |
+
"recall_queries": 919,
|
| 131 |
+
"useful": 623,
|
| 132 |
+
"retained": 964,
|
| 133 |
+
"excluded_positions": 0,
|
| 134 |
+
"mean_useful": 0.6462655601659751,
|
| 135 |
+
"mean_retained": 1.0
|
| 136 |
+
},
|
| 137 |
+
"3": {
|
| 138 |
+
"scored_queries": 970,
|
| 139 |
+
"excluded_queries": 0,
|
| 140 |
+
"hit": 0.7969072164948454,
|
| 141 |
+
"precision": 0.635237439779766,
|
| 142 |
+
"macro_precision": 0.6355670103092783,
|
| 143 |
+
"ndcg": 0.5429708859076539,
|
| 144 |
+
"known_positive_recall": 0.05057307077004017,
|
| 145 |
+
"recall_queries": 925,
|
| 146 |
+
"useful": 1846,
|
| 147 |
+
"retained": 2906,
|
| 148 |
+
"excluded_positions": 4,
|
| 149 |
+
"mean_useful": 1.9030927835051545,
|
| 150 |
+
"mean_retained": 2.995876288659794
|
| 151 |
+
},
|
| 152 |
+
"5": {
|
| 153 |
+
"scored_queries": 970,
|
| 154 |
+
"excluded_queries": 0,
|
| 155 |
+
"hit": 0.8422680412371134,
|
| 156 |
+
"precision": 0.6308613922743235,
|
| 157 |
+
"macro_precision": 0.63106529209622,
|
| 158 |
+
"ndcg": 0.5504066702414289,
|
| 159 |
+
"known_positive_recall": 0.08121602068063129,
|
| 160 |
+
"recall_queries": 925,
|
| 161 |
+
"useful": 3054,
|
| 162 |
+
"retained": 4841,
|
| 163 |
+
"excluded_positions": 9,
|
| 164 |
+
"mean_useful": 3.1484536082474226,
|
| 165 |
+
"mean_retained": 4.990721649484536
|
| 166 |
+
},
|
| 167 |
+
"10": {
|
| 168 |
+
"scored_queries": 970,
|
| 169 |
+
"excluded_queries": 0,
|
| 170 |
+
"hit": 0.8824742268041237,
|
| 171 |
+
"precision": 0.6132933636551582,
|
| 172 |
+
"macro_precision": 0.613993618065783,
|
| 173 |
+
"ndcg": 0.5604826042172213,
|
| 174 |
+
"known_positive_recall": 0.15004173661513098,
|
| 175 |
+
"recall_queries": 925,
|
| 176 |
+
"useful": 5933,
|
| 177 |
+
"retained": 9674,
|
| 178 |
+
"excluded_positions": 26,
|
| 179 |
+
"mean_useful": 6.116494845360824,
|
| 180 |
+
"mean_retained": 9.97319587628866
|
| 181 |
+
}
|
| 182 |
+
},
|
| 183 |
+
"256": {
|
| 184 |
+
"1": {
|
| 185 |
+
"scored_queries": 964,
|
| 186 |
+
"excluded_queries": 6,
|
| 187 |
+
"hit": 0.6037344398340249,
|
| 188 |
+
"precision": 0.6037344398340249,
|
| 189 |
+
"macro_precision": 0.6037344398340249,
|
| 190 |
+
"ndcg": 0.5127445168938944,
|
| 191 |
+
"known_positive_recall": 0.015976956236435802,
|
| 192 |
+
"recall_queries": 919,
|
| 193 |
+
"useful": 582,
|
| 194 |
+
"retained": 964,
|
| 195 |
+
"excluded_positions": 0,
|
| 196 |
+
"mean_useful": 0.6037344398340249,
|
| 197 |
+
"mean_retained": 1.0
|
| 198 |
+
},
|
| 199 |
+
"3": {
|
| 200 |
+
"scored_queries": 970,
|
| 201 |
+
"excluded_queries": 0,
|
| 202 |
+
"hit": 0.7608247422680412,
|
| 203 |
+
"precision": 0.5955249569707401,
|
| 204 |
+
"macro_precision": 0.5962199312714777,
|
| 205 |
+
"ndcg": 0.5111829833045054,
|
| 206 |
+
"known_positive_recall": 0.046233615834183193,
|
| 207 |
+
"recall_queries": 925,
|
| 208 |
+
"useful": 1730,
|
| 209 |
+
"retained": 2905,
|
| 210 |
+
"excluded_positions": 5,
|
| 211 |
+
"mean_useful": 1.7835051546391754,
|
| 212 |
+
"mean_retained": 2.9948453608247423
|
| 213 |
+
},
|
| 214 |
+
"5": {
|
| 215 |
+
"scored_queries": 970,
|
| 216 |
+
"excluded_queries": 0,
|
| 217 |
+
"hit": 0.8216494845360824,
|
| 218 |
+
"precision": 0.5925696594427244,
|
| 219 |
+
"macro_precision": 0.5929381443298969,
|
| 220 |
+
"ndcg": 0.5174515106874455,
|
| 221 |
+
"known_positive_recall": 0.0744111972114428,
|
| 222 |
+
"recall_queries": 925,
|
| 223 |
+
"useful": 2871,
|
| 224 |
+
"retained": 4845,
|
| 225 |
+
"excluded_positions": 5,
|
| 226 |
+
"mean_useful": 2.95979381443299,
|
| 227 |
+
"mean_retained": 4.994845360824742
|
| 228 |
+
},
|
| 229 |
+
"10": {
|
| 230 |
+
"scored_queries": 970,
|
| 231 |
+
"excluded_queries": 0,
|
| 232 |
+
"hit": 0.8731958762886598,
|
| 233 |
+
"precision": 0.5818876497315159,
|
| 234 |
+
"macro_precision": 0.5823911798396334,
|
| 235 |
+
"ndcg": 0.5291576970298347,
|
| 236 |
+
"known_positive_recall": 0.1419199356602385,
|
| 237 |
+
"recall_queries": 925,
|
| 238 |
+
"useful": 5635,
|
| 239 |
+
"retained": 9684,
|
| 240 |
+
"excluded_positions": 16,
|
| 241 |
+
"mean_useful": 5.809278350515464,
|
| 242 |
+
"mean_retained": 9.983505154639175
|
| 243 |
+
}
|
| 244 |
+
},
|
| 245 |
+
"128": {
|
| 246 |
+
"1": {
|
| 247 |
+
"scored_queries": 964,
|
| 248 |
+
"excluded_queries": 6,
|
| 249 |
+
"hit": 0.5726141078838174,
|
| 250 |
+
"precision": 0.5726141078838174,
|
| 251 |
+
"macro_precision": 0.5726141078838174,
|
| 252 |
+
"ndcg": 0.47288085358624776,
|
| 253 |
+
"known_positive_recall": 0.015019269226427342,
|
| 254 |
+
"recall_queries": 919,
|
| 255 |
+
"useful": 552,
|
| 256 |
+
"retained": 964,
|
| 257 |
+
"excluded_positions": 0,
|
| 258 |
+
"mean_useful": 0.5726141078838174,
|
| 259 |
+
"mean_retained": 1.0
|
| 260 |
+
},
|
| 261 |
+
"3": {
|
| 262 |
+
"scored_queries": 970,
|
| 263 |
+
"excluded_queries": 0,
|
| 264 |
+
"hit": 0.7402061855670103,
|
| 265 |
+
"precision": 0.5779310344827586,
|
| 266 |
+
"macro_precision": 0.5780068728522336,
|
| 267 |
+
"ndcg": 0.4818931326910878,
|
| 268 |
+
"known_positive_recall": 0.044531751738185646,
|
| 269 |
+
"recall_queries": 925,
|
| 270 |
+
"useful": 1676,
|
| 271 |
+
"retained": 2900,
|
| 272 |
+
"excluded_positions": 10,
|
| 273 |
+
"mean_useful": 1.7278350515463918,
|
| 274 |
+
"mean_retained": 2.9896907216494846
|
| 275 |
+
},
|
| 276 |
+
"5": {
|
| 277 |
+
"scored_queries": 970,
|
| 278 |
+
"excluded_queries": 0,
|
| 279 |
+
"hit": 0.8041237113402062,
|
| 280 |
+
"precision": 0.574591351127664,
|
| 281 |
+
"macro_precision": 0.5746048109965636,
|
| 282 |
+
"ndcg": 0.48751134015621145,
|
| 283 |
+
"known_positive_recall": 0.07159737689223461,
|
| 284 |
+
"recall_queries": 925,
|
| 285 |
+
"useful": 2777,
|
| 286 |
+
"retained": 4833,
|
| 287 |
+
"excluded_positions": 17,
|
| 288 |
+
"mean_useful": 2.8628865979381444,
|
| 289 |
+
"mean_retained": 4.982474226804124
|
| 290 |
+
},
|
| 291 |
+
"10": {
|
| 292 |
+
"scored_queries": 970,
|
| 293 |
+
"excluded_queries": 0,
|
| 294 |
+
"hit": 0.8649484536082475,
|
| 295 |
+
"precision": 0.5659733002173238,
|
| 296 |
+
"macro_precision": 0.5665320733104239,
|
| 297 |
+
"ndcg": 0.5035498696905354,
|
| 298 |
+
"known_positive_recall": 0.13623796718702827,
|
| 299 |
+
"recall_queries": 925,
|
| 300 |
+
"useful": 5469,
|
| 301 |
+
"retained": 9663,
|
| 302 |
+
"excluded_positions": 37,
|
| 303 |
+
"mean_useful": 5.638144329896908,
|
| 304 |
+
"mean_retained": 9.961855670103093
|
| 305 |
+
}
|
| 306 |
+
}
|
| 307 |
+
},
|
| 308 |
+
"corpus_passages": 82719
|
| 309 |
+
}
|
| 310 |
+
}
|
evaluation/dimensions.json
CHANGED
|
The diff for this file is too large to render.
See raw diff
|
|
|
evaluation/figures.py
CHANGED
|
@@ -63,7 +63,7 @@ def reranker(summary: dict) -> str:
|
|
| 63 |
|
| 64 |
|
| 65 |
def retriever(summary: dict) -> str:
|
| 66 |
-
cutoffs =
|
| 67 |
x = list(range(len(cutoffs)))
|
| 68 |
panels = []
|
| 69 |
legend = [('Upstream Gemma', UPSTREAM, None), ('Daecore v2', FIT, None)]
|
|
|
|
| 63 |
|
| 64 |
|
| 65 |
def retriever(summary: dict) -> str:
|
| 66 |
+
cutoffs = sorted(map(int, summary['models']['upstream']))
|
| 67 |
x = list(range(len(cutoffs)))
|
| 68 |
panels = []
|
| 69 |
legend = [('Upstream Gemma', UPSTREAM, None), ('Daecore v2', FIT, None)]
|
evaluation/metrics.py
CHANGED
|
@@ -143,7 +143,13 @@ def summarize_dimensions(data: dict) -> dict:
|
|
| 143 |
panels[name] = {'queries': len(panel), 'ndcg@10': {
|
| 144 |
str(w): mean(row['ndcg@10'][str(w)] for row in panel) for w in widths
|
| 145 |
}}
|
| 146 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 147 |
|
| 148 |
|
| 149 |
def _summarize_ranked_grades(data: dict, *, include_selected: bool) -> dict:
|
|
@@ -155,6 +161,12 @@ def _summarize_ranked_grades(data: dict, *, include_selected: bool) -> dict:
|
|
| 155 |
rows = data['rows']
|
| 156 |
if not rows or len({row['id'] for row in rows}) != len(rows):
|
| 157 |
raise ValueError('Ranked rows require unique nonempty query identities')
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 158 |
# Validate before excluding fully abstained prefixes. An omitted query
|
| 159 |
# must never conceal a malformed grade, model map or selected depth.
|
| 160 |
for row in rows:
|
|
@@ -163,9 +175,9 @@ def _summarize_ranked_grades(data: dict, *, include_selected: bool) -> dict:
|
|
| 163 |
raise ValueError('Reference grade counts must cover grades zero through three')
|
| 164 |
for model in data['model_order']:
|
| 165 |
ranked, excluded = row['ranked_grades'][model], row['excluded'][model]
|
| 166 |
-
if (len(ranked) !=
|
| 167 |
or any(type(x) is not bool for x in excluded)):
|
| 168 |
-
raise ValueError('Ranked rows
|
| 169 |
if include_selected:
|
| 170 |
depth = row['selected_depth'][model]
|
| 171 |
if type(depth) is not int or not 3 <= depth <= 20:
|
|
@@ -174,7 +186,9 @@ def _summarize_ranked_grades(data: dict, *, include_selected: bool) -> dict:
|
|
| 174 |
for grade, drop in zip(ranked, excluded, strict=True)):
|
| 175 |
raise ValueError('Only declared abstentions may lack grades')
|
| 176 |
output = {'queries': len(rows), 'models': {}}
|
| 177 |
-
requested_cutoffs = (
|
|
|
|
|
|
|
| 178 |
for model in data['model_order']:
|
| 179 |
cutoffs = {}
|
| 180 |
for cutoff in requested_cutoffs:
|
|
|
|
| 143 |
panels[name] = {'queries': len(panel), 'ndcg@10': {
|
| 144 |
str(w): mean(row['ndcg@10'][str(w)] for row in panel) for w in widths
|
| 145 |
}}
|
| 146 |
+
result = {'queries': len(rows), 'vector_bytes_fp32': {str(w): w * 4 for w in widths}, 'public': panels}
|
| 147 |
+
if 'private' in data:
|
| 148 |
+
private = data['private']
|
| 149 |
+
if private['model_order'] != [str(w) for w in widths]:
|
| 150 |
+
raise ValueError('Private width rankings must cover the same ordered widths')
|
| 151 |
+
result['private'] = summarize_record('retriever', private)
|
| 152 |
+
return result
|
| 153 |
|
| 154 |
|
| 155 |
def _summarize_ranked_grades(data: dict, *, include_selected: bool) -> dict:
|
|
|
|
| 161 |
rows = data['rows']
|
| 162 |
if not rows or len({row['id'] for row in rows}) != len(rows):
|
| 163 |
raise ValueError('Ranked rows require unique nonempty query identities')
|
| 164 |
+
# Published promotion and serving v1 records contain twenty positions.
|
| 165 |
+
# Extended dense/serving records declare their measured depth explicitly.
|
| 166 |
+
depth_limit = data.get('ranking_depth', 20)
|
| 167 |
+
allowed_depths = (20, 50) if include_selected else (10, 20, 50)
|
| 168 |
+
if type(depth_limit) is not int or depth_limit not in allowed_depths:
|
| 169 |
+
raise ValueError('Unsupported measured ranking depth')
|
| 170 |
# Validate before excluding fully abstained prefixes. An omitted query
|
| 171 |
# must never conceal a malformed grade, model map or selected depth.
|
| 172 |
for row in rows:
|
|
|
|
| 175 |
raise ValueError('Reference grade counts must cover grades zero through three')
|
| 176 |
for model in data['model_order']:
|
| 177 |
ranked, excluded = row['ranked_grades'][model], row['excluded'][model]
|
| 178 |
+
if (len(ranked) != depth_limit or len(excluded) != depth_limit
|
| 179 |
or any(type(x) is not bool for x in excluded)):
|
| 180 |
+
raise ValueError('Ranked rows must cover exactly the declared ranking depth')
|
| 181 |
if include_selected:
|
| 182 |
depth = row['selected_depth'][model]
|
| 183 |
if type(depth) is not int or not 3 <= depth <= 20:
|
|
|
|
| 186 |
for grade, drop in zip(ranked, excluded, strict=True)):
|
| 187 |
raise ValueError('Only declared abstentions may lack grades')
|
| 188 |
output = {'queries': len(rows), 'models': {}}
|
| 189 |
+
requested_cutoffs = tuple(k for k in CUTOFFS if k <= depth_limit)
|
| 190 |
+
if include_selected:
|
| 191 |
+
requested_cutoffs += ('selected',)
|
| 192 |
for model in data['model_order']:
|
| 193 |
cutoffs = {}
|
| 194 |
for cutoff in requested_cutoffs:
|
evaluation/retriever-summary.json
CHANGED
|
@@ -8,9 +8,9 @@
|
|
| 8 |
"hit": 0.534020618556701,
|
| 9 |
"precision": 0.534020618556701,
|
| 10 |
"macro_precision": 0.534020618556701,
|
| 11 |
-
"ndcg": 0.
|
| 12 |
-
"known_positive_recall": 0.
|
| 13 |
-
"recall_queries":
|
| 14 |
"useful": 518,
|
| 15 |
"retained": 970,
|
| 16 |
"excluded_positions": 0,
|
|
@@ -23,9 +23,9 @@
|
|
| 23 |
"hit": 0.7298969072164948,
|
| 24 |
"precision": 0.5082530949105915,
|
| 25 |
"macro_precision": 0.5082474226804123,
|
| 26 |
-
"ndcg": 0.
|
| 27 |
-
"known_positive_recall": 0.
|
| 28 |
-
"recall_queries":
|
| 29 |
"useful": 1478,
|
| 30 |
"retained": 2908,
|
| 31 |
"excluded_positions": 2,
|
|
@@ -38,9 +38,9 @@
|
|
| 38 |
"hit": 0.7948453608247422,
|
| 39 |
"precision": 0.49948379103861246,
|
| 40 |
"macro_precision": 0.4996219931271478,
|
| 41 |
-
"ndcg": 0.
|
| 42 |
-
"known_positive_recall": 0.
|
| 43 |
-
"recall_queries":
|
| 44 |
"useful": 2419,
|
| 45 |
"retained": 4843,
|
| 46 |
"excluded_positions": 7,
|
|
@@ -53,9 +53,9 @@
|
|
| 53 |
"hit": 0.8484536082474227,
|
| 54 |
"precision": 0.4739561802397685,
|
| 55 |
"macro_precision": 0.47407543773523153,
|
| 56 |
-
"ndcg": 0.
|
| 57 |
-
"known_positive_recall": 0.
|
| 58 |
-
"recall_queries":
|
| 59 |
"useful": 4586,
|
| 60 |
"retained": 9676,
|
| 61 |
"excluded_positions": 24,
|
|
@@ -68,14 +68,29 @@
|
|
| 68 |
"hit": 0.8876288659793814,
|
| 69 |
"precision": 0.4393610421836228,
|
| 70 |
"macro_precision": 0.44005992535723476,
|
| 71 |
-
"ndcg": 0.
|
| 72 |
-
"known_positive_recall": 0.
|
| 73 |
-
"recall_queries":
|
| 74 |
"useful": 8499,
|
| 75 |
"retained": 19344,
|
| 76 |
"excluded_positions": 56,
|
| 77 |
"mean_useful": 8.761855670103094,
|
| 78 |
"mean_retained": 19.942268041237114
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 79 |
}
|
| 80 |
},
|
| 81 |
"finetuned": {
|
|
@@ -85,9 +100,9 @@
|
|
| 85 |
"hit": 0.6474226804123712,
|
| 86 |
"precision": 0.6474226804123712,
|
| 87 |
"macro_precision": 0.6474226804123712,
|
| 88 |
-
"ndcg": 0.
|
| 89 |
-
"known_positive_recall": 0.
|
| 90 |
-
"recall_queries":
|
| 91 |
"useful": 628,
|
| 92 |
"retained": 970,
|
| 93 |
"excluded_positions": 0,
|
|
@@ -100,9 +115,9 @@
|
|
| 100 |
"hit": 0.7938144329896907,
|
| 101 |
"precision": 0.6308009625300791,
|
| 102 |
"macro_precision": 0.6309278350515464,
|
| 103 |
-
"ndcg": 0.
|
| 104 |
-
"known_positive_recall": 0.
|
| 105 |
-
"recall_queries":
|
| 106 |
"useful": 1835,
|
| 107 |
"retained": 2909,
|
| 108 |
"excluded_positions": 1,
|
|
@@ -115,9 +130,9 @@
|
|
| 115 |
"hit": 0.8474226804123711,
|
| 116 |
"precision": 0.6266005782734407,
|
| 117 |
"macro_precision": 0.6269759450171821,
|
| 118 |
-
"ndcg": 0.
|
| 119 |
-
"known_positive_recall": 0.
|
| 120 |
-
"recall_queries":
|
| 121 |
"useful": 3034,
|
| 122 |
"retained": 4842,
|
| 123 |
"excluded_positions": 8,
|
|
@@ -130,9 +145,9 @@
|
|
| 130 |
"hit": 0.8814432989690721,
|
| 131 |
"precision": 0.6117999586691465,
|
| 132 |
"macro_precision": 0.6123367697594502,
|
| 133 |
-
"ndcg": 0.
|
| 134 |
-
"known_positive_recall": 0.
|
| 135 |
-
"recall_queries":
|
| 136 |
"useful": 5921,
|
| 137 |
"retained": 9678,
|
| 138 |
"excluded_positions": 22,
|
|
@@ -145,14 +160,29 @@
|
|
| 145 |
"hit": 0.9051546391752577,
|
| 146 |
"precision": 0.5810720008270016,
|
| 147 |
"macro_precision": 0.5815832221977094,
|
| 148 |
-
"ndcg": 0.
|
| 149 |
-
"known_positive_recall": 0.
|
| 150 |
-
"recall_queries":
|
| 151 |
"useful": 11242,
|
| 152 |
"retained": 19347,
|
| 153 |
"excluded_positions": 53,
|
| 154 |
"mean_useful": 11.589690721649484,
|
| 155 |
"mean_retained": 19.945360824742266
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 156 |
}
|
| 157 |
}
|
| 158 |
},
|
|
|
|
| 8 |
"hit": 0.534020618556701,
|
| 9 |
"precision": 0.534020618556701,
|
| 10 |
"macro_precision": 0.534020618556701,
|
| 11 |
+
"ndcg": 0.4592047128129602,
|
| 12 |
+
"known_positive_recall": 0.014444554022554491,
|
| 13 |
+
"recall_queries": 925,
|
| 14 |
"useful": 518,
|
| 15 |
"retained": 970,
|
| 16 |
"excluded_positions": 0,
|
|
|
|
| 23 |
"hit": 0.7298969072164948,
|
| 24 |
"precision": 0.5082530949105915,
|
| 25 |
"macro_precision": 0.5082474226804123,
|
| 26 |
+
"ndcg": 0.4413745306026501,
|
| 27 |
+
"known_positive_recall": 0.03938697535513605,
|
| 28 |
+
"recall_queries": 925,
|
| 29 |
"useful": 1478,
|
| 30 |
"retained": 2908,
|
| 31 |
"excluded_positions": 2,
|
|
|
|
| 38 |
"hit": 0.7948453608247422,
|
| 39 |
"precision": 0.49948379103861246,
|
| 40 |
"macro_precision": 0.4996219931271478,
|
| 41 |
+
"ndcg": 0.43462885647387695,
|
| 42 |
+
"known_positive_recall": 0.06172226483192968,
|
| 43 |
+
"recall_queries": 925,
|
| 44 |
"useful": 2419,
|
| 45 |
"retained": 4843,
|
| 46 |
"excluded_positions": 7,
|
|
|
|
| 53 |
"hit": 0.8484536082474227,
|
| 54 |
"precision": 0.4739561802397685,
|
| 55 |
"macro_precision": 0.47407543773523153,
|
| 56 |
+
"ndcg": 0.43629894440575245,
|
| 57 |
+
"known_positive_recall": 0.11180689389376004,
|
| 58 |
+
"recall_queries": 925,
|
| 59 |
"useful": 4586,
|
| 60 |
"retained": 9676,
|
| 61 |
"excluded_positions": 24,
|
|
|
|
| 68 |
"hit": 0.8876288659793814,
|
| 69 |
"precision": 0.4393610421836228,
|
| 70 |
"macro_precision": 0.44005992535723476,
|
| 71 |
+
"ndcg": 0.4525661889832219,
|
| 72 |
+
"known_positive_recall": 0.19611550928151497,
|
| 73 |
+
"recall_queries": 925,
|
| 74 |
"useful": 8499,
|
| 75 |
"retained": 19344,
|
| 76 |
"excluded_positions": 56,
|
| 77 |
"mean_useful": 8.761855670103094,
|
| 78 |
"mean_retained": 19.942268041237114
|
| 79 |
+
},
|
| 80 |
+
"50": {
|
| 81 |
+
"scored_queries": 970,
|
| 82 |
+
"excluded_queries": 0,
|
| 83 |
+
"hit": 0.931958762886598,
|
| 84 |
+
"precision": 0.38598560079443894,
|
| 85 |
+
"macro_precision": 0.3867178488702632,
|
| 86 |
+
"ndcg": 0.5107049932805754,
|
| 87 |
+
"known_positive_recall": 0.41118693689165037,
|
| 88 |
+
"recall_queries": 925,
|
| 89 |
+
"useful": 18657,
|
| 90 |
+
"retained": 48336,
|
| 91 |
+
"excluded_positions": 164,
|
| 92 |
+
"mean_useful": 19.234020618556702,
|
| 93 |
+
"mean_retained": 49.83092783505155
|
| 94 |
}
|
| 95 |
},
|
| 96 |
"finetuned": {
|
|
|
|
| 100 |
"hit": 0.6474226804123712,
|
| 101 |
"precision": 0.6474226804123712,
|
| 102 |
"macro_precision": 0.6474226804123712,
|
| 103 |
+
"ndcg": 0.5312714776632302,
|
| 104 |
+
"known_positive_recall": 0.017869431761799028,
|
| 105 |
+
"recall_queries": 925,
|
| 106 |
"useful": 628,
|
| 107 |
"retained": 970,
|
| 108 |
"excluded_positions": 0,
|
|
|
|
| 115 |
"hit": 0.7938144329896907,
|
| 116 |
"precision": 0.6308009625300791,
|
| 117 |
"macro_precision": 0.6309278350515464,
|
| 118 |
+
"ndcg": 0.5406982958852574,
|
| 119 |
+
"known_positive_recall": 0.05034823759106709,
|
| 120 |
+
"recall_queries": 925,
|
| 121 |
"useful": 1835,
|
| 122 |
"retained": 2909,
|
| 123 |
"excluded_positions": 1,
|
|
|
|
| 130 |
"hit": 0.8474226804123711,
|
| 131 |
"precision": 0.6266005782734407,
|
| 132 |
"macro_precision": 0.6269759450171821,
|
| 133 |
+
"ndcg": 0.5453498298950709,
|
| 134 |
+
"known_positive_recall": 0.08055790386567772,
|
| 135 |
+
"recall_queries": 925,
|
| 136 |
"useful": 3034,
|
| 137 |
"retained": 4842,
|
| 138 |
"excluded_positions": 8,
|
|
|
|
| 145 |
"hit": 0.8814432989690721,
|
| 146 |
"precision": 0.6117999586691465,
|
| 147 |
"macro_precision": 0.6123367697594502,
|
| 148 |
+
"ndcg": 0.5574064002423099,
|
| 149 |
+
"known_positive_recall": 0.15074270958466607,
|
| 150 |
+
"recall_queries": 925,
|
| 151 |
"useful": 5921,
|
| 152 |
"retained": 9678,
|
| 153 |
"excluded_positions": 22,
|
|
|
|
| 160 |
"hit": 0.9051546391752577,
|
| 161 |
"precision": 0.5810720008270016,
|
| 162 |
"macro_precision": 0.5815832221977094,
|
| 163 |
+
"ndcg": 0.5818739813983459,
|
| 164 |
+
"known_positive_recall": 0.27556154002694083,
|
| 165 |
+
"recall_queries": 925,
|
| 166 |
"useful": 11242,
|
| 167 |
"retained": 19347,
|
| 168 |
"excluded_positions": 53,
|
| 169 |
"mean_useful": 11.589690721649484,
|
| 170 |
"mean_retained": 19.945360824742266
|
| 171 |
+
},
|
| 172 |
+
"50": {
|
| 173 |
+
"scored_queries": 970,
|
| 174 |
+
"excluded_queries": 0,
|
| 175 |
+
"hit": 0.9443298969072165,
|
| 176 |
+
"precision": 0.5045207208325575,
|
| 177 |
+
"macro_precision": 0.5051647365119746,
|
| 178 |
+
"ndcg": 0.6449716439245425,
|
| 179 |
+
"known_positive_recall": 0.5791516627126307,
|
| 180 |
+
"recall_queries": 925,
|
| 181 |
+
"useful": 24385,
|
| 182 |
+
"retained": 48333,
|
| 183 |
+
"excluded_positions": 167,
|
| 184 |
+
"mean_useful": 25.13917525773196,
|
| 185 |
+
"mean_retained": 49.827835051546394
|
| 186 |
}
|
| 187 |
}
|
| 188 |
},
|
evaluation/retriever.json
CHANGED
|
The diff for this file is too large to render.
See raw diff
|
|
|
evaluation/serving-summary.json
CHANGED
|
@@ -8,9 +8,9 @@
|
|
| 8 |
"hit": 0.7233160621761658,
|
| 9 |
"precision": 0.7233160621761658,
|
| 10 |
"macro_precision": 0.7233160621761658,
|
| 11 |
-
"ndcg": 0.
|
| 12 |
-
"known_positive_recall": 0.
|
| 13 |
-
"recall_queries":
|
| 14 |
"useful": 698,
|
| 15 |
"retained": 965,
|
| 16 |
"excluded_positions": 0,
|
|
@@ -23,9 +23,9 @@
|
|
| 23 |
"hit": 0.8536082474226804,
|
| 24 |
"precision": 0.7157640565712314,
|
| 25 |
"macro_precision": 0.7166666666666667,
|
| 26 |
-
"ndcg": 0.
|
| 27 |
-
"known_positive_recall": 0.
|
| 28 |
-
"recall_queries":
|
| 29 |
"useful": 2075,
|
| 30 |
"retained": 2899,
|
| 31 |
"excluded_positions": 11,
|
|
@@ -38,9 +38,9 @@
|
|
| 38 |
"hit": 0.8845360824742268,
|
| 39 |
"precision": 0.7028311634635255,
|
| 40 |
"macro_precision": 0.7031786941580757,
|
| 41 |
-
"ndcg": 0.
|
| 42 |
-
"known_positive_recall": 0.
|
| 43 |
-
"recall_queries":
|
| 44 |
"useful": 3401,
|
| 45 |
"retained": 4839,
|
| 46 |
"excluded_positions": 11,
|
|
@@ -53,9 +53,9 @@
|
|
| 53 |
"hit": 0.8989690721649485,
|
| 54 |
"precision": 0.67180070291503,
|
| 55 |
"macro_precision": 0.6722631320569465,
|
| 56 |
-
"ndcg": 0.
|
| 57 |
-
"known_positive_recall": 0.
|
| 58 |
-
"recall_queries":
|
| 59 |
"useful": 6499,
|
| 60 |
"retained": 9674,
|
| 61 |
"excluded_positions": 26,
|
|
@@ -68,24 +68,39 @@
|
|
| 68 |
"hit": 0.9103092783505154,
|
| 69 |
"precision": 0.61794500723589,
|
| 70 |
"macro_precision": 0.6184956360634603,
|
| 71 |
-
"ndcg": 0.
|
| 72 |
-
"known_positive_recall": 0.
|
| 73 |
-
"recall_queries":
|
| 74 |
"useful": 11956,
|
| 75 |
"retained": 19348,
|
| 76 |
"excluded_positions": 52,
|
| 77 |
"mean_useful": 12.32577319587629,
|
| 78 |
"mean_retained": 19.94639175257732
|
| 79 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 80 |
"selected": {
|
| 81 |
"scored_queries": 970,
|
| 82 |
"excluded_queries": 0,
|
| 83 |
"hit": 0.8618556701030928,
|
| 84 |
"precision": 0.8033770583310076,
|
| 85 |
"macro_precision": 0.7040325511660611,
|
| 86 |
-
"ndcg": 0.
|
| 87 |
-
"known_positive_recall": 0.
|
| 88 |
-
"recall_queries":
|
| 89 |
"useful": 5757,
|
| 90 |
"retained": 7166,
|
| 91 |
"excluded_positions": 17,
|
|
@@ -100,9 +115,9 @@
|
|
| 100 |
"hit": 0.7243523316062176,
|
| 101 |
"precision": 0.7243523316062176,
|
| 102 |
"macro_precision": 0.7243523316062176,
|
| 103 |
-
"ndcg": 0.
|
| 104 |
-
"known_positive_recall": 0.
|
| 105 |
-
"recall_queries":
|
| 106 |
"useful": 699,
|
| 107 |
"retained": 965,
|
| 108 |
"excluded_positions": 0,
|
|
@@ -115,9 +130,9 @@
|
|
| 115 |
"hit": 0.8525773195876288,
|
| 116 |
"precision": 0.7157640565712314,
|
| 117 |
"macro_precision": 0.7166666666666667,
|
| 118 |
-
"ndcg": 0.
|
| 119 |
-
"known_positive_recall": 0.
|
| 120 |
-
"recall_queries":
|
| 121 |
"useful": 2075,
|
| 122 |
"retained": 2899,
|
| 123 |
"excluded_positions": 11,
|
|
@@ -130,9 +145,9 @@
|
|
| 130 |
"hit": 0.8845360824742268,
|
| 131 |
"precision": 0.7023563455973543,
|
| 132 |
"macro_precision": 0.702766323024055,
|
| 133 |
-
"ndcg": 0.
|
| 134 |
-
"known_positive_recall": 0.
|
| 135 |
-
"recall_queries":
|
| 136 |
"useful": 3398,
|
| 137 |
"retained": 4838,
|
| 138 |
"excluded_positions": 12,
|
|
@@ -145,9 +160,9 @@
|
|
| 145 |
"hit": 0.8989690721649485,
|
| 146 |
"precision": 0.6723514211886304,
|
| 147 |
"macro_precision": 0.6727589592538046,
|
| 148 |
-
"ndcg": 0.
|
| 149 |
-
"known_positive_recall": 0.
|
| 150 |
-
"recall_queries":
|
| 151 |
"useful": 6505,
|
| 152 |
"retained": 9675,
|
| 153 |
"excluded_positions": 25,
|
|
@@ -160,24 +175,39 @@
|
|
| 160 |
"hit": 0.9103092783505154,
|
| 161 |
"precision": 0.6178416373785404,
|
| 162 |
"macro_precision": 0.6183925432799551,
|
| 163 |
-
"ndcg": 0.
|
| 164 |
-
"known_positive_recall": 0.
|
| 165 |
-
"recall_queries":
|
| 166 |
"useful": 11954,
|
| 167 |
"retained": 19348,
|
| 168 |
"excluded_positions": 52,
|
| 169 |
"mean_useful": 12.323711340206186,
|
| 170 |
"mean_retained": 19.94639175257732
|
| 171 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 172 |
"selected": {
|
| 173 |
"scored_queries": 970,
|
| 174 |
"excluded_queries": 0,
|
| 175 |
"hit": 0.8608247422680413,
|
| 176 |
"precision": 0.8027855153203343,
|
| 177 |
"macro_precision": 0.7038533836490732,
|
| 178 |
-
"ndcg": 0.
|
| 179 |
-
"known_positive_recall": 0.
|
| 180 |
-
"recall_queries":
|
| 181 |
"useful": 5764,
|
| 182 |
"retained": 7180,
|
| 183 |
"excluded_positions": 17,
|
|
|
|
| 8 |
"hit": 0.7233160621761658,
|
| 9 |
"precision": 0.7233160621761658,
|
| 10 |
"macro_precision": 0.7233160621761658,
|
| 11 |
+
"ndcg": 0.6458425857389588,
|
| 12 |
+
"known_positive_recall": 0.021209009540392988,
|
| 13 |
+
"recall_queries": 920,
|
| 14 |
"useful": 698,
|
| 15 |
"retained": 965,
|
| 16 |
"excluded_positions": 0,
|
|
|
|
| 23 |
"hit": 0.8536082474226804,
|
| 24 |
"precision": 0.7157640565712314,
|
| 25 |
"macro_precision": 0.7166666666666667,
|
| 26 |
+
"ndcg": 0.6374467700901159,
|
| 27 |
+
"known_positive_recall": 0.0589083117861917,
|
| 28 |
+
"recall_queries": 925,
|
| 29 |
"useful": 2075,
|
| 30 |
"retained": 2899,
|
| 31 |
"excluded_positions": 11,
|
|
|
|
| 38 |
"hit": 0.8845360824742268,
|
| 39 |
"precision": 0.7028311634635255,
|
| 40 |
"macro_precision": 0.7031786941580757,
|
| 41 |
+
"ndcg": 0.632593817782834,
|
| 42 |
+
"known_positive_recall": 0.09351524011264825,
|
| 43 |
+
"recall_queries": 925,
|
| 44 |
"useful": 3401,
|
| 45 |
"retained": 4839,
|
| 46 |
"excluded_positions": 11,
|
|
|
|
| 53 |
"hit": 0.8989690721649485,
|
| 54 |
"precision": 0.67180070291503,
|
| 55 |
"macro_precision": 0.6722631320569465,
|
| 56 |
+
"ndcg": 0.6315462046841925,
|
| 57 |
+
"known_positive_recall": 0.17039667688880922,
|
| 58 |
+
"recall_queries": 925,
|
| 59 |
"useful": 6499,
|
| 60 |
"retained": 9674,
|
| 61 |
"excluded_positions": 26,
|
|
|
|
| 68 |
"hit": 0.9103092783505154,
|
| 69 |
"precision": 0.61794500723589,
|
| 70 |
"macro_precision": 0.6184956360634603,
|
| 71 |
+
"ndcg": 0.6416731640651591,
|
| 72 |
+
"known_positive_recall": 0.29717756138114526,
|
| 73 |
+
"recall_queries": 925,
|
| 74 |
"useful": 11956,
|
| 75 |
"retained": 19348,
|
| 76 |
"excluded_positions": 52,
|
| 77 |
"mean_useful": 12.32577319587629,
|
| 78 |
"mean_retained": 19.94639175257732
|
| 79 |
},
|
| 80 |
+
"50": {
|
| 81 |
+
"scored_queries": 970,
|
| 82 |
+
"excluded_queries": 0,
|
| 83 |
+
"hit": 0.9309278350515464,
|
| 84 |
+
"precision": 0.4043216890119198,
|
| 85 |
+
"macro_precision": 0.4044939351610387,
|
| 86 |
+
"ndcg": 0.6051209510262134,
|
| 87 |
+
"known_positive_recall": 0.4598077844640693,
|
| 88 |
+
"recall_queries": 925,
|
| 89 |
+
"useful": 19572,
|
| 90 |
+
"retained": 48407,
|
| 91 |
+
"excluded_positions": 93,
|
| 92 |
+
"mean_useful": 20.177319587628865,
|
| 93 |
+
"mean_retained": 49.904123711340205
|
| 94 |
+
},
|
| 95 |
"selected": {
|
| 96 |
"scored_queries": 970,
|
| 97 |
"excluded_queries": 0,
|
| 98 |
"hit": 0.8618556701030928,
|
| 99 |
"precision": 0.8033770583310076,
|
| 100 |
"macro_precision": 0.7040325511660611,
|
| 101 |
+
"ndcg": 0.6359990281117442,
|
| 102 |
+
"known_positive_recall": 0.128026594631235,
|
| 103 |
+
"recall_queries": 925,
|
| 104 |
"useful": 5757,
|
| 105 |
"retained": 7166,
|
| 106 |
"excluded_positions": 17,
|
|
|
|
| 115 |
"hit": 0.7243523316062176,
|
| 116 |
"precision": 0.7243523316062176,
|
| 117 |
"macro_precision": 0.7243523316062176,
|
| 118 |
+
"ndcg": 0.6461386627189736,
|
| 119 |
+
"known_positive_recall": 0.021228078953055077,
|
| 120 |
+
"recall_queries": 920,
|
| 121 |
"useful": 699,
|
| 122 |
"retained": 965,
|
| 123 |
"excluded_positions": 0,
|
|
|
|
| 130 |
"hit": 0.8525773195876288,
|
| 131 |
"precision": 0.7157640565712314,
|
| 132 |
"macro_precision": 0.7166666666666667,
|
| 133 |
+
"ndcg": 0.6373247400469291,
|
| 134 |
+
"known_positive_recall": 0.05893952330659241,
|
| 135 |
+
"recall_queries": 925,
|
| 136 |
"useful": 2075,
|
| 137 |
"retained": 2899,
|
| 138 |
"excluded_positions": 11,
|
|
|
|
| 145 |
"hit": 0.8845360824742268,
|
| 146 |
"precision": 0.7023563455973543,
|
| 147 |
"macro_precision": 0.702766323024055,
|
| 148 |
+
"ndcg": 0.6321994825191731,
|
| 149 |
+
"known_positive_recall": 0.09344587015679974,
|
| 150 |
+
"recall_queries": 925,
|
| 151 |
"useful": 3398,
|
| 152 |
"retained": 4838,
|
| 153 |
"excluded_positions": 12,
|
|
|
|
| 160 |
"hit": 0.8989690721649485,
|
| 161 |
"precision": 0.6723514211886304,
|
| 162 |
"macro_precision": 0.6727589592538046,
|
| 163 |
+
"ndcg": 0.631891212140629,
|
| 164 |
+
"known_positive_recall": 0.17053838380216288,
|
| 165 |
+
"recall_queries": 925,
|
| 166 |
"useful": 6505,
|
| 167 |
"retained": 9675,
|
| 168 |
"excluded_positions": 25,
|
|
|
|
| 175 |
"hit": 0.9103092783505154,
|
| 176 |
"precision": 0.6178416373785404,
|
| 177 |
"macro_precision": 0.6183925432799551,
|
| 178 |
+
"ndcg": 0.6416203428445276,
|
| 179 |
+
"known_positive_recall": 0.2971249330393808,
|
| 180 |
+
"recall_queries": 925,
|
| 181 |
"useful": 11954,
|
| 182 |
"retained": 19348,
|
| 183 |
"excluded_positions": 52,
|
| 184 |
"mean_useful": 12.323711340206186,
|
| 185 |
"mean_retained": 19.94639175257732
|
| 186 |
},
|
| 187 |
+
"50": {
|
| 188 |
+
"scored_queries": 970,
|
| 189 |
+
"excluded_queries": 0,
|
| 190 |
+
"hit": 0.9309278350515464,
|
| 191 |
+
"precision": 0.4043216890119198,
|
| 192 |
+
"macro_precision": 0.4044939351610387,
|
| 193 |
+
"ndcg": 0.6050905860403181,
|
| 194 |
+
"known_positive_recall": 0.4598077844640693,
|
| 195 |
+
"recall_queries": 925,
|
| 196 |
+
"useful": 19572,
|
| 197 |
+
"retained": 48407,
|
| 198 |
+
"excluded_positions": 93,
|
| 199 |
+
"mean_useful": 20.177319587628865,
|
| 200 |
+
"mean_retained": 49.904123711340205
|
| 201 |
+
},
|
| 202 |
"selected": {
|
| 203 |
"scored_queries": 970,
|
| 204 |
"excluded_queries": 0,
|
| 205 |
"hit": 0.8608247422680413,
|
| 206 |
"precision": 0.8027855153203343,
|
| 207 |
"macro_precision": 0.7038533836490732,
|
| 208 |
+
"ndcg": 0.6358997417721292,
|
| 209 |
+
"known_positive_recall": 0.12825382711378852,
|
| 210 |
+
"recall_queries": 925,
|
| 211 |
"useful": 5764,
|
| 212 |
"retained": 7180,
|
| 213 |
"excluded_positions": 17,
|
evaluation/serving.json
CHANGED
|
The diff for this file is too large to render.
See raw diff
|
|
|
figures/gemma-comparison.svg
CHANGED
|
|
|
|
publication-manifest.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
| 1 |
{
|
| 2 |
"schema": "daecore.retrieval-publication-manifest",
|
| 3 |
-
"tool_sha256": "
|
| 4 |
"model_id": "embeddinggemma-300m-memory-ft-v2",
|
| 5 |
"public_repo": "Daecore/embeddinggemma-300m-memory-ft-v2",
|
| 6 |
"role": "dense-retriever",
|
|
@@ -12,7 +12,7 @@
|
|
| 12 |
"export_receipt_sha256": "e24e1692f7e03c9af8040bb3f6e310a061d305f485861007f1f37d36808b20d8",
|
| 13 |
"serving_sha256": "8c088b1635875d5406e148782b303cfc5a961138600783fb4842b413b98a98c3",
|
| 14 |
"source_model_tree_sha256": "30a7b9f1c42b1fcee29a6729742685d5b744e8fa87d13cefab2fd2d49ea9150e",
|
| 15 |
-
"staged_at": "2026-09-
|
| 16 |
"files": {
|
| 17 |
"added_tokens.json": {
|
| 18 |
"sha256": "50b2f405ba56a26d4913fd772089992252d7f942123cc0a034d96424221ba946",
|
|
@@ -80,8 +80,8 @@
|
|
| 80 |
"binding": "packaging record"
|
| 81 |
},
|
| 82 |
"evaluation/README.md": {
|
| 83 |
-
"sha256": "
|
| 84 |
-
"size":
|
| 85 |
"binding": "packaging record"
|
| 86 |
},
|
| 87 |
"evaluation/serving-qualification.json": {
|
|
@@ -90,18 +90,23 @@
|
|
| 90 |
"binding": "packaging record"
|
| 91 |
},
|
| 92 |
"evaluation/metrics.py": {
|
| 93 |
-
"sha256": "
|
| 94 |
-
"size":
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 95 |
"binding": "packaging record"
|
| 96 |
},
|
| 97 |
"evaluation/retriever-summary.json": {
|
| 98 |
-
"sha256": "
|
| 99 |
-
"size":
|
| 100 |
"binding": "packaging record"
|
| 101 |
},
|
| 102 |
"evaluation/dimensions-summary.json": {
|
| 103 |
-
"sha256": "
|
| 104 |
-
"size":
|
| 105 |
"binding": "packaging record"
|
| 106 |
},
|
| 107 |
"evaluation/promotion-summary.json": {
|
|
@@ -110,18 +115,18 @@
|
|
| 110 |
"binding": "packaging record"
|
| 111 |
},
|
| 112 |
"evaluation/serving-summary.json": {
|
| 113 |
-
"sha256": "
|
| 114 |
-
"size":
|
| 115 |
"binding": "packaging record"
|
| 116 |
},
|
| 117 |
"evaluation/retriever.json": {
|
| 118 |
-
"sha256": "
|
| 119 |
-
"size":
|
| 120 |
"binding": "packaging record"
|
| 121 |
},
|
| 122 |
"evaluation/dimensions.json": {
|
| 123 |
-
"sha256": "
|
| 124 |
-
"size":
|
| 125 |
"binding": "packaging record"
|
| 126 |
},
|
| 127 |
"evaluation/promotion.json": {
|
|
@@ -130,13 +135,13 @@
|
|
| 130 |
"binding": "packaging record"
|
| 131 |
},
|
| 132 |
"evaluation/serving.json": {
|
| 133 |
-
"sha256": "
|
| 134 |
-
"size":
|
| 135 |
"binding": "packaging record"
|
| 136 |
},
|
| 137 |
"evaluation/figures.py": {
|
| 138 |
-
"sha256": "
|
| 139 |
-
"size":
|
| 140 |
"binding": "packaging record"
|
| 141 |
},
|
| 142 |
"evaluation/svg_figures.py": {
|
|
@@ -145,15 +150,15 @@
|
|
| 145 |
"binding": "packaging record"
|
| 146 |
},
|
| 147 |
"figures/gemma-comparison.svg": {
|
| 148 |
-
"sha256": "
|
| 149 |
-
"size":
|
| 150 |
"binding": "packaging record"
|
| 151 |
},
|
| 152 |
"README.md": {
|
| 153 |
-
"sha256": "
|
| 154 |
-
"size":
|
| 155 |
"binding": "model card with upload-relative links",
|
| 156 |
-
"source_sha256": "
|
| 157 |
}
|
| 158 |
}
|
| 159 |
}
|
|
|
|
| 1 |
{
|
| 2 |
"schema": "daecore.retrieval-publication-manifest",
|
| 3 |
+
"tool_sha256": "46dcb4b17750d5a4ba3d44b1236f33e9e5657673b648ae6cad48fe4ac4f147f8",
|
| 4 |
"model_id": "embeddinggemma-300m-memory-ft-v2",
|
| 5 |
"public_repo": "Daecore/embeddinggemma-300m-memory-ft-v2",
|
| 6 |
"role": "dense-retriever",
|
|
|
|
| 12 |
"export_receipt_sha256": "e24e1692f7e03c9af8040bb3f6e310a061d305f485861007f1f37d36808b20d8",
|
| 13 |
"serving_sha256": "8c088b1635875d5406e148782b303cfc5a961138600783fb4842b413b98a98c3",
|
| 14 |
"source_model_tree_sha256": "30a7b9f1c42b1fcee29a6729742685d5b744e8fa87d13cefab2fd2d49ea9150e",
|
| 15 |
+
"staged_at": "2026-09-30T01:53:48+00:00",
|
| 16 |
"files": {
|
| 17 |
"added_tokens.json": {
|
| 18 |
"sha256": "50b2f405ba56a26d4913fd772089992252d7f942123cc0a034d96424221ba946",
|
|
|
|
| 80 |
"binding": "packaging record"
|
| 81 |
},
|
| 82 |
"evaluation/README.md": {
|
| 83 |
+
"sha256": "390b5e5d525ae6c0b2c3370050af1cc14b4574127bb258db911fb0afed670494",
|
| 84 |
+
"size": 9703,
|
| 85 |
"binding": "packaging record"
|
| 86 |
},
|
| 87 |
"evaluation/serving-qualification.json": {
|
|
|
|
| 90 |
"binding": "packaging record"
|
| 91 |
},
|
| 92 |
"evaluation/metrics.py": {
|
| 93 |
+
"sha256": "bbccf29f34db24aefc5630adcaf9630e5f214010902dbe1df0032618aa329a45",
|
| 94 |
+
"size": 18329,
|
| 95 |
+
"binding": "packaging record"
|
| 96 |
+
},
|
| 97 |
+
"evaluation/chunking.md": {
|
| 98 |
+
"sha256": "cda82fca3a467f02e06625a607fc8a5bfc9b16d48f4fe225d7f2282330d4955e",
|
| 99 |
+
"size": 3509,
|
| 100 |
"binding": "packaging record"
|
| 101 |
},
|
| 102 |
"evaluation/retriever-summary.json": {
|
| 103 |
+
"sha256": "b320430acd9938d08e31eb00693dd9f775e9890b17427f86577c9bbd91dd1bba",
|
| 104 |
+
"size": 6071,
|
| 105 |
"binding": "packaging record"
|
| 106 |
},
|
| 107 |
"evaluation/dimensions-summary.json": {
|
| 108 |
+
"sha256": "3b160deb874b98e835b92b15070afe36e548b1ce692a7ef817af77026e5064f8",
|
| 109 |
+
"size": 9753,
|
| 110 |
"binding": "packaging record"
|
| 111 |
},
|
| 112 |
"evaluation/promotion-summary.json": {
|
|
|
|
| 115 |
"binding": "packaging record"
|
| 116 |
},
|
| 117 |
"evaluation/serving-summary.json": {
|
| 118 |
+
"sha256": "cd87757b53482e2ec9450a3c64f31a32afad2710900fbcbd7885828929aa9535",
|
| 119 |
+
"size": 7035,
|
| 120 |
"binding": "packaging record"
|
| 121 |
},
|
| 122 |
"evaluation/retriever.json": {
|
| 123 |
+
"sha256": "c209c96fd17144a6f7d2c7174bc27db4b346288910d085e2b614286b80ae4559",
|
| 124 |
+
"size": 1156315,
|
| 125 |
"binding": "packaging record"
|
| 126 |
},
|
| 127 |
"evaluation/dimensions.json": {
|
| 128 |
+
"sha256": "3d9b49d6bda19fe9f9b7f8859ce550defa230757fef85ee2bc1ff175f6f4fe4b",
|
| 129 |
+
"size": 1063026,
|
| 130 |
"binding": "packaging record"
|
| 131 |
},
|
| 132 |
"evaluation/promotion.json": {
|
|
|
|
| 135 |
"binding": "packaging record"
|
| 136 |
},
|
| 137 |
"evaluation/serving.json": {
|
| 138 |
+
"sha256": "f4870a5118400152d4a78c7fd7c96df4df73901bfe0c590fdcd00c2af477a4be",
|
| 139 |
+
"size": 1184937,
|
| 140 |
"binding": "packaging record"
|
| 141 |
},
|
| 142 |
"evaluation/figures.py": {
|
| 143 |
+
"sha256": "6bc89020363df7a0cfcdebaa4e8be5b37385ab97b2b7bbf3f60cb9b46b2ac63e",
|
| 144 |
+
"size": 5408,
|
| 145 |
"binding": "packaging record"
|
| 146 |
},
|
| 147 |
"evaluation/svg_figures.py": {
|
|
|
|
| 150 |
"binding": "packaging record"
|
| 151 |
},
|
| 152 |
"figures/gemma-comparison.svg": {
|
| 153 |
+
"sha256": "4393db8afc6de32877a66b863708e396a1503d753932680dc7e4e590c4ccd513",
|
| 154 |
+
"size": 14507,
|
| 155 |
"binding": "packaging record"
|
| 156 |
},
|
| 157 |
"README.md": {
|
| 158 |
+
"sha256": "b5fe1ab8b6d3821cf5ee42e9ffbb2d88d8509233d27834382ce63a385f5d60a4",
|
| 159 |
+
"size": 16388,
|
| 160 |
"binding": "model card with upload-relative links",
|
| 161 |
+
"source_sha256": "58c1418eb860b7e8d2db5b84085780bb665ffbd92d06224a59cff700cebe0f02"
|
| 162 |
}
|
| 163 |
}
|
| 164 |
}
|