tnh0527's picture
Update ettin-150m-memory-reranker-ft-v1 documentation
6c038cf verified
|
Raw History Blame Contribute Delete
7.9 kB
# Ettin reranker evaluation
These records support the [model card's](../README.md) ranking results and compare Ettin's CUDA and Vulkan serving paths. They contain anonymized relevance grades and scores, without private queries or passage text.
The [chunking guide](chunking.md#retrieval-model-data) explains the synthetic
source documents and prepared chunks that form the private candidate pools.
| File | Contents |
|---|---|
| `reranker.json` | Grades and upstream and fine-tuned scores for 970 fixed pools of 50 passages |
| `serving.json` | Reviewed top-50 hybrid rankings from CUDA and Vulkan on the same 970 queries |
| `reranker-summary.json`, `serving-summary.json` | The summaries that `metrics.py` reproduces |
| `serving-qualification.json` | Serving checks, public reranking results and tested limits |
| `metrics.py`, `figures.py`, `svg_figures.py` | Metric and figure code |
## Reproduce the tables
From this directory, run with Python 3.11 or later:
```sh
python metrics.py --directory .
python figures.py --output ../figures
```
No packages or network access are needed. The outputs match `reranker-summary.json` and `serving-summary.json`; the second command redraws the card's comparison figure. This recomputes saved results, rather than running inference again. The shared pipeline table uses `serving-summary.json`'s `answerable` section; its outer `models` section preserves all-query coverage.
## Ranking quality
`reranker.json` holds the exact upstream and fine-tuned model identities and their rankings on the same 970 pools of 50 passages. All 48,500 pairs were judged; 873 pools contain useful evidence. Both models ran in PyTorch FP16 with the same 1,153-token pair construction, Transformers 5.2.0 and Sentence Transformers 5.5.1. Fresh inference on three complete pools checked the fine-tune's saved scores. This comparison measures the trained models, not ONNX latency.
The model card presents results on the 873 answerable pools. The records also report all-query Hit and precision. Conditional nDCG and recall use the answerable pools. Grades 2 and 3 count as useful; nDCG uses gain `2**grade - 1` with logarithmic rank discount. Ties retain candidate order.
The card shows cutoffs 1, 3, 5, 10 and 20, plus the full pool at 50. Daecore reranks up to 50 candidates and returns a selected prefix of 3–20. Recall here is relative to useful passages in the candidate pool. At 50, answerable-pool Hit and recall are 100% and precision is 42.68% for both models; these are properties of the shared pool. The 97 pools without useful evidence are excluded from the ranking-quality comparison because no ordering can make them answerable. The saved all-query metrics describe candidate coverage, not Ettin's conditional ranking quality. nDCG@50 remains sensitive to order: 0.7891 upstream and 0.8753 for the fine-tune.
Random ordering is an exact within-pool reference: expected precision is R/N and Hit@k is `1 - C(N-R, k)/C(N, k)`, for N candidates and R useful passages. It already reaches 97.30% Hit@20 on answerable pools, so early precision and graded ranking are more informative here.
The FiQA and SciFact results in `serving-qualification.json` rerank fixed pools of 50 candidates retrieved by upstream Gemma, with official relevance judgments. The fine-tune loses quality on FiQA and slightly improves SciFact. These are reranking results, not dense retrieval scores.
## Query coverage
The shared 970-query panel contains 620 generated-source questions, 200
questions written with potentially absent evidence, and 150 project-document
questions. These are natural-language questions. In the pipeline replay, 483
use the original query alone; 487 multipart queries also use two deterministic
subquestions. This exercises hybrid retrieval but does not provide a separately
balanced terse-query or paraphrase benchmark. Ettin trained on 294 answerable
queries from an earlier lexical supplement. Its separate held-out artifacts
are no longer available, so those historical checks cannot be recomputed and
are not part of the comparison reported here.
## CUDA and Vulkan
Both model cards share the CUDA fixed-depth table from `serving.json`,
using **925 queries with known grade-2/3 evidence** in a shared reference
across dense and hybrid retrieval. These queries have **122,231 graded
query–passage pairs**, within a full reference of 127,665 pairs across 970
queries. Membership comes from the reference, so pipeline retrieval misses
remain included. This is a different population and candidate set from
Ettin's historical 873-answerable-pool table.
Depths 3–50 use all 925 queries. Depth 1 uses the same 920 for both providers;
five are omitted because a first result was ungradable. Fixed cutoffs do not
evaluate the selector's choice of a 3–20-item prefix.
At 50, the hybrid pool reaches **97.62% Hit and 45.98% known-positive recall**
on the shared answerable subset. This is candidate coverage against the wider
reference. Ettin's standalone fixed-pool Hit and recall instead reach 100%
by construction on its answerable pools.
`serving.json` compares CUDA FP16 and Vulkan FP32 on 970 candidate pools retrieved with [Gemma v2](https://huggingface.co/Daecore/embeddinggemma-300m-memory-ft-v2) and BM25. Ettin's learned weights and the score mapping are unchanged. This is a retrieval-pipeline check on fixed inputs; it does not measure a change in Gemma.
| Metric | CUDA FP16 | Vulkan FP32 |
|---|---:|---:|
| Hit@1 (920 queries) | 75.87% | 75.98% |
| Hit@3 | 89.51% | 89.41% |
| Hit@5 | 92.76% | 92.76% |
| Hit@10 | 94.27% | 94.27% |
| Hit@20 | 95.46% | 95.46% |
| Hit@50 | 97.62% | 97.62% |
| nDCG@10 | 0.6378 | 0.6381 |
| Selected-prefix precision | 81.88% | 81.82% |
The top-50 completion shares its expanded reference with Gemma's dense and
smaller-width evaluations: 127,665 graded query–passage pairs. It adds 40,231 numeric
grades and 222 abstentions from 40,453 pairs across those panels. Two
independent GPT-6.1 Sol judges and a blind adjudicator handled disagreements;
the primary assistant reviewed 349 complete pairs, including every abstention
and final quotation flag. Sixteen quotation repairs changed no grades. These
are model judgments without human review.
Existing grades and rankings are unchanged. The earlier reference expansion
changed recall and ideal DCG; the card now also conditions quality on the
shared 925 known-answerable queries. Each provider omits 52 abstained positions
at rank 20 and 93 at rank 50, with no backfill. Precision pools retained
positions and nDCG reindexes them. The standalone fixed-pool Ettin comparison
and public benchmark labels are unchanged.
### All-query coverage
The saved records also preserve all 970 source queries, including 45 with no
known useful evidence. This separate audit view uses the full 127,665-pair
reference; it does not replace the answerable-only card table.
| Retrieval path | Queries | Hit@3 | Hit@50 | nDCG@10 |
|---|---:|---:|---:|---:|
| Gemma v2 + BM25 + Ettin, CUDA | 970 | 85.36% | 93.09% | 0.6315 |
### Serving latency
On 20 matched pools replayed twice, second-pass Vulkan reranking took 1.56 s median and 2.11 s at the 95th percentile, against 1.92 s and 2.70 s for the previous DirectML graph. Measurements used one Windows x64 machine with an RTX 3060 Ti (8 GB). They do not establish performance on other GPUs or Linux. `serving-qualification.json` records the checks and their limits.
## Limits
Some questions omit a project or subject, making relevance ambiguous. The private panel is mostly generated and model-judged, has no human reference, and was reused during development. Training-seed variation is unmeasured. A reranker cannot recover passages absent from its candidates, and several useful passages may repeat one fact. Passage-level scores do not establish unique-fact coverage, answer completeness or agent success.