Ettin-150M memory reranker
An English cross-encoder that reorders search results so the most useful passages come first. Give it a query and candidate passages from notes, procedures, decision records or technical documentation; it returns one relevance score per pair. Place it after a keyword or embedding search when the first few results matter most. It is a fine-tune of Ettin Reranker 150M and runs locally through ONNX Runtime.
On Daecore's evaluation data, fine-tuning raised precision@3 from 59.76% to 80.03% and nDCG@10 from 0.5519 to 0.7407 over the upstream checkpoint, measured on 873 answerable pools of 50 candidates. Public results are mixed: FiQA reranking falls below upstream, while SciFact rises slightly.
This revision adds a Vulkan FP32 graph, model.vulkan.onnx; the learned weights and the CUDA FP16 graph are unchanged. It omits the DirectML graph, which remains in the previous revision.
- Query limit: 128 content tokens, excluding special tokens
- Passage limit: 1,022 content tokens, excluding special tokens
- Output: one relevance score per pair; higher scores rank first
- Runtimes: ONNX Runtime 1.24.4 with CUDA FP16 (
model.onnx) or Vulkan FP32 (model.vulkan.onnx); the FP32 graph was also checked for correctness on CPU - License: Apache-2.0
Quick start
Install the CUDA build of ONNX Runtime and download the package. Keep model.onnx, model.onnx.data and the tokenizer from the same package together.
pip install "onnxruntime-gpu[cuda,cudnn]==1.24.4" tokenizers numpy huggingface_hub
hf download Daecore/ettin-150m-memory-reranker-ft-v1 --local-dir downloaded-model
This CUDA example scores two passages:
import numpy as np
import onnxruntime as ort
from pathlib import Path
from tokenizers import Tokenizer
root = Path("downloaded-model") # The exact downloaded serving package.
tokenizer = Tokenizer.from_file(f"{root}/tokenizer.json")
tokenizer.no_padding()
tokenizer.no_truncation()
ort.preload_dlls()
session = ort.InferenceSession(
f"{root}/model.onnx", providers=[("CUDAExecutionProvider", {"use_tf32": "0"})]
)
if "CUDAExecutionProvider" not in session.get_providers():
raise RuntimeError("CUDA did not initialize; check the ONNX Runtime CUDA dependencies.")
query = "What must happen before a database migration?"
passages = [
"Take a verified backup before applying the migration.",
"The dashboard theme uses a blue background.",
]
query_tokens = tokenizer.encode(query, add_special_tokens=False)
if len(query_tokens.ids) > 128:
raise ValueError("Query exceeds the qualified 128-token content limit.")
scores = []
for passage in passages:
passage_tokens = tokenizer.encode(passage, add_special_tokens=False)
passage_tokens.truncate(1022)
pair = tokenizer.post_process(query_tokens, passage_tokens, add_special_tokens=True)
inputs = {
"input_ids": np.array([pair.ids], dtype=np.int64),
"attention_mask": np.array([pair.attention_mask], dtype=np.int64),
}
scores.append(float(session.run(None, inputs)[0].reshape(-1)[0]))
print(sorted(zip(passages, scores), key=lambda item: item[1], reverse=True))
Higher scores rank first; raw scores are not probabilities. The input envelope is 128 query-content tokens plus 1,022 passage-content tokens and three special tokens, a maximum of 1,153 tokens. The example rejects an oversized query and truncates the passage, matching the tested input policy. Without CUDA, use the FP32 graph model.vulkan.onnx with Vulkan or the CPU provider, as described under Runtime and limits.
Evaluation
Ettin alone: ranking a fixed candidate pool
Upstream and fine-tuned Ettin scored the same 970 pools of 50 candidates, with all 48,500 pairs judged. The table uses the 873 answerable pools that contain at least one useful passage, so it measures ordering after retrieval. βUpstreamβ is the exact starting checkpoint without Daecore fine-tuning. The headline panel uses natural-language questions; it does not report separate terse-query or paraphrase robustness scores.
Ettin scores up to 50 candidates; Daecore returns a selected prefix of 3β20. Each cell below reads upstream β Daecore fine-tune.
| Cutoff | Hit | Precision | Pool recall | nDCG |
|---|---|---|---|---|
| 3 | 83.16% β 92.55% | 59.76% β 80.03% | 10.03% β 14.28% | 0.5180 β 0.7198 |
| 5 | 89.46% β 94.04% | 58.58% β 77.94% | 15.80% β 22.34% | 0.5264 β 0.7212 |
| 10 | 93.81% β 96.91% | 56.74% β 74.18% | 29.07% β 39.95% | 0.5519 β 0.7407 |
| 20 | 98.17% β 98.97% | 54.46% β 67.34% | 53.78% β 67.88% | 0.6154 β 0.7869 |
| 50 (full pool) | 100.00% β 100.00% | 42.68% β 42.68% | 100.00% β 100.00% | 0.7891 β 0.8753 |
Grades 2 and 3 count as useful. Hit asks whether at least one useful passage appears, precision measures the useful share of returned positions, and nDCG rewards higher grades near the top. Pool recall measures the fraction of useful passages found within these 50 candidates, not across the full corpus. At 50, both models have returned the same pool: Hit, precision and recall are identical, while nDCG still measures ordering.
Many pools contain several useful passages, so random ordering already reaches 97.30% Hit@20; early precision and graded ranking are more informative than deep Hit. The comparison runs both models in PyTorch FP16 with matched tokenization on reused development inputs. The evaluation companion includes all-query results, model identities and metric code.
Ettin alone: public reranking
Both models rerank the same 50 candidates that upstream Gemma retrieved for each query, scored with official relevance judgments:
| Dataset | Queries | Upstream Ettin | Daecore fine-tune |
|---|---|---|---|
| FiQA | 648 | 0.4860 | 0.4527 |
| SciFact | 300 | 0.7487 | 0.7544 |
These are nDCG@10 scores for reranking fixed candidate pools, not dense Gemma scores. Fine-tuning lowered FiQA and slightly raised SciFact. Later continuation recipes did not recover FiQA while keeping Daecore ranking quality, so the original learned weights remain.
Full pipeline: Gemma v2 + BM25 + Ettin
This separate replay measures the complete retrieval path: semantic and lexical search, reciprocal-rank fusion, then Ettin reranking. It covers all 970 development queries over 82,719 passages, including queries with no known useful evidence. The table reports fixed return depths on CUDA FP16 and is shared by the Gemma and Ettin cards; it is not either model's standalone score.
| Return depth | Hit | Precision | Known-positive recall | nDCG |
|---|---|---|---|---|
| 3 | 85.36% | 71.58% | 9.52% | 0.6577 |
| 5 | 88.45% | 70.28% | 15.38% | 0.6555 |
| 10 | 89.90% | 67.18% | 28.03% | 0.6611 |
| 20 | 91.03% | 61.79% | 48.85% | 0.6841 |
Grades 2β3 count as useful. Hit and nDCG average all 970 queries; precision pools the retained positions. Recall averages the 895 queries with known positives and measures coverage of judged passages, not of every useful passage in the corpus. A shared judging-exclusion set removes 52 positions from the top 20 without backfilling. Daecore's selector chooses a prefix of 3β20; these fixed-depth scores do not evaluate that choice. The evaluation companion provides the records, exclusions and CUDA/Vulkan comparison; the pipeline overview explains the retrieval path.
Training data and objective
Training passages come from generated organizational documents and one person's project documentation, much of it AI-written. Alongside full questions, a dedicated lexical supplement contributes 294 answerable training queries covering short keywords, exact lookups, identifier-bearing questions, pasted fragments and shorthand or typos. Model judges assigned grades from 0 to 3, and several passages can be useful for one query. ListNet learns ordering within each candidate list, and a pairwise term emphasizes distinctions between neighboring grades. Lists rotate through the full labeled pools, so useful, partly useful and unhelpful passages stay in the comparison.
The private passages came from frozen versions of Daecore's structural chunker, which follows document headings, carries section context and handles tables and code blocks. Earlier parser outputs remain in the corpus. Ettin learns relevance among these prepared passages; it does not choose their boundaries. Its tokenizer and pair-length limit are applied afterward. A split that separates a fact from its context, or a candidate omitted by first-stage retrieval, can limit the final result even when reranking works well.
Training recipe and exposure
| Item | Value |
|---|---|
| Starting model | cross-encoder/ettin-reranker-150m-v1 |
| Training queries | 2,882 with ranking supervision; 30 no-positive queries excluded from the ranking loss |
| Unique queryβpassage pairs | 144,100 |
| Unique passage texts | 35,006 |
| Supervision | Four relevance grades, 0β3; grades 2β3 count as useful |
| Objective | ListNet with an adjacent-grade pairwise term weighted 0.2 |
| Adapter | LoRA+, rank 16, alpha 32; merged for serving |
| Learning rates | 5e-5 for LoRA A and the scoring head; 8e-4 for LoRA B |
| Training steps | 2,882 scheduled; 2,881 applied updates, accumulating four lists of 16 per step |
| Exposure | 11,528 lists; 184,448 candidate presentations |
Runtime and limits
model.onnx is the CUDA FP16 graph and model.vulkan.onnx the Vulkan FP32 graph; both read model.onnx.data. The Vulkan derivation rewrites bounded mask operations and keeps the learned weights. The FP32 graph's results were also checked for correctness on the CPU execution provider; no CPU latency guidance is given.
Vulkan requires onnxruntime==1.24.4 and the native WebGPU plugin onnxruntime-ep-webgpu==0.4.0, registered explicitly. Add the plugin device to the session options with dawnBackendType=Vulkan, and create the session without a providers list; the plugin name alone does not select Vulkan. serving.vulkan.json records the tested provider options.
On 970 Daecore candidate pools, CUDA and Vulkan produced the same Hit@5, Hit@10 and Hit@20; Hit@3 differed by one query and nDCG@10 by 0.0004. The paths are not bit-exact, so close scores can swap order. Checks used Windows x64 with an NVIDIA RTX 3060 Ti (8 GB); other GPUs and Linux are untested, even where Vulkan is available.
Inside Daecore, Ettin reorders up to 50 deduplicated BM25 and Gemma candidates; the retrieval pipeline overview describes that path and result selection.
The task-specific panel is mostly generated and model-labeled, with no human reference, and was reused during development; training-seed variation is unmeasured. A reranker cannot recover a passage missing from its candidates, and several relevant passages may repeat one fact. These results do not establish unique-evidence coverage or downstream agent success.
License
Apache-2.0. The package includes upstream attribution and a modification notice.
- Downloads last month
- 69
Model tree for Daecore/ettin-150m-memory-reranker-ft-v1
Base model
jhu-clsp/ettin-encoder-150m