Ettin-150M memory reranker

An English cross-encoder that reorders search results so the most useful passages come first. Give it a query and candidate passages from notes, procedures, decision records or technical documentation; it returns one relevance score per pair. Place it after a keyword or embedding search when the first few results matter most. It is a fine-tune of Ettin Reranker 150M and runs locally through ONNX Runtime.

On Daecore's evaluation data, fine-tuning raised precision@3 from 59.76% to 80.03% and nDCG@10 from 0.5519 to 0.7407 over the upstream checkpoint, measured on 873 answerable pools of 50 candidates. Public results are mixed: FiQA reranking falls below upstream, while SciFact rises slightly.

This revision adds a Vulkan FP32 graph, model.vulkan.onnx; the learned weights and the CUDA FP16 graph are unchanged. It omits the DirectML graph, which remains in the previous revision.

  • Query limit: 128 content tokens, excluding special tokens
  • Passage limit: 1,022 content tokens, excluding special tokens
  • Output: one relevance score per pair; higher scores rank first
  • Runtimes: ONNX Runtime 1.24.4 with CUDA FP16 (model.onnx) or Vulkan FP32 (model.vulkan.onnx); the FP32 graph was also checked for correctness on CPU
  • License: Apache-2.0

Quick start

Install the CUDA build of ONNX Runtime and download the package. Keep model.onnx, model.onnx.data and the tokenizer from the same package together.

pip install "onnxruntime-gpu[cuda,cudnn]==1.24.4" tokenizers numpy huggingface_hub
hf download Daecore/ettin-150m-memory-reranker-ft-v1 --local-dir downloaded-model

This CUDA example scores two passages:

import numpy as np
import onnxruntime as ort
from pathlib import Path
from tokenizers import Tokenizer

root = Path("downloaded-model")  # The exact downloaded serving package.
tokenizer = Tokenizer.from_file(f"{root}/tokenizer.json")
tokenizer.no_padding()
tokenizer.no_truncation()
ort.preload_dlls()
session = ort.InferenceSession(
    f"{root}/model.onnx", providers=[("CUDAExecutionProvider", {"use_tf32": "0"})]
)
if "CUDAExecutionProvider" not in session.get_providers():
    raise RuntimeError("CUDA did not initialize; check the ONNX Runtime CUDA dependencies.")
query = "What must happen before a database migration?"
passages = [
    "Take a verified backup before applying the migration.",
    "The dashboard theme uses a blue background.",
]
query_tokens = tokenizer.encode(query, add_special_tokens=False)
if len(query_tokens.ids) > 128:
    raise ValueError("Query exceeds the qualified 128-token content limit.")
scores = []
for passage in passages:
    passage_tokens = tokenizer.encode(passage, add_special_tokens=False)
    passage_tokens.truncate(1022)
    pair = tokenizer.post_process(query_tokens, passage_tokens, add_special_tokens=True)
    inputs = {
        "input_ids": np.array([pair.ids], dtype=np.int64),
        "attention_mask": np.array([pair.attention_mask], dtype=np.int64),
    }
    scores.append(float(session.run(None, inputs)[0].reshape(-1)[0]))
print(sorted(zip(passages, scores), key=lambda item: item[1], reverse=True))

Higher scores rank first; raw scores are not probabilities. The input envelope is 128 query-content tokens plus 1,022 passage-content tokens and three special tokens, a maximum of 1,153 tokens. The example rejects an oversized query and truncates the passage, matching the tested input policy. Without CUDA, use the FP32 graph model.vulkan.onnx with Vulkan or the CPU provider, as described under Runtime and limits.

Evaluation

Ettin alone: ranking a fixed candidate pool

Upstream and fine-tuned Ettin scored the same 970 pools of 50 candidates, with all 48,500 pairs judged. The table uses the 873 answerable pools that contain at least one useful passage, so it measures ordering after retrieval. β€œUpstream” is the exact starting checkpoint without Daecore fine-tuning. The headline panel uses natural-language questions; it does not report separate terse-query or paraphrase robustness scores.

Ettin scores up to 50 candidates; Daecore returns a selected prefix of 3–20. Each cell below reads upstream β†’ Daecore fine-tune.

Cutoff Hit Precision Pool recall nDCG
3 83.16% β†’ 92.55% 59.76% β†’ 80.03% 10.03% β†’ 14.28% 0.5180 β†’ 0.7198
5 89.46% β†’ 94.04% 58.58% β†’ 77.94% 15.80% β†’ 22.34% 0.5264 β†’ 0.7212
10 93.81% β†’ 96.91% 56.74% β†’ 74.18% 29.07% β†’ 39.95% 0.5519 β†’ 0.7407
20 98.17% β†’ 98.97% 54.46% β†’ 67.34% 53.78% β†’ 67.88% 0.6154 β†’ 0.7869
50 (full pool) 100.00% β†’ 100.00% 42.68% β†’ 42.68% 100.00% β†’ 100.00% 0.7891 β†’ 0.8753

Ettin ordering quality across return depths

Grades 2 and 3 count as useful. Hit asks whether at least one useful passage appears, precision measures the useful share of returned positions, and nDCG rewards higher grades near the top. Pool recall measures the fraction of useful passages found within these 50 candidates, not across the full corpus. At 50, both models have returned the same pool: Hit, precision and recall are identical, while nDCG still measures ordering.

Many pools contain several useful passages, so random ordering already reaches 97.30% Hit@20; early precision and graded ranking are more informative than deep Hit. The comparison runs both models in PyTorch FP16 with matched tokenization on reused development inputs. The evaluation companion includes all-query results, model identities and metric code.

Ettin alone: public reranking

Both models rerank the same 50 candidates that upstream Gemma retrieved for each query, scored with official relevance judgments:

Dataset Queries Upstream Ettin Daecore fine-tune
FiQA 648 0.4860 0.4527
SciFact 300 0.7487 0.7544

These are nDCG@10 scores for reranking fixed candidate pools, not dense Gemma scores. Fine-tuning lowered FiQA and slightly raised SciFact. Later continuation recipes did not recover FiQA while keeping Daecore ranking quality, so the original learned weights remain.

Full pipeline: Gemma v2 + BM25 + Ettin

This separate replay measures the complete retrieval path: semantic and lexical search, reciprocal-rank fusion, then Ettin reranking. It covers all 970 development queries over 82,719 passages, including queries with no known useful evidence. The table reports fixed return depths on CUDA FP16 and is shared by the Gemma and Ettin cards; it is not either model's standalone score.

Return depth Hit Precision Known-positive recall nDCG
3 85.36% 71.58% 9.52% 0.6577
5 88.45% 70.28% 15.38% 0.6555
10 89.90% 67.18% 28.03% 0.6611
20 91.03% 61.79% 48.85% 0.6841

Grades 2–3 count as useful. Hit and nDCG average all 970 queries; precision pools the retained positions. Recall averages the 895 queries with known positives and measures coverage of judged passages, not of every useful passage in the corpus. A shared judging-exclusion set removes 52 positions from the top 20 without backfilling. Daecore's selector chooses a prefix of 3–20; these fixed-depth scores do not evaluate that choice. The evaluation companion provides the records, exclusions and CUDA/Vulkan comparison; the pipeline overview explains the retrieval path.

Training data and objective

Training passages come from generated organizational documents and one person's project documentation, much of it AI-written. Alongside full questions, a dedicated lexical supplement contributes 294 answerable training queries covering short keywords, exact lookups, identifier-bearing questions, pasted fragments and shorthand or typos. Model judges assigned grades from 0 to 3, and several passages can be useful for one query. ListNet learns ordering within each candidate list, and a pairwise term emphasizes distinctions between neighboring grades. Lists rotate through the full labeled pools, so useful, partly useful and unhelpful passages stay in the comparison.

The private passages came from frozen versions of Daecore's structural chunker, which follows document headings, carries section context and handles tables and code blocks. Earlier parser outputs remain in the corpus. Ettin learns relevance among these prepared passages; it does not choose their boundaries. Its tokenizer and pair-length limit are applied afterward. A split that separates a fact from its context, or a candidate omitted by first-stage retrieval, can limit the final result even when reranking works well.

Training recipe and exposure
Item Value
Starting model cross-encoder/ettin-reranker-150m-v1
Training queries 2,882 with ranking supervision; 30 no-positive queries excluded from the ranking loss
Unique query–passage pairs 144,100
Unique passage texts 35,006
Supervision Four relevance grades, 0–3; grades 2–3 count as useful
Objective ListNet with an adjacent-grade pairwise term weighted 0.2
Adapter LoRA+, rank 16, alpha 32; merged for serving
Learning rates 5e-5 for LoRA A and the scoring head; 8e-4 for LoRA B
Training steps 2,882 scheduled; 2,881 applied updates, accumulating four lists of 16 per step
Exposure 11,528 lists; 184,448 candidate presentations

Runtime and limits

model.onnx is the CUDA FP16 graph and model.vulkan.onnx the Vulkan FP32 graph; both read model.onnx.data. The Vulkan derivation rewrites bounded mask operations and keeps the learned weights. The FP32 graph's results were also checked for correctness on the CPU execution provider; no CPU latency guidance is given.

Vulkan requires onnxruntime==1.24.4 and the native WebGPU plugin onnxruntime-ep-webgpu==0.4.0, registered explicitly. Add the plugin device to the session options with dawnBackendType=Vulkan, and create the session without a providers list; the plugin name alone does not select Vulkan. serving.vulkan.json records the tested provider options.

On 970 Daecore candidate pools, CUDA and Vulkan produced the same Hit@5, Hit@10 and Hit@20; Hit@3 differed by one query and nDCG@10 by 0.0004. The paths are not bit-exact, so close scores can swap order. Checks used Windows x64 with an NVIDIA RTX 3060 Ti (8 GB); other GPUs and Linux are untested, even where Vulkan is available.

Inside Daecore, Ettin reorders up to 50 deduplicated BM25 and Gemma candidates; the retrieval pipeline overview describes that path and result selection.

The task-specific panel is mostly generated and model-labeled, with no human reference, and was reused during development; training-seed variation is unmeasured. A reranker cannot recover a passage missing from its candidates, and several relevant passages may repeat one fact. These results do not establish unique-evidence coverage or downstream agent success.

License

Apache-2.0. The package includes upstream attribution and a modification notice.

Downloads last month
69
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for Daecore/ettin-150m-memory-reranker-ft-v1

Finetuned
(2)
this model