Publish ettin-150m-memory-reranker-ft-v1 (qualified ONNX export and complete attribution)
Browse files- MODIFICATIONS.md +11 -8
- README.md +56 -8
- model.directml.onnx +3 -0
- publication-manifest.json +17 -6
- serving.directml.json +25 -0
MODIFICATIONS.md
CHANGED
|
@@ -14,20 +14,23 @@ What changed:
|
|
| 14 |
with rank-16 LoRA+. The adapter was merged into the model weights.
|
| 15 |
Relevance grades were frontier-model judgments under a frozen protocol,
|
| 16 |
not human annotations.
|
| 17 |
-
- **Export.** The merged model was exported to ONNX
|
| 18 |
`input_ids` and `attention_mask` and emitting one score per query-passage
|
| 19 |
-
pair
|
| 20 |
-
|
| 21 |
-
|
| 22 |
-
|
| 23 |
-
|
|
|
|
|
|
|
| 24 |
|
| 25 |
Files modified or generated by Daecore:
|
| 26 |
|
| 27 |
- `model.onnx`: the generated ONNX graph of the fine-tuned reranker;
|
| 28 |
-
- `model.
|
|
|
|
| 29 |
- `config.json`: the export configuration of the modified reranker;
|
| 30 |
-
- `serving.json`: the
|
| 31 |
|
| 32 |
Unchanged: `tokenizer.json` and `tokenizer_config.json` preserve the upstream
|
| 33 |
tokenizer contract.
|
|
|
|
| 14 |
with rank-16 LoRA+. The adapter was merged into the model weights.
|
| 15 |
Relevance grades were frontier-model judgments under a frozen protocol,
|
| 16 |
not human annotations.
|
| 17 |
+
- **Export.** The merged model was exported to ONNX graphs taking
|
| 18 |
`input_ids` and `attention_mask` and emitting one score per query-passage
|
| 19 |
+
pair. CUDA uses FP16. DirectML widens the same stored weights and floating
|
| 20 |
+
calculations to FP32, and uses compatible inferred reshapes. Both graphs
|
| 21 |
+
share one external parameter file; the CUDA graph and weights are unchanged.
|
| 22 |
+
- **Serving contracts.** `serving.json` and `serving.directml.json` use the
|
| 23 |
+
same descriptor format for each graph: 1,153 total pair tokens and batch
|
| 24 |
+
ceilings of 16 on CUDA and 8 on DirectML. CPU handles permitted control
|
| 25 |
+
operations, not a fallback execution of the whole reranker.
|
| 26 |
|
| 27 |
Files modified or generated by Daecore:
|
| 28 |
|
| 29 |
- `model.onnx`: the generated ONNX graph of the fine-tuned reranker;
|
| 30 |
+
- `model.directml.onnx`: the FP32 calculation graph for DirectML;
|
| 31 |
+
- `model.onnx.data`: the fine-tuned parameters shared by both graphs;
|
| 32 |
- `config.json`: the export configuration of the modified reranker;
|
| 33 |
+
- `serving.json` and `serving.directml.json`: the runtime and sequence contracts.
|
| 34 |
|
| 35 |
Unchanged: `tokenizer.json` and `tokenizer_config.json` preserve the upstream
|
| 36 |
tokenizer contract.
|
README.md
CHANGED
|
@@ -40,8 +40,8 @@ The limitations section spells out each.
|
|
| 40 |
| Input | one query and one passage, 1,153 tokens total |
|
| 41 |
| Pool | a fixed 50-passage candidate set; the model reorders it and never adds to it |
|
| 42 |
| Output | one relevance score per pair, used for ordering |
|
| 43 |
-
| Serving precision | FP16
|
| 44 |
-
| Deployment state |
|
| 45 |
|
| 46 |
## Intended use
|
| 47 |
|
|
@@ -126,7 +126,50 @@ sits at rank 47. On that panel the export reranked a pool in 0.584 s median
|
|
| 126 |
and 0.912 s at the 95th percentile, and the Torch FP16 model in 0.532 s and
|
| 127 |
0.777 s.
|
| 128 |
|
| 129 |
-
##
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 130 |
|
| 131 |
1. **Model-judged relevance.** No human grades; session agreement 0.858 and
|
| 132 |
kappa 0.682 on a 419-query audit. Scores measure agreement with that
|
|
@@ -142,14 +185,16 @@ and 0.912 s at the 95th percentile, and the Torch FP16 model in 0.532 s and
|
|
| 142 |
5. **First-stage ceiling.** 112 of 553 development queries and 95 of 990 gate
|
| 143 |
queries have no useful passage in any top-50 pool; ranks 51 to 200 were
|
| 144 |
never judged.
|
| 145 |
-
6. **
|
| 146 |
-
|
|
|
|
|
|
|
| 147 |
7. **The reproducible parity read fails the frozen policy.** The recorded
|
| 148 |
development-panel pass cannot be re-run, and the 2026-09-04 recheck on the
|
| 149 |
training partition fails two checks at exact FP16 ties. The operator
|
| 150 |
accepted the recorded qualification with this disclosure rather than
|
| 151 |
-
exporting an FP32 graph or amending the policy
|
| 152 |
-
|
| 153 |
|
| 154 |
## Reproducibility
|
| 155 |
|
|
@@ -161,7 +206,10 @@ and 0.912 s at the 95th percentile, and the Torch FP16 model in 0.532 s and
|
|
| 161 |
- Recipe id `final-listnet-adjacent-loraplus16-full-coverage-w16-lr5e5`;
|
| 162 |
registry manifest `../model_registry/ettin-150m-memory-reranker-ft-v1.json`;
|
| 163 |
experiment log `../history/README.md`.
|
| 164 |
-
-
|
|
|
|
|
|
|
|
|
|
| 165 |
- Build stack: Transformers 5.2.0, Sentence Transformers 5.5.1, PyTorch
|
| 166 |
2.10.0; export opset 18.
|
| 167 |
|
|
|
|
| 40 |
| Input | one query and one passage, 1,153 tokens total |
|
| 41 |
| Pool | a fixed 50-passage candidate set; the model reorders it and never adds to it |
|
| 42 |
| Output | one relevance score per pair, used for ordering |
|
| 43 |
+
| Serving precision | FP16 on CUDA; FP32 calculations on DirectML, sharing the same stored weights |
|
| 44 |
+
| Deployment state | CUDA release plus an approved DirectML graph; installation is a separate update |
|
| 45 |
|
| 46 |
## Intended use
|
| 47 |
|
|
|
|
| 126 |
and 0.912 s at the 95th percentile, and the Torch FP16 model in 0.532 s and
|
| 127 |
0.777 s.
|
| 128 |
|
| 129 |
+
### DirectML comparison
|
| 130 |
+
|
| 131 |
+
The DirectML package adds `model.directml.onnx` alongside the unchanged CUDA
|
| 132 |
+
`model.onnx`. Both read the same `model.onnx.data`; no second parameter file
|
| 133 |
+
or training framework is needed at runtime. DirectML uses full-precision
|
| 134 |
+
calculations because the half-precision graph had larger score differences.
|
| 135 |
+
CUDA keeps its faster half-precision graph. Each provider plans batches from
|
| 136 |
+
available memory and pair length, with measured ceilings of 16 rows for CUDA
|
| 137 |
+
and 8 for DirectML.
|
| 138 |
+
|
| 139 |
+
On a current 970-query replay (48,500 fixed-pool pairs), DirectML completed
|
| 140 |
+
without a failure: 2.231 seconds median, 3.247 seconds at the 95th percentile,
|
| 141 |
+
and 4.286 seconds maximum on an RTX 3060 Ti. No call exceeded ten seconds.
|
| 142 |
+
This panel retains its per-query evidence; 2,498 pairs are unjudged, and their
|
| 143 |
+
grades remain unknown. It is separate from the lost historical gate evidence.
|
| 144 |
+
|
| 145 |
+
Against the published CUDA graph, nDCG@10 changed by +0.000017. Eight first
|
| 146 |
+
choices and 36 top-ten sets changed. Both versions returned known-useful
|
| 147 |
+
evidence for 816 of 970 queries; DirectML returned 5,528 known-useful passages
|
| 148 |
+
versus 5,524. The unchanged calibration was applied for comparison only; its
|
| 149 |
+
DirectML registration is not implied by these results.
|
| 150 |
+
|
| 151 |
+
Two strict cross-precision checks fail and were explicitly accepted:
|
| 152 |
+
Hit@1 and Hit@5 each lose one query, and one passage moves five positions
|
| 153 |
+
against a limit of four. The affected useful passages remain in the returned
|
| 154 |
+
sets; the five-position move is rank 33 to 28, outside the returned three.
|
| 155 |
+
On the 482 fully judged pools, no Hit metric decreases.
|
| 156 |
+
|
| 157 |
+
Running the same FP32 graph on CUDA for all 165 queries with changed
|
| 158 |
+
deliveries or relevant ranking differences isolates the provider error:
|
| 159 |
+
8,250 pairs have a 95th-percentile score difference of 0.00000954 and a
|
| 160 |
+
maximum of 0.000149. One pair exceeds the diagnostic 0.0001 bound; one full
|
| 161 |
+
ordering changes and no first choice changes. The two Hit losses and the
|
| 162 |
+
five-position move also occur on CUDA FP32, identifying precision rather
|
| 163 |
+
than DirectML as their cause.
|
| 164 |
+
|
| 165 |
+
On the supported sequential build policy, concurrent embedding and recall
|
| 166 |
+
kept reranking below 3.2 seconds in three rounds. Eight queued recalls across
|
| 167 |
+
four threads completed within 9.8 seconds. Forcing all three GPU models to
|
| 168 |
+
remain resident on this 8 GB card exceeded the memory budget and failed the
|
| 169 |
+
ten-second limit; those forced-overload timings are not a passing result.
|
| 170 |
+
Worker-kill recovery completed for reranking, embedding, and classification.
|
| 171 |
+
|
| 172 |
+
## Limitations, ranked
|
| 173 |
|
| 174 |
1. **Model-judged relevance.** No human grades; session agreement 0.858 and
|
| 175 |
kappa 0.682 on a 419-query audit. Scores measure agreement with that
|
|
|
|
| 185 |
5. **First-stage ceiling.** 112 of 553 development queries and 95 of 990 gate
|
| 186 |
queries have no useful passage in any top-50 pool; ranks 51 to 200 were
|
| 187 |
never judged.
|
| 188 |
+
6. **Provider limits.** CPU reranking is not qualified. DirectML has been
|
| 189 |
+
measured on an NVIDIA RTX 3060 Ti, not AMD or Intel hardware, and its
|
| 190 |
+
cross-precision policy exceptions above were accepted with disclosure.
|
| 191 |
+
Hosts without a qualified reranker serve fusion order with a typed notice.
|
| 192 |
7. **The reproducible parity read fails the frozen policy.** The recorded
|
| 193 |
development-panel pass cannot be re-run, and the 2026-09-04 recheck on the
|
| 194 |
training partition fails two checks at exact FP16 ties. The operator
|
| 195 |
accepted the recorded qualification with this disclosure rather than
|
| 196 |
+
exporting an FP32 graph or amending the policy. The current FP32 provider
|
| 197 |
+
comparison is separate evidence and does not erase that historical result.
|
| 198 |
|
| 199 |
## Reproducibility
|
| 200 |
|
|
|
|
| 206 |
- Recipe id `final-listnet-adjacent-loraplus16-full-coverage-w16-lr5e5`;
|
| 207 |
registry manifest `../model_registry/ettin-150m-memory-reranker-ft-v1.json`;
|
| 208 |
experiment log `../history/README.md`.
|
| 209 |
+
- CUDA graph SHA-256 `7df4a3b3cd89d216c016abd7d169f0fd215e9f79c06a2ffed57bac1abf556073`.
|
| 210 |
+
- DirectML graph SHA-256 `05abdc03628b599712a993381919aef1e22fcc5ad6384608abc2e069c402e369`.
|
| 211 |
+
- Shared weights SHA-256 `78ddab681c2f506d62bbcd475fec5c10c62d867ed60d97a0170828efaff5bdde`.
|
| 212 |
+
- CUDA serving-contract SHA-256 `ce05a6a810007372c6de0f4c56dc94f652e1d28f2eeddf1878b64a170b4c48c2`.
|
| 213 |
- Build stack: Transformers 5.2.0, Sentence Transformers 5.5.1, PyTorch
|
| 214 |
2.10.0; export opset 18.
|
| 215 |
|
model.directml.onnx
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:05abdc03628b599712a993381919aef1e22fcc5ad6384608abc2e069c402e369
|
| 3 |
+
size 2242802
|
publication-manifest.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
| 1 |
{
|
| 2 |
"schema": "daecore.retrieval-publication-manifest",
|
| 3 |
-
"
|
| 4 |
"model_id": "ettin-150m-memory-reranker-ft-v1",
|
| 5 |
"public_repo": "Daecore/ettin-150m-memory-reranker-ft-v1",
|
| 6 |
"role": "cross-encoder-reranker",
|
|
@@ -9,15 +9,21 @@
|
|
| 9 |
"revision": "025501c4e0f9bbeb4c5b198318e0089ff061cc14",
|
| 10 |
"tree_sha256": "e0edfb38cc495f23507de50233d0965f4e72ad0872e70e6f4e8fc889d394ea01"
|
| 11 |
},
|
| 12 |
-
"export_receipt_sha256": "
|
| 13 |
"serving_sha256": "ce05a6a810007372c6de0f4c56dc94f652e1d28f2eeddf1878b64a170b4c48c2",
|
| 14 |
"source_model_tree_sha256": "bd8028b5dbc6ab2cd2ab9a6019346de7b25536f50022617194bcb8503d4467ad",
|
|
|
|
| 15 |
"files": {
|
| 16 |
"config.json": {
|
| 17 |
"sha256": "a897acffc99f97238bb5bc0e1bccd6b7fe8fdfebe1450d4db3a9ab026cce6795",
|
| 18 |
"size": 2020,
|
| 19 |
"binding": "export receipt (file hash)"
|
| 20 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 21 |
"model.onnx": {
|
| 22 |
"sha256": "7df4a3b3cd89d216c016abd7d169f0fd215e9f79c06a2ffed57bac1abf556073",
|
| 23 |
"size": 2231565,
|
|
@@ -28,6 +34,11 @@
|
|
| 28 |
"size": 299270144,
|
| 29 |
"binding": "export receipt (file hash)"
|
| 30 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 31 |
"serving.json": {
|
| 32 |
"sha256": "381c2fba69bd65cd1e90eac10b4048f1bcaac35f77bac358709194e05432f742",
|
| 33 |
"size": 704,
|
|
@@ -44,8 +55,8 @@
|
|
| 44 |
"binding": "export receipt (file hash)"
|
| 45 |
},
|
| 46 |
"README.md": {
|
| 47 |
-
"sha256": "
|
| 48 |
-
"size":
|
| 49 |
"binding": "packaging record"
|
| 50 |
},
|
| 51 |
"LICENSE": {
|
|
@@ -59,8 +70,8 @@
|
|
| 59 |
"binding": "packaging record"
|
| 60 |
},
|
| 61 |
"MODIFICATIONS.md": {
|
| 62 |
-
"sha256": "
|
| 63 |
-
"size":
|
| 64 |
"binding": "packaging record"
|
| 65 |
}
|
| 66 |
}
|
|
|
|
| 1 |
{
|
| 2 |
"schema": "daecore.retrieval-publication-manifest",
|
| 3 |
+
"tool_sha256": "2369a55f4fa3f007f800834880b6e4a65624c91589b469f48133ebc058c3d2ef",
|
| 4 |
"model_id": "ettin-150m-memory-reranker-ft-v1",
|
| 5 |
"public_repo": "Daecore/ettin-150m-memory-reranker-ft-v1",
|
| 6 |
"role": "cross-encoder-reranker",
|
|
|
|
| 9 |
"revision": "025501c4e0f9bbeb4c5b198318e0089ff061cc14",
|
| 10 |
"tree_sha256": "e0edfb38cc495f23507de50233d0965f4e72ad0872e70e6f4e8fc889d394ea01"
|
| 11 |
},
|
| 12 |
+
"export_receipt_sha256": "4faf07d568cc5da6400e9aa2645d73b2a2562cb720742c64106b2a70f43b12ba",
|
| 13 |
"serving_sha256": "ce05a6a810007372c6de0f4c56dc94f652e1d28f2eeddf1878b64a170b4c48c2",
|
| 14 |
"source_model_tree_sha256": "bd8028b5dbc6ab2cd2ab9a6019346de7b25536f50022617194bcb8503d4467ad",
|
| 15 |
+
"staged_at": "2026-09-11T14:57:19+00:00",
|
| 16 |
"files": {
|
| 17 |
"config.json": {
|
| 18 |
"sha256": "a897acffc99f97238bb5bc0e1bccd6b7fe8fdfebe1450d4db3a9ab026cce6795",
|
| 19 |
"size": 2020,
|
| 20 |
"binding": "export receipt (file hash)"
|
| 21 |
},
|
| 22 |
+
"model.directml.onnx": {
|
| 23 |
+
"sha256": "05abdc03628b599712a993381919aef1e22fcc5ad6384608abc2e069c402e369",
|
| 24 |
+
"size": 2242802,
|
| 25 |
+
"binding": "export receipt (file hash)"
|
| 26 |
+
},
|
| 27 |
"model.onnx": {
|
| 28 |
"sha256": "7df4a3b3cd89d216c016abd7d169f0fd215e9f79c06a2ffed57bac1abf556073",
|
| 29 |
"size": 2231565,
|
|
|
|
| 34 |
"size": 299270144,
|
| 35 |
"binding": "export receipt (file hash)"
|
| 36 |
},
|
| 37 |
+
"serving.directml.json": {
|
| 38 |
+
"sha256": "7a43e53663efc6e82db9b29a13fa3775a35d73bd6731e1163601f8f36748925f",
|
| 39 |
+
"size": 711,
|
| 40 |
+
"binding": "export receipt (file hash)"
|
| 41 |
+
},
|
| 42 |
"serving.json": {
|
| 43 |
"sha256": "381c2fba69bd65cd1e90eac10b4048f1bcaac35f77bac358709194e05432f742",
|
| 44 |
"size": 704,
|
|
|
|
| 55 |
"binding": "export receipt (file hash)"
|
| 56 |
},
|
| 57 |
"README.md": {
|
| 58 |
+
"sha256": "19b97e33bddbcb587446bbbb164cd091fc2cb534911cf0aafb0efc8013bda68c",
|
| 59 |
+
"size": 11806,
|
| 60 |
"binding": "packaging record"
|
| 61 |
},
|
| 62 |
"LICENSE": {
|
|
|
|
| 70 |
"binding": "packaging record"
|
| 71 |
},
|
| 72 |
"MODIFICATIONS.md": {
|
| 73 |
+
"sha256": "1f8c677387fb72ef31b7a107e866868a13821b6282debebc6f8150eac9866b18",
|
| 74 |
+
"size": 1958,
|
| 75 |
"binding": "packaging record"
|
| 76 |
}
|
| 77 |
}
|
serving.directml.json
ADDED
|
@@ -0,0 +1,25 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"external_data": "model.onnx.data",
|
| 3 |
+
"geometry": {
|
| 4 |
+
"dynamic_batch": true,
|
| 5 |
+
"dynamic_sequence": true,
|
| 6 |
+
"max_batch": 8,
|
| 7 |
+
"max_sequence": 1153
|
| 8 |
+
},
|
| 9 |
+
"graph": "model.directml.onnx",
|
| 10 |
+
"inputs": [
|
| 11 |
+
"input_ids",
|
| 12 |
+
"attention_mask"
|
| 13 |
+
],
|
| 14 |
+
"output": "scores",
|
| 15 |
+
"precision": "float32",
|
| 16 |
+
"provider_chain": [
|
| 17 |
+
"DmlExecutionProvider",
|
| 18 |
+
"CPUExecutionProvider"
|
| 19 |
+
],
|
| 20 |
+
"role": "cross-encoder-reranker",
|
| 21 |
+
"schema": "daecore.retrieval-onnx-serving.v1",
|
| 22 |
+
"sha256": "b83ffa3a44f19d145ecb87e36a56da290b25def03db267e33c1c57a7a46dda76",
|
| 23 |
+
"source_model_tree_sha256": "bd8028b5dbc6ab2cd2ab9a6019346de7b25536f50022617194bcb8503d4467ad",
|
| 24 |
+
"torch_required_at_runtime": false
|
| 25 |
+
}
|