tnh0527 commited on
Commit
2e1d6d0
·
verified ·
1 Parent(s): bc82ea2

Publish ettin-150m-memory-reranker-ft-v1 (qualified ONNX export and complete attribution)

Browse files
MODIFICATIONS.md CHANGED
@@ -14,20 +14,23 @@ What changed:
14
  with rank-16 LoRA+. The adapter was merged into the model weights.
15
  Relevance grades were frontier-model judgments under a frozen protocol,
16
  not human annotations.
17
- - **Export.** The merged model was exported to ONNX as one graph taking
18
  `input_ids` and `attention_mask` and emitting one score per query-passage
19
- pair, in FP16 for the CUDA realization, with the parameters stored as
20
- external data.
21
- - **Serving contract.** `serving.json` records the runtime and sequence
22
- geometry the graph was qualified under: 1,153 total pair tokens, batch 16,
23
- CUDA first with CPU as the floor.
 
 
24
 
25
  Files modified or generated by Daecore:
26
 
27
  - `model.onnx`: the generated ONNX graph of the fine-tuned reranker;
28
- - `model.onnx.data`: the fine-tuned parameters used by that graph;
 
29
  - `config.json`: the export configuration of the modified reranker;
30
- - `serving.json`: the Daecore runtime and sequence-geometry contract.
31
 
32
  Unchanged: `tokenizer.json` and `tokenizer_config.json` preserve the upstream
33
  tokenizer contract.
 
14
  with rank-16 LoRA+. The adapter was merged into the model weights.
15
  Relevance grades were frontier-model judgments under a frozen protocol,
16
  not human annotations.
17
+ - **Export.** The merged model was exported to ONNX graphs taking
18
  `input_ids` and `attention_mask` and emitting one score per query-passage
19
+ pair. CUDA uses FP16. DirectML widens the same stored weights and floating
20
+ calculations to FP32, and uses compatible inferred reshapes. Both graphs
21
+ share one external parameter file; the CUDA graph and weights are unchanged.
22
+ - **Serving contracts.** `serving.json` and `serving.directml.json` use the
23
+ same descriptor format for each graph: 1,153 total pair tokens and batch
24
+ ceilings of 16 on CUDA and 8 on DirectML. CPU handles permitted control
25
+ operations, not a fallback execution of the whole reranker.
26
 
27
  Files modified or generated by Daecore:
28
 
29
  - `model.onnx`: the generated ONNX graph of the fine-tuned reranker;
30
+ - `model.directml.onnx`: the FP32 calculation graph for DirectML;
31
+ - `model.onnx.data`: the fine-tuned parameters shared by both graphs;
32
  - `config.json`: the export configuration of the modified reranker;
33
+ - `serving.json` and `serving.directml.json`: the runtime and sequence contracts.
34
 
35
  Unchanged: `tokenizer.json` and `tokenizer_config.json` preserve the upstream
36
  tokenizer contract.
README.md CHANGED
@@ -40,8 +40,8 @@ The limitations section spells out each.
40
  | Input | one query and one passage, 1,153 tokens total |
41
  | Pool | a fixed 50-passage candidate set; the model reorders it and never adds to it |
42
  | Output | one relevance score per pair, used for ordering |
43
- | Serving precision | FP16 through ONNX Runtime with the CUDA overlay |
44
- | Deployment state | protected gate passed; package qualified; not yet published or activated |
45
 
46
  ## Intended use
47
 
@@ -126,7 +126,50 @@ sits at rank 47. On that panel the export reranked a pool in 0.584 s median
126
  and 0.912 s at the 95th percentile, and the Torch FP16 model in 0.532 s and
127
  0.777 s.
128
 
129
- ## Limitations, ranked
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
130
 
131
  1. **Model-judged relevance.** No human grades; session agreement 0.858 and
132
  kappa 0.682 on a 419-query audit. Scores measure agreement with that
@@ -142,14 +185,16 @@ and 0.912 s at the 95th percentile, and the Torch FP16 model in 0.532 s and
142
  5. **First-stage ceiling.** 112 of 553 development queries and 95 of 990 gate
143
  queries have no useful passage in any top-50 pool; ranks 51 to 200 were
144
  never judged.
145
- 6. **No CPU or DirectML realization is qualified.** Hosts without the CUDA
146
- provider serve the fused first-stage order with a typed notice.
 
 
147
  7. **The reproducible parity read fails the frozen policy.** The recorded
148
  development-panel pass cannot be re-run, and the 2026-09-04 recheck on the
149
  training partition fails two checks at exact FP16 ties. The operator
150
  accepted the recorded qualification with this disclosure rather than
151
- exporting an FP32 graph or amending the policy; an FP32 export and re-run
152
- would settle it.
153
 
154
  ## Reproducibility
155
 
@@ -161,7 +206,10 @@ and 0.912 s at the 95th percentile, and the Torch FP16 model in 0.532 s and
161
  - Recipe id `final-listnet-adjacent-loraplus16-full-coverage-w16-lr5e5`;
162
  registry manifest `../model_registry/ettin-150m-memory-reranker-ft-v1.json`;
163
  experiment log `../history/README.md`.
164
- - Serving graph SHA-256 `ce05a6a810007372c6de0f4c56dc94f652e1d28f2eeddf1878b64a170b4c48c2`.
 
 
 
165
  - Build stack: Transformers 5.2.0, Sentence Transformers 5.5.1, PyTorch
166
  2.10.0; export opset 18.
167
 
 
40
  | Input | one query and one passage, 1,153 tokens total |
41
  | Pool | a fixed 50-passage candidate set; the model reorders it and never adds to it |
42
  | Output | one relevance score per pair, used for ordering |
43
+ | Serving precision | FP16 on CUDA; FP32 calculations on DirectML, sharing the same stored weights |
44
+ | Deployment state | CUDA release plus an approved DirectML graph; installation is a separate update |
45
 
46
  ## Intended use
47
 
 
126
  and 0.912 s at the 95th percentile, and the Torch FP16 model in 0.532 s and
127
  0.777 s.
128
 
129
+ ### DirectML comparison
130
+
131
+ The DirectML package adds `model.directml.onnx` alongside the unchanged CUDA
132
+ `model.onnx`. Both read the same `model.onnx.data`; no second parameter file
133
+ or training framework is needed at runtime. DirectML uses full-precision
134
+ calculations because the half-precision graph had larger score differences.
135
+ CUDA keeps its faster half-precision graph. Each provider plans batches from
136
+ available memory and pair length, with measured ceilings of 16 rows for CUDA
137
+ and 8 for DirectML.
138
+
139
+ On a current 970-query replay (48,500 fixed-pool pairs), DirectML completed
140
+ without a failure: 2.231 seconds median, 3.247 seconds at the 95th percentile,
141
+ and 4.286 seconds maximum on an RTX 3060 Ti. No call exceeded ten seconds.
142
+ This panel retains its per-query evidence; 2,498 pairs are unjudged, and their
143
+ grades remain unknown. It is separate from the lost historical gate evidence.
144
+
145
+ Against the published CUDA graph, nDCG@10 changed by +0.000017. Eight first
146
+ choices and 36 top-ten sets changed. Both versions returned known-useful
147
+ evidence for 816 of 970 queries; DirectML returned 5,528 known-useful passages
148
+ versus 5,524. The unchanged calibration was applied for comparison only; its
149
+ DirectML registration is not implied by these results.
150
+
151
+ Two strict cross-precision checks fail and were explicitly accepted:
152
+ Hit@1 and Hit@5 each lose one query, and one passage moves five positions
153
+ against a limit of four. The affected useful passages remain in the returned
154
+ sets; the five-position move is rank 33 to 28, outside the returned three.
155
+ On the 482 fully judged pools, no Hit metric decreases.
156
+
157
+ Running the same FP32 graph on CUDA for all 165 queries with changed
158
+ deliveries or relevant ranking differences isolates the provider error:
159
+ 8,250 pairs have a 95th-percentile score difference of 0.00000954 and a
160
+ maximum of 0.000149. One pair exceeds the diagnostic 0.0001 bound; one full
161
+ ordering changes and no first choice changes. The two Hit losses and the
162
+ five-position move also occur on CUDA FP32, identifying precision rather
163
+ than DirectML as their cause.
164
+
165
+ On the supported sequential build policy, concurrent embedding and recall
166
+ kept reranking below 3.2 seconds in three rounds. Eight queued recalls across
167
+ four threads completed within 9.8 seconds. Forcing all three GPU models to
168
+ remain resident on this 8 GB card exceeded the memory budget and failed the
169
+ ten-second limit; those forced-overload timings are not a passing result.
170
+ Worker-kill recovery completed for reranking, embedding, and classification.
171
+
172
+ ## Limitations, ranked
173
 
174
  1. **Model-judged relevance.** No human grades; session agreement 0.858 and
175
  kappa 0.682 on a 419-query audit. Scores measure agreement with that
 
185
  5. **First-stage ceiling.** 112 of 553 development queries and 95 of 990 gate
186
  queries have no useful passage in any top-50 pool; ranks 51 to 200 were
187
  never judged.
188
+ 6. **Provider limits.** CPU reranking is not qualified. DirectML has been
189
+ measured on an NVIDIA RTX 3060 Ti, not AMD or Intel hardware, and its
190
+ cross-precision policy exceptions above were accepted with disclosure.
191
+ Hosts without a qualified reranker serve fusion order with a typed notice.
192
  7. **The reproducible parity read fails the frozen policy.** The recorded
193
  development-panel pass cannot be re-run, and the 2026-09-04 recheck on the
194
  training partition fails two checks at exact FP16 ties. The operator
195
  accepted the recorded qualification with this disclosure rather than
196
+ exporting an FP32 graph or amending the policy. The current FP32 provider
197
+ comparison is separate evidence and does not erase that historical result.
198
 
199
  ## Reproducibility
200
 
 
206
  - Recipe id `final-listnet-adjacent-loraplus16-full-coverage-w16-lr5e5`;
207
  registry manifest `../model_registry/ettin-150m-memory-reranker-ft-v1.json`;
208
  experiment log `../history/README.md`.
209
+ - CUDA graph SHA-256 `7df4a3b3cd89d216c016abd7d169f0fd215e9f79c06a2ffed57bac1abf556073`.
210
+ - DirectML graph SHA-256 `05abdc03628b599712a993381919aef1e22fcc5ad6384608abc2e069c402e369`.
211
+ - Shared weights SHA-256 `78ddab681c2f506d62bbcd475fec5c10c62d867ed60d97a0170828efaff5bdde`.
212
+ - CUDA serving-contract SHA-256 `ce05a6a810007372c6de0f4c56dc94f652e1d28f2eeddf1878b64a170b4c48c2`.
213
  - Build stack: Transformers 5.2.0, Sentence Transformers 5.5.1, PyTorch
214
  2.10.0; export opset 18.
215
 
model.directml.onnx ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:05abdc03628b599712a993381919aef1e22fcc5ad6384608abc2e069c402e369
3
+ size 2242802
publication-manifest.json CHANGED
@@ -1,6 +1,6 @@
1
  {
2
  "schema": "daecore.retrieval-publication-manifest",
3
- "staged_at": "2026-09-07T00:27:07+00:00",
4
  "model_id": "ettin-150m-memory-reranker-ft-v1",
5
  "public_repo": "Daecore/ettin-150m-memory-reranker-ft-v1",
6
  "role": "cross-encoder-reranker",
@@ -9,15 +9,21 @@
9
  "revision": "025501c4e0f9bbeb4c5b198318e0089ff061cc14",
10
  "tree_sha256": "e0edfb38cc495f23507de50233d0965f4e72ad0872e70e6f4e8fc889d394ea01"
11
  },
12
- "export_receipt_sha256": "0b9698df82a890a1688ab23dca2e32182fbccd31f6b460deea0222678d8ce99c",
13
  "serving_sha256": "ce05a6a810007372c6de0f4c56dc94f652e1d28f2eeddf1878b64a170b4c48c2",
14
  "source_model_tree_sha256": "bd8028b5dbc6ab2cd2ab9a6019346de7b25536f50022617194bcb8503d4467ad",
 
15
  "files": {
16
  "config.json": {
17
  "sha256": "a897acffc99f97238bb5bc0e1bccd6b7fe8fdfebe1450d4db3a9ab026cce6795",
18
  "size": 2020,
19
  "binding": "export receipt (file hash)"
20
  },
 
 
 
 
 
21
  "model.onnx": {
22
  "sha256": "7df4a3b3cd89d216c016abd7d169f0fd215e9f79c06a2ffed57bac1abf556073",
23
  "size": 2231565,
@@ -28,6 +34,11 @@
28
  "size": 299270144,
29
  "binding": "export receipt (file hash)"
30
  },
 
 
 
 
 
31
  "serving.json": {
32
  "sha256": "381c2fba69bd65cd1e90eac10b4048f1bcaac35f77bac358709194e05432f742",
33
  "size": 704,
@@ -44,8 +55,8 @@
44
  "binding": "export receipt (file hash)"
45
  },
46
  "README.md": {
47
- "sha256": "7b2135ffa768c9d919fd8d71ab47fcb7971ba7bfc9e7f33b00183b533e43bb9b",
48
- "size": 8755,
49
  "binding": "packaging record"
50
  },
51
  "LICENSE": {
@@ -59,8 +70,8 @@
59
  "binding": "packaging record"
60
  },
61
  "MODIFICATIONS.md": {
62
- "sha256": "128a860497b8e31ea659025420a0ca7861b2481bf3d58fda1ad9fe4bcea7710c",
63
- "size": 1639,
64
  "binding": "packaging record"
65
  }
66
  }
 
1
  {
2
  "schema": "daecore.retrieval-publication-manifest",
3
+ "tool_sha256": "2369a55f4fa3f007f800834880b6e4a65624c91589b469f48133ebc058c3d2ef",
4
  "model_id": "ettin-150m-memory-reranker-ft-v1",
5
  "public_repo": "Daecore/ettin-150m-memory-reranker-ft-v1",
6
  "role": "cross-encoder-reranker",
 
9
  "revision": "025501c4e0f9bbeb4c5b198318e0089ff061cc14",
10
  "tree_sha256": "e0edfb38cc495f23507de50233d0965f4e72ad0872e70e6f4e8fc889d394ea01"
11
  },
12
+ "export_receipt_sha256": "4faf07d568cc5da6400e9aa2645d73b2a2562cb720742c64106b2a70f43b12ba",
13
  "serving_sha256": "ce05a6a810007372c6de0f4c56dc94f652e1d28f2eeddf1878b64a170b4c48c2",
14
  "source_model_tree_sha256": "bd8028b5dbc6ab2cd2ab9a6019346de7b25536f50022617194bcb8503d4467ad",
15
+ "staged_at": "2026-09-11T14:57:19+00:00",
16
  "files": {
17
  "config.json": {
18
  "sha256": "a897acffc99f97238bb5bc0e1bccd6b7fe8fdfebe1450d4db3a9ab026cce6795",
19
  "size": 2020,
20
  "binding": "export receipt (file hash)"
21
  },
22
+ "model.directml.onnx": {
23
+ "sha256": "05abdc03628b599712a993381919aef1e22fcc5ad6384608abc2e069c402e369",
24
+ "size": 2242802,
25
+ "binding": "export receipt (file hash)"
26
+ },
27
  "model.onnx": {
28
  "sha256": "7df4a3b3cd89d216c016abd7d169f0fd215e9f79c06a2ffed57bac1abf556073",
29
  "size": 2231565,
 
34
  "size": 299270144,
35
  "binding": "export receipt (file hash)"
36
  },
37
+ "serving.directml.json": {
38
+ "sha256": "7a43e53663efc6e82db9b29a13fa3775a35d73bd6731e1163601f8f36748925f",
39
+ "size": 711,
40
+ "binding": "export receipt (file hash)"
41
+ },
42
  "serving.json": {
43
  "sha256": "381c2fba69bd65cd1e90eac10b4048f1bcaac35f77bac358709194e05432f742",
44
  "size": 704,
 
55
  "binding": "export receipt (file hash)"
56
  },
57
  "README.md": {
58
+ "sha256": "19b97e33bddbcb587446bbbb164cd091fc2cb534911cf0aafb0efc8013bda68c",
59
+ "size": 11806,
60
  "binding": "packaging record"
61
  },
62
  "LICENSE": {
 
70
  "binding": "packaging record"
71
  },
72
  "MODIFICATIONS.md": {
73
+ "sha256": "1f8c677387fb72ef31b7a107e866868a13821b6282debebc6f8150eac9866b18",
74
+ "size": 1958,
75
  "binding": "packaging record"
76
  }
77
  }
serving.directml.json ADDED
@@ -0,0 +1,25 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "external_data": "model.onnx.data",
3
+ "geometry": {
4
+ "dynamic_batch": true,
5
+ "dynamic_sequence": true,
6
+ "max_batch": 8,
7
+ "max_sequence": 1153
8
+ },
9
+ "graph": "model.directml.onnx",
10
+ "inputs": [
11
+ "input_ids",
12
+ "attention_mask"
13
+ ],
14
+ "output": "scores",
15
+ "precision": "float32",
16
+ "provider_chain": [
17
+ "DmlExecutionProvider",
18
+ "CPUExecutionProvider"
19
+ ],
20
+ "role": "cross-encoder-reranker",
21
+ "schema": "daecore.retrieval-onnx-serving.v1",
22
+ "sha256": "b83ffa3a44f19d145ecb87e36a56da290b25def03db267e33c1c57a7a46dda76",
23
+ "source_model_tree_sha256": "bd8028b5dbc6ab2cd2ab9a6019346de7b25536f50022617194bcb8503d4467ad",
24
+ "torch_required_at_runtime": false
25
+ }