alexwengg commited on
Commit
62ffd36
·
verified ·
1 Parent(s): 8b0c36f

Compare the full published synthetic test and report failed numerical gates

Browse files

All 24,370 decisions: PyTorch and both unchanged Core ML exports score 24,359 correct (99.9549%) with 100% selected-index agreement. Strict parity fails: 11 probability errors above 0.005 per export (max 0.0204874) and one sum outside the Swift guard. Original median 1.003 ms; ANE-gather 1.052 ms on M5 Pro. Publish exact manifest and complete per-row trace; no normalization or tolerance changes. Historical browser reproduction now pins its exact source commit.

README.md CHANGED
@@ -36,6 +36,7 @@ not load Qwen weights.
36
  | `preprocessing.py` | Upstream-compatible byte encoding and input validation |
37
  | `conversion.json` | Architecture, conversion versions, and portable package hashes |
38
  | `assets.lock.json` | Pinned upstream source, model, and demo hashes |
 
39
  | `checksums.json` | SHA-256 of each distributed file except this checksum file |
40
  | `reports/` | Per-row parity, original Cua metrics, and compute-placement report |
41
 
@@ -142,8 +143,64 @@ operation counts are not a measurement of time spent on each processor.
142
  These results verify conversion on three demo forms and three PDFs. They do not
143
  establish generalization, live GUI completion rates, or production safety.
144
  The original demo is not redistributed here; its exact revision and SHA-256 are
145
- in the asset lock. No training or larger-corpus evaluation was performed for
146
- this conversion.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
147
 
148
  ## ANE profile
149
 
@@ -249,13 +306,13 @@ The pinned model and dataset cards declare MIT. See [LICENSE](LICENSE),
249
 
250
  [![Actual CUA-driven browser forms](demo/browser-demo.gif)](demo/browser-demo.mp4)
251
 
252
- The [native Swift browser demo](https://github.com/FluidInference/FluidAudio/tree/codex/cua-s1-forms/Examples/CuaS1FormsDemo)
253
  loads both variants into independent WKWebViews. It reads actual DOM labels,
254
  roles and state, asks the model for a choice, applies compatible fill/check
255
  actions, dispatches events, and independently verifies the resulting DOM.
256
  Source values are user-entered or supplied by the original public examples.
257
  HTML contains controls, not source values or expected choices. The
258
- [recording](https://github.com/FluidInference/FluidAudio/blob/codex/cua-s1-forms/Examples/CuaS1FormsDemo/browser-demo.mp4)
259
  shows patient, job and insurance forms: **100/100 original decisions** across
260
  both models, with event-count, stale-observation and explicit-click checks.
261
  Full actual contexts/candidates/actions are in [browser-validation.json](reports/browser-validation.json).
@@ -272,9 +329,12 @@ raw samples, exact model hashes, per-form statistics, and load/first-call costs.
272
  The earlier compute-plan counts still apply to these unchanged artifacts;
273
  no utilization, energy saving or held-out accuracy claim is made.
274
 
275
- Reproduce from the FluidAudio branch:
 
276
 
277
  ```bash
 
 
278
  Examples/CuaS1FormsDemo/run.sh --browser
279
  swift run --package-path Examples/CuaS1FormsDemo -c release CuaS1FormsDemo \
280
  --benchmark --report /absolute/path/to/variant-comparison.json \
@@ -283,5 +343,5 @@ swift run --package-path Examples/CuaS1FormsDemo -c release CuaS1FormsDemo \
283
 
284
  Both packages are fetched by pinned revisions and verified hashes, or supplied
285
  with `--model /path/to/original.mlpackage --ane-model /path/to/ane-gather.mlpackage`.
286
- The [demo README](https://github.com/FluidInference/FluidAudio/tree/codex/cua-s1-forms/Examples/CuaS1FormsDemo#matched-swift-benchmark)
287
  also documents real-browser recording and its separate validation trace.
 
36
  | `preprocessing.py` | Upstream-compatible byte encoding and input validation |
37
  | `conversion.json` | Architecture, conversion versions, and portable package hashes |
38
  | `assets.lock.json` | Pinned upstream source, model, and demo hashes |
39
+ | `synthetic-test.lock.json` | Full published synthetic test revision, count, size, and SHA-256 |
40
  | `checksums.json` | SHA-256 of each distributed file except this checksum file |
41
  | `reports/` | Per-row parity, original Cua metrics, and compute-placement report |
42
 
 
143
  These results verify conversion on three demo forms and three PDFs. They do not
144
  establish generalization, live GUI completion rates, or production safety.
145
  The original demo is not redistributed here; its exact revision and SHA-256 are
146
+ in the asset lock. No training was performed. The separate full synthetic test below evaluates
147
+ the unchanged artifacts and exposes numerical failures outside this demo.
148
+
149
+ ## Full published synthetic test
150
+
151
+ Evaluated the complete published synthetic [`test.jsonl`](https://huggingface.co/datasets/cua-ai/cua-s1-forms/blob/8273f34778b99ac2e12d9f6e7d57dad99ae20845/test.jsonl):
152
+ **24,370 decisions across 1,040 episode seeds**, with no excluded rows or truncated
153
+ inputs. The dataset revision is `8273f34778b99ac2e12d9f6e7d57dad99ae20845`; its
154
+ SHA-256 is `d63a7e0db195d4d20154a40b2f8dd09ce3bb65487a158c638da5c609d4475e7c`.
155
+ The upstream [model card](https://huggingface.co/cua-ai/cua-s1-forms/blob/f54adbf447f4ca6ec259f529ee3f2e3e09f8cc71/README.md) reports 99.95% on an approximately 15,000-row
156
+ synthetic test. Our result rounds to that accuracy, but the released file contains
157
+ 24,370 rows; this does not reconstruct the card's unspecified smaller manifest.
158
+
159
+ | Model / backend | Correct decisions | Top-1 accuracy | Median call | p95 call |
160
+ | --- | ---: | ---: | ---: | ---: |
161
+ | Upstream PyTorch / CPU | 24,359 / 24,370 | 99.9549% | 1.787 ms | 3.104 ms |
162
+ | Original FP16 Core ML / CPU + ANE | 24,359 / 24,370 | 99.9549% | 1.003 ms | 1.133 ms |
163
+ | ANE-gather FP16 Core ML / CPU + ANE | 24,359 / 24,370 | 99.9549% | 1.052 ms | 1.177 ms |
164
+
165
+ Both Core ML exports select the same option as PyTorch on **all 24,370 rows**.
166
+ All three share the same 11 errors: choosing `fill` when the label is `skip`.
167
+ The original Cua evaluator counts these as wrong/unsafe actions; this offline
168
+ benchmark executes no actions. All 9,802 fill, 816 check, and 1,040 click labels
169
+ are correct; skip accuracy is 12,701/12,712. Higher ANE placement is approximately
170
+ **4.8% slower** by median here, so the original remains the default.
171
+
172
+ **Strict numerical conversion parity fails for both exports.** Each has 11 rows
173
+ above the original 0.005 absolute probability-error limit, with a maximum error
174
+ of **0.0204874**. One further row (zero-based index 19270) has a live-probability
175
+ sum of **0.99893665**. That lies outside the current Swift manager's 0.001
176
+ normalization guard, so this output would be rejected despite its correct argmax.
177
+ These are raw model-output accuracy results, not a claim that every row succeeds
178
+ through the Swift API. No probabilities were renormalized and no tolerances or
179
+ model weights were changed. The earlier 196-row demo passed its numerical gates;
180
+ the larger split exposes failures that demo did not cover.
181
+
182
+ Measured September 19, 2026 on **Apple M5 Pro, 24 GB, macOS 27.0 (26A428)**,
183
+ Python 3.11.11, PyTorch 2.7.0, coremltools 9.0. Batch size 1, three warmup rows per
184
+ model, one timed pass over the whole split; both Core ML models remain loaded
185
+ and alternate AB/BA order by row. PyTorch uses two CPU threads and one inter-op
186
+ thread, with the Transformer fast path disabled. Timers cover PyTorch
187
+ forward + softmax or synchronous Core ML prediction, excluding encoding,
188
+ validation, loading, UI, and network. These compare deployment backends, not
189
+ algorithms on equal hardware, and differ from the separate Swift timings at the end of this card.
190
+
191
+ Upstream describes this synthetic split as disjoint from training/validation by
192
+ form signature; those signatures were not independently re-audited here. No
193
+ training, validation inference, test-based tuning, or hosted Jev/API comparison
194
+ was performed. This measures supplied-option classification, not unseen real-world
195
+ GUI completion or document extraction.
196
+
197
+ [Full report](reports/synthetic-test.json), [complete compressed per-row trace](reports/synthetic-test-decisions.jsonl.gz),
198
+ and [test manifest](synthetic-test.lock.json) retain the exact protocol, hashes,
199
+ paired decisions, action metrics, and all failed numerical checks. Reproduce with
200
+ `benchmark-synthetic.py --require-parity` in the [Mobius toolkit](https://github.com/FluidInference/mobius/tree/codex/cua-s1-forms/models/computer-use/cua-s1-forms/coreml#full-published-synthetic-test).
201
+ That command exits 1 for the recorded numerical failures; it does not normalize
202
+ scores or weaken the original gates. Treat these artifacts as under review
203
+ until the numerical failures are resolved and validated.
204
 
205
  ## ANE profile
206
 
 
306
 
307
  [![Actual CUA-driven browser forms](demo/browser-demo.gif)](demo/browser-demo.mp4)
308
 
309
+ The [native Swift browser demo](https://github.com/FluidInference/FluidAudio/tree/7f9eb92b0af8594c4e048a9e57f697340aacfa67/Examples/CuaS1FormsDemo)
310
  loads both variants into independent WKWebViews. It reads actual DOM labels,
311
  roles and state, asks the model for a choice, applies compatible fill/check
312
  actions, dispatches events, and independently verifies the resulting DOM.
313
  Source values are user-entered or supplied by the original public examples.
314
  HTML contains controls, not source values or expected choices. The
315
+ [recording](https://github.com/FluidInference/FluidAudio/blob/7f9eb92b0af8594c4e048a9e57f697340aacfa67/Examples/CuaS1FormsDemo/browser-demo.mp4)
316
  shows patient, job and insurance forms: **100/100 original decisions** across
317
  both models, with event-count, stale-observation and explicit-click checks.
318
  Full actual contexts/candidates/actions are in [browser-validation.json](reports/browser-validation.json).
 
329
  The earlier compute-plan counts still apply to these unchanged artifacts;
330
  no utilization, energy saving or held-out accuracy claim is made.
331
 
332
+ Reproduce from FluidAudio commit `7f9eb92b0af8594c4e048a9e57f697340aacfa67` (the example is
333
+ retained in history and is not part of the current library PR):
334
 
335
  ```bash
336
+ git worktree add --detach /tmp/cua-s1-browser-repro 7f9eb92b0af8594c4e048a9e57f697340aacfa67
337
+ cd /tmp/cua-s1-browser-repro
338
  Examples/CuaS1FormsDemo/run.sh --browser
339
  swift run --package-path Examples/CuaS1FormsDemo -c release CuaS1FormsDemo \
340
  --benchmark --report /absolute/path/to/variant-comparison.json \
 
343
 
344
  Both packages are fetched by pinned revisions and verified hashes, or supplied
345
  with `--model /path/to/original.mlpackage --ane-model /path/to/ane-gather.mlpackage`.
346
+ The [demo README](https://github.com/FluidInference/FluidAudio/tree/7f9eb92b0af8594c4e048a9e57f697340aacfa67/Examples/CuaS1FormsDemo#matched-swift-benchmark)
347
  also documents real-browser recording and its separate validation trace.
checksums.json CHANGED
@@ -1,7 +1,7 @@
1
  {
2
  "LICENSE": "c0779290c1d4783169aa3dbfb55feb505e563ef8a004bbf55298ceffcfbda8d9",
3
  "NOTICES.md": "027c72741eaa695e60d7b6cebd3666372c81d443a96b673b72a897cf432ddfd0",
4
- "README.md": "96d395d4792f8d9b86878dbf37c63243a31b67adc8f14fbe7792b28f4cdb657a",
5
  "UPSTREAM-THIRD-PARTY-NOTICES.md": "4091e69b45c8cc97e30a066fbbd56148dbef66ab716432c048d2333c9c464213",
6
  "ane-gather/conversion.json": "4d569a4020f1e9711da3f7fff0d11acef1812c4652414cf64728c77bc2bae4c4",
7
  "ane-gather/cua_s1_forms_fp16_options32.mlmodelc/analytics/coremldata.bin": "526dccb87bdc036df1e8add0b50dd6540db534808fc2f8caa97cd897b4fc1fa3",
@@ -33,6 +33,9 @@
33
  "reports/browser-validation.json": "299641c7b9f6cc2464cfcd984b4ef4445385e4710bc15abae7cee8ba262331b6",
34
  "reports/swift-ane-validation.json": "ed323e6d6619fb7698d3546599d6450b745bc8cc6100c02e13f7ad4d59859a3a",
35
  "reports/swift-variant-comparison.json": "593053d7011a1c6d6881b130c3408099d0fcc09c714c4064d4f853f0dc63b651",
 
 
36
  "reports/upstream-metrics.json": "7d6207e430504dc0baec96d8180cb03f8aadc5a6e016b53c4cf03a80baf169ab",
37
- "reports/verification.json": "9a5d22ec4b447f0710c12e03c69d7ed802158bb94e9a3c6b059731730ced9f30"
 
38
  }
 
1
  {
2
  "LICENSE": "c0779290c1d4783169aa3dbfb55feb505e563ef8a004bbf55298ceffcfbda8d9",
3
  "NOTICES.md": "027c72741eaa695e60d7b6cebd3666372c81d443a96b673b72a897cf432ddfd0",
4
+ "README.md": "5739754098ca5a480397693115f079e3e500c2a06abb822f152cf7cbdc9943ac",
5
  "UPSTREAM-THIRD-PARTY-NOTICES.md": "4091e69b45c8cc97e30a066fbbd56148dbef66ab716432c048d2333c9c464213",
6
  "ane-gather/conversion.json": "4d569a4020f1e9711da3f7fff0d11acef1812c4652414cf64728c77bc2bae4c4",
7
  "ane-gather/cua_s1_forms_fp16_options32.mlmodelc/analytics/coremldata.bin": "526dccb87bdc036df1e8add0b50dd6540db534808fc2f8caa97cd897b4fc1fa3",
 
33
  "reports/browser-validation.json": "299641c7b9f6cc2464cfcd984b4ef4445385e4710bc15abae7cee8ba262331b6",
34
  "reports/swift-ane-validation.json": "ed323e6d6619fb7698d3546599d6450b745bc8cc6100c02e13f7ad4d59859a3a",
35
  "reports/swift-variant-comparison.json": "593053d7011a1c6d6881b130c3408099d0fcc09c714c4064d4f853f0dc63b651",
36
+ "reports/synthetic-test-decisions.jsonl.gz": "b6f6edc9de34fb37b10847a1f9568b530ce2b8e1bd241f51bf2e637f489d6554",
37
+ "reports/synthetic-test.json": "2ae85cfeb270b4f31be4f1d9082f4c73b3b9b084d6f1ddb661712a87d46b0f59",
38
  "reports/upstream-metrics.json": "7d6207e430504dc0baec96d8180cb03f8aadc5a6e016b53c4cf03a80baf169ab",
39
+ "reports/verification.json": "9a5d22ec4b447f0710c12e03c69d7ed802158bb94e9a3c6b059731730ced9f30",
40
+ "synthetic-test.lock.json": "e582b951ecc07809a0b06a2bc0790827b90a86ce57c70d5b6177ef04f0b4f1ca"
41
  }
reports/synthetic-test-decisions.jsonl.gz ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:b6f6edc9de34fb37b10847a1f9568b530ce2b8e1bd241f51bf2e637f489d6554
3
+ size 2231147
reports/synthetic-test.json ADDED
@@ -0,0 +1,845 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "purpose": "Full published synthetic test; no fitting, filtering, resampling or test-based tuning",
3
+ "created_utc": "2026-09-19T21:10:18.821257+00:00",
4
+ "dataset": {
5
+ "dataset_repository": "cua-ai/cua-s1-forms",
6
+ "dataset_revision": "8273f34778b99ac2e12d9f6e7d57dad99ae20845",
7
+ "path": "artifacts/test.jsonl",
8
+ "url": "https://huggingface.co/datasets/cua-ai/cua-s1-forms/resolve/8273f34778b99ac2e12d9f6e7d57dad99ae20845/test.jsonl",
9
+ "sha256": "d63a7e0db195d4d20154a40b2f8dd09ce3bb65487a158c638da5c609d4475e7c",
10
+ "rows": 24370,
11
+ "bytes": 18710924,
12
+ "upstream_split_claim": "Synthetic test, form-signature-disjoint from train/validation. Train and validation signatures have not been independently re-audited here.",
13
+ "model_card_claim": {
14
+ "synthetic_test_top1": 0.9995,
15
+ "approximate_decisions": 15000,
16
+ "note": "The released file has 24370 rows. No matching full result ledger or exact 15000-row manifest is supplied by the linked model card."
17
+ }
18
+ },
19
+ "manifest_sha256": "e582b951ecc07809a0b06a2bc0790827b90a86ce57c70d5b6177ef04f0b4f1ca",
20
+ "input_census": {
21
+ "rows": 24370,
22
+ "unique_episode_seeds": 1040,
23
+ "maxima": {
24
+ "context_bytes": 217,
25
+ "option_bytes": 92,
26
+ "max_options": 32
27
+ },
28
+ "per_action": {
29
+ "check": 816,
30
+ "click": 1040,
31
+ "fill": 9802,
32
+ "skip": 12712
33
+ },
34
+ "option_count_histogram": {
35
+ "7": 17,
36
+ "8": 46,
37
+ "9": 80,
38
+ "10": 227,
39
+ "11": 361,
40
+ "12": 589,
41
+ "13": 888,
42
+ "14": 1254,
43
+ "15": 1513,
44
+ "16": 1761,
45
+ "17": 2373,
46
+ "18": 2014,
47
+ "19": 2277,
48
+ "20": 2506,
49
+ "21": 2396,
50
+ "22": 1818,
51
+ "23": 1422,
52
+ "24": 907,
53
+ "25": 688,
54
+ "26": 524,
55
+ "27": 483,
56
+ "28": 89,
57
+ "29": 33,
58
+ "30": 29,
59
+ "31": 35,
60
+ "32": 40
61
+ },
62
+ "excluded_rows": 0,
63
+ "truncated_contexts": 0,
64
+ "truncated_options": 0
65
+ },
66
+ "model_revision": "f54adbf447f4ca6ec259f529ee3f2e3e09f8cc71",
67
+ "source_revision": "83f142c4290a0f7d9ed545ae8532858c6e4f8145",
68
+ "conversions": {
69
+ "baseline": {
70
+ "model": "cua_s1_forms_fp16_options32.mlpackage",
71
+ "precision": "float16",
72
+ "minimum_target": "iOS17/macOS14",
73
+ "limits": {
74
+ "context_bytes": 224,
75
+ "option_bytes": 96,
76
+ "max_options": 32
77
+ },
78
+ "model_config": {
79
+ "context_tokens": 224,
80
+ "encoder": "tinyx",
81
+ "heads": 4,
82
+ "hf_model": "Qwen/Qwen2.5-0.5B",
83
+ "layers": 2,
84
+ "option_tokens": 96,
85
+ "rank": 128,
86
+ "width": 128
87
+ },
88
+ "parameters": 706048,
89
+ "assets_lock_sha256": "8e65ad70af6bb814b571cdcfe828ba4bc339147da2d2211cbeac416163ef18ba",
90
+ "model_revision": "f54adbf447f4ca6ec259f529ee3f2e3e09f8cc71",
91
+ "source_revision": "83f142c4290a0f7d9ed545ae8532858c6e4f8145",
92
+ "trace_row": 0,
93
+ "trace_dataset_revision": "8273f34778b99ac2e12d9f6e7d57dad99ae20845",
94
+ "export_seconds": 0.5374167499830946,
95
+ "python": "3.11.11",
96
+ "torch": "2.7.0",
97
+ "coremltools": "9.0",
98
+ "package_files": {
99
+ "Data/com.apple.CoreML/model.mlmodel": "70485fc18cbb21785df833cbddddc0b5b59acb00d22394b76e55307e2c135dd0",
100
+ "Data/com.apple.CoreML/weights/weight.bin": "4da9259f798e44f5a1b50769ee1916fd3747c4d723dd9997b516c7fe238c7895",
101
+ "Manifest.json": "2bc0f5f62337b27fb6b0ecde248f1e3dc269e1ba4b65516aaeede2a60e293dcc"
102
+ }
103
+ },
104
+ "ane-gather": {
105
+ "model": "cua_s1_forms_fp16_options32.mlpackage",
106
+ "precision": "float16",
107
+ "optimization": "ane-gather",
108
+ "minimum_target": "iOS17/macOS14",
109
+ "limits": {
110
+ "context_bytes": 224,
111
+ "option_bytes": 96,
112
+ "max_options": 32
113
+ },
114
+ "model_config": {
115
+ "context_tokens": 224,
116
+ "encoder": "tinyx",
117
+ "heads": 4,
118
+ "hf_model": "Qwen/Qwen2.5-0.5B",
119
+ "layers": 2,
120
+ "option_tokens": 96,
121
+ "rank": 128,
122
+ "width": 128
123
+ },
124
+ "parameters": 706048,
125
+ "assets_lock_sha256": "8e65ad70af6bb814b571cdcfe828ba4bc339147da2d2211cbeac416163ef18ba",
126
+ "model_revision": "f54adbf447f4ca6ec259f529ee3f2e3e09f8cc71",
127
+ "source_revision": "83f142c4290a0f7d9ed545ae8532858c6e4f8145",
128
+ "trace_row": 0,
129
+ "trace_dataset_revision": "8273f34778b99ac2e12d9f6e7d57dad99ae20845",
130
+ "export_seconds": 0.6771400420111604,
131
+ "python": "3.11.11",
132
+ "torch": "2.7.0",
133
+ "coremltools": "9.0",
134
+ "package_files": {
135
+ "Data/com.apple.CoreML/model.mlmodel": "de18e313c3b625e35d008ed8b6b24108edc6df7bf2fae9533af03519eca11b63",
136
+ "Data/com.apple.CoreML/weights/weight.bin": "4da9259f798e44f5a1b50769ee1916fd3747c4d723dd9997b516c7fe238c7895",
137
+ "Manifest.json": "38f812a04eb2322080634ba788c61df362336168466ca67e5549058d346ac793"
138
+ }
139
+ }
140
+ },
141
+ "environment": {
142
+ "machine": "Apple M5 Pro",
143
+ "memory_bytes": 25769803776,
144
+ "os": "27.0",
145
+ "os_build": "26A428",
146
+ "python": "3.11.11",
147
+ "torch": "2.7.0",
148
+ "coremltools": "9.0",
149
+ "torch_cpu_threads": 2,
150
+ "torch_interop_threads": 1,
151
+ "torch_mha_fastpath": false
152
+ },
153
+ "protocol": {
154
+ "batch_size": 1,
155
+ "warmups_per_model": 3,
156
+ "timed_passes": 1,
157
+ "coreml_compute_units": "CPU_AND_NE",
158
+ "coreml_order": "AB on even rows, BA on odd rows",
159
+ "timing": "PyTorch CPU forward+softmax / synchronous Core ML predict; excludes encoding, validation, GUI and network",
160
+ "preprocessing": "Upstream variable-length collator; independent fixed-shape Core ML encoder shared by both exports",
161
+ "load_note": "Process/model creation can use existing system caches; not a cold-start experiment",
162
+ "output_audit": "Imperfect finite probability sums are retained unchanged and fail the original normalization gate",
163
+ "calibration_note": "NLL/ECE use raw emitted scores; no renormalization or calibration was applied"
164
+ },
165
+ "thresholds": {
166
+ "argmax_agreement": 1.0,
167
+ "max_abs_probability_error": 0.005,
168
+ "allow_accuracy_loss": false,
169
+ "probability_sum_atol": 0.001,
170
+ "probability_sum_rtol": 1e-05
171
+ },
172
+ "host_preprocessing_ms": {
173
+ "upstream_collator_ms": {
174
+ "count": 24370,
175
+ "median_ms": 0.09608300752006471,
176
+ "p95_ms": 0.14412500750040635,
177
+ "min_ms": 0.04137499490752816,
178
+ "max_ms": 0.96245898748748
179
+ },
180
+ "shared_coreml_encoding_ms": {
181
+ "count": 24370,
182
+ "median_ms": 0.028999987989664078,
183
+ "p95_ms": 0.044917018385604024,
184
+ "min_ms": 0.012375006917864084,
185
+ "max_ms": 0.3988750104326755
186
+ }
187
+ },
188
+ "results": {
189
+ "upstream_pytorch": {
190
+ "correct": 24359,
191
+ "rows": 24370,
192
+ "accuracy": 0.999548625359048,
193
+ "model_call_latency_ms": {
194
+ "count": 24370,
195
+ "median_ms": 1.786978988093324,
196
+ "p95_ms": 3.103799853124656,
197
+ "min_ms": 1.019332994474098,
198
+ "max_ms": 15.003458975115791
199
+ },
200
+ "load_ms": 8.14687501406297,
201
+ "first_call_ms": 5.291624984238297,
202
+ "calibration": {
203
+ "nll": 0.0021302032038990836,
204
+ "ece": 0.00035450826133425167,
205
+ "equal_width_bins": 15,
206
+ "bins": [
207
+ {
208
+ "bin": 7,
209
+ "count": 2,
210
+ "accuracy": 0.5,
211
+ "mean_confidence": 0.5216051042079926
212
+ },
213
+ {
214
+ "bin": 8,
215
+ "count": 5,
216
+ "accuracy": 0.8,
217
+ "mean_confidence": 0.575071656703949
218
+ },
219
+ {
220
+ "bin": 9,
221
+ "count": 7,
222
+ "accuracy": 0.7142857142857143,
223
+ "mean_confidence": 0.6404258779117039
224
+ },
225
+ {
226
+ "bin": 10,
227
+ "count": 4,
228
+ "accuracy": 1.0,
229
+ "mean_confidence": 0.6963506639003754
230
+ },
231
+ {
232
+ "bin": 11,
233
+ "count": 11,
234
+ "accuracy": 0.9090909090909091,
235
+ "mean_confidence": 0.7664791941642761
236
+ },
237
+ {
238
+ "bin": 12,
239
+ "count": 21,
240
+ "accuracy": 0.9523809523809523,
241
+ "mean_confidence": 0.8310647464933849
242
+ },
243
+ {
244
+ "bin": 13,
245
+ "count": 22,
246
+ "accuracy": 0.9545454545454546,
247
+ "mean_confidence": 0.9173552746122534
248
+ },
249
+ {
250
+ "bin": 14,
251
+ "count": 24298,
252
+ "accuracy": 0.9998353773973166,
253
+ "mean_confidence": 0.999802232897422
254
+ }
255
+ ],
256
+ "nll_probability_floor": 1.1754943508222875e-38
257
+ },
258
+ "upstream_action_metrics": {
259
+ "examples": 24370,
260
+ "accuracy": 0.999548625359048,
261
+ "coverage": 0.4788264259335248,
262
+ "abstention_rate": 0.5211735740664751,
263
+ "selective_accuracy": 0.9990573313908647,
264
+ "wrong_action_rate": 0.00045137464095199015,
265
+ "wrong_target_rate": 0.0,
266
+ "unsafe_action_rate": 0.00045137464095199015,
267
+ "counts": {
268
+ "total": 24370,
269
+ "correct": 24359,
270
+ "abstained": 12701,
271
+ "acted": 11669,
272
+ "acted_correct": 11658,
273
+ "wrong_action": 11,
274
+ "unsafe_action": 11
275
+ },
276
+ "per_action": {
277
+ "abstain": {
278
+ "accuracy": 0.9991346758967904,
279
+ "abstention_rate": 0.9991346758967904,
280
+ "examples": 12712
281
+ },
282
+ "check": {
283
+ "accuracy": 1.0,
284
+ "abstention_rate": 0.0,
285
+ "examples": 816
286
+ },
287
+ "click": {
288
+ "accuracy": 1.0,
289
+ "abstention_rate": 0.0,
290
+ "examples": 1040
291
+ },
292
+ "fill": {
293
+ "accuracy": 1.0,
294
+ "abstention_rate": 0.0,
295
+ "examples": 9802
296
+ }
297
+ }
298
+ },
299
+ "macro_action_accuracy": 0.9997836689741976,
300
+ "wrong_rows": [
301
+ 137,
302
+ 3216,
303
+ 6631,
304
+ 6748,
305
+ 7280,
306
+ 10703,
307
+ 12255,
308
+ 13480,
309
+ 15214,
310
+ 21022,
311
+ 23671
312
+ ]
313
+ },
314
+ "baseline": {
315
+ "correct": 24359,
316
+ "rows": 24370,
317
+ "accuracy": 0.999548625359048,
318
+ "model_call_latency_ms": {
319
+ "count": 24370,
320
+ "median_ms": 1.0031874990090728,
321
+ "p95_ms": 1.1327079846523702,
322
+ "min_ms": 0.8830419974401593,
323
+ "max_ms": 3.726124996319413
324
+ },
325
+ "load_ms": 612.4412920034956,
326
+ "first_call_ms": 3.2538340019527823,
327
+ "calibration": {
328
+ "nll": 0.002108235388127147,
329
+ "ece": 0.000331678613561794,
330
+ "equal_width_bins": 15,
331
+ "bins": [
332
+ {
333
+ "bin": 7,
334
+ "count": 2,
335
+ "accuracy": 0.5,
336
+ "mean_confidence": 0.5224609375
337
+ },
338
+ {
339
+ "bin": 8,
340
+ "count": 5,
341
+ "accuracy": 0.8,
342
+ "mean_confidence": 0.57607421875
343
+ },
344
+ {
345
+ "bin": 9,
346
+ "count": 6,
347
+ "accuracy": 0.8333333333333334,
348
+ "mean_confidence": 0.637451171875
349
+ },
350
+ {
351
+ "bin": 10,
352
+ "count": 5,
353
+ "accuracy": 0.8,
354
+ "mean_confidence": 0.69248046875
355
+ },
356
+ {
357
+ "bin": 11,
358
+ "count": 12,
359
+ "accuracy": 0.9166666666666666,
360
+ "mean_confidence": 0.7687581380208334
361
+ },
362
+ {
363
+ "bin": 12,
364
+ "count": 20,
365
+ "accuracy": 0.95,
366
+ "mean_confidence": 0.8322265625
367
+ },
368
+ {
369
+ "bin": 13,
370
+ "count": 22,
371
+ "accuracy": 0.9545454545454546,
372
+ "mean_confidence": 0.9176136363636364
373
+ },
374
+ {
375
+ "bin": 14,
376
+ "count": 24298,
377
+ "accuracy": 0.9998353773973166,
378
+ "mean_confidence": 0.9998245660008025
379
+ }
380
+ ],
381
+ "nll_probability_floor": 1.1754943508222875e-38
382
+ },
383
+ "upstream_action_metrics": {
384
+ "examples": 24370,
385
+ "accuracy": 0.999548625359048,
386
+ "coverage": 0.4788264259335248,
387
+ "abstention_rate": 0.5211735740664751,
388
+ "selective_accuracy": 0.9990573313908647,
389
+ "wrong_action_rate": 0.00045137464095199015,
390
+ "wrong_target_rate": 0.0,
391
+ "unsafe_action_rate": 0.00045137464095199015,
392
+ "counts": {
393
+ "total": 24370,
394
+ "correct": 24359,
395
+ "abstained": 12701,
396
+ "acted": 11669,
397
+ "acted_correct": 11658,
398
+ "wrong_action": 11,
399
+ "unsafe_action": 11
400
+ },
401
+ "per_action": {
402
+ "abstain": {
403
+ "accuracy": 0.9991346758967904,
404
+ "abstention_rate": 0.9991346758967904,
405
+ "examples": 12712
406
+ },
407
+ "check": {
408
+ "accuracy": 1.0,
409
+ "abstention_rate": 0.0,
410
+ "examples": 816
411
+ },
412
+ "click": {
413
+ "accuracy": 1.0,
414
+ "abstention_rate": 0.0,
415
+ "examples": 1040
416
+ },
417
+ "fill": {
418
+ "accuracy": 1.0,
419
+ "abstention_rate": 0.0,
420
+ "examples": 9802
421
+ }
422
+ }
423
+ },
424
+ "macro_action_accuracy": 0.9997836689741976,
425
+ "wrong_rows": [
426
+ 137,
427
+ 3216,
428
+ 6631,
429
+ 6748,
430
+ 7280,
431
+ 10703,
432
+ 12255,
433
+ 13480,
434
+ 15214,
435
+ 21022,
436
+ 23671
437
+ ],
438
+ "paired_with_upstream": {
439
+ "rows": 24370,
440
+ "argmax_agreement": 24370,
441
+ "agreement_rate": 1.0,
442
+ "reference_correct_candidate_wrong": 0,
443
+ "reference_wrong_candidate_correct": 0,
444
+ "both_wrong": 11,
445
+ "disagreement_rows": []
446
+ },
447
+ "max_abs_probability_error": 0.020487427711486816,
448
+ "mean_row_max_abs_probability_error": 4.673297645567538e-05,
449
+ "probability_tolerance_violation_rows": [
450
+ 137,
451
+ 884,
452
+ 2945,
453
+ 3216,
454
+ 8706,
455
+ 9193,
456
+ 12725,
457
+ 13480,
458
+ 13768,
459
+ 15698,
460
+ 23992
461
+ ],
462
+ "conversion_parity_passed": false,
463
+ "output_validation_failures": {
464
+ "19270": [
465
+ "Live option probabilities do not sum to one"
466
+ ]
467
+ }
468
+ },
469
+ "ane-gather": {
470
+ "correct": 24359,
471
+ "rows": 24370,
472
+ "accuracy": 0.999548625359048,
473
+ "model_call_latency_ms": {
474
+ "count": 24370,
475
+ "median_ms": 1.0515624890103936,
476
+ "p95_ms": 1.1766475494368933,
477
+ "min_ms": 0.9340409887954593,
478
+ "max_ms": 2.8469589888118207
479
+ },
480
+ "load_ms": 570.0647499761544,
481
+ "first_call_ms": 1.2669999850913882,
482
+ "calibration": {
483
+ "nll": 0.002108235388127147,
484
+ "ece": 0.000331678613561794,
485
+ "equal_width_bins": 15,
486
+ "bins": [
487
+ {
488
+ "bin": 7,
489
+ "count": 2,
490
+ "accuracy": 0.5,
491
+ "mean_confidence": 0.5224609375
492
+ },
493
+ {
494
+ "bin": 8,
495
+ "count": 5,
496
+ "accuracy": 0.8,
497
+ "mean_confidence": 0.57607421875
498
+ },
499
+ {
500
+ "bin": 9,
501
+ "count": 6,
502
+ "accuracy": 0.8333333333333334,
503
+ "mean_confidence": 0.637451171875
504
+ },
505
+ {
506
+ "bin": 10,
507
+ "count": 5,
508
+ "accuracy": 0.8,
509
+ "mean_confidence": 0.69248046875
510
+ },
511
+ {
512
+ "bin": 11,
513
+ "count": 12,
514
+ "accuracy": 0.9166666666666666,
515
+ "mean_confidence": 0.7687581380208334
516
+ },
517
+ {
518
+ "bin": 12,
519
+ "count": 20,
520
+ "accuracy": 0.95,
521
+ "mean_confidence": 0.8322265625
522
+ },
523
+ {
524
+ "bin": 13,
525
+ "count": 22,
526
+ "accuracy": 0.9545454545454546,
527
+ "mean_confidence": 0.9176136363636364
528
+ },
529
+ {
530
+ "bin": 14,
531
+ "count": 24298,
532
+ "accuracy": 0.9998353773973166,
533
+ "mean_confidence": 0.9998245660008025
534
+ }
535
+ ],
536
+ "nll_probability_floor": 1.1754943508222875e-38
537
+ },
538
+ "upstream_action_metrics": {
539
+ "examples": 24370,
540
+ "accuracy": 0.999548625359048,
541
+ "coverage": 0.4788264259335248,
542
+ "abstention_rate": 0.5211735740664751,
543
+ "selective_accuracy": 0.9990573313908647,
544
+ "wrong_action_rate": 0.00045137464095199015,
545
+ "wrong_target_rate": 0.0,
546
+ "unsafe_action_rate": 0.00045137464095199015,
547
+ "counts": {
548
+ "total": 24370,
549
+ "correct": 24359,
550
+ "abstained": 12701,
551
+ "acted": 11669,
552
+ "acted_correct": 11658,
553
+ "wrong_action": 11,
554
+ "unsafe_action": 11
555
+ },
556
+ "per_action": {
557
+ "abstain": {
558
+ "accuracy": 0.9991346758967904,
559
+ "abstention_rate": 0.9991346758967904,
560
+ "examples": 12712
561
+ },
562
+ "check": {
563
+ "accuracy": 1.0,
564
+ "abstention_rate": 0.0,
565
+ "examples": 816
566
+ },
567
+ "click": {
568
+ "accuracy": 1.0,
569
+ "abstention_rate": 0.0,
570
+ "examples": 1040
571
+ },
572
+ "fill": {
573
+ "accuracy": 1.0,
574
+ "abstention_rate": 0.0,
575
+ "examples": 9802
576
+ }
577
+ }
578
+ },
579
+ "macro_action_accuracy": 0.9997836689741976,
580
+ "wrong_rows": [
581
+ 137,
582
+ 3216,
583
+ 6631,
584
+ 6748,
585
+ 7280,
586
+ 10703,
587
+ 12255,
588
+ 13480,
589
+ 15214,
590
+ 21022,
591
+ 23671
592
+ ],
593
+ "paired_with_upstream": {
594
+ "rows": 24370,
595
+ "argmax_agreement": 24370,
596
+ "agreement_rate": 1.0,
597
+ "reference_correct_candidate_wrong": 0,
598
+ "reference_wrong_candidate_correct": 0,
599
+ "both_wrong": 11,
600
+ "disagreement_rows": []
601
+ },
602
+ "max_abs_probability_error": 0.020487427711486816,
603
+ "mean_row_max_abs_probability_error": 4.673297645567538e-05,
604
+ "probability_tolerance_violation_rows": [
605
+ 137,
606
+ 884,
607
+ 2945,
608
+ 3216,
609
+ 8706,
610
+ 9193,
611
+ 12725,
612
+ 13480,
613
+ 13768,
614
+ 15698,
615
+ 23992
616
+ ],
617
+ "conversion_parity_passed": false,
618
+ "output_validation_failures": {
619
+ "19270": [
620
+ "Live option probabilities do not sum to one"
621
+ ]
622
+ }
623
+ }
624
+ },
625
+ "failures": [
626
+ {
627
+ "row": 137,
628
+ "context": "TASK fill the form from the document, then submit\nFORM Lakeside Vet - Pet Owner Registration - Microsoft Edge\nELEMENT Edit \"Full name\" value=\"\"",
629
+ "gold_option": "skip",
630
+ "selected_options": {
631
+ "upstream_pytorch": "fill Last: Novak",
632
+ "baseline": "fill Last: Novak",
633
+ "ane-gather": "fill Last: Novak"
634
+ }
635
+ },
636
+ {
637
+ "row": 884,
638
+ "context": "TASK fill the form from the document, then submit\nFORM Acme Robotics - Job Application\nELEMENT Edit \"Health plan\" value=\"\"",
639
+ "gold_option": "skip",
640
+ "selected_options": {
641
+ "upstream_pytorch": "skip",
642
+ "baseline": "skip",
643
+ "ane-gather": "skip"
644
+ }
645
+ },
646
+ {
647
+ "row": 2945,
648
+ "context": "TASK fill the form from the document, then submit\nFORM Metro Utilities - New Service Request\nELEMENT Edit \"Organization\" value=\"\"",
649
+ "gold_option": "skip",
650
+ "selected_options": {
651
+ "upstream_pytorch": "skip",
652
+ "baseline": "skip",
653
+ "ane-gather": "skip"
654
+ }
655
+ },
656
+ {
657
+ "row": 3216,
658
+ "context": "TASK fill the form from the document, then submit\nFORM Northwind Clinic - New Patient Registration - Google Chrome\nELEMENT Edit \"Email\" value=\"oivanova66@yahoo.com\"",
659
+ "gold_option": "skip",
660
+ "selected_options": {
661
+ "upstream_pytorch": "fill Email address: oivanova66@yahoo.com",
662
+ "baseline": "fill Email address: oivanova66@yahoo.com",
663
+ "ane-gather": "fill Email address: oivanova66@yahoo.com"
664
+ }
665
+ },
666
+ {
667
+ "row": 6631,
668
+ "context": "TASK fill the form from the document, then submit\nFORM Blue Sun Freight - Vendor Onboarding - Google Chrome\nELEMENT Edit \"Given name\" value=\"Priya\"",
669
+ "gold_option": "skip",
670
+ "selected_options": {
671
+ "upstream_pytorch": "fill Physician: Dr. Fischer",
672
+ "baseline": "fill Physician: Dr. Fischer",
673
+ "ane-gather": "fill Physician: Dr. Fischer"
674
+ }
675
+ },
676
+ {
677
+ "row": 6748,
678
+ "context": "TASK fill the form from the document, then submit\nFORM JevBrowser\nELEMENT Edit \"State / Province\" value=\"NE\"",
679
+ "gold_option": "skip",
680
+ "selected_options": {
681
+ "upstream_pytorch": "fill ST: NE",
682
+ "baseline": "fill ST: NE",
683
+ "ane-gather": "fill ST: NE"
684
+ }
685
+ },
686
+ {
687
+ "row": 7280,
688
+ "context": "TASK fill the form from the document, then submit\nFORM \nELEMENT Edit \"Given name\" value=\"Tomas\" hint=\"First\"",
689
+ "gold_option": "skip",
690
+ "selected_options": {
691
+ "upstream_pytorch": "fill First name: Tomas",
692
+ "baseline": "fill First name: Tomas",
693
+ "ane-gather": "fill First name: Tomas"
694
+ }
695
+ },
696
+ {
697
+ "row": 8706,
698
+ "context": "TASK fill the form from the document, then submit\nFORM Northwind Clinic - New Patient Registration - Mozilla Firefox\nELEMENT CheckBox \"I authorize verification of the information provided\" unchecked",
699
+ "gold_option": "check",
700
+ "selected_options": {
701
+ "upstream_pytorch": "check",
702
+ "baseline": "check",
703
+ "ane-gather": "check"
704
+ }
705
+ },
706
+ {
707
+ "row": 9193,
708
+ "context": "TASK fill the form from the document, then submit\nFORM Blue Sun Freight - Vendor Onboarding\nELEMENT Edit \"DL number\" value=\"\"",
709
+ "gold_option": "skip",
710
+ "selected_options": {
711
+ "upstream_pytorch": "skip",
712
+ "baseline": "skip",
713
+ "ane-gather": "skip"
714
+ }
715
+ },
716
+ {
717
+ "row": 10703,
718
+ "context": "TASK fill the form from the document, then submit\nFORM Apex Gym - Member Enrollment - JevBrowser\nELEMENT Edit \"Last\" value=\"Silva\" hint=\"Family name\"",
719
+ "gold_option": "skip",
720
+ "selected_options": {
721
+ "upstream_pytorch": "fill Family name: Silva",
722
+ "baseline": "fill Family name: Silva",
723
+ "ane-gather": "fill Family name: Silva"
724
+ }
725
+ },
726
+ {
727
+ "row": 12255,
728
+ "context": "TASK fill the form from the document, then submit\nFORM Globex Insurance - Auto Claim Form - JevBrowser\nELEMENT Edit \"Taxpayer identification number\" value=\"\"",
729
+ "gold_option": "skip",
730
+ "selected_options": {
731
+ "upstream_pytorch": "fill National ID: 813-20-3361",
732
+ "baseline": "fill National ID: 813-20-3361",
733
+ "ane-gather": "fill National ID: 813-20-3361"
734
+ }
735
+ },
736
+ {
737
+ "row": 12725,
738
+ "context": "TASK fill the form from the document, then submit\nFORM JevBrowser\nELEMENT Edit \"Company\" value=\"\"",
739
+ "gold_option": "skip",
740
+ "selected_options": {
741
+ "upstream_pytorch": "skip",
742
+ "baseline": "skip",
743
+ "ane-gather": "skip"
744
+ }
745
+ },
746
+ {
747
+ "row": 13480,
748
+ "context": "TASK fill the form from the document, then submit\nFORM Evergreen Rentals - Tenant Application\nELEMENT Edit \"Current employer\" value=\"\"",
749
+ "gold_option": "skip",
750
+ "selected_options": {
751
+ "upstream_pytorch": "fill Company reg: REG3878384",
752
+ "baseline": "fill Company reg: REG3878384",
753
+ "ane-gather": "fill Company reg: REG3878384"
754
+ }
755
+ },
756
+ {
757
+ "row": 13768,
758
+ "context": "TASK fill the form from the document, then submit\nFORM Northwind Clinic - New Patient Registration\nELEMENT Edit \"Legal name\" value=\"\"",
759
+ "gold_option": "skip",
760
+ "selected_options": {
761
+ "upstream_pytorch": "skip",
762
+ "baseline": "skip",
763
+ "ane-gather": "skip"
764
+ }
765
+ },
766
+ {
767
+ "row": 15214,
768
+ "context": "TASK fill the form from the document, then submit\nFORM Riverside Dental - Patient Intake\nELEMENT Edit \"Relation to applicant\" value=\"\"",
769
+ "gold_option": "skip",
770
+ "selected_options": {
771
+ "upstream_pytorch": "fill Reg plate: WN9-8147",
772
+ "baseline": "fill Reg plate: WN9-8147",
773
+ "ane-gather": "fill Reg plate: WN9-8147"
774
+ }
775
+ },
776
+ {
777
+ "row": 15698,
778
+ "context": "TASK fill the form from the document, then submit\nFORM Orbit Telecom - Account Transfer - Mozilla Firefox\nELEMENT CheckBox \"I am legally authorized to work in this country\" checked",
779
+ "gold_option": "skip",
780
+ "selected_options": {
781
+ "upstream_pytorch": "skip",
782
+ "baseline": "skip",
783
+ "ane-gather": "skip"
784
+ }
785
+ },
786
+ {
787
+ "row": 19270,
788
+ "context": "TASK fill the form from the document, then submit\nFORM JevBrowser\nELEMENT Edit \"Plate number\" value=\"old\"",
789
+ "gold_option": "fill License plate: SJ3-1115",
790
+ "selected_options": {
791
+ "upstream_pytorch": "fill License plate: SJ3-1115",
792
+ "baseline": "fill License plate: SJ3-1115",
793
+ "ane-gather": "fill License plate: SJ3-1115"
794
+ }
795
+ },
796
+ {
797
+ "row": 21022,
798
+ "context": "TASK fill the form from the document, then submit\nFORM Harbor University - Graduate Application - JevBrowser\nELEMENT Edit \"State\" value=\"NE\"",
799
+ "gold_option": "skip",
800
+ "selected_options": {
801
+ "upstream_pytorch": "fill National ID: 794-03-1212",
802
+ "baseline": "fill National ID: 794-03-1212",
803
+ "ane-gather": "fill National ID: 794-03-1212"
804
+ }
805
+ },
806
+ {
807
+ "row": 23671,
808
+ "context": "TASK fill the form from the document, then submit\nFORM \nELEMENT Edit \"Email\" value=\"lchen6@yahoo.com\"",
809
+ "gold_option": "skip",
810
+ "selected_options": {
811
+ "upstream_pytorch": "fill Email: lchen6@yahoo.com",
812
+ "baseline": "fill Email: lchen6@yahoo.com",
813
+ "ane-gather": "fill Email: lchen6@yahoo.com"
814
+ }
815
+ },
816
+ {
817
+ "row": 23992,
818
+ "context": "TASK fill the form from the document, then submit\nFORM Apex Gym - Member Enrollment - Microsoft Edge\nELEMENT Edit \"Total years experience\" value=\"\"",
819
+ "gold_option": "skip",
820
+ "selected_options": {
821
+ "upstream_pytorch": "skip",
822
+ "baseline": "skip",
823
+ "ane-gather": "skip"
824
+ }
825
+ }
826
+ ],
827
+ "trace": {
828
+ "file": "synthetic-test-decisions.jsonl.gz",
829
+ "sha256": "b6f6edc9de34fb37b10847a1f9568b530ce2b8e1bd241f51bf2e637f489d6554",
830
+ "rows": 24370
831
+ },
832
+ "upstream_evaluator_sha256": "49dccdeb4e2b456187cdc4491a6d78fabac6a45384272f8e3ec802a58872aa6e",
833
+ "harness_sha256": {
834
+ "benchmark-synthetic.py": "96b75133b38735594e7b5483cec0f39a5f9868045d0e7438434df8af698215fa",
835
+ "synthetic_test.py": "e2b68725236e6827eb422480c8cc8aa189b2490c1594650d8f2928771d1bc626",
836
+ "score-report.py": "c78e45caa86fb59f9a56ef18cc6a35020ea6e281a6a23d23618f56da85a0d95e"
837
+ },
838
+ "conversion_parity_passed": false,
839
+ "limitations": [
840
+ "The release file has 24370 rows; the model card claims approximately 15000. This is not a reconstruction of that unspecified manifest.",
841
+ "Upstream describes form-disjoint synthetic data; this is not an unseen real-world GUI or task-completion benchmark.",
842
+ "No fresh hosted Jev calls are included. Published hosted-Jev percentages are not a matched baseline for this run.",
843
+ "Local timings use CPU for PyTorch and CPU+ANE for Core ML; they are backend/deployment measurements, not equal-hardware algorithm comparisons."
844
+ ]
845
+ }
synthetic-test.lock.json ADDED
@@ -0,0 +1,15 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "dataset_repository": "cua-ai/cua-s1-forms",
3
+ "dataset_revision": "8273f34778b99ac2e12d9f6e7d57dad99ae20845",
4
+ "path": "artifacts/test.jsonl",
5
+ "url": "https://huggingface.co/datasets/cua-ai/cua-s1-forms/resolve/8273f34778b99ac2e12d9f6e7d57dad99ae20845/test.jsonl",
6
+ "sha256": "d63a7e0db195d4d20154a40b2f8dd09ce3bb65487a158c638da5c609d4475e7c",
7
+ "rows": 24370,
8
+ "bytes": 18710924,
9
+ "upstream_split_claim": "Synthetic test, form-signature-disjoint from train/validation. Train and validation signatures have not been independently re-audited here.",
10
+ "model_card_claim": {
11
+ "synthetic_test_top1": 0.9995,
12
+ "approximate_decisions": 15000,
13
+ "note": "The released file has 24370 rows. No matching full result ledger or exact 15000-row manifest is supplied by the linked model card."
14
+ }
15
+ }