File size: 7,895 Bytes
c87b915
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
---
license: mit
library_name: coreml
pipeline_tag: text-classification
base_model: cua-ai/cua-s1-forms
base_model_relation: quantized
datasets:
- cua-ai/cua-s1-forms
tags:
- coreml
- apple-silicon
- computer-use
- classification
- cua-s1
- fp16
---

# CUA-S1-FORMS — Core ML

FP16 Core ML conversion of [Cua's CUA-S1-FORMS](https://huggingface.co/cua-ai/cua-s1-forms),
a **706,048-parameter** specialist that selects among supplied form actions.
The portable package is **1,511,163 bytes (1.51 MB)**. No text generation, KV cache,
or external tokenizer is required.

The model uses small Transformer encoders over UTF-8 bytes and an attention
readout. It is a classifier, not an autoregressive LLM. The inherited checkpoint
configuration contains an unused `hf_model` value; this `tinyx` checkpoint does
not load Qwen weights.

## Files

| File | Purpose |
| --- | --- |
| `cua_s1_forms_fp16_options32.mlpackage/` | Portable model; compile locally or add to Xcode |
| `cua_s1_forms_fp16_options32.mlmodelc/` | Compiled bundle for FluidAudio's model loader |
| `preprocessing.py` | Upstream-compatible byte encoding and input validation |
| `conversion.json` | Architecture, conversion versions, and portable package hashes |
| `assets.lock.json` | Pinned upstream source, model, and demo hashes |
| `checksums.json` | SHA-256 of each distributed file except this checksum file |
| `reports/` | Per-row parity, original Cua metrics, and compute-placement report |

The model targets **iOS 17/macOS 14 or newer**. Runtime validation used an Apple
silicon Mac. Use the portable package for local compilation on other supported
systems; iPhone performance and compatibility of the precompiled bundle across
older OS versions have not been measured.

## Python usage

Install `coremltools==9.0`, `numpy==1.26.4`, and `huggingface_hub` on macOS.
Download this repository, then run from its directory:

```python
import coremltools as ct
from preprocessing import InputLimits, prepare_inputs

model = ct.models.MLModel(
    "cua_s1_forms_fp16_options32.mlpackage",
    compute_units=ct.ComputeUnit.CPU_AND_NE,
)
options = ["fill E-mail: person@example.com", "check", "click", "skip"]
inputs = prepare_inputs(
    'TASK fill the form from the document, then submit\n'
    'FORM Contact details\nELEMENT Edit "Email address" value=""',
    options,
    InputLimits(),
)
probabilities = model.predict(inputs)["probabilities"][0, :len(options)]
print(options[int(probabilities.argmax())], probabilities)
```

During review, download the model PR's revision (for example `--revision refs/pr/1`)
with `hf download FluidInference/cua-s1-forms-coreml --local-dir ./cua-coreml`.
After merge, the default `main` revision contains the artifacts.

## Swift usage

The proposed [FluidAudio integration](https://github.com/FluidInference/FluidAudio/tree/codex/cua-s1-forms)
provides `CuaS1FormsManager`:

```swift
import FluidAudio
import Foundation

let manager = try await CuaS1FormsManager.load(
    from: URL(fileURLWithPath: "/models/cua_s1_forms_fp16_options32.mlpackage"))
let decision = try await manager.score(
    context: "TASK fill the form from the document, then submit\nFORM Contact details\nELEMENT Edit \"Email address\" value=\"\"",
    options: ["fill E-mail: person@example.com", "check", "click", "skip"])
print(decision.selectedOption, decision.probabilities)
```

After the model and Swift PRs land, `try await CuaS1FormsManager.load()` downloads
and caches the compiled artifact automatically.

## Tensor interface

| Name | Type | Shape |
| --- | --- | --- |
| `context_ids` | int32 | `[1, 224]` |
| `option_ids` | int32 | `[1, 32, 96]` |
| `option_mask` | int32 | `[1, 32]` |
| `logits` | float32 output | `[1, 32]` |
| `probabilities` | float32 output | `[1, 32]` |

Encode UTF-8 bytes plus one, pad with zero, and truncate by bytes at 224 for
context and 96 per option. Supply a nonempty context and 2–32 nonempty options.
Set option-mask entries to one for supplied options and zero for padding.
Padded logits are `-10000`; padded probabilities are zero. Inputs above 32
options must be rejected or use a separately exported larger-capacity model.
The example helper rejects overflow rather than dropping choices.

## Conversion verification

On the complete pinned **196-row upstream demo**, PyTorch and Core ML both
selected **196/196 labeled options correctly**, and matched each other's selected
option on every row. The unmodified upstream Cua evaluator reports 36 fill,
4 check, 6 click, and 150 skip decisions, with zero wrong actions, wrong targets,
or unsafe actions on these saved predictions. It counts `skip` as abstention,
so its 23.47% coverage corresponds to 46 actionable decisions.

| Check | Core ML `ALL` | Core ML `CPU_AND_NE` |
| --- | ---: | ---: |
| Selected options matching PyTorch | 196 / 196 | 196 / 196 |
| Maximum absolute probability error | 0.003099 | 0.002336 |
| Warm model-call median | 1.85 ms | 0.90 ms |
| Warm model-call p95 | 2.49 ms | 0.94 ms |

Measured September 19, 2026 on Apple M5 Pro, 24 GB, macOS 27.0, using
Python 3.11.11, PyTorch 2.7.0, and coremltools 9.0. Timing is exploratory and
includes Python call overhead; model loading, encoding, document extraction,
UI observation, and action execution are excluded. It is not an optimized
PyTorch/MPS speed comparison.

The conversion gates require 100% selected-option agreement, no accuracy loss,
maximum absolute probability error ≤ 0.005, finite outputs, normalized live
probabilities, and zero probability for padding. The FP32 export adapter differs
from the unmodified PyTorch reference by at most 0.00000113. Six reversed-option
checks also pass on each Core ML configuration. The checked-in placement report
counts 149 Neural Engine operations, 24 CPU operations, and zero GPU operations;
operation counts are not a measurement of time spent on each processor.

These results verify conversion on three demo forms and three PDFs. They do not
establish generalization, live GUI completion rates, or production safety.
The original demo is not redistributed here; its exact revision and SHA-256 are
in the asset lock. No training or larger-corpus evaluation was performed for
this conversion.

## Application responsibilities

The application must extract document entities, describe UI elements, build
candidate actions, and validate and order the selected actions. Submission and
other effects require application authorization. Scores are not calibrated
confidence guarantees. Text outside the byte limits is truncated, and arbitrary
new forms and languages require their own evaluation.

## Source, reproduction, and license

The [Mobius conversion toolkit](https://github.com/FluidInference/mobius/tree/codex/cua-s1-forms/models/computer-use/cua-s1-forms/coreml)
contains the conversion code, lockfile, original reference implementation and
evaluator, tests, and full reproduction instructions. Adaptations are limited
to export-compatible masking, a floating-point clamp constant, finite padded
logits, and disabling the fused PyTorch Transformer fast path during tracing.
All trained layers and checkpoint tensors are retained; internal compute and
weights are converted to FP16.

- Model: [`cua-ai/cua-s1-forms` at `f54adbf447f4ca6ec259f529ee3f2e3e09f8cc71`](https://huggingface.co/cua-ai/cua-s1-forms/tree/f54adbf447f4ca6ec259f529ee3f2e3e09f8cc71).
- Demo: [`cua-ai/cua-s1-forms` at `8273f34778b99ac2e12d9f6e7d57dad99ae20845`](https://huggingface.co/datasets/cua-ai/cua-s1-forms/tree/8273f34778b99ac2e12d9f6e7d57dad99ae20845).
- Code: [`trycua/cua` at `83f142c4290a0f7d9ed545ae8532858c6e4f8145`](https://github.com/trycua/cua/tree/83f142c4290a0f7d9ed545ae8532858c6e4f8145/libs/cua-s1).

The pinned model and dataset cards declare MIT. See [LICENSE](LICENSE),
[NOTICES.md](NOTICES.md), and the preserved
[upstream third-party notices](UPSTREAM-THIRD-PARTY-NOTICES.md).