cua-s1-forms-coreml / README.md
alexwengg's picture
Add optional ANE-gather variant with 98.2 percent ANE placement
f66dd2a verified
|
Raw History Blame
12.1 kB
metadata
license: mit
library_name: coreml
pipeline_tag: text-classification
base_model: cua-ai/cua-s1-forms
base_model_relation: quantized
datasets:
  - cua-ai/cua-s1-forms
tags:
  - coreml
  - apple-silicon
  - computer-use
  - classification
  - cua-s1
  - fp16

CUA-S1-FORMS — Core ML

FP16 Core ML conversion of Cua's CUA-S1-FORMS, a 706,048-parameter specialist that selects among supplied form actions. The portable package is 1,511,163 bytes (1.51 MB). No text generation, KV cache, or external tokenizer is required.

The model uses small Transformer encoders over UTF-8 bytes and an attention readout. It is a classifier, not an autoregressive LLM. The inherited checkpoint configuration contains an unused hf_model value; this tinyx checkpoint does not load Qwen weights.

Files

File Purpose
cua_s1_forms_fp16_options32.mlpackage/ Portable model; compile locally or add to Xcode
cua_s1_forms_fp16_options32.mlmodelc/ Compiled bundle for FluidAudio's model loader
preprocessing.py Upstream-compatible byte encoding and input validation
conversion.json Architecture, conversion versions, and portable package hashes
assets.lock.json Pinned upstream source, model, and demo hashes
checksums.json SHA-256 of each distributed file except this checksum file
reports/ Per-row parity, original Cua metrics, and compute-placement report

The model targets iOS 17/macOS 14 or newer. Runtime validation used an Apple silicon Mac. Use the portable package for local compilation on other supported systems; iPhone performance and compatibility of the precompiled bundle across older OS versions have not been measured.

Python usage

Install coremltools==9.0, numpy==1.26.4, and huggingface_hub on macOS. Download this repository, then run from its directory:

import coremltools as ct
from preprocessing import InputLimits, prepare_inputs

model = ct.models.MLModel(
    "cua_s1_forms_fp16_options32.mlpackage",
    compute_units=ct.ComputeUnit.CPU_AND_NE,
)
options = ["fill E-mail: person@example.com", "check", "click", "skip"]
inputs = prepare_inputs(
    'TASK fill the form from the document, then submit\n'
    'FORM Contact details\nELEMENT Edit "Email address" value=""',
    options,
    InputLimits(),
)
probabilities = model.predict(inputs)["probabilities"][0, :len(options)]
print(options[int(probabilities.argmax())], probabilities)

During review, download the model PR's revision (for example --revision refs/pr/1) with hf download FluidInference/cua-s1-forms-coreml --local-dir ./cua-coreml. After merge, the default main revision contains the artifacts.

Swift usage

The proposed FluidAudio integration provides CuaS1FormsManager:

import FluidAudio
import Foundation

let manager = try await CuaS1FormsManager.load(
    from: URL(fileURLWithPath: "/models/cua_s1_forms_fp16_options32.mlpackage"))
let decision = try await manager.score(
    context: "TASK fill the form from the document, then submit\nFORM Contact details\nELEMENT Edit \"Email address\" value=\"\"",
    options: ["fill E-mail: person@example.com", "check", "click", "skip"])
print(decision.selectedOption, decision.probabilities)

After the model and Swift PRs land, try await CuaS1FormsManager.load() downloads and caches the compiled artifact automatically.

Tensor interface

Name Type Shape
context_ids int32 [1, 224]
option_ids int32 [1, 32, 96]
option_mask int32 [1, 32]
logits float32 output [1, 32]
probabilities float32 output [1, 32]

Encode UTF-8 bytes plus one, pad with zero, and truncate by bytes at 224 for context and 96 per option. Supply a nonempty context and 2–32 nonempty options. Set option-mask entries to one for supplied options and zero for padding. Padded logits are -10000; padded probabilities are zero. Inputs above 32 options must be rejected or use a separately exported larger-capacity model. The example helper rejects overflow rather than dropping choices.

Conversion verification

On the complete pinned 196-row upstream demo, PyTorch and Core ML both selected 196/196 labeled options correctly, and matched each other's selected option on every row. The unmodified upstream Cua evaluator reports 36 fill, 4 check, 6 click, and 150 skip decisions, with zero wrong actions, wrong targets, or unsafe actions on these saved predictions. It counts skip as abstention, so its 23.47% coverage corresponds to 46 actionable decisions.

Check Core ML ALL Core ML CPU_AND_NE
Selected options matching PyTorch 196 / 196 196 / 196
Maximum absolute probability error 0.003099 0.002336
Warm model-call median 1.85 ms 0.90 ms
Warm model-call p95 2.49 ms 0.94 ms

Measured September 19, 2026 on Apple M5 Pro, 24 GB, macOS 27.0, using Python 3.11.11, PyTorch 2.7.0, and coremltools 9.0. Timing is exploratory and includes Python call overhead; model loading, encoding, document extraction, UI observation, and action execution are excluded. It is not an optimized PyTorch/MPS speed comparison.

The conversion gates require 100% selected-option agreement, no accuracy loss, maximum absolute probability error ≤ 0.005, finite outputs, normalized live probabilities, and zero probability for padding. The FP32 export adapter differs from the unmodified PyTorch reference by at most 0.00000113. Six reversed-option checks also pass on each Core ML configuration. The checked-in placement report counts 149 Neural Engine operations, 24 CPU operations, and zero GPU operations; operation counts are not a measurement of time spent on each processor.

These results verify conversion on three demo forms and three PDFs. They do not establish generalization, live GUI completion rates, or production safety. The original demo is not redistributed here; its exact revision and SHA-256 are in the asset lock. No training or larger-corpus evaluation was performed for this conversion.

ANE profile

A separate September 19 profile uses the same portable-package hashes on the M5 Pro, with real demo rows 0, 68, and 130 (27, 21, and 19 options). Each policy runs two warmup passes and ten timed passes, totaling 30 timed predictions. All 120 timed predictions select the correct labels.

Policy CPU ops GPU ops ANE ops Warm p50 Warm p95
CPU_ONLY 173 0 0 1.527 ms 1.602 ms
CPU_AND_GPU 0 173 0 0.929 ms 2.380 ms
CPU_AND_NE 24 0 149 0.929 ms 0.973 ms
ALL 0 173 0 0.912 ms 1.229 ms

CPU_AND_NE assigns 86.1% of operations to ANE; ALL chooses the GPU on this Mac. CPU fallbacks cover integer/mask preparation and embedding gathers. These are public MLComputePlan preferred-device assignments, not measurements of utilization, energy, or time spent on each device. No Instruments runtime trace was captured.

ANE model loading took 566.8 ms, followed by a 1.65 ms first prediction, with system caches retained. These are not first-install cold-start numbers. Warm timing includes Python model-call overhead and excludes encoding, Swift/UI work, and animation. This three-row timing manifest differs from the full conversion parity run above; no weights or graph were changed.

See reports/ane-profile.json for all operation assignments, individual timings, hashes, and the protocol; reports/ane-fallback.json records rejection reasons. Reproduce with uv run --frozen python profile-coreml.py in the Mobius conversion directory.

Optional higher-ANE variant

The ane-gather/ directory contains an alternative portable package and compiled bundle with the same int32 inputs and float32 outputs and all trained weights. The variant uses shared float16 mask inputs and unsigned 16-bit embedding indices to eliminate negative-index correction and place the gathers on ANE. Valid byte IDs 0–256 remain exact. It is 1,509,491 bytes as a portable package.

On this M5 Pro, the scheduler plan is 162 ANE operations and 3 CPU input casts (98.2% ANE) for both CPU_AND_NE and ALL. The default model has 149 ANE and 24 CPU operations (86.1%) under CPU_AND_NE. Counts are not runtime or energy shares, and host byte encoding still runs outside the model.

The optional variant passes 196/196 decisions against upstream on ALL and CPU_AND_NE, with maximum probability error 0.002336 under the unchanged 0.005 tolerance. 28 Python regression tests pass, including all byte-ID boundaries, full option capacity, truncation, and reordered choices. The Swift manager independently passes all 196 reference decisions, compiled-cache loading, and concurrent/reordered requests; the native demo passes its three-form checks.

A matched same-process ABBA comparison uses three real inputs and 60 timed calls per model after warmup:

Artifact CPU ops ANE ops Warm p50 Warm p95
Root/default 24 149 0.915 ms 0.968 ms
ane-gather/ 3 162 0.970 ms 0.988 ms

Higher ANE placement is about 6% slower in this local comparison, so the root/default artifact remains unchanged. No energy or CPU-time saving is claimed. The original input names, dtypes, shapes, and byte encoding still apply. Load ane-gather/cua_s1_forms_fp16_options32.mlpackage with the existing Python or Swift APIs, or pass its local path to the Swift demo's --model argument.

Reports: parity, Swift validation, compute plans, fallbacks, and matched comparison. Reproduce with uv run --frozen python convert-coreml.py --optimization ane-gather --output-dir build/ane-gather in the Mobius conversion directory.

Application responsibilities

The application must extract document entities, describe UI elements, build candidate actions, and validate and order the selected actions. Submission and other effects require application authorization. Scores are not calibrated confidence guarantees. Text outside the byte limits is truncated, and arbitrary new forms and languages require their own evaluation.

Source, reproduction, and license

The Mobius conversion toolkit contains the conversion code, lockfile, original reference implementation and evaluator, tests, and full reproduction instructions. Adaptations are limited to export-compatible masking, a floating-point clamp constant, finite padded logits, and disabling the fused PyTorch Transformer fast path during tracing. All trained layers and checkpoint tensors are retained; internal compute and weights are converted to FP16.

The pinned model and dataset cards declare MIT. See LICENSE, NOTICES.md, and the preserved upstream third-party notices.