|
Download README.md from FluidInference/cua-s1-forms-coreml: direct link, hf CLI and curl.
- Browser
- Download file 7.9 kB
-
https://huggingface.co/FluidInference/cua-s1-forms-coreml/resolve/c87b915d302bdbff644709408d8fbd8c8effe894/README.md
- Command line
-
hf download hf://FluidInference/cua-s1-forms-coreml@c87b915d302bdbff644709408d8fbd8c8effe894/README.md
-
curl -L -o README.md https://huggingface.co/FluidInference/cua-s1-forms-coreml/resolve/c87b915d302bdbff644709408d8fbd8c8effe894/README.md
7.9 kB
| license: mit | |
| library_name: coreml | |
| pipeline_tag: text-classification | |
| base_model: cua-ai/cua-s1-forms | |
| base_model_relation: quantized | |
| datasets: | |
| - cua-ai/cua-s1-forms | |
| tags: | |
| - coreml | |
| - apple-silicon | |
| - computer-use | |
| - classification | |
| - cua-s1 | |
| - fp16 | |
| # CUA-S1-FORMS — Core ML | |
| FP16 Core ML conversion of [Cua's CUA-S1-FORMS](https://huggingface.co/cua-ai/cua-s1-forms), | |
| a **706,048-parameter** specialist that selects among supplied form actions. | |
| The portable package is **1,511,163 bytes (1.51 MB)**. No text generation, KV cache, | |
| or external tokenizer is required. | |
| The model uses small Transformer encoders over UTF-8 bytes and an attention | |
| readout. It is a classifier, not an autoregressive LLM. The inherited checkpoint | |
| configuration contains an unused `hf_model` value; this `tinyx` checkpoint does | |
| not load Qwen weights. | |
| ## Files | |
| | File | Purpose | | |
| | --- | --- | | |
| | `cua_s1_forms_fp16_options32.mlpackage/` | Portable model; compile locally or add to Xcode | | |
| | `cua_s1_forms_fp16_options32.mlmodelc/` | Compiled bundle for FluidAudio's model loader | | |
| | `preprocessing.py` | Upstream-compatible byte encoding and input validation | | |
| | `conversion.json` | Architecture, conversion versions, and portable package hashes | | |
| | `assets.lock.json` | Pinned upstream source, model, and demo hashes | | |
| | `checksums.json` | SHA-256 of each distributed file except this checksum file | | |
| | `reports/` | Per-row parity, original Cua metrics, and compute-placement report | | |
| The model targets **iOS 17/macOS 14 or newer**. Runtime validation used an Apple | |
| silicon Mac. Use the portable package for local compilation on other supported | |
| systems; iPhone performance and compatibility of the precompiled bundle across | |
| older OS versions have not been measured. | |
| ## Python usage | |
| Install `coremltools==9.0`, `numpy==1.26.4`, and `huggingface_hub` on macOS. | |
| Download this repository, then run from its directory: | |
| ```python | |
| import coremltools as ct | |
| from preprocessing import InputLimits, prepare_inputs | |
| model = ct.models.MLModel( | |
| "cua_s1_forms_fp16_options32.mlpackage", | |
| compute_units=ct.ComputeUnit.CPU_AND_NE, | |
| ) | |
| options = ["fill E-mail: person@example.com", "check", "click", "skip"] | |
| inputs = prepare_inputs( | |
| 'TASK fill the form from the document, then submit\n' | |
| 'FORM Contact details\nELEMENT Edit "Email address" value=""', | |
| options, | |
| InputLimits(), | |
| ) | |
| probabilities = model.predict(inputs)["probabilities"][0, :len(options)] | |
| print(options[int(probabilities.argmax())], probabilities) | |
| ``` | |
| During review, download the model PR's revision (for example `--revision refs/pr/1`) | |
| with `hf download FluidInference/cua-s1-forms-coreml --local-dir ./cua-coreml`. | |
| After merge, the default `main` revision contains the artifacts. | |
| ## Swift usage | |
| The proposed [FluidAudio integration](https://github.com/FluidInference/FluidAudio/tree/codex/cua-s1-forms) | |
| provides `CuaS1FormsManager`: | |
| ```swift | |
| import FluidAudio | |
| import Foundation | |
| let manager = try await CuaS1FormsManager.load( | |
| from: URL(fileURLWithPath: "/models/cua_s1_forms_fp16_options32.mlpackage")) | |
| let decision = try await manager.score( | |
| context: "TASK fill the form from the document, then submit\nFORM Contact details\nELEMENT Edit \"Email address\" value=\"\"", | |
| options: ["fill E-mail: person@example.com", "check", "click", "skip"]) | |
| print(decision.selectedOption, decision.probabilities) | |
| ``` | |
| After the model and Swift PRs land, `try await CuaS1FormsManager.load()` downloads | |
| and caches the compiled artifact automatically. | |
| ## Tensor interface | |
| | Name | Type | Shape | | |
| | --- | --- | --- | | |
| | `context_ids` | int32 | `[1, 224]` | | |
| | `option_ids` | int32 | `[1, 32, 96]` | | |
| | `option_mask` | int32 | `[1, 32]` | | |
| | `logits` | float32 output | `[1, 32]` | | |
| | `probabilities` | float32 output | `[1, 32]` | | |
| Encode UTF-8 bytes plus one, pad with zero, and truncate by bytes at 224 for | |
| context and 96 per option. Supply a nonempty context and 2–32 nonempty options. | |
| Set option-mask entries to one for supplied options and zero for padding. | |
| Padded logits are `-10000`; padded probabilities are zero. Inputs above 32 | |
| options must be rejected or use a separately exported larger-capacity model. | |
| The example helper rejects overflow rather than dropping choices. | |
| ## Conversion verification | |
| On the complete pinned **196-row upstream demo**, PyTorch and Core ML both | |
| selected **196/196 labeled options correctly**, and matched each other's selected | |
| option on every row. The unmodified upstream Cua evaluator reports 36 fill, | |
| 4 check, 6 click, and 150 skip decisions, with zero wrong actions, wrong targets, | |
| or unsafe actions on these saved predictions. It counts `skip` as abstention, | |
| so its 23.47% coverage corresponds to 46 actionable decisions. | |
| | Check | Core ML `ALL` | Core ML `CPU_AND_NE` | | |
| | --- | ---: | ---: | | |
| | Selected options matching PyTorch | 196 / 196 | 196 / 196 | | |
| | Maximum absolute probability error | 0.003099 | 0.002336 | | |
| | Warm model-call median | 1.85 ms | 0.90 ms | | |
| | Warm model-call p95 | 2.49 ms | 0.94 ms | | |
| Measured September 19, 2026 on Apple M5 Pro, 24 GB, macOS 27.0, using | |
| Python 3.11.11, PyTorch 2.7.0, and coremltools 9.0. Timing is exploratory and | |
| includes Python call overhead; model loading, encoding, document extraction, | |
| UI observation, and action execution are excluded. It is not an optimized | |
| PyTorch/MPS speed comparison. | |
| The conversion gates require 100% selected-option agreement, no accuracy loss, | |
| maximum absolute probability error ≤ 0.005, finite outputs, normalized live | |
| probabilities, and zero probability for padding. The FP32 export adapter differs | |
| from the unmodified PyTorch reference by at most 0.00000113. Six reversed-option | |
| checks also pass on each Core ML configuration. The checked-in placement report | |
| counts 149 Neural Engine operations, 24 CPU operations, and zero GPU operations; | |
| operation counts are not a measurement of time spent on each processor. | |
| These results verify conversion on three demo forms and three PDFs. They do not | |
| establish generalization, live GUI completion rates, or production safety. | |
| The original demo is not redistributed here; its exact revision and SHA-256 are | |
| in the asset lock. No training or larger-corpus evaluation was performed for | |
| this conversion. | |
| ## Application responsibilities | |
| The application must extract document entities, describe UI elements, build | |
| candidate actions, and validate and order the selected actions. Submission and | |
| other effects require application authorization. Scores are not calibrated | |
| confidence guarantees. Text outside the byte limits is truncated, and arbitrary | |
| new forms and languages require their own evaluation. | |
| ## Source, reproduction, and license | |
| The [Mobius conversion toolkit](https://github.com/FluidInference/mobius/tree/codex/cua-s1-forms/models/computer-use/cua-s1-forms/coreml) | |
| contains the conversion code, lockfile, original reference implementation and | |
| evaluator, tests, and full reproduction instructions. Adaptations are limited | |
| to export-compatible masking, a floating-point clamp constant, finite padded | |
| logits, and disabling the fused PyTorch Transformer fast path during tracing. | |
| All trained layers and checkpoint tensors are retained; internal compute and | |
| weights are converted to FP16. | |
| - Model: [`cua-ai/cua-s1-forms` at `f54adbf447f4ca6ec259f529ee3f2e3e09f8cc71`](https://huggingface.co/cua-ai/cua-s1-forms/tree/f54adbf447f4ca6ec259f529ee3f2e3e09f8cc71). | |
| - Demo: [`cua-ai/cua-s1-forms` at `8273f34778b99ac2e12d9f6e7d57dad99ae20845`](https://huggingface.co/datasets/cua-ai/cua-s1-forms/tree/8273f34778b99ac2e12d9f6e7d57dad99ae20845). | |
| - Code: [`trycua/cua` at `83f142c4290a0f7d9ed545ae8532858c6e4f8145`](https://github.com/trycua/cua/tree/83f142c4290a0f7d9ed545ae8532858c6e4f8145/libs/cua-s1). | |
| The pinned model and dataset cards declare MIT. See [LICENSE](LICENSE), | |
| [NOTICES.md](NOTICES.md), and the preserved | |
| [upstream third-party notices](UPSTREAM-THIRD-PARTY-NOTICES.md). | |