--- license: mit tags: - onnx - computer-use - form-filling - webgpu - wasm - cua base_model: cua-ai/cua-s1-forms --- # CUA-S1 Forms ONNX ONNX conversion of the original [`cua-ai/cua-s1-forms`](https://huggingface.co/cua-ai/cua-s1-forms) model for local and browser inference experiments. Original project: https://github.com/trycua/cua/tree/main/libs/cua-s1 This repository only provides a deployment-format conversion of the original trained checkpoint. The architecture, weights, training, evaluation, and research belong to the original CUA-S1 project. ## Why ONNX? CUA-S1 Forms is a small specialist model for form-oriented computer-use tasks. Instead of generating text token-by-token, it scores a bounded set of candidate actions in a single forward pass. This makes it suitable for lightweight local inference using: - ONNX Runtime - ONNX Runtime Web - WebAssembly - WebGPU No server-side inference is required. ## Model The original model uses: - byte-level input encoding - 2-layer Transformer encoder - width 128 - 4 attention heads - option-attention scoring - ~706K parameters Please refer to the original model card for the authoritative architecture, training details, evaluation results, limitations, and safety guidance: https://huggingface.co/cua-ai/cua-s1-forms ## Files ```text cua-s1-forms.onnx cua-s1-forms-web.json ``` ## Input contract ```text context_ids int64 [B, 224] context_mask bool [B, 224] option_ids int64 [B, N, 96] option_token_mask bool [B, N, 96] option_mask bool [B, N] ``` Output: ```text logits float32 [B, N] ``` Where: ```text B = batch size N = candidate options ``` ## Byte encoding CUA-S1 uses UTF-8 bytes directly: ```python byte_id = byte + 1 ``` `0` is reserved for padding. This means browser implementations can use JavaScript `TextEncoder` without requiring a tokenizer model. ## Python example ```python import onnxruntime as ort session = ort.InferenceSession( "cua-s1-forms.onnx", providers=["CPUExecutionProvider"] ) logits = session.run( ["logits"], { "context_ids": context_ids, "context_mask": context_mask, "option_ids": option_ids, "option_token_mask": option_token_mask, "option_mask": option_mask, } )[0] ``` ## Preliminary CPU result Google Colab, ONNX Runtime CPU, batch 1, 8 candidate options: ```text Median latency 15.0 ms Mean latency 16.8 ms P95 24.3 ms Throughput ~60 decisions/sec ``` These are deployment measurements for this ONNX conversion, not official CUA-S1 benchmark results. Browser WASM and WebGPU tests will be added separately. ## Intended browser deployment ```text Browser | v TextEncoder | v ONNX Runtime Web | +--> WASM | +--> WebGPU | v CUA-S1 Forms | v Action scores ``` ## Limitations All limitations of the original model still apply. CUA-S1 Forms is a specialist form-oriented model, not a general-purpose computer-use agent. Refer to the original model card for detailed scope, limitations, evaluation, and safety guidance. ## Attribution Original model: https://huggingface.co/cua-ai/cua-s1-forms Original source: https://github.com/trycua/cua/tree/main/libs/cua-s1 ONNX conversion: https://huggingface.co/yasserrmd/cua-s1-forms-onnx Converted for ONNX / browser deployment experiments by Mohamed Yasser. This repository is not an official CUA-AI release. ## License MIT, following the original model repository.