f-trycua commited on
Commit
249c395
·
verified ·
1 Parent(s): 1f93fd0

Add model card and Apache-2.0 license

Browse files

Adds a model card with YAML metadata (license: apache-2.0, library_name: cua-s1). This checkpoint was trained from scratch, so the card declares no base_model. The license matches cua-ai/cua-s1-4b-0.2.

The card documents intended and out-of-scope use, the text and multimodal (frozen SigLIP) layout, a training summary, measured cua-bench-s1 results, the pinned revision (1f93fd0f), and safetensors loading with cua_s1.nano. It also notes that the jev-use chooser does not support this checkpoint. Sources: trycua/cua libs/cua-s1/MODEL_CARD.md and README.md (trycua/cua#4186).

Only README.md changes; the weights do not.

Files changed (1) hide show
  1. README.md +174 -0
README.md ADDED
@@ -0,0 +1,174 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ library_name: cua-s1
4
+ tags:
5
+ - computer-use
6
+ - gui-agent
7
+ - option-attention
8
+ - cua-s1
9
+ - safetensors
10
+ language:
11
+ - en
12
+ ---
13
+
14
+ # cua-s1-nano-0.1
15
+
16
+ `cua-s1-nano-0.1` is a small research classifier for closed-option
17
+ computer-use decisions. It has about 855,000 trainable parameters and was
18
+ trained from scratch. Given a screen state and a closed set of candidate
19
+ (element, action) options, it scores every option in one parallel forward
20
+ pass and selects the highest-scoring option for each element.
21
+
22
+ This checkpoint belongs to the [Cua-S1](https://github.com/trycua/cua/tree/main/libs/cua-s1)
23
+ research family. It is not a general-purpose assistant. Do not treat a result
24
+ on one task family as evidence of reliability outside it.
25
+
26
+ ## Files
27
+
28
+ | Path | Contents |
29
+ | --- | --- |
30
+ | `text/model.safetensors`, `text/config.json` | Text checkpoint: a byte-level context encoder over an accessibility-tree excerpt for each element |
31
+ | `multimodal/model.safetensors`, `multimodal/config.json` | Multimodal checkpoint: a trainable projection over frozen `google/siglip-base-patch16-224` features from a screenshot crop for each element |
32
+
33
+ Each checkpoint has 855,296 parameters. Tensors are stored only in
34
+ safetensors, and the JSON sidecar holds the architecture config and a SHA-256
35
+ signature over the tensors. This repository contains no pickle files and does
36
+ not redistribute the SigLIP backbone.
37
+
38
+ ## Intended use
39
+
40
+ - Research on closed-option computer-use GUI decisions (element and action
41
+ selection) with a very small, fast architecture.
42
+ - Comparing a from-scratch specialist against larger checkpoints, such as
43
+ [`cua-s1-4b-0.2`](https://huggingface.co/cua-ai/cua-s1-4b-0.2), on
44
+ [`cua-bench-s1`](https://github.com/trycua/cua/tree/main/libs/cua-bench-s1).
45
+ - Studying task-family generalization, including out-of-domain checks.
46
+
47
+ ### Out of scope
48
+
49
+ - General-purpose or open-ended computer operation.
50
+ - Unsupervised operation on production accounts or sensitive data.
51
+ - Actions with financial, legal, medical, employment, safety, or other
52
+ high-impact consequences.
53
+ - Bypassing access controls, consent, rate limits, or service policies.
54
+ - Treating a selected option as proof that the action is correct, safe, or
55
+ succeeded.
56
+
57
+ ## How it works
58
+
59
+ An option-attention head uses each candidate option as a query against the
60
+ element's context tokens and produces one logit per option. Options are
61
+ always encoded as text by a shared byte-level option encoder. The context
62
+ depends on the modality:
63
+
64
+ - **Text**: a small trainable byte-level transformer over a rendered
65
+ accessibility-tree excerpt for the element. It needs no extra dependencies.
66
+ - **Multimodal**: a frozen vision backbone over a screenshot crop of the
67
+ element, followed by a small trainable projection. This checkpoint's config
68
+ selects `siglip` (`google/siglip-base-patch16-224`).
69
+
70
+ Scoring is one parallel forward pass, not autoregressive generation. Latency
71
+ is in the low milliseconds per task on a GPU and under 100 ms on a CPU.
72
+
73
+ ## Training summary
74
+
75
+ The model was trained from scratch; it has no base model. It was trained on
76
+ the `data/v1` split of the `cua-bench-s1` core GUI task families:
77
+ `form_filling`, `login_auth`, `consent_checkbox`, `multi_step_submit`,
78
+ `pagination`, and `search_filter`. `cua-bench-s1` builds these families with
79
+ its synthetic generator and its converters for
80
+ [AndroidControl](https://github.com/google-research/google-research/tree/master/android_control)
81
+ (Apache-2.0) and GUI-360 (MIT). For exact provenance, see
82
+ [Data sources and provenance](https://github.com/trycua/cua/tree/main/libs/cua-bench-s1#data-sources-and-provenance).
83
+ The two-stage training recipe (base training, then an optional fine-tune) is
84
+ [`libs/cua-s1/training/train_nano.py`](https://github.com/trycua/cua/blob/main/libs/cua-s1/training/train_nano.py).
85
+ No dataset is distributed with this checkpoint.
86
+
87
+ ## Evaluation
88
+
89
+ These are measured results from the `cua-bench-s1`
90
+ [Results section](https://github.com/trycua/cua/tree/main/libs/cua-bench-s1#results):
91
+
92
+ - Held-out cross-dataset text split: task-level accuracy of 0.000 to 0.286,
93
+ depending on family. The per-element text context carries no goal string.
94
+ - Same-distribution multimodal split (the split it was trained on): 1.000.
95
+ - `chess`, on the same 15-position subset used for every model: 0.000 task
96
+ accuracy and 0.227 element accuracy in both modalities.
97
+ - `game_control`: 0.000 task accuracy and 0.333 element accuracy. Neither
98
+ result survives chance correction.
99
+ - `safety_gate`: 0.000 zero-shot, and 1.000 after fine-tuning on that
100
+ family's own split.
101
+ - External `general_decision` benchmark: no result, because the model cannot
102
+ produce that benchmark's text-only decision shape.
103
+
104
+ The strong same-distribution result and weak cross-dataset result mean that
105
+ this checkpoint mostly reflects its training distribution. It is not a
106
+ general GUI decision model.
107
+
108
+ ## How to run
109
+
110
+ Pin the download to the revision the Cua-S1 documentation was verified
111
+ against. The weights are unchanged at later revisions of this repository
112
+ that change only documentation.
113
+
114
+ | Artifact | Revision |
115
+ | --- | --- |
116
+ | `cua-ai/cua-s1-nano-0.1` (weights) | `1f93fd0fdcbe33740334948f967dff9f6c8e9f34` |
117
+
118
+ From a checkout of [`trycua/cua`](https://github.com/trycua/cua):
119
+
120
+ ```bash
121
+ uv sync --project libs/cua-s1/python # text modality
122
+ uv sync --project libs/cua-s1/python --extra nano-vision # adds the multimodal backbone
123
+ libs/cua-s1/python/.venv/bin/hf download cua-ai/cua-s1-nano-0.1 \
124
+ --revision 1f93fd0fdcbe33740334948f967dff9f6c8e9f34 \
125
+ --local-dir "$HOME/cua-s1-models/cua-s1-nano-0.1"
126
+ ```
127
+
128
+ ```python
129
+ from pathlib import Path
130
+ from cua_s1.nano import load_nano_checkpoint, select_device
131
+
132
+ root = Path.home() / "cua-s1-models" / "cua-s1-nano-0.1"
133
+ # Validates the format, version, and SHA-256 tensor signature before returning.
134
+ model, collator, config = load_nano_checkpoint(root / "text", select_device("auto"))
135
+ ```
136
+
137
+ The multimodal checkpoint loads the same way from `root / "multimodal"`. On
138
+ first use, it downloads the frozen `google/siglip-base-patch16-224` backbone
139
+ through Transformers.
140
+
141
+ ### jev-use chooser
142
+
143
+ The [jev-use closed-candidate chooser](https://github.com/trycua/cua/blob/main/libs/cua-driver/examples/jev-use/decision-models.md)
144
+ supports only the Cua-S1-4B adapters. It fails setup for `cua-s1-nano-0.1`,
145
+ and no Cua Driver integration runs this checkpoint yet. To drive the jev-use
146
+ flow with a Cua-S1 model, use
147
+ [`cua-s1-4b-0.2`](https://huggingface.co/cua-ai/cua-s1-4b-0.2) and follow
148
+ [Get the weights and run inference](https://github.com/trycua/cua/tree/main/libs/cua-s1#get-the-weights-and-run-inference).
149
+ Use this checkpoint through `cua_s1.nano` directly, or through
150
+ `cua-bench-s1`'s model adapters.
151
+
152
+ ## Limitations
153
+
154
+ - It needs a closed, pre-enumerated option set for each element and cannot
155
+ propose an action outside that set.
156
+ - It scores each element independently, so it cannot compare which element
157
+ to act on next across a screen.
158
+ - Its small size trades capacity for speed.
159
+ - Multimodal behavior inherits the frozen backbone's limitations and depends
160
+ on screenshot-crop quality.
161
+ - `game_control` and `chess` are out-of-domain checks only.
162
+ - No weights-backed CI or canonical Cua Driver desktop E2E covers this
163
+ checkpoint.
164
+
165
+ Run it in an isolated environment with least-privilege credentials and
166
+ bounded actions. Verify outcomes independently, and require human
167
+ confirmation before consequential or irreversible actions.
168
+
169
+ ## License
170
+
171
+ The weights in this repository are licensed under Apache-2.0. The frozen
172
+ `google/siglip-base-patch16-224` backbone used by the multimodal checkpoint
173
+ is not redistributed here: its own license governs it. The Cua-S1 source code
174
+ is MIT-licensed in [`trycua/cua`](https://github.com/trycua/cua).