Duplicate from cua-ai/cua-s1-forms
Browse filesCo-authored-by: Dillon DuPont <ddupont@users.noreply.huggingface.co>
- .gitattributes +35 -0
- README.md +125 -0
- cua-s1-forms.json +44 -0
- cua-s1-forms.pt +3 -0
- cua-s1-forms.safetensors +3 -0
.gitattributes
ADDED
|
@@ -0,0 +1,35 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
*.7z filter=lfs diff=lfs merge=lfs -text
|
| 2 |
+
*.arrow filter=lfs diff=lfs merge=lfs -text
|
| 3 |
+
*.bin filter=lfs diff=lfs merge=lfs -text
|
| 4 |
+
*.bz2 filter=lfs diff=lfs merge=lfs -text
|
| 5 |
+
*.ckpt filter=lfs diff=lfs merge=lfs -text
|
| 6 |
+
*.ftz filter=lfs diff=lfs merge=lfs -text
|
| 7 |
+
*.gz filter=lfs diff=lfs merge=lfs -text
|
| 8 |
+
*.h5 filter=lfs diff=lfs merge=lfs -text
|
| 9 |
+
*.joblib filter=lfs diff=lfs merge=lfs -text
|
| 10 |
+
*.lfs.* filter=lfs diff=lfs merge=lfs -text
|
| 11 |
+
*.mlmodel filter=lfs diff=lfs merge=lfs -text
|
| 12 |
+
*.model filter=lfs diff=lfs merge=lfs -text
|
| 13 |
+
*.msgpack filter=lfs diff=lfs merge=lfs -text
|
| 14 |
+
*.npy filter=lfs diff=lfs merge=lfs -text
|
| 15 |
+
*.npz filter=lfs diff=lfs merge=lfs -text
|
| 16 |
+
*.onnx filter=lfs diff=lfs merge=lfs -text
|
| 17 |
+
*.ot filter=lfs diff=lfs merge=lfs -text
|
| 18 |
+
*.parquet filter=lfs diff=lfs merge=lfs -text
|
| 19 |
+
*.pb filter=lfs diff=lfs merge=lfs -text
|
| 20 |
+
*.pickle filter=lfs diff=lfs merge=lfs -text
|
| 21 |
+
*.pkl filter=lfs diff=lfs merge=lfs -text
|
| 22 |
+
*.pt filter=lfs diff=lfs merge=lfs -text
|
| 23 |
+
*.pth filter=lfs diff=lfs merge=lfs -text
|
| 24 |
+
*.rar filter=lfs diff=lfs merge=lfs -text
|
| 25 |
+
*.safetensors filter=lfs diff=lfs merge=lfs -text
|
| 26 |
+
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
| 27 |
+
*.tar.* filter=lfs diff=lfs merge=lfs -text
|
| 28 |
+
*.tar filter=lfs diff=lfs merge=lfs -text
|
| 29 |
+
*.tflite filter=lfs diff=lfs merge=lfs -text
|
| 30 |
+
*.tgz filter=lfs diff=lfs merge=lfs -text
|
| 31 |
+
*.wasm filter=lfs diff=lfs merge=lfs -text
|
| 32 |
+
*.xz filter=lfs diff=lfs merge=lfs -text
|
| 33 |
+
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
+
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
+
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
README.md
ADDED
|
@@ -0,0 +1,125 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: mit
|
| 3 |
+
tags:
|
| 4 |
+
- jev
|
| 5 |
+
- system-one
|
| 6 |
+
- computer-use
|
| 7 |
+
- form-filling
|
| 8 |
+
- option-attention
|
| 9 |
+
base_model: []
|
| 10 |
+
pipeline_tag: other
|
| 11 |
+
---
|
| 12 |
+
|
| 13 |
+
# cua-s1-forms
|
| 14 |
+
|
| 15 |
+
A small, jev-like ("System One") one-pass option scorer for GUI form filling, trained to
|
| 16 |
+
work as the decision layer behind [cua-driver](https://github.com/trycua/cua/tree/main/libs/cua-driver).
|
| 17 |
+
|
| 18 |
+
Unlike an autoregressive LLM, this model does not generate text. Given a UI element and a
|
| 19 |
+
list of typed options (one option per document entity, plus `check` / `click` / `skip`), it
|
| 20 |
+
returns one probability per option in a single forward pass — the same input/output
|
| 21 |
+
contract as TypeSafe's [Jev](https://typesafe.ai/blog/introducing-system-one-models-and-jev).
|
| 22 |
+
Every actionable element on a form is scored independently and in parallel in one batch;
|
| 23 |
+
execution order (fills, then checkboxes, then the one submit click) is decided by
|
| 24 |
+
downstream code, not the model.
|
| 25 |
+
|
| 26 |
+
Full writeup, training code, synthetic data generator and live Cua Driver integration:
|
| 27 |
+
https://github.com/trycua/cua/tree/main/libs/cua-s1.
|
| 28 |
+
|
| 29 |
+
## Architecture
|
| 30 |
+
|
| 31 |
+
- Byte-level embedding + 2-layer Transformer encoder (width 128, 4 heads) over the context
|
| 32 |
+
and, separately, over each option's text
|
| 33 |
+
- jevlike's `AttentionHead`: each option becomes a query against the context tokens,
|
| 34 |
+
producing an attended context vector, then a shared dot product turns each
|
| 35 |
+
(option, attended-context) pair into one logit; softmax over the live option count
|
| 36 |
+
- 706,048 trainable parameters, 2.8 MB checkpoint (`state_dict` + `config` + training
|
| 37 |
+
history + best validation metrics)
|
| 38 |
+
|
| 39 |
+
## Input / output
|
| 40 |
+
|
| 41 |
+
Context (one per element, byte-truncated to 224 bytes):
|
| 42 |
+
```
|
| 43 |
+
TASK fill the form from the document, then submit
|
| 44 |
+
FORM Northwind Clinic - New Patient Registration
|
| 45 |
+
ELEMENT Edit "Phone number" value=""
|
| 46 |
+
```
|
| 47 |
+
Options (one per document entity, plus the three fixed actions, byte-truncated to 96 bytes
|
| 48 |
+
each): `fill Tel: (503) 555-0142`, `fill DOB: 03/14/1987`, ..., `check`, `click`, `skip`.
|
| 49 |
+
|
| 50 |
+
Output: one probability per option. The executor picks the argmax, looks up the entity by
|
| 51 |
+
index if the action is `fill`, and orders the resulting actions before sending them to
|
| 52 |
+
cua-driver (`set_value` / `click`).
|
| 53 |
+
|
| 54 |
+
## Training
|
| 55 |
+
|
| 56 |
+
- 10,000 synthetic episodes (`cua_s1/synth.py`): random form (2–16 fields from a 55-concept
|
| 57 |
+
catalogue with form-label/document-label synonyms), random person, random document with
|
| 58 |
+
distractor entities and forced look-alike confuser pairs (e.g. `email` vs `street`,
|
| 59 |
+
`phone` vs `emergency contact phone`, `state` vs `university`), random window-title
|
| 60 |
+
suffixes and 20% title dropout
|
| 61 |
+
- Splits are disjoint by exact form field signature — a test form's field set never appears
|
| 62 |
+
in training
|
| 63 |
+
- AdamW, cosine schedule with warmup, 6 epochs, batch size 128, cross-entropy over the live
|
| 64 |
+
option count
|
| 65 |
+
|
| 66 |
+
## Results (see the repo's `docs/RESULTS.md` for the full ladder)
|
| 67 |
+
|
| 68 |
+
| split | top-1 | notes |
|
| 69 |
+
| --- | ---: | --- |
|
| 70 |
+
| synthetic test (form-disjoint, ~15k decisions) | 99.95% | hard confuser pairs forced in |
|
| 71 |
+
| real demo eval (3 real forms + 3 real PDFs, 196 decisions, nothing synthetic) | 100% | |
|
| 72 |
+
| shuffled-context control | 37% | confirms the model reads the element, not option statistics |
|
| 73 |
+
|
| 74 |
+
Head-to-head against the real hosted Jev API (`jev-latest`, zero fine-tuning, same task):
|
| 75 |
+
99.7% for this model vs 83.6% for hosted Jev overall; 96% for hosted Jev on decisions that
|
| 76 |
+
require real judgment (fill vs check vs click) and 74% on recognizing an already-filled
|
| 77 |
+
field as a no-op — a convention this model was trained on and hosted Jev was not. Full
|
| 78 |
+
numbers in the repo.
|
| 79 |
+
|
| 80 |
+
## Files
|
| 81 |
+
|
| 82 |
+
- `cua-s1-forms.safetensors` + `cua-s1-forms.json` — the checkpoint in the format
|
| 83 |
+
`cua_s1.checkpoint` expects: tensors only in safetensors, everything else (architecture
|
| 84 |
+
config, a SHA-256 signature over the tensors, free-form metadata) in a plain JSON
|
| 85 |
+
sidecar. This is the format to use; `cua_s1`'s own loader rejects pickled `.pt`/`.pth`
|
| 86 |
+
files by design (arbitrary pickle is a code-execution risk for a public checkpoint).
|
| 87 |
+
- `cua-s1-forms.pt` — the original PyTorch pickle checkpoint, kept only for anyone still
|
| 88 |
+
loading it directly with `torch.load(..., weights_only=False)` outside `cua_s1`. New code
|
| 89 |
+
should use the safetensors pair above.
|
| 90 |
+
|
| 91 |
+
Both encode the exact same weights; converted with a script that reimplements
|
| 92 |
+
`cua_s1.checkpoint.save_checkpoint_files`'s exact document/signature format, and verified to
|
| 93 |
+
produce bit-for-bit identical model output against the original `.pt`.
|
| 94 |
+
|
| 95 |
+
## Usage
|
| 96 |
+
|
| 97 |
+
```python
|
| 98 |
+
from pathlib import Path
|
| 99 |
+
from huggingface_hub import hf_hub_download
|
| 100 |
+
from cua_s1.model import load_checkpoint, select_device
|
| 101 |
+
|
| 102 |
+
repo = "cua-ai/cua-s1-forms"
|
| 103 |
+
weights = Path(hf_hub_download(repo, "cua-s1-forms.safetensors"))
|
| 104 |
+
hf_hub_download(repo, "cua-s1-forms.json", local_dir=weights.parent) # sits next to the weights
|
| 105 |
+
|
| 106 |
+
# validates format, version and SHA-256 tensor signature before returning
|
| 107 |
+
model, collator, config = load_checkpoint(weights, select_device("auto"))
|
| 108 |
+
```
|
| 109 |
+
|
| 110 |
+
See [`cua_s1/planner.py`](https://github.com/trycua/cua/tree/main/libs/cua-s1/python/src/cua_s1/planner.py)
|
| 111 |
+
for the full snapshot → score → order → execute loop against a live Cua Driver session.
|
| 112 |
+
|
| 113 |
+
## Limitations
|
| 114 |
+
|
| 115 |
+
- Only ever chooses among entities a PDF/document extractor already found as `Label: value`
|
| 116 |
+
pairs; it cannot invent a value.
|
| 117 |
+
- Trained entirely on synthetic forms plus a small (196-decision) real eval; not validated
|
| 118 |
+
on arbitrary real-world forms outside the demo set.
|
| 119 |
+
- Byte-level encoder, English-centric label vocabulary.
|
| 120 |
+
- Not calibrated with TypeSafe's RLCD method — this is an independent research checkpoint,
|
| 121 |
+
not a reproduction of Jev.
|
| 122 |
+
|
| 123 |
+
## License
|
| 124 |
+
|
| 125 |
+
MIT.
|
cua-s1-forms.json
ADDED
|
@@ -0,0 +1,44 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"config": {
|
| 3 |
+
"context_tokens": 224,
|
| 4 |
+
"encoder": "tinyx",
|
| 5 |
+
"heads": 4,
|
| 6 |
+
"hf_model": "Qwen/Qwen2.5-0.5B",
|
| 7 |
+
"layers": 2,
|
| 8 |
+
"option_tokens": 96,
|
| 9 |
+
"rank": 128,
|
| 10 |
+
"width": 128
|
| 11 |
+
},
|
| 12 |
+
"format": "cua-s1",
|
| 13 |
+
"format_version": 1,
|
| 14 |
+
"metadata": {
|
| 15 |
+
"best_validation": {
|
| 16 |
+
"ece": 0.00014795374386267213,
|
| 17 |
+
"examples": 22054,
|
| 18 |
+
"nll": 0.0019850827802381913,
|
| 19 |
+
"per_action": {
|
| 20 |
+
"check": {
|
| 21 |
+
"acc": 1.0,
|
| 22 |
+
"n": 835
|
| 23 |
+
},
|
| 24 |
+
"click": {
|
| 25 |
+
"acc": 1.0,
|
| 26 |
+
"n": 939
|
| 27 |
+
},
|
| 28 |
+
"fill": {
|
| 29 |
+
"acc": 0.9998858968507531,
|
| 30 |
+
"n": 8764
|
| 31 |
+
},
|
| 32 |
+
"skip": {
|
| 33 |
+
"acc": 0.9989579715178881,
|
| 34 |
+
"n": 11516
|
| 35 |
+
}
|
| 36 |
+
},
|
| 37 |
+
"rows_per_second": 2589.071654630599,
|
| 38 |
+
"top1": 0.9994105100631714
|
| 39 |
+
},
|
| 40 |
+
"converted_by": "scripts/convert_to_safetensors.py",
|
| 41 |
+
"source_checkpoint": "jevform-best.pt"
|
| 42 |
+
},
|
| 43 |
+
"state_signature": "0c661c669bc22189c81637553dcd105fe6178c8a8de85f6f6307c3c635403bfa"
|
| 44 |
+
}
|
cua-s1-forms.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:f5077f0c9baf6b5fc10f21512e1aa15207a395598416a6ffdd95f0d3dd5ab8df
|
| 3 |
+
size 2840436
|
cua-s1-forms.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:05954c1caf51c2fb6c13ea4acbfc88a2e7653dea192252bb51dc89e76a356ddc
|
| 3 |
+
size 2828784
|