|
Download README.md from precisit/one-pass-sv-forms: direct link, hf CLI and curl.
- Browser
- Download file 8.94 kB
-
https://huggingface.co/precisit/one-pass-sv-forms/resolve/adce86401213ba6ee8cfccadeb8bc4bf192092f2/README.md
- Command line
-
hf download hf://precisit/one-pass-sv-forms@adce86401213ba6ee8cfccadeb8bc4bf192092f2/README.md
-
curl -L -o README.md https://huggingface.co/precisit/one-pass-sv-forms/resolve/adce86401213ba6ee8cfccadeb8bc4bf192092f2/README.md
8.94 kB
| license: mit | |
| language: | |
| - sv | |
| library_name: coreml | |
| pipeline_tag: text-classification | |
| tags: | |
| - jev | |
| - system-one | |
| - one-pass-scorer | |
| - computer-use | |
| - form-filling | |
| - swedish | |
| - coreml | |
| - apple-silicon | |
| base_model: [] | |
| datasets: | |
| - precisit/one-pass-sv-forms-synthetic | |
| # One-Pass SV-Forms (SV0) — a Swedish one-pass form specialist (research checkpoint) | |
| **This is a research checkpoint, not a product.** It is a 706 048-parameter model that scores a | |
| supplied list of options in **one forward pass** — the "Jev"/System One shape — trained from | |
| scratch on **synthetic** Swedish forms. It generates no text, needs no tokenizer, and answers in | |
| about a millisecond on an Apple device (Core ML: 1.4 MB fp16, 787 KB int8). | |
| | | | | |
| | --- | --- | | |
| | Task | given a UI element (role, label, state) plus the entities extracted from a document, pick one option: `fyll <entitet>` / `kryssa` / `klicka` / `hoppa över` | | |
| | Architecture | byte-level embeddings + 2-layer Transformer encoder (width 128, 4 heads) + option-attention head; `tinyx` | | |
| | Parameters | 706 048 | | |
| | Context / option budget | 224 bytes of context, 96 bytes per option, **up to 40 options** | | |
| | Training data | 10 000 synthetic Swedish form episodes (234 921 decisions); **no real form, person or customer data** | | |
| | Licence | MIT (weights and code) | | |
| | Lineage | independent implementation of the same interface as [Cua's CUA-S1](https://huggingface.co/cua-ai/cua-s1-forms) and [jevlike](https://github.com/vinnylarouge/jevlike); trained from scratch, not a fine-tune of either | | |
| ## Measured results (our harness, Apple M4) | |
| | Run | decisions | top-1 | majority-class baseline | ECE | silently skipped a required fill | | |
| | --- | ---: | ---: | ---: | ---: | ---: | | |
| | Held-out synthetic test (form-signature disjoint) | 2 310 | **99.29 %** | 47.36 % | 0.0014 | **1** | | |
| | Hand-written out-of-distribution demo | 50 | **100.00 %** | 64.00 % | — | 0 | | |
| | Shuffled-context control (same test rows) | 2 310 | 31.99 % | 47.36 % | — | — | | |
| The shuffled-context control is the number to look at first: rotating contexts between rows drops | |
| the model to 31.99 %, well below the majority baseline, so it is reading the element label and the | |
| document rather than exploiting option statistics. Without that control a 100 % demo set means | |
| nothing. | |
| For comparison, the released English checkpoint of the same family scores **21.4 %** on these | |
| Swedish rows — below the majority baseline, at its own shuffled-control floor. The language was | |
| the barrier, not the interface. | |
| ### Data-size curve (same 4-epoch schedule, same seed, corrected generator) | |
| | Episodes | Train decisions | Test top-1 | Silent skips | ECE | Training time (M4) | | |
| | ---: | ---: | ---: | ---: | ---: | ---: | | |
| | 900 | 21 305 | 77.10 % | 127 | 0.0216 | 32 min | | |
| | 4 000 | 94 405 | 98.94 % | 6 | 0.0014 | 56 min | | |
| | **10 000** | **234 921** | **99.29 %** | **1** | 0.0014 | 163 min | | |
| This checkpoint is the 10 000-episode point. The curve is published because it is the useful | |
| part: at 900 episodes the same pipeline sits at 77 %, so a single small run would have been | |
| misread as a statement about the method rather than about the data. | |
| ## Changelog | |
| - **v2 (2026-09-21)** — this release. Two fixes and a scale-up, all three visible above: | |
| 1. **Corpus correctness.** The Swedish catalogue's `personnummer` generator emitted nine digits | |
| formatted `DDDDDDDD-DD` and `organisationsnummer` emitted `DDDDD-DDDDD`; both are now the | |
| real Swedish shape `YYMMDD-XXXX`. As a side effect the models can no longer identify those | |
| two concepts from the *shape of the value* alone — the bug had been an accidental | |
| giveaway — so per-action accuracy on `fill` is the honest kind now. | |
| 2. **Scale.** 900 → 10 000 episodes, which moved top-1 from 77 % to 99.29 % and silent skips | |
| from 127 to 1. | |
| 3. **Correction of our own first explanation.** An intermediate 900-episode run on the | |
| corrected corpus scored 83.02 % (v1) versus 77.10 % (corrected), and we initially attributed | |
| that six-point drop to the removed giveaway. The 4 000- and 10 000-episode points show the | |
| dominant factor was data size, not corpus correctness. The measurement is recorded here | |
| because we published the wrong explanation for part of a working day. | |
| - **v1 (2026-09-21)** — first release: 900 synthetic episodes, 83.02 % top-1, 7.5 % silent skips, | |
| 787 KB int8. Superseded; kept in the history of this repository rather than hidden. | |
| ## What it is NOT | |
| - **Not trained or tested on real Swedish forms.** The corpus is synthetic by construction: | |
| fictional names, `.invalid` e-mail domains, generated personnummer-shaped identifiers that pass | |
| their checksum but belong to nobody. The only non-synthetic evidence is a 50-decision set we | |
| wrote ourselves. A small set built from two of our own shipped forms (Kanslist, Pratsam) is the | |
| next measurement, not a claim in this card. | |
| - **Not an agent.** It does not generate text, cannot invent a value it was not given, and does | |
| not decide execution order — the option list and the sequence come from ordinary code around it. | |
| - **Not safe unsupervised.** One required fill out of 29 839 was still answered "skip": a silent | |
| failure. Any real integration must check outcomes outside the model (fail closed, dry run, one | |
| submit, human review before consequential actions). | |
| - **Not tuned for throughput.** One decision per `predict` call is what the latency numbers | |
| describe; batching is untested. | |
| ## Core ML export (same checkpoint, coremltools 9.0, Apple M4) | |
| | Variant | package | top-1 | argmax parity vs PyTorch | median latency | p95 | | |
| | --- | ---: | ---: | ---: | ---: | ---: | | |
| | fp16, CPU + ANE | 1.44 MB | 99.29 % | 0.99950 (15 / 29 839) | 1.314 ms | 1.412 ms | | |
| | fp16, CPU only | 1.44 MB | 99.28 % | — | 1.727 ms | 1.929 ms | | |
| | **int8, CPU + ANE** | **787 KB** | **99.29 %** | 0.99956 (13 / 29 839) | 1.314 ms | 1.411 ms | | |
| | int4, CPU + ANE | 481 KB | **49.92 %** | 0.49675 (15 015 / 29 839) | 1.984 ms | 2.115 ms | | |
| **int8 is the variant to use; int4 does not work for this checkpoint** and the collapsed variant is | |
| published as evidence rather than omitted. Two honest notes about it: | |
| - On the *converged English* checkpoint of this family, the same 4-bit palettisation cost 14 | |
| decisions in 24 367. On both of ours it destroys the model (83 % → 51 % on v1, 99.29 % → 49.92 % | |
| here) and the 4-bit graph does not compile for the Neural Engine (`ANECCompile() FAILED`). | |
| Whatever the difference is, it is not training level: we first assumed quantisation headroom | |
| tracked convergence, and the well-trained checkpoint falsifies that. We do not know the exact | |
| quantisation recipe behind the English release, so we report the contradiction instead of | |
| explaining it away. | |
| - fp16 parity is 99.95 %, not 100 %, on 29 839 rows. Measure parity per checkpoint; do not inherit | |
| someone else's 1.000000. | |
| ## Usage | |
| ```python | |
| from pathlib import Path | |
| from huggingface_hub import hf_hub_download | |
| weights = Path(hf_hub_download("precisit/one-pass-sv-forms", "sv0-forms.safetensors")) | |
| hf_hub_download("precisit/one-pass-sv-forms", "sv0-forms.json", local_dir=weights.parent) | |
| # then either use the Core ML packages (no Python runtime needed, see example.py) or the PyTorch | |
| # code and the full recipe in the toolkit repository: | |
| # https://github.com/precisit/one-pass-specialists | |
| ``` | |
| The context string and options must be built exactly as in training (byte ids = UTF-8 byte + 1, | |
| zero padded; `UPPGIFT fyll i formuläret från dokumentet och skicka sedan in` / `FORM <titel>` / | |
| `ELEMENT <roll> "<etikett>" value="…"`). A mismatch there is the most likely cause of poor output — | |
| the generator, the catalogue and the scoring harness are in the toolkit repository. | |
| ## Building your own | |
| [`precisit/one-pass-specialists`](https://github.com/precisit/one-pass-specialists) is the public | |
| toolkit that produced this checkpoint: catalogue, synthetic generator with form-disjoint splits, | |
| trainer, evaluation (including the shuffled-context control and the silent-skip count), Core ML | |
| export with a bit-identity proof before anything is written, and the SV0 recipe with all receipts. | |
| The synthetic corpus itself is published as | |
| [`precisit/one-pass-sv-forms-synthetic`](https://huggingface.co/datasets/precisit/one-pass-sv-forms-synthetic). | |
| ## Licence and attribution | |
| MIT. The architecture, training loop and evaluation metrics are vendored unmodified from Cua's | |
| MIT-licensed [`libs/cua-s1`](https://github.com/trycua/cua/tree/main/libs/cua-s1) (which credits | |
| `jevlike` for the option-attention head design). The Swedish catalogue, synthetic generator, | |
| training runs, measurements and Core ML tooling are Precisit's. **No TypeSafe AI code, weights, | |
| data or API output was used** — "Jev" and "System One" are their names for a similar interface, and | |
| this is an independent implementation. See `THIRD_PARTY_NOTICES.md`. | |