--- license: mit language: - sv library_name: coreml pipeline_tag: text-classification tags: - jev - system-one - one-pass-scorer - computer-use - form-filling - swedish - coreml - apple-silicon base_model: [] datasets: - precisit/one-pass-sv-forms-synthetic --- # One-Pass SV-Forms (SV0) — a Swedish one-pass form specialist (research checkpoint) **This is a research checkpoint, not a product.** It is a 706 048-parameter model that scores a supplied list of options in **one forward pass** — the "Jev"/System One shape — trained from scratch on **synthetic** Swedish forms. It generates no text, needs no tokenizer, and answers in about a millisecond on an Apple device (Core ML: 1.4 MB fp16, 787 KB int8). | | | | --- | --- | | Task | given a UI element (role, label, state) plus the entities extracted from a document, pick one option: `fyll ` / `kryssa` / `klicka` / `hoppa över` | | Architecture | byte-level embeddings + 2-layer Transformer encoder (width 128, 4 heads) + option-attention head; `tinyx` | | Parameters | 706 048 | | Context / option budget | 224 bytes of context, 96 bytes per option, **up to 40 options** | | Training data | 10 000 synthetic Swedish form episodes (234 921 decisions); **no real form, person or customer data** | | Licence | MIT (weights and code) | | Lineage | independent implementation of the same interface as [Cua's CUA-S1](https://huggingface.co/cua-ai/cua-s1-forms) and [jevlike](https://github.com/vinnylarouge/jevlike); trained from scratch, not a fine-tune of either | ## Measured results (our harness, Apple M4) | Run | decisions | top-1 | majority-class baseline | ECE | silently skipped a required fill | | --- | ---: | ---: | ---: | ---: | ---: | | Held-out synthetic test (form-signature disjoint) | 2 310 | **99.29 %** | 47.36 % | 0.0014 | **1** | | Hand-written out-of-distribution demo | 50 | **100.00 %** | 64.00 % | — | 0 | | Shuffled-context control (same test rows) | 2 310 | 31.99 % | 47.36 % | — | — | The shuffled-context control is the number to look at first: rotating contexts between rows drops the model to 31.99 %, well below the majority baseline, so it is reading the element label and the document rather than exploiting option statistics. Without that control a 100 % demo set means nothing. For comparison, the released English checkpoint of the same family scores **21.4 %** on these Swedish rows — below the majority baseline, at its own shuffled-control floor. The language was the barrier, not the interface. ### Data-size curve (same 4-epoch schedule, same seed, corrected generator) | Episodes | Train decisions | Test top-1 | Silent skips | ECE | Training time (M4) | | ---: | ---: | ---: | ---: | ---: | ---: | | 900 | 21 305 | 77.10 % | 127 | 0.0216 | 32 min | | 4 000 | 94 405 | 98.94 % | 6 | 0.0014 | 56 min | | **10 000** | **234 921** | **99.29 %** | **1** | 0.0014 | 163 min | This checkpoint is the 10 000-episode point. The curve is published because it is the useful part: at 900 episodes the same pipeline sits at 77 %, so a single small run would have been misread as a statement about the method rather than about the data. ## Changelog - **v2 (2026-09-21)** — this release. Two fixes and a scale-up, all three visible above: 1. **Corpus correctness.** The Swedish catalogue's `personnummer` generator emitted nine digits formatted `DDDDDDDD-DD` and `organisationsnummer` emitted `DDDDD-DDDDD`; both are now the real Swedish shape `YYMMDD-XXXX`. As a side effect the models can no longer identify those two concepts from the *shape of the value* alone — the bug had been an accidental giveaway — so per-action accuracy on `fill` is the honest kind now. 2. **Scale.** 900 → 10 000 episodes, which moved top-1 from 77 % to 99.29 % and silent skips from 127 to 1. 3. **Correction of our own first explanation.** An intermediate 900-episode run on the corrected corpus scored 83.02 % (v1) versus 77.10 % (corrected), and we initially attributed that six-point drop to the removed giveaway. The 4 000- and 10 000-episode points show the dominant factor was data size, not corpus correctness. The measurement is recorded here because we published the wrong explanation for part of a working day. - **v1 (2026-09-21)** — first release: 900 synthetic episodes, 83.02 % top-1, 7.5 % silent skips, 787 KB int8. Superseded; kept in the history of this repository rather than hidden. ## What it is NOT - **Not trained or tested on real Swedish forms.** The corpus is synthetic by construction: fictional names, `.invalid` e-mail domains, generated personnummer-shaped identifiers that pass their checksum but belong to nobody. The only non-synthetic evidence is a 50-decision set we wrote ourselves. A small set built from two of our own shipped forms (Kanslist, Pratsam) is the next measurement, not a claim in this card. - **Not an agent.** It does not generate text, cannot invent a value it was not given, and does not decide execution order — the option list and the sequence come from ordinary code around it. - **Not safe unsupervised.** One required fill out of 29 839 was still answered "skip": a silent failure. Any real integration must check outcomes outside the model (fail closed, dry run, one submit, human review before consequential actions). - **Not tuned for throughput.** One decision per `predict` call is what the latency numbers describe; batching is untested. ## Core ML export (same checkpoint, coremltools 9.0, Apple M4) | Variant | package | top-1 | argmax parity vs PyTorch | median latency | p95 | | --- | ---: | ---: | ---: | ---: | ---: | | fp16, CPU + ANE | 1.44 MB | 99.29 % | 0.99950 (15 / 29 839) | 1.314 ms | 1.412 ms | | fp16, CPU only | 1.44 MB | 99.28 % | — | 1.727 ms | 1.929 ms | | **int8, CPU + ANE** | **787 KB** | **99.29 %** | 0.99956 (13 / 29 839) | 1.314 ms | 1.411 ms | | int4, CPU + ANE | 481 KB | **49.92 %** | 0.49675 (15 015 / 29 839) | 1.984 ms | 2.115 ms | **int8 is the variant to use; int4 does not work for this checkpoint** and the collapsed variant is published as evidence rather than omitted. Two honest notes about it: - On the *converged English* checkpoint of this family, the same 4-bit palettisation cost 14 decisions in 24 367. On both of ours it destroys the model (83 % → 51 % on v1, 99.29 % → 49.92 % here) and the 4-bit graph does not compile for the Neural Engine (`ANECCompile() FAILED`). Whatever the difference is, it is not training level: we first assumed quantisation headroom tracked convergence, and the well-trained checkpoint falsifies that. We do not know the exact quantisation recipe behind the English release, so we report the contradiction instead of explaining it away. - fp16 parity is 99.95 %, not 100 %, on 29 839 rows. Measure parity per checkpoint; do not inherit someone else's 1.000000. ## Usage ```python from pathlib import Path from huggingface_hub import hf_hub_download weights = Path(hf_hub_download("precisit/one-pass-sv-forms", "sv0-forms.safetensors")) hf_hub_download("precisit/one-pass-sv-forms", "sv0-forms.json", local_dir=weights.parent) # then either use the Core ML packages (no Python runtime needed, see example.py) or the PyTorch # code and the full recipe in the toolkit repository: # https://github.com/precisit/one-pass-specialists ``` The context string and options must be built exactly as in training (byte ids = UTF-8 byte + 1, zero padded; `UPPGIFT fyll i formuläret från dokumentet och skicka sedan in` / `FORM ` / `ELEMENT "" value="…"`). A mismatch there is the most likely cause of poor output — the generator, the catalogue and the scoring harness are in the toolkit repository. ## Building your own [`precisit/one-pass-specialists`](https://github.com/precisit/one-pass-specialists) is the public toolkit that produced this checkpoint: catalogue, synthetic generator with form-disjoint splits, trainer, evaluation (including the shuffled-context control and the silent-skip count), Core ML export with a bit-identity proof before anything is written, and the SV0 recipe with all receipts. The synthetic corpus itself is published as [`precisit/one-pass-sv-forms-synthetic`](https://huggingface.co/datasets/precisit/one-pass-sv-forms-synthetic). ## Licence and attribution MIT. The architecture, training loop and evaluation metrics are vendored unmodified from Cua's MIT-licensed [`libs/cua-s1`](https://github.com/trycua/cua/tree/main/libs/cua-s1) (which credits `jevlike` for the option-attention head design). The Swedish catalogue, synthetic generator, training runs, measurements and Core ML tooling are Precisit's. **No TypeSafe AI code, weights, data or API output was used** — "Jev" and "System One" are their names for a similar interface, and this is an independent implementation. See `THIRD_PARTY_NOTICES.md`.