--- license: mit language: - sv library_name: coreml pipeline_tag: text-classification tags: - jev - system-one - one-pass-scorer - computer-use - form-filling - swedish - coreml - apple-silicon base_model: [] --- # One-Pass SV-Forms (SV0) — a Swedish one-pass form specialist (research checkpoint) **This is a research checkpoint, not a product.** It is a 706 048-parameter model that scores a supplied list of options in **one forward pass** — the "Jev"/System One shape — trained on **synthetic** Swedish forms by a synthetic-data recipe we are publishing alongside it. It generates no text, needs no tokenizer, and runs in ~1 ms per decision on an Apple laptop or phone (Core ML package: 1.4 MB fp16, 787 KB int8). | | | | --- | --- | | Task | given a UI element (role, label, state) plus the entities extracted from a document, pick one option: `fyll ` / `kryssa` / `klicka` / `hoppa över` | | Architecture | byte-level embeddings + 2-layer Transformer encoder (width 128, 4 heads) + option-attention head; `tinyx` | | Parameters | 706 048 (2.83 MB fp16 checkpoint) | | Context / option budget | 224 bytes of context, 96 bytes per option, **up to 40 options** | | Training data | 900 synthetic Swedish form episodes (21 306 decisions), generated locally; **no real form, person or customer data** | | Licence | MIT (weights and code) | | Lineage | independent implementation of the same contract as [Cua's CUA-S1](https://huggingface.co/cua-ai/cua-s1-forms) and [jevlike](https://github.com/vinnylarouge/jevlike); trained from scratch, not a fine-tune of either | ## Measured results (our harness, Apple M4) | Run | decisions | top-1 | majority-class baseline | ECE | silently skipped a required fill | | --- | ---: | ---: | ---: | ---: | ---: | | Held-out synthetic test (form-signature disjoint) | 2 315 | **83.02 %** | 50.45 % | 0.017 | **76 (7.5 % of fills)** | | Hand-written out-of-distribution Swedish demo | 50 | **86.00 %** | 64.00 % | 0.092 | 3 (of 32 fills) | | Shuffled-context control (test) | 2 315 | 34.08 % | 50.45 % | 0.497 | 264 | Per-action accuracy on the test split: `check` 98.3 %, `click` 93.7 %, `skip` 88.0 %, `fill` 75.5 %. The shuffled-context control is the interesting one: rotating contexts between rows drops the model below the majority baseline, i.e. it really reads the element label and the document rather than exploiting option statistics. For comparison, the released English checkpoint of the same family scores **21.4 %** on exactly these Swedish rows — below the majority baseline, at its own shuffled-control floor. The language was the barrier, not the contract. Core ML export (same checkpoint, converted with coremltools 9.0): | Variant | package | top-1 | argmax parity vs PyTorch | median latency | p95 | | --- | ---: | ---: | ---: | ---: | ---: | | fp16 (CPU + ANE) | 1.4 MB | 83.00 % | 0.99870 | 1.31 ms | 1.46 ms | | **int8 (CPU + ANE)** | **787 KB** | **83.09 %** | 0.99611 | 1.32 ms | 1.40 ms | | int4 (CPU + ANE) | 481 KB | 50.87 % | 0.48702 | 1.97 ms | 2.10 ms | | PyTorch reference | 2.83 MB | 83.02 % | — | — | — | We publish the int4 result because it is the most useful thing in this table: the same 4-bit palettisation costs 0.06 pp on a converged model and destroys this one, and it also fails ANE compilation. **Quantisation headroom is a property of the training run, not of the architecture.** Use int8. `int4-weights/` ships the collapsed variant for reproducibility only. ## What it is NOT - **Not trained or tested on real Swedish forms.** The corpus is synthetic (`Label: value` document entities, `.invalid` e-mail domains, fictional names, locally generated personnummer-shaped strings with no link to any real person) and the only non-synthetic evidence is a 50-decision set we wrote ourselves. - **Not a general-purpose assistant or an autonomous agent.** It does not generate text, cannot invent a value it was not given, and does not decide execution order. - **Not finished.** Validation was still improving when the run stopped (48 % → 63 % → 79 % → 83 % top-1 over four epochs), so treat 83 % as a floor for this recipe, not a ceiling. - **Not safe to run unsupervised on real data.** 7.5 % of required fills are answered "skip" — a required field that silently stays empty. Any real integration must verify outcomes outside the model (fail-closed execution, dry run, one submit, human review before consequential actions) — see the runtime contract in Cua's `planner.py` for a good pattern. ## Intended use Research on bounded, high-volume Swedish interface workflows where the option set is supplied by deterministic code: filling forms from an extracted document, triaging/routing where the choices are known in advance, or as a **criteria-decision** layer where a calibrated distribution per question is more useful than generated prose. Also as a worked example of how to build and measure a language-local one-pass specialist. ## Usage ```python from pathlib import Path from huggingface_hub import hf_hub_download # the checkpoint format is .safetensors + .json (architecture + hashes) weights = Path(hf_hub_download("precisit/one-pass-sv-forms", "sv0-forms.safetensors")) hf_hub_download("precisit/one-pass-sv-forms", "sv0-forms.json", local_dir=weights.parent) # then either use the Core ML packages (no Python runtime needed) or the PyTorch code in # the toolkit repository: https://github.com/precisit/one-pass-specialists — the Core ML # packages below live in *this* repository and need no PyTorch. ``` The context string and options must be built exactly as in training (byte ids = UTF-8 byte + 1, zero-padded; `UPPGIFT fyll i formuläret från dokumentet och skicka sedan in` / `FORM ` / `ELEMENT <role> "<label>" value="…"`). A mismatch there is the most likely cause of poor output — the recipe and the generator are in the toolkit repository. ## Building your own The toolkit that produced this checkpoint is public: [`precisit/one-pass-specialists`](https://github.com/precisit/one-pass-specialists). The interesting work is the **catalogue** — the concepts, the labels they appear under, and their value formats — because everything above the trainer is language- and vertical-neutral. The recipe, the corpus manifest and both result files for this checkpoint are in that repository under `examples/sv-forms/`. ## Licence and attribution MIT. The architecture, training loop and evaluation metrics come from Cua's MIT-licensed `libs/cua-s1` (see `THIRD_PARTY_NOTICES.md`); the option-attention head design is credited there to `jevlike` (MIT). The Swedish catalogue, synthetic generator, training run, measurements and Core ML export are ours — **Precisit AB, 2026**. No TypeSafe AI code, weights or data was used; "Jev" is their product and this is an independent implementation of a similar interface.