File size: 6,925 Bytes
c8d55fb 1b5f74c c8d55fb 1b5f74c c8d55fb | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 | ---
license: mit
language:
- sv
library_name: coreml
pipeline_tag: text-classification
tags:
- jev
- system-one
- one-pass-scorer
- computer-use
- form-filling
- swedish
- coreml
- apple-silicon
base_model: []
---
# One-Pass SV-Forms (SV0) — a Swedish one-pass form specialist (research checkpoint)
**This is a research checkpoint, not a product.** It is a 706 048-parameter model that scores a
supplied list of options in **one forward pass** — the "Jev"/System One shape — trained on
**synthetic** Swedish forms by a synthetic-data recipe we are publishing alongside it. It
generates no text, needs no tokenizer, and runs in ~1 ms per decision on an Apple laptop or
phone (Core ML package: 1.4 MB fp16, 787 KB int8).
| | |
| --- | --- |
| Task | given a UI element (role, label, state) plus the entities extracted from a document, pick one option: `fyll <entity>` / `kryssa` / `klicka` / `hoppa över` |
| Architecture | byte-level embeddings + 2-layer Transformer encoder (width 128, 4 heads) + option-attention head; `tinyx` |
| Parameters | 706 048 (2.83 MB fp16 checkpoint) |
| Context / option budget | 224 bytes of context, 96 bytes per option, **up to 40 options** |
| Training data | 900 synthetic Swedish form episodes (21 306 decisions), generated locally; **no real form, person or customer data** |
| Licence | MIT (weights and code) |
| Lineage | independent implementation of the same contract as [Cua's CUA-S1](https://huggingface.co/cua-ai/cua-s1-forms) and [jevlike](https://github.com/vinnylarouge/jevlike); trained from scratch, not a fine-tune of either |
## Measured results (our harness, Apple M4)
| Run | decisions | top-1 | majority-class baseline | ECE | silently skipped a required fill |
| --- | ---: | ---: | ---: | ---: | ---: |
| Held-out synthetic test (form-signature disjoint) | 2 315 | **83.02 %** | 50.45 % | 0.017 | **76 (7.5 % of fills)** |
| Hand-written out-of-distribution Swedish demo | 50 | **86.00 %** | 64.00 % | 0.092 | 3 (of 32 fills) |
| Shuffled-context control (test) | 2 315 | 34.08 % | 50.45 % | 0.497 | 264 |
Per-action accuracy on the test split: `check` 98.3 %, `click` 93.7 %, `skip` 88.0 %,
`fill` 75.5 %. The shuffled-context control is the interesting one: rotating contexts between
rows drops the model below the majority baseline, i.e. it really reads the element label and the
document rather than exploiting option statistics.
For comparison, the released English checkpoint of the same family scores **21.4 %** on exactly
these Swedish rows — below the majority baseline, at its own shuffled-control floor. The
language was the barrier, not the contract.
Core ML export (same checkpoint, converted with coremltools 9.0):
| Variant | package | top-1 | argmax parity vs PyTorch | median latency | p95 |
| --- | ---: | ---: | ---: | ---: | ---: |
| fp16 (CPU + ANE) | 1.4 MB | 83.00 % | 0.99870 | 1.31 ms | 1.46 ms |
| **int8 (CPU + ANE)** | **787 KB** | **83.09 %** | 0.99611 | 1.32 ms | 1.40 ms |
| int4 (CPU + ANE) | 481 KB | 50.87 % | 0.48702 | 1.97 ms | 2.10 ms |
| PyTorch reference | 2.83 MB | 83.02 % | — | — | — |
We publish the int4 result because it is the most useful thing in this table: the same 4-bit
palettisation costs 0.06 pp on a converged model and destroys this one, and it also fails ANE
compilation. **Quantisation headroom is a property of the training run, not of the architecture.**
Use int8. `int4-weights/` ships the collapsed variant for reproducibility only.
## What it is NOT
- **Not trained or tested on real Swedish forms.** The corpus is synthetic (`Label: value`
document entities, `.invalid` e-mail domains, fictional names, locally generated
personnummer-shaped strings with no link to any real person) and the only non-synthetic
evidence is a 50-decision set we wrote ourselves.
- **Not a general-purpose assistant or an autonomous agent.** It does not generate text, cannot
invent a value it was not given, and does not decide execution order.
- **Not finished.** Validation was still improving when the run stopped (48 % → 63 % → 79 % →
83 % top-1 over four epochs), so treat 83 % as a floor for this recipe, not a ceiling.
- **Not safe to run unsupervised on real data.** 7.5 % of required fills are answered "skip" —
a required field that silently stays empty. Any real integration must verify outcomes outside
the model (fail-closed execution, dry run, one submit, human review before consequential
actions) — see the runtime contract in Cua's `planner.py` for a good pattern.
## Intended use
Research on bounded, high-volume Swedish interface workflows where the option set is supplied by
deterministic code: filling forms from an extracted document, triaging/routing where the choices
are known in advance, or as a **criteria-decision** layer where a calibrated distribution per
question is more useful than generated prose. Also as a worked example of how to build and
measure a language-local one-pass specialist.
## Usage
```python
from pathlib import Path
from huggingface_hub import hf_hub_download
# the checkpoint format is <name>.safetensors + <name>.json (architecture + hashes)
weights = Path(hf_hub_download("precisit/one-pass-sv-forms", "sv0-forms.safetensors"))
hf_hub_download("precisit/one-pass-sv-forms", "sv0-forms.json", local_dir=weights.parent)
# then either use the Core ML packages (no Python runtime needed) or the PyTorch code in
# the toolkit repository: https://github.com/precisit/one-pass-specialists — the Core ML
# packages below live in *this* repository and need no PyTorch.
```
The context string and options must be built exactly as in training (byte ids = UTF-8 byte + 1,
zero-padded; `UPPGIFT fyll i formuläret från dokumentet och skicka sedan in` / `FORM <title>` /
`ELEMENT <role> "<label>" value="…"`). A mismatch there is the most likely cause of poor output —
the recipe and the generator are in the toolkit repository.
## Building your own
The toolkit that produced this checkpoint is public: [`precisit/one-pass-specialists`](https://github.com/precisit/one-pass-specialists). The interesting work is the **catalogue** — the concepts, the labels they appear under, and their value formats — because everything above the trainer is language- and vertical-neutral. The recipe, the corpus manifest and both result files for this checkpoint are in that repository under `examples/sv-forms/`.
## Licence and attribution
MIT. The architecture, training loop and evaluation metrics come from Cua's MIT-licensed
`libs/cua-s1` (see `THIRD_PARTY_NOTICES.md`); the option-attention head design is credited there
to `jevlike` (MIT). The Swedish catalogue, synthetic generator, training run, measurements and
Core ML export are ours — **Precisit AB, 2026**. No TypeSafe AI code, weights or data was used;
"Jev" is their product and this is an independent implementation of a similar interface.
|