one-pass-sv-forms / README.md
maglun's picture
Upload README.md with huggingface_hub
1b5f74c verified
|
Raw History Blame
6.93 kB
metadata
license: mit
language:
  - sv
library_name: coreml
pipeline_tag: text-classification
tags:
  - jev
  - system-one
  - one-pass-scorer
  - computer-use
  - form-filling
  - swedish
  - coreml
  - apple-silicon
base_model: []

One-Pass SV-Forms (SV0) — a Swedish one-pass form specialist (research checkpoint)

This is a research checkpoint, not a product. It is a 706 048-parameter model that scores a supplied list of options in one forward pass — the "Jev"/System One shape — trained on synthetic Swedish forms by a synthetic-data recipe we are publishing alongside it. It generates no text, needs no tokenizer, and runs in ~1 ms per decision on an Apple laptop or phone (Core ML package: 1.4 MB fp16, 787 KB int8).

Task given a UI element (role, label, state) plus the entities extracted from a document, pick one option: fyll <entity> / kryssa / klicka / hoppa över
Architecture byte-level embeddings + 2-layer Transformer encoder (width 128, 4 heads) + option-attention head; tinyx
Parameters 706 048 (2.83 MB fp16 checkpoint)
Context / option budget 224 bytes of context, 96 bytes per option, up to 40 options
Training data 900 synthetic Swedish form episodes (21 306 decisions), generated locally; no real form, person or customer data
Licence MIT (weights and code)
Lineage independent implementation of the same contract as Cua's CUA-S1 and jevlike; trained from scratch, not a fine-tune of either

Measured results (our harness, Apple M4)

Run decisions top-1 majority-class baseline ECE silently skipped a required fill
Held-out synthetic test (form-signature disjoint) 2 315 83.02 % 50.45 % 0.017 76 (7.5 % of fills)
Hand-written out-of-distribution Swedish demo 50 86.00 % 64.00 % 0.092 3 (of 32 fills)
Shuffled-context control (test) 2 315 34.08 % 50.45 % 0.497 264

Per-action accuracy on the test split: check 98.3 %, click 93.7 %, skip 88.0 %, fill 75.5 %. The shuffled-context control is the interesting one: rotating contexts between rows drops the model below the majority baseline, i.e. it really reads the element label and the document rather than exploiting option statistics.

For comparison, the released English checkpoint of the same family scores 21.4 % on exactly these Swedish rows — below the majority baseline, at its own shuffled-control floor. The language was the barrier, not the contract.

Core ML export (same checkpoint, converted with coremltools 9.0):

Variant package top-1 argmax parity vs PyTorch median latency p95
fp16 (CPU + ANE) 1.4 MB 83.00 % 0.99870 1.31 ms 1.46 ms
int8 (CPU + ANE) 787 KB 83.09 % 0.99611 1.32 ms 1.40 ms
int4 (CPU + ANE) 481 KB 50.87 % 0.48702 1.97 ms 2.10 ms
PyTorch reference 2.83 MB 83.02 % — — —

We publish the int4 result because it is the most useful thing in this table: the same 4-bit palettisation costs 0.06 pp on a converged model and destroys this one, and it also fails ANE compilation. Quantisation headroom is a property of the training run, not of the architecture. Use int8. int4-weights/ ships the collapsed variant for reproducibility only.

What it is NOT

  • Not trained or tested on real Swedish forms. The corpus is synthetic (Label: value document entities, .invalid e-mail domains, fictional names, locally generated personnummer-shaped strings with no link to any real person) and the only non-synthetic evidence is a 50-decision set we wrote ourselves.
  • Not a general-purpose assistant or an autonomous agent. It does not generate text, cannot invent a value it was not given, and does not decide execution order.
  • Not finished. Validation was still improving when the run stopped (48 % → 63 % → 79 % → 83 % top-1 over four epochs), so treat 83 % as a floor for this recipe, not a ceiling.
  • Not safe to run unsupervised on real data. 7.5 % of required fills are answered "skip" — a required field that silently stays empty. Any real integration must verify outcomes outside the model (fail-closed execution, dry run, one submit, human review before consequential actions) — see the runtime contract in Cua's planner.py for a good pattern.

Intended use

Research on bounded, high-volume Swedish interface workflows where the option set is supplied by deterministic code: filling forms from an extracted document, triaging/routing where the choices are known in advance, or as a criteria-decision layer where a calibrated distribution per question is more useful than generated prose. Also as a worked example of how to build and measure a language-local one-pass specialist.

Usage

from pathlib import Path
from huggingface_hub import hf_hub_download

# the checkpoint format is <name>.safetensors + <name>.json (architecture + hashes)
weights = Path(hf_hub_download("precisit/one-pass-sv-forms", "sv0-forms.safetensors"))
hf_hub_download("precisit/one-pass-sv-forms", "sv0-forms.json", local_dir=weights.parent)
# then either use the Core ML packages (no Python runtime needed) or the PyTorch code in
# the toolkit repository: https://github.com/precisit/one-pass-specialists — the Core ML
# packages below live in *this* repository and need no PyTorch.

The context string and options must be built exactly as in training (byte ids = UTF-8 byte + 1, zero-padded; UPPGIFT fyll i formuläret från dokumentet och skicka sedan in / FORM <title> / ELEMENT <role> "<label>" value="…"). A mismatch there is the most likely cause of poor output — the recipe and the generator are in the toolkit repository.

Building your own

The toolkit that produced this checkpoint is public: precisit/one-pass-specialists. The interesting work is the catalogue — the concepts, the labels they appear under, and their value formats — because everything above the trainer is language- and vertical-neutral. The recipe, the corpus manifest and both result files for this checkpoint are in that repository under examples/sv-forms/.

Licence and attribution

MIT. The architecture, training loop and evaluation metrics come from Cua's MIT-licensed libs/cua-s1 (see THIRD_PARTY_NOTICES.md); the option-attention head design is credited there to jevlike (MIT). The Swedish catalogue, synthetic generator, training run, measurements and Core ML export are ours — Precisit AB, 2026. No TypeSafe AI code, weights or data was used; "Jev" is their product and this is an independent implementation of a similar interface.