one-pass-sv-forms / README.md
maglun's picture
v2: corrected generator, 10 000-episode checkpoint, changelog, Core ML packages
3f1f3aa verified
|
Raw History Blame
8.94 kB
metadata
license: mit
language:
  - sv
library_name: coreml
pipeline_tag: text-classification
tags:
  - jev
  - system-one
  - one-pass-scorer
  - computer-use
  - form-filling
  - swedish
  - coreml
  - apple-silicon
base_model: []
datasets:
  - precisit/one-pass-sv-forms-synthetic

One-Pass SV-Forms (SV0) — a Swedish one-pass form specialist (research checkpoint)

This is a research checkpoint, not a product. It is a 706 048-parameter model that scores a supplied list of options in one forward pass — the "Jev"/System One shape — trained from scratch on synthetic Swedish forms. It generates no text, needs no tokenizer, and answers in about a millisecond on an Apple device (Core ML: 1.4 MB fp16, 787 KB int8).

Task given a UI element (role, label, state) plus the entities extracted from a document, pick one option: fyll <entitet> / kryssa / klicka / hoppa över
Architecture byte-level embeddings + 2-layer Transformer encoder (width 128, 4 heads) + option-attention head; tinyx
Parameters 706 048
Context / option budget 224 bytes of context, 96 bytes per option, up to 40 options
Training data 10 000 synthetic Swedish form episodes (234 921 decisions); no real form, person or customer data
Licence MIT (weights and code)
Lineage independent implementation of the same interface as Cua's CUA-S1 and jevlike; trained from scratch, not a fine-tune of either

Measured results (our harness, Apple M4)

Run decisions top-1 majority-class baseline ECE silently skipped a required fill
Held-out synthetic test (form-signature disjoint) 2 310 99.29 % 47.36 % 0.0014 1
Hand-written out-of-distribution demo 50 100.00 % 64.00 % — 0
Shuffled-context control (same test rows) 2 310 31.99 % 47.36 % — —

The shuffled-context control is the number to look at first: rotating contexts between rows drops the model to 31.99 %, well below the majority baseline, so it is reading the element label and the document rather than exploiting option statistics. Without that control a 100 % demo set means nothing.

For comparison, the released English checkpoint of the same family scores 21.4 % on these Swedish rows — below the majority baseline, at its own shuffled-control floor. The language was the barrier, not the interface.

Data-size curve (same 4-epoch schedule, same seed, corrected generator)

Episodes Train decisions Test top-1 Silent skips ECE Training time (M4)
900 21 305 77.10 % 127 0.0216 32 min
4 000 94 405 98.94 % 6 0.0014 56 min
10 000 234 921 99.29 % 1 0.0014 163 min

This checkpoint is the 10 000-episode point. The curve is published because it is the useful part: at 900 episodes the same pipeline sits at 77 %, so a single small run would have been misread as a statement about the method rather than about the data.

Changelog

  • v2 (2026-09-21) — this release. Two fixes and a scale-up, all three visible above:
    1. Corpus correctness. The Swedish catalogue's personnummer generator emitted nine digits formatted DDDDDDDD-DD and organisationsnummer emitted DDDDD-DDDDD; both are now the real Swedish shape YYMMDD-XXXX. As a side effect the models can no longer identify those two concepts from the shape of the value alone — the bug had been an accidental giveaway — so per-action accuracy on fill is the honest kind now.
    2. Scale. 900 → 10 000 episodes, which moved top-1 from 77 % to 99.29 % and silent skips from 127 to 1.
    3. Correction of our own first explanation. An intermediate 900-episode run on the corrected corpus scored 83.02 % (v1) versus 77.10 % (corrected), and we initially attributed that six-point drop to the removed giveaway. The 4 000- and 10 000-episode points show the dominant factor was data size, not corpus correctness. The measurement is recorded here because we published the wrong explanation for part of a working day.
  • v1 (2026-09-21) — first release: 900 synthetic episodes, 83.02 % top-1, 7.5 % silent skips, 787 KB int8. Superseded; kept in the history of this repository rather than hidden.

What it is NOT

  • Not trained or tested on real Swedish forms. The corpus is synthetic by construction: fictional names, .invalid e-mail domains, generated personnummer-shaped identifiers that pass their checksum but belong to nobody. The only non-synthetic evidence is a 50-decision set we wrote ourselves. A small set built from two of our own shipped forms (Kanslist, Pratsam) is the next measurement, not a claim in this card.
  • Not an agent. It does not generate text, cannot invent a value it was not given, and does not decide execution order — the option list and the sequence come from ordinary code around it.
  • Not safe unsupervised. One required fill out of 29 839 was still answered "skip": a silent failure. Any real integration must check outcomes outside the model (fail closed, dry run, one submit, human review before consequential actions).
  • Not tuned for throughput. One decision per predict call is what the latency numbers describe; batching is untested.

Core ML export (same checkpoint, coremltools 9.0, Apple M4)

Variant package top-1 argmax parity vs PyTorch median latency p95
fp16, CPU + ANE 1.44 MB 99.29 % 0.99950 (15 / 29 839) 1.314 ms 1.412 ms
fp16, CPU only 1.44 MB 99.28 % — 1.727 ms 1.929 ms
int8, CPU + ANE 787 KB 99.29 % 0.99956 (13 / 29 839) 1.314 ms 1.411 ms
int4, CPU + ANE 481 KB 49.92 % 0.49675 (15 015 / 29 839) 1.984 ms 2.115 ms

int8 is the variant to use; int4 does not work for this checkpoint and the collapsed variant is published as evidence rather than omitted. Two honest notes about it:

  • On the converged English checkpoint of this family, the same 4-bit palettisation cost 14 decisions in 24 367. On both of ours it destroys the model (83 % → 51 % on v1, 99.29 % → 49.92 % here) and the 4-bit graph does not compile for the Neural Engine (ANECCompile() FAILED). Whatever the difference is, it is not training level: we first assumed quantisation headroom tracked convergence, and the well-trained checkpoint falsifies that. We do not know the exact quantisation recipe behind the English release, so we report the contradiction instead of explaining it away.
  • fp16 parity is 99.95 %, not 100 %, on 29 839 rows. Measure parity per checkpoint; do not inherit someone else's 1.000000.

Usage

from pathlib import Path
from huggingface_hub import hf_hub_download

weights = Path(hf_hub_download("precisit/one-pass-sv-forms", "sv0-forms.safetensors"))
hf_hub_download("precisit/one-pass-sv-forms", "sv0-forms.json", local_dir=weights.parent)
# then either use the Core ML packages (no Python runtime needed, see example.py) or the PyTorch
# code and the full recipe in the toolkit repository:
# https://github.com/precisit/one-pass-specialists

The context string and options must be built exactly as in training (byte ids = UTF-8 byte + 1, zero padded; UPPGIFT fyll i formuläret från dokumentet och skicka sedan in / FORM <titel> / ELEMENT <roll> "<etikett>" value="…"). A mismatch there is the most likely cause of poor output — the generator, the catalogue and the scoring harness are in the toolkit repository.

Building your own

precisit/one-pass-specialists is the public toolkit that produced this checkpoint: catalogue, synthetic generator with form-disjoint splits, trainer, evaluation (including the shuffled-context control and the silent-skip count), Core ML export with a bit-identity proof before anything is written, and the SV0 recipe with all receipts. The synthetic corpus itself is published as precisit/one-pass-sv-forms-synthetic.

Licence and attribution

MIT. The architecture, training loop and evaluation metrics are vendored unmodified from Cua's MIT-licensed libs/cua-s1 (which credits jevlike for the option-attention head design). The Swedish catalogue, synthetic generator, training runs, measurements and Core ML tooling are Precisit's. No TypeSafe AI code, weights, data or API output was used — "Jev" and "System One" are their names for a similar interface, and this is an independent implementation. See THIRD_PARTY_NOTICES.md.