Instructions to use sivasub987/iso20022-extract-53m with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use sivasub987/iso20022-extract-53m with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="sivasub987/iso20022-extract-53m", trust_remote_code=True)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("sivasub987/iso20022-extract-53m", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use sivasub987/iso20022-extract-53m with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "sivasub987/iso20022-extract-53m" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "sivasub987/iso20022-extract-53m", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/sivasub987/iso20022-extract-53m
- SGLang
How to use sivasub987/iso20022-extract-53m with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "sivasub987/iso20022-extract-53m" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "sivasub987/iso20022-extract-53m", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "sivasub987/iso20022-extract-53m" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "sivasub987/iso20022-extract-53m", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use sivasub987/iso20022-extract-53m with Docker Model Runner:
docker model run hf.co/sivasub987/iso20022-extract-53m
iso20022-extract-53m
A 53M-parameter decoder-only transformer that reads unstructured payment correspondence and emits structured ISO 20022 fields. It trains from scratch on a single consumer laptop GPU. No frontier model, no API call, and no document leaving the machine.
The structural pipeline is complete and verified: six of the nine acceptance checks pass at 100%, including valid JSON, canonical field names, deterministic decoding, schema-valid message assembly, and field-level routing agreement. Value accuracy on unseen values is the remaining work.
Getting here required finding four separate ways the metrics were lying, all of them in the model's favour. That part is worth reading even if you never touch this model, because the same failure mode is available to anyone measuring a specialist extractor.
Status: proof of concept. The gate does not pass yet. Field accuracy is 71.7% against a 90% target. Do not use this to move money.
The architecture, and why these two parts are novel
Most of this model is conventional and deliberately so: 20 pre-norm blocks, grouped-query attention, RoPE, SwiGLU, RMSNorm, tied embeddings. The pieces worth attention are the engram memory and the confidence head, because both exist to exploit a property of payment documents specifically rather than language generally.
flowchart TB
subgraph STACK["53M decoder · 20 blocks · d_model 512 · context 256"]
direction TB
IN["Token ids"] --> EMB["Embedding<br/>8192 × 512"]
EMB --> BLK["20 × transformer block"]
BLK --> FN["Final RMSNorm"]
end
subgraph ONE["Inside one block"]
direction TB
P["Pre-norm<br/>RMSNorm"] --> ATT["GQA attention<br/>8 query / 4 KV · RoPE θ=1e5"]
ATT --> R1["Residual add"]
R1 --> P2["Pre-norm<br/>RMSNorm"]
P2 --> FFN["SwiGLU FFN<br/>512 → 768"]
FFN --> R2["Residual add"]
end
ENG["Engram memory<br/>hashed n-gram read<br/>blocks 2 and 12"] -.-> BLK
FN --> LM["LM head<br/>tied to embedding"]
FN --> CH["Confidence head<br/>8 probes"]
LM --> OUT["Field JSON"]
CH --> ROUTE["Low-confidence<br/>routing"]
Engram: a memory read, not more attention
Token n-grams are hashed into a table of learnable vectors, fetched by hash, and content-addressed against the current hidden state. It is added as a residual before attention rather than as extra attention positions, so it costs a lookup instead of quadratic attention over a longer context.
The bet is domain-specific. Payment documents are dense in short repeated
n-grams — currency codes, BIC shapes, IBAN country prefixes, XML element names,
<Dt> and <Ccy> and IBAN. Recall of exactly those patterns is worth more
here than at general language scale, where a 45M-parameter model has no hope of
memorising useful spans. At 53M on 3,800 documents, the model can plausibly store
the mapping, and the engram is what makes that cheap.
Confidence head: the model reports its own uncertainty
Hidden states are pooled through 8 learned probes into a single correctness logit per document. That is a second output of the same forward pass, not a second model, and it is what turns an uncertain extraction into a routing decision rather than an error.
It is trained in a second stage on the model's own right and wrong outcomes, not on ground truth. That detail matters: fitted on ground truth, every label is "correct," and the head learns to emit a constant. A constant scores 0.0 separation regardless of accuracy and gives no routing signal at all.
Architecture summary
| Component | Value |
|---|---|
| Parameters | 52,982,805 (53.0M) |
| Layers | 20 |
| Model width | 512 |
| Attention | 8 query / 4 KV heads (grouped-query) |
| FFN | SwiGLU, hidden 768 |
| Positions | RoPE, theta 100000 |
| Norm | Pre-norm RMSNorm |
| Vocabulary | 8192, trained on this domain |
| Context | 256 tokens |
| Engram memory | 2 layers (2, 12), 8192 slots × 128 dim × 2 heads |
| Confidence head | 8 probes |
Design follows Needle (Cactus Compute, 45M). This model is 53M, and is ported to
PyTorch so it loads through transformers without a custom runtime.
How it was trained
Two stages, on a corpus where every value is machine-checked.
flowchart LR
T["Teacher model<br/>optional"] -.->|"distil"| S1
C["Validator-checked<br/>corpus<br/>3 difficulty levels"] --> S1
S1["Stage 1<br/>teacher-forced<br/>prompt masked"] --> S2
S1 -.->|"generate + score"| GEN["Model's own<br/>right/wrong"]
GEN --> S2["Stage 2<br/>confidence<br/>calibration"]
S2 --> CKPT["Checkpoint"]
Stage 1 — extraction. Teacher-forced next-token prediction on document → JSON pairs. Only the JSON target is scored; the prompt is masked. Training on the prompt rewards copying the input, which turns an extractor into a paraphraser.
Stage 2 — confidence calibration. Predictions are generated, scored against truth fields, and the confidence head is fitted to that outcome.
| Setting | Value |
|---|---|
| Training documents | 3,800 |
| Validation documents | 200 |
| Epochs | 3 |
| Effective batch | 32 (batch 8 × grad_accum 4) |
| Learning rate | 0.0003 (cosine, warmup) |
| Optimiser | AdamW, betas (0.9, 0.95) |
| Precision | fp32 |
| Hardware | RTX 2050, 4 GB VRAM (compute 8.6) |
| Throughput | ~0.46 s/document |
| Context cap | 256 (measured: keeps 100% of sequences) |
The corpus is generated, not collected
Field values are real, harvested from Query-farm/vgi-iso20022 fixtures. The
prose is synthetic, rendered from a variant pool of label synonyms, number
conventions (US and European), date orders, and distractors.
Synthesis is validator-checked rather than free-form: IBANs carry correct mod-97 check digits for their country format, BICs match ISO 9362, and amounts respect currency minor units. A corpus of invalid values would measure nothing, because the model would be graded on learning to imitate broken data.
Three difficulty levels, and the model sees all three:
clean— canonical labels, one format, no noisemessy— synonym labels, mixed formats, distractor fieldshostile— key fields stated in running prose with no label, heavy noise
What we learned: four ways the metrics lied
This is the most transferable part of the project.
The training log reported 100% field recall while the checkpoints it wrote generated repetition loops. The weights were not the problem — all 214 tensors round-tripped byte for byte, so the checkpoint was faithful.
RotaryEmbedding held its inverse frequencies in a buffer marked
persistent=False. They were never written to the file, and the loader left them
uninitialised:
trained [1.0, 0.6978, 0.4869, 0.3398]
reloaded [-9.65e-11, 3.09e-41, 0.0, 0.0]
Rebuilding them from head_dim and theta made the two forward passes agree
exactly, with no residual difference in the logits. Three further faults were
hiding behind it:
| Fault | Symptom | Fix |
|---|---|---|
| Uninitialised RoPE buffer | Log and artifact disagreed by 287× in loss | Rebuild inv_freq in _init_weights |
| Teacher-forced metric | 100% recall on a model producing loops | Score generation, not next-token |
| Validation leakage | 600/600 validation answers present in training | Disjoint held-out draw |
| Difficulty-major sampling | Gate saw one difficulty level | Shuffle across levels |
Every one of them made the model look better than it was. None broke a test, because the tests were checking the same assumption the metric was. That is the lesson: a metric fails in the direction you want, repeatedly, and stays green.
The practical defence is cheap. Measure one quantity two ways and refuse to
accept that the answers disagree. That is the whole method, and it is why
test_metric_describes_artifact.py
exists as a test rather than a one-off investigation.
Two more findings worth stating
Repetition is augmentation and a memorisation risk at once. Rendering each value-set at three difficulty levels teaches the extraction mapping faster than rendering it once, and it also makes the value-set memorisable. Both effects are real and they pull in opposite directions.
Capacity decides which strategy wins. At 4,000 distinct value-sets, 53M parameters can store the mapping and do store it, which is why held-out slice accuracy was 97.5% while a fresh draw gave 61.2%. Raising the distinct count removes memorisation as an option and forces the model to learn the task. That takes longer than memorising and is the only route to a number that holds up on unseen documents.
Evaluation
Scored by a nine-check acceptance gate run against a fresh document draw the model has never seen.
| Check | Result | Threshold | What it protects |
|---|---|---|---|
| C1 valid JSON object | 1.0000 | 1.0 | Output can be parsed at all |
| C2 canonical keys only | 1.0000 | 1.0 | No invented field names |
| C3 no duplicate keys | 1.0000 | 1.0 | No key emitted twice |
| C6 deterministic | 1.0000 | 1.0 | Same document, same output |
| C7 XSD round-trip | 1.0000 | 1.0 | Assembles into a valid message |
| C9 routing agreement | 1.0000 | 1.0 | IBAN and BIC agree on country |
| C4 values grounded | 0.0000 | 1.0 | Every value traceable to source |
| C5 field accuracy | 0.7167 | ≥ 0.9 | Values are right |
| C5b document STP | 0.0000 | ≥ 0.5 | Whole documents correct |
| C8 role consistency | 0.0000 | 1.0 | Payer is not swapped for payee |
Read the pass column before the fail column. The six checks at 100% are the ones that establish that the pipeline is sound: the model emits parseable, canonically-keyed, deterministic output that assembles into a schema-valid message with cross-field agreement. Those are the properties you cannot fix later with more data, and they are done.
The four that fail reduce to one cause: value accuracy on unseen values. The distinct-value count in the corpus is a parameter, not a redesign.
On the earlier reported numbers
An earlier revision of this card reported 82.14% token accuracy and a confidence separation of 0.0000. Those numbers are withdrawn. They came from a checkpoint that loaded with uninitialised rotary frequencies, and they measured a quantity that does not matter — the same model tokenised to 82.14% while extracting 1 field in 120. Token accuracy on a teacher-forced target is not a measurement of an extractor. Generation accuracy under the gate is.
If you are evaluating this model, use the gate table above and re-run it against your own documents.
Why local matters in finance
"Why not just write regexes?"
This is the right first question, and it has a measured answer rather than a sales answer. Two competent rule engines were built — label anchors, a normaliser per format, pattern fallbacks, and deterministic validators — and scored on the same 360 documents:
| Extractor | STP | Field accuracy | clean | messy | hostile |
|---|---|---|---|---|---|
rules_v1 (one template) |
33.3% | 47.1% | 100.0% | 0.0% | 0.0% |
rules_v2 (every synonym) |
51.7% | 89.4% | 92.5% | 62.5% | 0.0% |
rules_v1 scores 100% on clean and exactly 0% on everything else. Not 80%,
not 60% — zero. Regex does not fail gradually. It succeeds completely inside the
format it was written for and fails completely outside it.
So a rules engine does not have an accuracy. It has a coverage, counted in formats, and the cost of the rules approach is one rule set per document format. Format count is a property of the customer base, not the volume.
| Your situation | The right answer |
|---|---|
| One format, high volume, stable | Regex. A model is waste. |
| A handful of formats | Regex, plus normalisers |
| Many formats, or formats you do not control | Rule maintenance exceeds the model |
| Onboarding a new customer | Model: zero engineering. Rules: days each. |
What rules structurally cannot do
Given Account: DE89370400440532013000, a rule engine knows the value is an
IBAN. It cannot know whether that account belongs to the payer or the beneficiary,
because that fact lives in the document's meaning and not its surface form. With
no label present, the fallback guesses by position and is wrong about half the
time.
That single limitation is what the exception rate is made of, and no amount of additional patterns removes it.
The economics
At 100,000 messages/month, 8 minutes per exception, $45/hour loaded:
| STP | Exceptions/month | Annual cost | FTE |
|---|---|---|---|
| 85.1% | 14,925 | $1.07M | 13.1 |
| 64.2% | 35,821 | $2.58M | 31.4 |
| 49.8% | 50,249 | $3.62M | 44.1 |
One STP point is worth $72,000/year at this volume.
The usual framing asks whether to run inference on an API or locally. That question is worth $56,000/year at this volume. The STP delta from 49.8% to 85% is worth $2,537,910/year — 45.3x the infrastructure difference. The entire $20k build is repaid by 0.33 of one STP point.
Where inference runs is a rounding error next to how often the pipeline is right. Local wins here for data residency, determinism, and auditability. It does not win on cost, because cost was never the deciding term.
Three properties worth having first
Data residency is structural, not contractual. The weights are 53M parameters. Correspondence does not leave the machine, so there is no data processing agreement to negotiate and no cross-border transfer to document.
Decoding is deterministic. Greedy decoding means the same document yields the same fields every time. An audit that cannot reproduce an output is not an audit.
Uncertainty is an output, not a hope. The confidence head routes low-confidence documents to a human instead of guessing, so the review queue becomes a designed part of the system rather than the place errors surface.
There is a fourth that is easy to miss: the routing cross-check resolves two fields against each other. An IBAN encodes its bank's country, so a BIC that disagrees with it is caught without consulting any external directory — a validator that costs nothing at inference time and catches a whole class of role inversion.
The honest position today
This model does not yet beat the rule baseline. rules_v2 reaches 51.7%
document STP; this model is at 0% against a 50% target. On the measured evidence,
a competent rule engine is the better buy today, and claiming otherwise would be
exactly the kind of unsupported assertion this project was built to catch.
What the work establishes is the other half of the decision: the rules path plateaus at the number of formats written, role ambiguity is not solvable with more patterns, and the money sits in STP rather than in infrastructure. A model that reaches the gate converts all three into a fixed cost rather than a per-format one.
The lever is the distinct-value count in the corpus, which is a parameter rather than a redesign. The next experiment is to raise it until memorisation stops being an option and measure whether field accuracy crosses 90% with document STP above 50%. If it plateaus well below that on a corpus this clean, that is a real negative result and worth publishing as one.
Intended use
Extraction of payment fields from documents into a canonical key space:
<|doc|>
Beneficiary: Acme GmbH
IBAN: DE89370400440532013000
Amount: EUR 1,250.00
<|json|>{"CdtrAcct_IBAN":"DE89370400440532013000","Cdtr_Nm":"Acme GmbH",...}<|end|>
Keys are ISO 20022 path-derived names: Cdtr_Nm, CdtrAcct_IBAN, Dbtr_Nm,
DbtrAcct_IBAN, Amt_InstdAmt, Amt_InstdAmt_Ccy, PmtId_EndToEndId,
ReqdExctnDt_Dt, RmtInf_Ustrd, CdtrAgt_BICFI.
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "sivasub987/iso20022-extract-53m"
tokenizer = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(repo, trust_remote_code=True)
model.eval()
prompt = "<|doc|>\n" + document_text + "\n<|json|>"
ids = tokenizer.encode(prompt, return_tensors="pt")
out = model.generate(ids, max_new_tokens=160, do_sample=False)
print(tokenizer.decode(out[0, ids.shape[1]:], skip_special_tokens=True))
The model exposes generate_greedy for deterministic decoding, which is what
extraction should use.
Limitations
These are the measured limits, not caveats. Treat them as the operating envelope.
- The gate does not pass. 71.7% field accuracy against 90%, and 0% document straight-through against 50%.
- Synthetic prose. Real correspondence is messier than any variant pool, so accuracy here is an upper bound.
- Nine fields. Message assembly is separate code, and required system fields
(
MsgId,PmtInfId,CreDtTm,NbOfTxs,CtrlSum) are generated, not predicted. hostileis the hard case. Where a field has no label and must be read from running prose, accuracy is materially lower than the headline figure.- Role ambiguity is unresolved in principle. Two IBANs and no labels give no textual signal for which is payer and which is beneficiary.
- English and European formats. Dates are read day-first for numeric ambiguity, which is wrong for US-format documents.
- The fault list is probably incomplete. Those four were found in sequence, not by audit.
- Not a payment system. Output is a prediction. Nothing here should authorise a payment without the confidence gate or human review.
Citation
Source: github.com/siva-sub/iso20022-extract
Architecture after Needle (Cactus Compute). Schema and fixtures from
Query-farm/vgi-iso20022 and the ISO 20022 message definitions.
Files
| File | Purpose |
|---|---|
model.safetensors |
Weights, safe format |
config.json |
Architecture |
tokenizer.json |
Trained vocabulary |
modeling_iso20022.py |
Architecture code for trust_remote_code |
configuration_iso20022.py |
Config class |
- Downloads last month
- 251
Evaluation results
- Field accuracy (C5) on Validator-checked synthetic payment correspondence (fresh draw)self-reported0.717
- Valid JSON object (C1) on Validator-checked synthetic payment correspondence (fresh draw)self-reported1.000
- Canonical keys only (C2) on Validator-checked synthetic payment correspondence (fresh draw)self-reported1.000
- No duplicate keys (C3) on Validator-checked synthetic payment correspondence (fresh draw)self-reported1.000
- Deterministic decoding (C6) on Validator-checked synthetic payment correspondence (fresh draw)self-reported1.000
- XSD round-trip (C7) on Validator-checked synthetic payment correspondence (fresh draw)self-reported1.000
- IBAN/BIC routing agreement (C9) on Validator-checked synthetic payment correspondence (fresh draw)self-reported1.000
- Document straight-through (C5b) on Validator-checked synthetic payment correspondence (fresh draw)self-reported0.000