iso20022-extract-53m

A 53M-parameter decoder-only transformer that reads unstructured payment correspondence and emits structured ISO 20022 fields. It trains from scratch on a single consumer laptop GPU. No frontier model, no API call, and no document leaving the machine.

The structural pipeline is complete and verified: six of the nine acceptance checks pass at 100%, including valid JSON, canonical field names, deterministic decoding, schema-valid message assembly, and field-level routing agreement. Value accuracy on unseen values is the remaining work.

Getting here required finding four separate ways the metrics were lying, all of them in the model's favour. That part is worth reading even if you never touch this model, because the same failure mode is available to anyone measuring a specialist extractor.

Status: proof of concept. The gate does not pass yet. Field accuracy is 71.7% against a 90% target. Do not use this to move money.


The architecture, and why these two parts are novel

Most of this model is conventional and deliberately so: 20 pre-norm blocks, grouped-query attention, RoPE, SwiGLU, RMSNorm, tied embeddings. The pieces worth attention are the engram memory and the confidence head, because both exist to exploit a property of payment documents specifically rather than language generally.

flowchart TB
    subgraph STACK["53M decoder · 20 blocks · d_model 512 · context 256"]
        direction TB
        IN["Token ids"] --> EMB["Embedding<br/>8192 × 512"]
        EMB --> BLK["20 × transformer block"]
        BLK --> FN["Final RMSNorm"]
    end

    subgraph ONE["Inside one block"]
        direction TB
        P["Pre-norm<br/>RMSNorm"] --> ATT["GQA attention<br/>8 query / 4 KV · RoPE θ=1e5"]
        ATT --> R1["Residual add"]
        R1 --> P2["Pre-norm<br/>RMSNorm"]
        P2 --> FFN["SwiGLU FFN<br/>512 → 768"]
        FFN --> R2["Residual add"]
    end

    ENG["Engram memory<br/>hashed n-gram read<br/>blocks 2 and 12"] -.-> BLK
    FN --> LM["LM head<br/>tied to embedding"]
    FN --> CH["Confidence head<br/>8 probes"]
    LM --> OUT["Field JSON"]
    CH --> ROUTE["Low-confidence<br/>routing"]

Engram: a memory read, not more attention

Token n-grams are hashed into a table of learnable vectors, fetched by hash, and content-addressed against the current hidden state. It is added as a residual before attention rather than as extra attention positions, so it costs a lookup instead of quadratic attention over a longer context.

The bet is domain-specific. Payment documents are dense in short repeated n-grams — currency codes, BIC shapes, IBAN country prefixes, XML element names, <Dt> and <Ccy> and IBAN. Recall of exactly those patterns is worth more here than at general language scale, where a 45M-parameter model has no hope of memorising useful spans. At 53M on 3,800 documents, the model can plausibly store the mapping, and the engram is what makes that cheap.

Confidence head: the model reports its own uncertainty

Hidden states are pooled through 8 learned probes into a single correctness logit per document. That is a second output of the same forward pass, not a second model, and it is what turns an uncertain extraction into a routing decision rather than an error.

It is trained in a second stage on the model's own right and wrong outcomes, not on ground truth. That detail matters: fitted on ground truth, every label is "correct," and the head learns to emit a constant. A constant scores 0.0 separation regardless of accuracy and gives no routing signal at all.

Architecture summary

Component Value
Parameters 52,982,805 (53.0M)
Layers 20
Model width 512
Attention 8 query / 4 KV heads (grouped-query)
FFN SwiGLU, hidden 768
Positions RoPE, theta 100000
Norm Pre-norm RMSNorm
Vocabulary 8192, trained on this domain
Context 256 tokens
Engram memory 2 layers (2, 12), 8192 slots × 128 dim × 2 heads
Confidence head 8 probes

Design follows Needle (Cactus Compute, 45M). This model is 53M, and is ported to PyTorch so it loads through transformers without a custom runtime.


How it was trained

Two stages, on a corpus where every value is machine-checked.

flowchart LR
    T["Teacher model<br/>optional"] -.->|"distil"| S1
    C["Validator-checked<br/>corpus<br/>3 difficulty levels"] --> S1
    S1["Stage 1<br/>teacher-forced<br/>prompt masked"] --> S2
    S1 -.->|"generate + score"| GEN["Model's own<br/>right/wrong"]
    GEN --> S2["Stage 2<br/>confidence<br/>calibration"]
    S2 --> CKPT["Checkpoint"]

Stage 1 — extraction. Teacher-forced next-token prediction on document → JSON pairs. Only the JSON target is scored; the prompt is masked. Training on the prompt rewards copying the input, which turns an extractor into a paraphraser.

Stage 2 — confidence calibration. Predictions are generated, scored against truth fields, and the confidence head is fitted to that outcome.

Setting Value
Training documents 3,800
Validation documents 200
Epochs 3
Effective batch 32 (batch 8 × grad_accum 4)
Learning rate 0.0003 (cosine, warmup)
Optimiser AdamW, betas (0.9, 0.95)
Precision fp32
Hardware RTX 2050, 4 GB VRAM (compute 8.6)
Throughput ~0.46 s/document
Context cap 256 (measured: keeps 100% of sequences)

The corpus is generated, not collected

Field values are real, harvested from Query-farm/vgi-iso20022 fixtures. The prose is synthetic, rendered from a variant pool of label synonyms, number conventions (US and European), date orders, and distractors.

Synthesis is validator-checked rather than free-form: IBANs carry correct mod-97 check digits for their country format, BICs match ISO 9362, and amounts respect currency minor units. A corpus of invalid values would measure nothing, because the model would be graded on learning to imitate broken data.

Three difficulty levels, and the model sees all three:

  • clean — canonical labels, one format, no noise
  • messy — synonym labels, mixed formats, distractor fields
  • hostile — key fields stated in running prose with no label, heavy noise

What we learned: four ways the metrics lied

This is the most transferable part of the project.

The training log reported 100% field recall while the checkpoints it wrote generated repetition loops. The weights were not the problem — all 214 tensors round-tripped byte for byte, so the checkpoint was faithful.

RotaryEmbedding held its inverse frequencies in a buffer marked persistent=False. They were never written to the file, and the loader left them uninitialised:

trained    [1.0, 0.6978, 0.4869, 0.3398]
reloaded   [-9.65e-11, 3.09e-41, 0.0, 0.0]

Rebuilding them from head_dim and theta made the two forward passes agree exactly, with no residual difference in the logits. Three further faults were hiding behind it:

Fault Symptom Fix
Uninitialised RoPE buffer Log and artifact disagreed by 287× in loss Rebuild inv_freq in _init_weights
Teacher-forced metric 100% recall on a model producing loops Score generation, not next-token
Validation leakage 600/600 validation answers present in training Disjoint held-out draw
Difficulty-major sampling Gate saw one difficulty level Shuffle across levels

Every one of them made the model look better than it was. None broke a test, because the tests were checking the same assumption the metric was. That is the lesson: a metric fails in the direction you want, repeatedly, and stays green.

The practical defence is cheap. Measure one quantity two ways and refuse to accept that the answers disagree. That is the whole method, and it is why test_metric_describes_artifact.py exists as a test rather than a one-off investigation.

Two more findings worth stating

Repetition is augmentation and a memorisation risk at once. Rendering each value-set at three difficulty levels teaches the extraction mapping faster than rendering it once, and it also makes the value-set memorisable. Both effects are real and they pull in opposite directions.

Capacity decides which strategy wins. At 4,000 distinct value-sets, 53M parameters can store the mapping and do store it, which is why held-out slice accuracy was 97.5% while a fresh draw gave 61.2%. Raising the distinct count removes memorisation as an option and forces the model to learn the task. That takes longer than memorising and is the only route to a number that holds up on unseen documents.


Evaluation

Scored by a nine-check acceptance gate run against a fresh document draw the model has never seen.

Check Result Threshold What it protects
C1 valid JSON object 1.0000 1.0 Output can be parsed at all
C2 canonical keys only 1.0000 1.0 No invented field names
C3 no duplicate keys 1.0000 1.0 No key emitted twice
C6 deterministic 1.0000 1.0 Same document, same output
C7 XSD round-trip 1.0000 1.0 Assembles into a valid message
C9 routing agreement 1.0000 1.0 IBAN and BIC agree on country
C4 values grounded 0.0000 1.0 Every value traceable to source
C5 field accuracy 0.7167 ≥ 0.9 Values are right
C5b document STP 0.0000 ≥ 0.5 Whole documents correct
C8 role consistency 0.0000 1.0 Payer is not swapped for payee

Read the pass column before the fail column. The six checks at 100% are the ones that establish that the pipeline is sound: the model emits parseable, canonically-keyed, deterministic output that assembles into a schema-valid message with cross-field agreement. Those are the properties you cannot fix later with more data, and they are done.

The four that fail reduce to one cause: value accuracy on unseen values. The distinct-value count in the corpus is a parameter, not a redesign.

On the earlier reported numbers

An earlier revision of this card reported 82.14% token accuracy and a confidence separation of 0.0000. Those numbers are withdrawn. They came from a checkpoint that loaded with uninitialised rotary frequencies, and they measured a quantity that does not matter — the same model tokenised to 82.14% while extracting 1 field in 120. Token accuracy on a teacher-forced target is not a measurement of an extractor. Generation accuracy under the gate is.

If you are evaluating this model, use the gate table above and re-run it against your own documents.


Why local matters in finance

"Why not just write regexes?"

This is the right first question, and it has a measured answer rather than a sales answer. Two competent rule engines were built — label anchors, a normaliser per format, pattern fallbacks, and deterministic validators — and scored on the same 360 documents:

Extractor STP Field accuracy clean messy hostile
rules_v1 (one template) 33.3% 47.1% 100.0% 0.0% 0.0%
rules_v2 (every synonym) 51.7% 89.4% 92.5% 62.5% 0.0%

rules_v1 scores 100% on clean and exactly 0% on everything else. Not 80%, not 60% — zero. Regex does not fail gradually. It succeeds completely inside the format it was written for and fails completely outside it.

So a rules engine does not have an accuracy. It has a coverage, counted in formats, and the cost of the rules approach is one rule set per document format. Format count is a property of the customer base, not the volume.

Your situation The right answer
One format, high volume, stable Regex. A model is waste.
A handful of formats Regex, plus normalisers
Many formats, or formats you do not control Rule maintenance exceeds the model
Onboarding a new customer Model: zero engineering. Rules: days each.

What rules structurally cannot do

Given Account: DE89370400440532013000, a rule engine knows the value is an IBAN. It cannot know whether that account belongs to the payer or the beneficiary, because that fact lives in the document's meaning and not its surface form. With no label present, the fallback guesses by position and is wrong about half the time.

That single limitation is what the exception rate is made of, and no amount of additional patterns removes it.

The economics

At 100,000 messages/month, 8 minutes per exception, $45/hour loaded:

STP Exceptions/month Annual cost FTE
85.1% 14,925 $1.07M 13.1
64.2% 35,821 $2.58M 31.4
49.8% 50,249 $3.62M 44.1

One STP point is worth $72,000/year at this volume.

The usual framing asks whether to run inference on an API or locally. That question is worth $56,000/year at this volume. The STP delta from 49.8% to 85% is worth $2,537,910/year — 45.3x the infrastructure difference. The entire $20k build is repaid by 0.33 of one STP point.

Where inference runs is a rounding error next to how often the pipeline is right. Local wins here for data residency, determinism, and auditability. It does not win on cost, because cost was never the deciding term.

Three properties worth having first

Data residency is structural, not contractual. The weights are 53M parameters. Correspondence does not leave the machine, so there is no data processing agreement to negotiate and no cross-border transfer to document.

Decoding is deterministic. Greedy decoding means the same document yields the same fields every time. An audit that cannot reproduce an output is not an audit.

Uncertainty is an output, not a hope. The confidence head routes low-confidence documents to a human instead of guessing, so the review queue becomes a designed part of the system rather than the place errors surface.

There is a fourth that is easy to miss: the routing cross-check resolves two fields against each other. An IBAN encodes its bank's country, so a BIC that disagrees with it is caught without consulting any external directory — a validator that costs nothing at inference time and catches a whole class of role inversion.

The honest position today

This model does not yet beat the rule baseline. rules_v2 reaches 51.7% document STP; this model is at 0% against a 50% target. On the measured evidence, a competent rule engine is the better buy today, and claiming otherwise would be exactly the kind of unsupported assertion this project was built to catch.

What the work establishes is the other half of the decision: the rules path plateaus at the number of formats written, role ambiguity is not solvable with more patterns, and the money sits in STP rather than in infrastructure. A model that reaches the gate converts all three into a fixed cost rather than a per-format one.

The lever is the distinct-value count in the corpus, which is a parameter rather than a redesign. The next experiment is to raise it until memorisation stops being an option and measure whether field accuracy crosses 90% with document STP above 50%. If it plateaus well below that on a corpus this clean, that is a real negative result and worth publishing as one.


Intended use

Extraction of payment fields from documents into a canonical key space:

<|doc|>
Beneficiary: Acme GmbH
IBAN: DE89370400440532013000
Amount: EUR 1,250.00
<|json|>{"CdtrAcct_IBAN":"DE89370400440532013000","Cdtr_Nm":"Acme GmbH",...}<|end|>

Keys are ISO 20022 path-derived names: Cdtr_Nm, CdtrAcct_IBAN, Dbtr_Nm, DbtrAcct_IBAN, Amt_InstdAmt, Amt_InstdAmt_Ccy, PmtId_EndToEndId, ReqdExctnDt_Dt, RmtInf_Ustrd, CdtrAgt_BICFI.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "sivasub987/iso20022-extract-53m"
tokenizer = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(repo, trust_remote_code=True)
model.eval()

prompt = "<|doc|>\n" + document_text + "\n<|json|>"
ids = tokenizer.encode(prompt, return_tensors="pt")
out = model.generate(ids, max_new_tokens=160, do_sample=False)
print(tokenizer.decode(out[0, ids.shape[1]:], skip_special_tokens=True))

The model exposes generate_greedy for deterministic decoding, which is what extraction should use.

Limitations

These are the measured limits, not caveats. Treat them as the operating envelope.

  • The gate does not pass. 71.7% field accuracy against 90%, and 0% document straight-through against 50%.
  • Synthetic prose. Real correspondence is messier than any variant pool, so accuracy here is an upper bound.
  • Nine fields. Message assembly is separate code, and required system fields (MsgId, PmtInfId, CreDtTm, NbOfTxs, CtrlSum) are generated, not predicted.
  • hostile is the hard case. Where a field has no label and must be read from running prose, accuracy is materially lower than the headline figure.
  • Role ambiguity is unresolved in principle. Two IBANs and no labels give no textual signal for which is payer and which is beneficiary.
  • English and European formats. Dates are read day-first for numeric ambiguity, which is wrong for US-format documents.
  • The fault list is probably incomplete. Those four were found in sequence, not by audit.
  • Not a payment system. Output is a prediction. Nothing here should authorise a payment without the confidence gate or human review.

Citation

Source: github.com/siva-sub/iso20022-extract

Architecture after Needle (Cactus Compute). Schema and fixtures from Query-farm/vgi-iso20022 and the ISO 20022 message definitions.

Files

File Purpose
model.safetensors Weights, safe format
config.json Architecture
tokenizer.json Trained vocabulary
modeling_iso20022.py Architecture code for trust_remote_code
configuration_iso20022.py Config class
Downloads last month
251
Safetensors
Model size
53M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Evaluation results

  • Field accuracy (C5) on Validator-checked synthetic payment correspondence (fresh draw)
    self-reported
    0.717
  • Valid JSON object (C1) on Validator-checked synthetic payment correspondence (fresh draw)
    self-reported
    1.000
  • Canonical keys only (C2) on Validator-checked synthetic payment correspondence (fresh draw)
    self-reported
    1.000
  • No duplicate keys (C3) on Validator-checked synthetic payment correspondence (fresh draw)
    self-reported
    1.000
  • Deterministic decoding (C6) on Validator-checked synthetic payment correspondence (fresh draw)
    self-reported
    1.000
  • XSD round-trip (C7) on Validator-checked synthetic payment correspondence (fresh draw)
    self-reported
    1.000
  • IBAN/BIC routing agreement (C9) on Validator-checked synthetic payment correspondence (fresh draw)
    self-reported
    1.000
  • Document straight-through (C5b) on Validator-checked synthetic payment correspondence (fresh draw)
    self-reported
    0.000