How to use from the
Use from the
Transformers library
# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("token-classification", model="mksl-ai/docuvision-layoutlmv3-large-funsd")
# Load model directly
from transformers import AutoProcessor, AutoModelForTokenClassification

processor = AutoProcessor.from_pretrained("mksl-ai/docuvision-layoutlmv3-large-funsd")
model = AutoModelForTokenClassification.from_pretrained("mksl-ai/docuvision-layoutlmv3-large-funsd", device_map="auto")
Quick Links

DocuVision LayoutLMv3-large (FUNSD, reading order) + key-value linker

These are the two learned models behind key-value extraction in my DocuVision pipeline:

  1. A LayoutLMv3-large entity tagger (repo root). It tags every word of a form as HEADER, QUESTION, ANSWER or O (7 BIO labels).
  2. A question → answer linker (linker/kv-linker.joblib). It decides which tagged question goes with which answer.

I fine-tuned the tagger on words sorted into reading order, not FUNSD's entity-by-entity order, because that's the order OCR returns at inference. It scores slightly lower on seqeval but higher on the metric that matters, end-to-end question/answer pair F1.

Results (FUNSD test split, 50 forms)

Precision Recall F1
ANSWER 0.829 0.888 0.858
QUESTION 0.825 0.837 0.831
HEADER 0.645 0.614 0.629
Micro avg (entity F1) 0.816 0.844 0.830

End-to-end key-value pair F1: 0.736 (tagger + linker on gold words). With Tesseract OCR on FUNSD's ~100-dpi faxes it drops to 0.43–0.46; OCR quality is the bottleneck there, not the models.

Ablations I ran

Base Boxes Word order lr / batch Time Entity F1 Pair F1
base word FUNSD 2e-5 / 4 3.9 min 0.824 0.651
base segment FUNSD 2e-5 / 4 4.0 min 0.814 0.659
base word reading 2e-5 / 4 4.1 min 0.812 0.701
large word FUNSD 1e-5 / 2 13.8 min 0.846 0.708
large word reading 1e-5 / 2 13.6 min 0.830 0.736

This repo is the last row.

Training details

Tagger. microsoft/layoutlmv3-large, 149 FUNSD training forms, a plain PyTorch loop: AdamW (lr 1e-5), batch size 2, 30 epochs, 10% linear warmup then linear decay, fp16 autocast, gradient clipping at 1.0, seed 42. Pages are split into 220-word windows so nothing is truncated. Word-level boxes (box_mode: "word" is stored in config.json). FUNSD has no validation split, so I fixed the 30-epoch schedule up front instead of picking the best epoch on the test set. 13.6 minutes on one RTX A5000. Per-epoch loss and metrics are in training_metrics.json.

Linker. A scikit-learn HistGradientBoostingClassifier (400 iterations, learning rate 0.05, 31 leaves) over 21 geometric and text features per candidate (question, answer) pair: offsets, overlaps, same-line, relative rank, trailing colon, digits, lengths, distance. Trained on all 75,504 candidate pairs from the training forms (3,049 positive). The decision threshold (0.35) was chosen by 5-fold GroupKFold cross-validation grouped by form; CV pair F1 0.853. At inference I also keep only the best question for each answer, which adds about 4 points end to end.

Usage

Tagger

from PIL import Image
from transformers import AutoProcessor, AutoModelForTokenClassification

repo = "mksl-ai/docuvision-layoutlmv3-large-funsd"
processor = AutoProcessor.from_pretrained(repo, apply_ocr=False)  # bring your own OCR
model = AutoModelForTokenClassification.from_pretrained(repo).eval()

image = Image.open("form.png").convert("RGB")
# words: list[str] from your OCR, in reading order
# boxes: list[[x0, y0, x1, y1]] per word, normalised to 0-1000
enc = processor(image, words, boxes=boxes, return_tensors="pt", truncation=True)
pred = model(**enc).logits.argmax(-1)[0]
labels = [model.config.id2label[i.item()] for i in pred]

Linker

import joblib
from huggingface_hub import hf_hub_download

bundle = joblib.load(hf_hub_download(repo, "linker/kv-linker.joblib"))
clf, threshold, feature_names = bundle["model"], bundle["threshold"], bundle["features"]

The linker expects the 21 features built by pair_features in docuvision/extraction/linker.py. The easiest way to use both models together is DocuVision itself: download this repo into models/layoutlmv3-funsd, put the joblib at models/kv-linker.joblib, and run the GPU profile.

About the joblib file: it's a pickle containing a dict {"model": HistGradientBoostingClassifier, "threshold": 0.35, "features": [...]}. Only load pickles from sources you trust. It was saved with scikit-learn 1.9.1 and joblib 1.6.0; other scikit-learn versions may warn or fail to load it.

Environment I trained with: transformers 4.57.6, torch 2.6.0 (CUDA 12.4), scikit-learn 1.9.1.

Limitations

  • English business forms only; FUNSD is 199 scanned forms, so layouts far from it (receipts, invoices with dense tables) will be weaker.
  • HEADER is the weak class (F1 0.63, only 127 test examples, and headers look a lot like questions).
  • The tagger needs good word boxes. On low-resolution scans, OCR errors dominate the final result.

License

The weights are CC BY-NC-SA 4.0, non-commercial: LayoutLMv3 is released under CC BY-NC-SA 4.0 and FUNSD is for non-commercial research use. The DocuVision code is MIT, but that doesn't extend to these weights.

Citation

If you use these, please cite the original work:

@inproceedings{huang2022layoutlmv3,
  title     = {LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking},
  author    = {Huang, Yupan and Lv, Tengchao and Cui, Lei and Lu, Yutong and Wei, Furu},
  booktitle = {ACM Multimedia},
  year      = {2022}
}

@inproceedings{jaume2019funsd,
  title     = {FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents},
  author    = {Jaume, Guillaume and Ekenel, Hazim Kemal and Thiran, Jean-Philippe},
  booktitle = {ICDAR-OST},
  year      = {2019}
}
Downloads last month
10
Safetensors
Model size
0.4B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mksl-ai/docuvision-layoutlmv3-large-funsd

Finetuned
(14)
this model

Dataset used to train mksl-ai/docuvision-layoutlmv3-large-funsd

Evaluation results