Instructions to use mksl-ai/docuvision-layoutlmv3-large-funsd with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use mksl-ai/docuvision-layoutlmv3-large-funsd with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="mksl-ai/docuvision-layoutlmv3-large-funsd")# Load model directly from transformers import AutoProcessor, AutoModelForTokenClassification processor = AutoProcessor.from_pretrained("mksl-ai/docuvision-layoutlmv3-large-funsd") model = AutoModelForTokenClassification.from_pretrained("mksl-ai/docuvision-layoutlmv3-large-funsd", device_map="auto") - Notebooks
- Google Colab
- Kaggle
DocuVision LayoutLMv3-large (FUNSD, reading order) + key-value linker
These are the two learned models behind key-value extraction in my DocuVision pipeline:
- A LayoutLMv3-large entity tagger (repo root). It tags every word of a
form as
HEADER,QUESTION,ANSWERorO(7 BIO labels). - A question → answer linker (
linker/kv-linker.joblib). It decides which tagged question goes with which answer.
I fine-tuned the tagger on words sorted into reading order, not FUNSD's entity-by-entity order, because that's the order OCR returns at inference. It scores slightly lower on seqeval but higher on the metric that matters, end-to-end question/answer pair F1.
Results (FUNSD test split, 50 forms)
| Precision | Recall | F1 | |
|---|---|---|---|
| ANSWER | 0.829 | 0.888 | 0.858 |
| QUESTION | 0.825 | 0.837 | 0.831 |
| HEADER | 0.645 | 0.614 | 0.629 |
| Micro avg (entity F1) | 0.816 | 0.844 | 0.830 |
End-to-end key-value pair F1: 0.736 (tagger + linker on gold words). With Tesseract OCR on FUNSD's ~100-dpi faxes it drops to 0.43–0.46; OCR quality is the bottleneck there, not the models.
Ablations I ran
| Base | Boxes | Word order | lr / batch | Time | Entity F1 | Pair F1 |
|---|---|---|---|---|---|---|
| base | word | FUNSD | 2e-5 / 4 | 3.9 min | 0.824 | 0.651 |
| base | segment | FUNSD | 2e-5 / 4 | 4.0 min | 0.814 | 0.659 |
| base | word | reading | 2e-5 / 4 | 4.1 min | 0.812 | 0.701 |
| large | word | FUNSD | 1e-5 / 2 | 13.8 min | 0.846 | 0.708 |
| large | word | reading | 1e-5 / 2 | 13.6 min | 0.830 | 0.736 |
This repo is the last row.
Training details
Tagger. microsoft/layoutlmv3-large, 149 FUNSD training forms, a plain
PyTorch loop: AdamW (lr 1e-5), batch size 2, 30 epochs, 10% linear warmup
then linear decay, fp16 autocast, gradient clipping at 1.0, seed 42. Pages
are split into 220-word windows so nothing is truncated. Word-level boxes
(box_mode: "word" is stored in config.json). FUNSD has no validation
split, so I fixed the 30-epoch schedule up front instead of picking the best
epoch on the test set. 13.6 minutes on one RTX A5000. Per-epoch loss and
metrics are in training_metrics.json.
Linker. A scikit-learn HistGradientBoostingClassifier (400 iterations,
learning rate 0.05, 31 leaves) over 21 geometric and text features per
candidate (question, answer) pair: offsets, overlaps, same-line, relative
rank, trailing colon, digits, lengths, distance. Trained on all 75,504
candidate pairs from the training forms (3,049 positive). The decision
threshold (0.35) was chosen by 5-fold GroupKFold cross-validation grouped
by form; CV pair F1 0.853. At inference I also keep only the best question
for each answer, which adds about 4 points end to end.
Usage
Tagger
from PIL import Image
from transformers import AutoProcessor, AutoModelForTokenClassification
repo = "mksl-ai/docuvision-layoutlmv3-large-funsd"
processor = AutoProcessor.from_pretrained(repo, apply_ocr=False) # bring your own OCR
model = AutoModelForTokenClassification.from_pretrained(repo).eval()
image = Image.open("form.png").convert("RGB")
# words: list[str] from your OCR, in reading order
# boxes: list[[x0, y0, x1, y1]] per word, normalised to 0-1000
enc = processor(image, words, boxes=boxes, return_tensors="pt", truncation=True)
pred = model(**enc).logits.argmax(-1)[0]
labels = [model.config.id2label[i.item()] for i in pred]
Linker
import joblib
from huggingface_hub import hf_hub_download
bundle = joblib.load(hf_hub_download(repo, "linker/kv-linker.joblib"))
clf, threshold, feature_names = bundle["model"], bundle["threshold"], bundle["features"]
The linker expects the 21 features built by pair_features in
docuvision/extraction/linker.py.
The easiest way to use both models together is DocuVision itself: download
this repo into models/layoutlmv3-funsd, put the joblib at
models/kv-linker.joblib, and run the GPU profile.
About the joblib file: it's a pickle containing a dict
{"model": HistGradientBoostingClassifier, "threshold": 0.35, "features": [...]}.
Only load pickles from sources you trust. It was saved with
scikit-learn 1.9.1 and joblib 1.6.0; other scikit-learn versions may
warn or fail to load it.
Environment I trained with: transformers 4.57.6, torch 2.6.0 (CUDA 12.4), scikit-learn 1.9.1.
Limitations
- English business forms only; FUNSD is 199 scanned forms, so layouts far from it (receipts, invoices with dense tables) will be weaker.
HEADERis the weak class (F1 0.63, only 127 test examples, and headers look a lot like questions).- The tagger needs good word boxes. On low-resolution scans, OCR errors dominate the final result.
License
The weights are CC BY-NC-SA 4.0, non-commercial: LayoutLMv3 is released under CC BY-NC-SA 4.0 and FUNSD is for non-commercial research use. The DocuVision code is MIT, but that doesn't extend to these weights.
Citation
If you use these, please cite the original work:
@inproceedings{huang2022layoutlmv3,
title = {LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking},
author = {Huang, Yupan and Lv, Tengchao and Cui, Lei and Lu, Yutong and Wei, Furu},
booktitle = {ACM Multimedia},
year = {2022}
}
@inproceedings{jaume2019funsd,
title = {FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents},
author = {Jaume, Guillaume and Ekenel, Hazim Kemal and Thiran, Jean-Philippe},
booktitle = {ICDAR-OST},
year = {2019}
}
- Downloads last month
- 10
Model tree for mksl-ai/docuvision-layoutlmv3-large-funsd
Base model
microsoft/layoutlmv3-largeDataset used to train mksl-ai/docuvision-layoutlmv3-large-funsd
Evaluation results
- Entity F1 (seqeval) on FUNSDtest set self-reported0.830
- Precision on FUNSDtest set self-reported0.816
- Recall on FUNSDtest set self-reported0.844