Instructions to use lexedit/herbert-polish-legal-ner with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use lexedit/herbert-polish-legal-ner with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="lexedit/herbert-polish-legal-ner")# Load model directly from transformers import AutoTokenizer, AutoModelForTokenClassification tokenizer = AutoTokenizer.from_pretrained("lexedit/herbert-polish-legal-ner") model = AutoModelForTokenClassification.from_pretrained("lexedit/herbert-polish-legal-ner", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Download examples/KNOWN_LIMITATIONS.md from lexedit/herbert-polish-legal-ner: direct link, hf CLI and curl.
- Browser
- Download file 2.23 kB
-
https://huggingface.co/lexedit/herbert-polish-legal-ner/resolve/main/examples/KNOWN_LIMITATIONS.md
- Command line
-
hf download hf://lexedit/herbert-polish-legal-ner/examples/KNOWN_LIMITATIONS.md
-
curl -L -o KNOWN_LIMITATIONS.md https://huggingface.co/lexedit/herbert-polish-legal-ner/resolve/main/examples/KNOWN_LIMITATIONS.md
Known limitations — real failure cases
This general-purpose variant is strongest on clean digital text. Its main weakness is heavily OCR-garbled names (scanned / photographed documents).
👉 For scanned / OCR'd documents, use the OCR-robust sibling
lexedit/herbert-polish-legal-ner-ocr— it catches several of the cases below that this model misses.
All inputs are synthetic — fictional names, no real or evaluation data.
Outputs are the model's actual predictions (quantized ONNX, recall-first PER
threshold 0.2). Reproduce with known_limitations.py.
Heavy OCR garble → name missed entirely
When OCR destroys most of a name's recognizable structure, this variant can fail to flag it at all:
| Input | Behind the garble | Model PER |
|---|---|---|
Pozwana ote tobodaaka wniosła sprzeciw. |
Bożena Łobodzińska | — (nothing) |
Powód go Hea stawił się osobiście. |
Igor Heliasz | — (nothing) |
Decyzję doręczono: ist yi daziak. |
Justyna Idziak | — (nothing) |
Wniosek złożył wiar Żoją w dniu 5 maja. |
Wiktor Żołądek | — (nothing) |
(The OCR-robust variant catches part of the last one — ['Żoją'] — instead of
missing it entirely.)
Partial detection → residue leak
| Input | Behind the garble | Model PER |
Problem |
|---|---|---|---|
WIESEAW KUE stawił się na rozprawie. |
Wiesław Kuc | ['WIE', 'AW KUE'] |
split into two spans → the gap (SE) leaks |
Other known behaviours
- Foreign / out-of-distribution names (e.g. German, Dutch) are less reliable — the model is tuned for Polish.
- Boundaries can occasionally over- or under-extend.
Mitigations (recommended)
- Use the OCR-robust variant for scanned documents (route by document type).
- Document-safety post-pass — propagate masking from any clean mention of a surname to its other inflected / OCR-variant mentions in the same document.
- Checksum regex for structured PII (PESEL, NIP, IBAN, …).
- Lower the threshold on noisy/scanned documents.
- Human review for high-stakes anonymisation.