Token Classification
Transformers
ONNX
Safetensors
Polish
bert
named-entity-recognition
ner
pii
pii-detection
anonymization
privacy
gdpr
polish
legal
legal-nlp
ocr-robust
herbert
Eval Results (legacy)
Instructions to use lexedit/herbert-polish-legal-ner-ocr with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use lexedit/herbert-polish-legal-ner-ocr with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="lexedit/herbert-polish-legal-ner-ocr")# Load model directly from transformers import AutoTokenizer, AutoModelForTokenClassification tokenizer = AutoTokenizer.from_pretrained("lexedit/herbert-polish-legal-ner-ocr") model = AutoModelForTokenClassification.from_pretrained("lexedit/herbert-polish-legal-ner-ocr", device_map="auto") - Notebooks
- Google Colab
- Kaggle
|
Download examples/KNOWN_LIMITATIONS.md from lexedit/herbert-polish-legal-ner-ocr: direct link, hf CLI and curl.
- Browser
- Download file 2.74 kB
-
https://huggingface.co/lexedit/herbert-polish-legal-ner-ocr/resolve/main/examples/KNOWN_LIMITATIONS.md
- Command line
-
hf download hf://lexedit/herbert-polish-legal-ner-ocr/examples/KNOWN_LIMITATIONS.md
-
curl -L -o KNOWN_LIMITATIONS.md https://huggingface.co/lexedit/herbert-polish-legal-ner-ocr/resolve/main/examples/KNOWN_LIMITATIONS.md
2.74 kB
Known limitations — real failure cases
The model is recall-tuned and OCR-robust, but not perfect. Below are honest, reproducible failure cases so you know where to add safeguards.
- All inputs are synthetic — fictional names, no real personal data and no evaluation/validation data is shared here.
- Outputs are the model's actual predictions (quantized ONNX, recall-first PER
threshold 0.2). Reproduce with
known_limitations.py. - The leak that matters is identity-level: if any mention of a person is missed (or only partially caught), that person can be re-identified.
1. Heavy OCR garble → name missed entirely
When OCR destroys most of a name's recognizable structure, the model can fail to flag it at all. This is the model's main remaining weakness.
| Input | Behind the garble | Model PER |
|---|---|---|
Pozwana ote tobodaaka wniosła sprzeciw. |
Bożena Łobodzińska | — (nothing) |
Powód go Hea stawił się osobiście. |
Igor Heliasz | — (nothing) |
Decyzję doręczono: ist yi daziak. |
Justyna Idziak | — (nothing) |
2. Partial detection → residue leak
The model catches part of a garbled name but not all of it, so a fragment of the real name stays visible after masking.
| Input | Behind the garble | Model PER |
Problem |
|---|---|---|---|
Wniosek złożył wiar Żoją w dniu 5 maja. |
Wiktor Żołądek | ['Żoją'] |
first name wiar left unmasked |
WIESEAW KUE stawił się na rozprawie. |
Wiesław Kuc | ['WIE', 'AW KUE'] |
split into two spans → the gap (SE) leaks |
Other known behaviours
- Foreign / out-of-distribution names (e.g. German, Dutch) are less reliable — the model is tuned for Polish.
- Boundaries can occasionally over- or under-extend (e.g. an adjacent honorific or a stray punctuation token captured with the name).
- On clean digital text this OCR-robust variant is slightly less precise than a non-OCR-augmented model (it is more eager on garbled-looking tokens).
Mitigations (recommended)
- Document-safety post-pass — once any clean mention of a surname is detected in a document, propagate masking to its other inflected / OCR-variant mentions (fuzzy match). This recovers many of the partial/residue cases above when a cleaner mention exists elsewhere in the same document.
- Checksum regex for structured PII (PESEL, NIP, IBAN, …) — deterministic and not subject to these OCR misses.
- Lower the threshold on noisy/scanned documents, or route them through OCR quality checks.
- Human review for high-stakes anonymisation — the model is a strong first pass, not a guarantee.