Token Classification
Transformers
ONNX
Safetensors
Polish
bert
named-entity-recognition
ner
pii
pii-detection
anonymization
privacy
gdpr
polish
legal
legal-nlp
herbert
Eval Results (legacy)
Instructions to use lexedit/herbert-polish-legal-ner with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use lexedit/herbert-polish-legal-ner with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="lexedit/herbert-polish-legal-ner")# Load model directly from transformers import AutoTokenizer, AutoModelForTokenClassification tokenizer = AutoTokenizer.from_pretrained("lexedit/herbert-polish-legal-ner") model = AutoModelForTokenClassification.from_pretrained("lexedit/herbert-polish-legal-ner", device_map="auto") - Notebooks
- Google Colab
- Kaggle
|
Download examples/KNOWN_LIMITATIONS.md from lexedit/herbert-polish-legal-ner: direct link, hf CLI and curl.
- Browser
- Download file 2.23 kB
-
https://huggingface.co/lexedit/herbert-polish-legal-ner/resolve/main/examples/KNOWN_LIMITATIONS.md
- Command line
-
hf download hf://lexedit/herbert-polish-legal-ner/examples/KNOWN_LIMITATIONS.md
-
curl -L -o KNOWN_LIMITATIONS.md https://huggingface.co/lexedit/herbert-polish-legal-ner/resolve/main/examples/KNOWN_LIMITATIONS.md
2.23 kB
| # Known limitations — real failure cases | |
| This general-purpose variant is strongest on **clean digital** text. Its main | |
| weakness is **heavily OCR-garbled names** (scanned / photographed documents). | |
| > 👉 For scanned / OCR'd documents, use the OCR-robust sibling | |
| > **[`lexedit/herbert-polish-legal-ner-ocr`](https://huggingface.co/lexedit/herbert-polish-legal-ner-ocr)** | |
| > — it catches several of the cases below that this model misses. | |
| All inputs are **synthetic** — fictional names, **no real or evaluation data**. | |
| Outputs are the model's **actual** predictions (quantized ONNX, recall-first PER | |
| threshold 0.2). Reproduce with [`known_limitations.py`](known_limitations.py). | |
| ## Heavy OCR garble → name missed entirely | |
| When OCR destroys most of a name's recognizable structure, this variant can fail to | |
| flag it at all: | |
| | Input | Behind the garble | Model `PER` | | |
| |---|---|---| | |
| | `Pozwana ote tobodaaka wniosła sprzeciw.` | *Bożena Łobodzińska* | — (nothing) | | |
| | `Powód go Hea stawił się osobiście.` | *Igor Heliasz* | — (nothing) | | |
| | `Decyzję doręczono: ist yi daziak.` | *Justyna Idziak* | — (nothing) | | |
| | `Wniosek złożył wiar Żoją w dniu 5 maja.` | *Wiktor Żołądek* | — (nothing) | | |
| *(The OCR-robust variant catches part of the last one — `['Żoją']` — instead of | |
| missing it entirely.)* | |
| ## Partial detection → residue leak | |
| | Input | Behind the garble | Model `PER` | Problem | | |
| |---|---|---|---| | |
| | `WIESEAW KUE stawił się na rozprawie.` | *Wiesław Kuc* | `['WIE', 'AW KUE']` | split into two spans → the gap (`SE`) leaks | | |
| ## Other known behaviours | |
| - **Foreign / out-of-distribution names** (e.g. German, Dutch) are less reliable — | |
| the model is tuned for Polish. | |
| - **Boundaries** can occasionally over- or under-extend. | |
| ## Mitigations (recommended) | |
| 1. **Use the OCR-robust variant for scanned documents** (route by document type). | |
| 2. **Document-safety post-pass** — propagate masking from any clean mention of a | |
| surname to its other inflected / OCR-variant mentions in the same document. | |
| 3. **Checksum regex** for structured PII (PESEL, NIP, IBAN, …). | |
| 4. **Lower the threshold on noisy/scanned documents.** | |
| 5. **Human review** for high-stakes anonymisation. | |