# Known limitations — real failure cases This general-purpose variant is strongest on **clean digital** text. Its main weakness is **heavily OCR-garbled names** (scanned / photographed documents). > 👉 For scanned / OCR'd documents, use the OCR-robust sibling > **[`lexedit/herbert-polish-legal-ner-ocr`](https://huggingface.co/lexedit/herbert-polish-legal-ner-ocr)** > — it catches several of the cases below that this model misses. All inputs are **synthetic** — fictional names, **no real or evaluation data**. Outputs are the model's **actual** predictions (quantized ONNX, recall-first PER threshold 0.2). Reproduce with [`known_limitations.py`](known_limitations.py). ## Heavy OCR garble → name missed entirely When OCR destroys most of a name's recognizable structure, this variant can fail to flag it at all: | Input | Behind the garble | Model `PER` | |---|---|---| | `Pozwana ote tobodaaka wniosła sprzeciw.` | *Bożena Łobodzińska* | — (nothing) | | `Powód go Hea stawił się osobiście.` | *Igor Heliasz* | — (nothing) | | `Decyzję doręczono: ist yi daziak.` | *Justyna Idziak* | — (nothing) | | `Wniosek złożył wiar Żoją w dniu 5 maja.` | *Wiktor Żołądek* | — (nothing) | *(The OCR-robust variant catches part of the last one — `['Żoją']` — instead of missing it entirely.)* ## Partial detection → residue leak | Input | Behind the garble | Model `PER` | Problem | |---|---|---|---| | `WIESEAW KUE stawił się na rozprawie.` | *Wiesław Kuc* | `['WIE', 'AW KUE']` | split into two spans → the gap (`SE`) leaks | ## Other known behaviours - **Foreign / out-of-distribution names** (e.g. German, Dutch) are less reliable — the model is tuned for Polish. - **Boundaries** can occasionally over- or under-extend. ## Mitigations (recommended) 1. **Use the OCR-robust variant for scanned documents** (route by document type). 2. **Document-safety post-pass** — propagate masking from any clean mention of a surname to its other inflected / OCR-variant mentions in the same document. 3. **Checksum regex** for structured PII (PESEL, NIP, IBAN, …). 4. **Lower the threshold on noisy/scanned documents.** 5. **Human review** for high-stakes anonymisation.