# Known limitations — real failure cases The model is **recall-tuned and OCR-robust, but not perfect**. Below are honest, reproducible failure cases so you know where to add safeguards. - All inputs are **synthetic** — fictional names, **no real personal data** and **no evaluation/validation data** is shared here. - Outputs are the model's **actual predictions** (quantized ONNX, recall-first PER threshold 0.2). Reproduce with [`known_limitations.py`](known_limitations.py). - The leak that matters is **identity-level**: if *any* mention of a person is missed (or only partially caught), that person can be re-identified. ## 1. Heavy OCR garble → name missed entirely When OCR destroys most of a name's recognizable structure, the model can fail to flag it at all. This is the model's **main remaining weakness**. | Input | Behind the garble | Model `PER` | |---|---|---| | `Pozwana ote tobodaaka wniosła sprzeciw.` | *Bożena Łobodzińska* | — (nothing) | | `Powód go Hea stawił się osobiście.` | *Igor Heliasz* | — (nothing) | | `Decyzję doręczono: ist yi daziak.` | *Justyna Idziak* | — (nothing) | ## 2. Partial detection → residue leak The model catches part of a garbled name but not all of it, so a fragment of the real name stays visible after masking. | Input | Behind the garble | Model `PER` | Problem | |---|---|---|---| | `Wniosek złożył wiar Żoją w dniu 5 maja.` | *Wiktor Żołądek* | `['Żoją']` | first name `wiar` left unmasked | | `WIESEAW KUE stawił się na rozprawie.` | *Wiesław Kuc* | `['WIE', 'AW KUE']` | split into two spans → the gap (`SE`) leaks | ## Other known behaviours - **Foreign / out-of-distribution names** (e.g. German, Dutch) are less reliable — the model is tuned for Polish. - **Boundaries** can occasionally over- or under-extend (e.g. an adjacent honorific or a stray punctuation token captured with the name). - On **clean digital text** this OCR-robust variant is slightly less precise than a non-OCR-augmented model (it is more eager on garbled-looking tokens). ## Mitigations (recommended) 1. **Document-safety post-pass** — once any clean mention of a surname is detected in a document, propagate masking to its other inflected / OCR-variant mentions (fuzzy match). This recovers many of the partial/residue cases above when a cleaner mention exists elsewhere in the same document. 2. **Checksum regex** for structured PII (PESEL, NIP, IBAN, …) — deterministic and not subject to these OCR misses. 3. **Lower the threshold on noisy/scanned documents**, or route them through OCR quality checks. 4. **Human review** for high-stakes anonymisation — the model is a strong first pass, not a guarantee.