herbert-polish-legal-ner / examples /KNOWN_LIMITATIONS.md
i008's picture
HerBERT Polish legal NER (general): weights + ONNX + card + examples
399294e verified
|
Raw History Blame Contribute Delete
2.23 kB
# Known limitations — real failure cases
This general-purpose variant is strongest on **clean digital** text. Its main
weakness is **heavily OCR-garbled names** (scanned / photographed documents).
> 👉 For scanned / OCR'd documents, use the OCR-robust sibling
> **[`lexedit/herbert-polish-legal-ner-ocr`](https://huggingface.co/lexedit/herbert-polish-legal-ner-ocr)**
> — it catches several of the cases below that this model misses.
All inputs are **synthetic** — fictional names, **no real or evaluation data**.
Outputs are the model's **actual** predictions (quantized ONNX, recall-first PER
threshold 0.2). Reproduce with [`known_limitations.py`](known_limitations.py).
## Heavy OCR garble → name missed entirely
When OCR destroys most of a name's recognizable structure, this variant can fail to
flag it at all:
| Input | Behind the garble | Model `PER` |
|---|---|---|
| `Pozwana ote tobodaaka wniosła sprzeciw.` | *Bożena Łobodzińska* | — (nothing) |
| `Powód go Hea stawił się osobiście.` | *Igor Heliasz* | — (nothing) |
| `Decyzję doręczono: ist yi daziak.` | *Justyna Idziak* | — (nothing) |
| `Wniosek złożył wiar Żoją w dniu 5 maja.` | *Wiktor Żołądek* | — (nothing) |
*(The OCR-robust variant catches part of the last one — `['Żoją']` — instead of
missing it entirely.)*
## Partial detection → residue leak
| Input | Behind the garble | Model `PER` | Problem |
|---|---|---|---|
| `WIESEAW KUE stawił się na rozprawie.` | *Wiesław Kuc* | `['WIE', 'AW KUE']` | split into two spans → the gap (`SE`) leaks |
## Other known behaviours
- **Foreign / out-of-distribution names** (e.g. German, Dutch) are less reliable —
the model is tuned for Polish.
- **Boundaries** can occasionally over- or under-extend.
## Mitigations (recommended)
1. **Use the OCR-robust variant for scanned documents** (route by document type).
2. **Document-safety post-pass** — propagate masking from any clean mention of a
surname to its other inflected / OCR-variant mentions in the same document.
3. **Checksum regex** for structured PII (PESEL, NIP, IBAN, …).
4. **Lower the threshold on noisy/scanned documents.**
5. **Human review** for high-stakes anonymisation.