herbert-polish-legal-ner-ocr / examples /KNOWN_LIMITATIONS.md
i008's picture
HerBERT Polish legal NER (OCR-robust): weights + ONNX + card + examples
8d8b211 verified
|
Raw History Blame Contribute Delete
2.74 kB

Known limitations — real failure cases

The model is recall-tuned and OCR-robust, but not perfect. Below are honest, reproducible failure cases so you know where to add safeguards.

  • All inputs are synthetic — fictional names, no real personal data and no evaluation/validation data is shared here.
  • Outputs are the model's actual predictions (quantized ONNX, recall-first PER threshold 0.2). Reproduce with known_limitations.py.
  • The leak that matters is identity-level: if any mention of a person is missed (or only partially caught), that person can be re-identified.

1. Heavy OCR garble → name missed entirely

When OCR destroys most of a name's recognizable structure, the model can fail to flag it at all. This is the model's main remaining weakness.

Input Behind the garble Model PER
Pozwana ote tobodaaka wniosła sprzeciw. Bożena Łobodzińska — (nothing)
Powód go Hea stawił się osobiście. Igor Heliasz — (nothing)
Decyzję doręczono: ist yi daziak. Justyna Idziak — (nothing)

2. Partial detection → residue leak

The model catches part of a garbled name but not all of it, so a fragment of the real name stays visible after masking.

Input Behind the garble Model PER Problem
Wniosek złożył wiar Żoją w dniu 5 maja. Wiktor Żołądek ['Żoją'] first name wiar left unmasked
WIESEAW KUE stawił się na rozprawie. Wiesław Kuc ['WIE', 'AW KUE'] split into two spans → the gap (SE) leaks

Other known behaviours

  • Foreign / out-of-distribution names (e.g. German, Dutch) are less reliable — the model is tuned for Polish.
  • Boundaries can occasionally over- or under-extend (e.g. an adjacent honorific or a stray punctuation token captured with the name).
  • On clean digital text this OCR-robust variant is slightly less precise than a non-OCR-augmented model (it is more eager on garbled-looking tokens).

Mitigations (recommended)

  1. Document-safety post-pass — once any clean mention of a surname is detected in a document, propagate masking to its other inflected / OCR-variant mentions (fuzzy match). This recovers many of the partial/residue cases above when a cleaner mention exists elsewhere in the same document.
  2. Checksum regex for structured PII (PESEL, NIP, IBAN, …) — deterministic and not subject to these OCR misses.
  3. Lower the threshold on noisy/scanned documents, or route them through OCR quality checks.
  4. Human review for high-stakes anonymisation — the model is a strong first pass, not a guarantee.