herbert-polish-legal-ner / examples /KNOWN_LIMITATIONS.md
i008's picture
HerBERT Polish legal NER (general): weights + ONNX + card + examples
399294e verified
|
Raw History Blame Contribute Delete
2.23 kB

Known limitations — real failure cases

This general-purpose variant is strongest on clean digital text. Its main weakness is heavily OCR-garbled names (scanned / photographed documents).

👉 For scanned / OCR'd documents, use the OCR-robust sibling lexedit/herbert-polish-legal-ner-ocr — it catches several of the cases below that this model misses.

All inputs are synthetic — fictional names, no real or evaluation data. Outputs are the model's actual predictions (quantized ONNX, recall-first PER threshold 0.2). Reproduce with known_limitations.py.

Heavy OCR garble → name missed entirely

When OCR destroys most of a name's recognizable structure, this variant can fail to flag it at all:

Input Behind the garble Model PER
Pozwana ote tobodaaka wniosła sprzeciw. Bożena Łobodzińska — (nothing)
Powód go Hea stawił się osobiście. Igor Heliasz — (nothing)
Decyzję doręczono: ist yi daziak. Justyna Idziak — (nothing)
Wniosek złożył wiar Żoją w dniu 5 maja. Wiktor Żołądek — (nothing)

(The OCR-robust variant catches part of the last one — ['Żoją'] — instead of missing it entirely.)

Partial detection → residue leak

Input Behind the garble Model PER Problem
WIESEAW KUE stawił się na rozprawie. Wiesław Kuc ['WIE', 'AW KUE'] split into two spans → the gap (SE) leaks

Other known behaviours

  • Foreign / out-of-distribution names (e.g. German, Dutch) are less reliable — the model is tuned for Polish.
  • Boundaries can occasionally over- or under-extend.

Mitigations (recommended)

  1. Use the OCR-robust variant for scanned documents (route by document type).
  2. Document-safety post-pass — propagate masking from any clean mention of a surname to its other inflected / OCR-variant mentions in the same document.
  3. Checksum regex for structured PII (PESEL, NIP, IBAN, …).
  4. Lower the threshold on noisy/scanned documents.
  5. Human review for high-stakes anonymisation.