Token Classification
Transformers
ONNX
Safetensors
Polish
bert
named-entity-recognition
ner
pii
pii-detection
anonymization
privacy
gdpr
polish
legal
legal-nlp
ocr-robust
herbert
Eval Results (legacy)
Instructions to use lexedit/herbert-polish-legal-ner-ocr with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use lexedit/herbert-polish-legal-ner-ocr with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="lexedit/herbert-polish-legal-ner-ocr")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForTokenClassification tokenizer = AutoTokenizer.from_pretrained("lexedit/herbert-polish-legal-ner-ocr") model = AutoModelForTokenClassification.from_pretrained("lexedit/herbert-polish-legal-ner-ocr", device_map="auto") - Notebooks
- Google Colab
- Kaggle
HerBERT Polish legal NER (OCR-robust): weights + ONNX + card + examples
Browse files- README.md +4 -0
- examples/STRONG_CASES.md +50 -0
- examples/strong_cases.py +81 -0
README.md
CHANGED
|
@@ -38,6 +38,10 @@ diacritics, confusable characters, fragmented surnames).
|
|
| 38 |
— an interactive, **fully client-side** Polish legal-document anonymisation demo
|
| 39 |
(same anonymiser family; the text never leaves your browser).
|
| 40 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 41 |
## Labels (29, BIO scheme)
|
| 42 |
|
| 43 |
`PER` (person) · `ORG` (organisation) · `LOC` (private address/location) ·
|
|
|
|
| 38 |
— an interactive, **fully client-side** Polish legal-document anonymisation demo
|
| 39 |
(same anonymiser family; the text never leaves your browser).
|
| 40 |
|
| 41 |
+
**📂 Example cases:** [`examples/STRONG_CASES.md`](examples/STRONG_CASES.md) (hard
|
| 42 |
+
inputs it handles — OCR garble, declension, ALL-CAPS) ·
|
| 43 |
+
[`examples/KNOWN_LIMITATIONS.md`](examples/KNOWN_LIMITATIONS.md) (where it still fails).
|
| 44 |
+
|
| 45 |
## Labels (29, BIO scheme)
|
| 46 |
|
| 47 |
`PER` (person) · `ORG` (organisation) · `LOC` (private address/location) ·
|
examples/STRONG_CASES.md
ADDED
|
@@ -0,0 +1,50 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Where the model shines — hard cases handled well
|
| 2 |
+
|
| 3 |
+
The flip side of [`KNOWN_LIMITATIONS.md`](KNOWN_LIMITATIONS.md): difficult inputs
|
| 4 |
+
the model gets **right**. These are exactly the cases a plain dictionary / regex
|
| 5 |
+
or a non-OCR-augmented model tends to miss.
|
| 6 |
+
|
| 7 |
+
All inputs are synthetic (fictional names). Outputs are the model's **actual**
|
| 8 |
+
predictions (quantized ONNX, recall-first PER threshold 0.2). Reproduce with
|
| 9 |
+
[`strong_cases.py`](strong_cases.py).
|
| 10 |
+
|
| 11 |
+
## OCR-garbled names (the variant's headline strength)
|
| 12 |
+
|
| 13 |
+
Names corrupted by optical character recognition — still recognised:
|
| 14 |
+
|
| 15 |
+
| Input | Behind the garble | Model |
|
| 16 |
+
|---|---|---|
|
| 17 |
+
| `Pozwany Damian Zotadek wniósł odpowiedź na pozew.` | *Damian Żołądek* (lost diacritics) | `PER 'Damian Zotadek'` |
|
| 18 |
+
| `Z powództwa Jiga Bękowiki przeciwko spółce.` | *Jaga Bętkowski* (heavy garble) | `PER 'Jiga Bękowiki'` |
|
| 19 |
+
| `Pełnomocnik: Ptek Gostoiski, adwokat.` | *Patryk Gostomski* (signature line) | `PER 'Ptek Gostoiski'` |
|
| 20 |
+
|
| 21 |
+
## Polish inflection (declension)
|
| 22 |
+
|
| 23 |
+
Polish surnames change form by grammatical case; the model tags the full inflected
|
| 24 |
+
surface, not just the nominative:
|
| 25 |
+
|
| 26 |
+
| Input | Case | Model |
|
| 27 |
+
|---|---|---|
|
| 28 |
+
| `Pozew skierowano przeciwko Jerzemu Sekule.` | dative | `PER 'Jerzemu Sekule'` |
|
| 29 |
+
| `Sprawa dotyczy akt Wojciecha Michalika.` | genitive | `PER 'Wojciecha Michalika'` |
|
| 30 |
+
|
| 31 |
+
## Tricky surface forms
|
| 32 |
+
|
| 33 |
+
| Input | Why it's hard | Model |
|
| 34 |
+
|---|---|---|
|
| 35 |
+
| `Powódka Anna Nowak-Kowalska złożyła wniosek.` | hyphenated double surname → one span | `PER 'Anna Nowak-Kowalska'` |
|
| 36 |
+
| `Stawiła się ALEKSANDRA ŚWIDERSKA, pełnomocnik.` | ALL-CAPS header/signature style | `PER 'ALEKSANDRA ŚWIDERSKA'` |
|
| 37 |
+
| `Z poważaniem,` / `Iga Mełech` / `radca prawny` | short female first + rare surname in a low-cue signature | `PER 'Iga Mełech'` |
|
| 38 |
+
|
| 39 |
+
## Multiple types + public/private distinction
|
| 40 |
+
|
| 41 |
+
```
|
| 42 |
+
Jan Kowalski, zam. ul. Słoneczna 5 w Krakowie, PESEL 02070803628.
|
| 43 |
+
PER 'Jan Kowalski'
|
| 44 |
+
LOC 'ul. Słoneczna 5' ← private address (masked)
|
| 45 |
+
LOC_PUB 'Krakowie' ← public city (kept by default)
|
| 46 |
+
ID '02070803628' ← national id (PESEL)
|
| 47 |
+
```
|
| 48 |
+
|
| 49 |
+
The model keeps the public place (`Krakowie`) visible while masking the private
|
| 50 |
+
street address (`ul. Słoneczna 5`) — so the legal context survives anonymisation.
|
examples/strong_cases.py
ADDED
|
@@ -0,0 +1,81 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
#!/usr/bin/env python3
|
| 2 |
+
"""Hard cases the model handles WELL — the flip side of KNOWN_LIMITATIONS.md.
|
| 3 |
+
|
| 4 |
+
All inputs are synthetic (fictional names). Outputs are the model's actual
|
| 5 |
+
predictions (quantized ONNX, recall-first PER threshold 0.2). Reproduce:
|
| 6 |
+
python examples/strong_cases.py
|
| 7 |
+
"""
|
| 8 |
+
import json
|
| 9 |
+
from pathlib import Path
|
| 10 |
+
|
| 11 |
+
import numpy as np
|
| 12 |
+
import onnxruntime as ort
|
| 13 |
+
from transformers import AutoTokenizer
|
| 14 |
+
|
| 15 |
+
ROOT = Path(__file__).resolve().parent.parent
|
| 16 |
+
PER_THRESHOLD = 0.2
|
| 17 |
+
|
| 18 |
+
tok = AutoTokenizer.from_pretrained(str(ROOT))
|
| 19 |
+
cfg = json.load(open(ROOT / "config.json", encoding="utf-8"))
|
| 20 |
+
id2label = {int(k): v for k, v in cfg["id2label"].items()}
|
| 21 |
+
label2id = cfg["label2id"]
|
| 22 |
+
sess = ort.InferenceSession(str(ROOT / "onnx" / "model_quantized.onnx"))
|
| 23 |
+
in_names = {i.name for i in sess.get_inputs()}
|
| 24 |
+
|
| 25 |
+
|
| 26 |
+
def softmax(x):
|
| 27 |
+
e = np.exp(x - x.max(-1, keepdims=True)); return e / e.sum(-1, keepdims=True)
|
| 28 |
+
|
| 29 |
+
|
| 30 |
+
def predict(text):
|
| 31 |
+
enc = tok(text, return_offsets_mapping=True, return_tensors="np", truncation=True, max_length=512)
|
| 32 |
+
offs = enc["offset_mapping"][0]
|
| 33 |
+
feeds = {"input_ids": enc["input_ids"].astype(np.int64),
|
| 34 |
+
"attention_mask": enc["attention_mask"].astype(np.int64)}
|
| 35 |
+
if "token_type_ids" in in_names:
|
| 36 |
+
feeds["token_type_ids"] = np.zeros_like(enc["input_ids"], dtype=np.int64)
|
| 37 |
+
probs = softmax(sess.run(None, feeds)[0][0])
|
| 38 |
+
ids = probs.argmax(-1)
|
| 39 |
+
pb, pi = label2id["B-PER"], label2id["I-PER"]
|
| 40 |
+
spans, cur = [], None
|
| 41 |
+
for i, (s, e) in enumerate(offs):
|
| 42 |
+
if s == e:
|
| 43 |
+
if cur: spans.append(cur); cur = None
|
| 44 |
+
continue
|
| 45 |
+
lab = id2label[int(ids[i])]
|
| 46 |
+
if lab == "O" and probs[i, pb] + probs[i, pi] >= PER_THRESHOLD:
|
| 47 |
+
lab = "I-PER" if (cur and cur[2] == "PER") else "B-PER"
|
| 48 |
+
if lab == "O":
|
| 49 |
+
if cur: spans.append(cur); cur = None
|
| 50 |
+
continue
|
| 51 |
+
tag, et = lab.split("-", 1)
|
| 52 |
+
if tag == "B" or cur is None or cur[2] != et:
|
| 53 |
+
if cur: spans.append(cur)
|
| 54 |
+
cur = [int(s), int(e), et]
|
| 55 |
+
else:
|
| 56 |
+
cur[1] = int(e)
|
| 57 |
+
if cur: spans.append(cur)
|
| 58 |
+
return [{"type": t, "start": a, "end": b, "text": text[a:b]} for a, b, t in spans]
|
| 59 |
+
|
| 60 |
+
|
| 61 |
+
CASES = [
|
| 62 |
+
("Pozwany Damian Zotadek wniósł odpowiedź na pozew.", "OCR: missing diacritics (Żołądek)"),
|
| 63 |
+
("Z powództwa Jiga Bękowiki przeciwko spółce.", "OCR: heavily garbled name (Jaga Bętkowski)"),
|
| 64 |
+
("Pełnomocnik: Ptek Gostoiski, adwokat.", "OCR garble in a signature line (Patryk Gostomski)"),
|
| 65 |
+
("Pozew skierowano przeciwko Jerzemu Sekule.", "inflected (dative) Polish name"),
|
| 66 |
+
("Sprawa dotyczy akt Wojciecha Michalika.", "inflected (genitive) Polish name"),
|
| 67 |
+
("Powódka Anna Nowak-Kowalska złożyła wniosek.", "hyphenated double surname (single span)"),
|
| 68 |
+
("Stawiła się ALEKSANDRA ŚWIDERSKA, pełnomocnik.", "ALL-CAPS name (header/signature style)"),
|
| 69 |
+
("Z poważaniem,\nIga Mełech\nradca prawny", "short female first + rare surname in a signature"),
|
| 70 |
+
("Jan Kowalski, zam. ul. Słoneczna 5 w Krakowie, PESEL 02070803628.",
|
| 71 |
+
"private address (LOC) vs public city (LOC_PUB) + national ID"),
|
| 72 |
+
]
|
| 73 |
+
|
| 74 |
+
if __name__ == "__main__":
|
| 75 |
+
print(f"PER threshold = {PER_THRESHOLD}\n")
|
| 76 |
+
for text, note in CASES:
|
| 77 |
+
print(f"# {note}")
|
| 78 |
+
print(f" input : {text.replace(chr(10), ' / ')}")
|
| 79 |
+
for e in predict(text):
|
| 80 |
+
print(f" {e['type']:8} {e['text']!r}")
|
| 81 |
+
print()
|