i008 commited on
Commit
eb952ad
·
verified ·
1 Parent(s): 984dca7

HerBERT Polish legal NER (OCR-robust): weights + ONNX + card + examples

Browse files
Files changed (3) hide show
  1. README.md +4 -0
  2. examples/STRONG_CASES.md +50 -0
  3. examples/strong_cases.py +81 -0
README.md CHANGED
@@ -38,6 +38,10 @@ diacritics, confusable characters, fragmented surnames).
38
  — an interactive, **fully client-side** Polish legal-document anonymisation demo
39
  (same anonymiser family; the text never leaves your browser).
40
 
 
 
 
 
41
  ## Labels (29, BIO scheme)
42
 
43
  `PER` (person) · `ORG` (organisation) · `LOC` (private address/location) ·
 
38
  — an interactive, **fully client-side** Polish legal-document anonymisation demo
39
  (same anonymiser family; the text never leaves your browser).
40
 
41
+ **📂 Example cases:** [`examples/STRONG_CASES.md`](examples/STRONG_CASES.md) (hard
42
+ inputs it handles — OCR garble, declension, ALL-CAPS) ·
43
+ [`examples/KNOWN_LIMITATIONS.md`](examples/KNOWN_LIMITATIONS.md) (where it still fails).
44
+
45
  ## Labels (29, BIO scheme)
46
 
47
  `PER` (person) · `ORG` (organisation) · `LOC` (private address/location) ·
examples/STRONG_CASES.md ADDED
@@ -0,0 +1,50 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Where the model shines — hard cases handled well
2
+
3
+ The flip side of [`KNOWN_LIMITATIONS.md`](KNOWN_LIMITATIONS.md): difficult inputs
4
+ the model gets **right**. These are exactly the cases a plain dictionary / regex
5
+ or a non-OCR-augmented model tends to miss.
6
+
7
+ All inputs are synthetic (fictional names). Outputs are the model's **actual**
8
+ predictions (quantized ONNX, recall-first PER threshold 0.2). Reproduce with
9
+ [`strong_cases.py`](strong_cases.py).
10
+
11
+ ## OCR-garbled names (the variant's headline strength)
12
+
13
+ Names corrupted by optical character recognition — still recognised:
14
+
15
+ | Input | Behind the garble | Model |
16
+ |---|---|---|
17
+ | `Pozwany Damian Zotadek wniósł odpowiedź na pozew.` | *Damian Żołądek* (lost diacritics) | `PER 'Damian Zotadek'` |
18
+ | `Z powództwa Jiga Bękowiki przeciwko spółce.` | *Jaga Bętkowski* (heavy garble) | `PER 'Jiga Bękowiki'` |
19
+ | `Pełnomocnik: Ptek Gostoiski, adwokat.` | *Patryk Gostomski* (signature line) | `PER 'Ptek Gostoiski'` |
20
+
21
+ ## Polish inflection (declension)
22
+
23
+ Polish surnames change form by grammatical case; the model tags the full inflected
24
+ surface, not just the nominative:
25
+
26
+ | Input | Case | Model |
27
+ |---|---|---|
28
+ | `Pozew skierowano przeciwko Jerzemu Sekule.` | dative | `PER 'Jerzemu Sekule'` |
29
+ | `Sprawa dotyczy akt Wojciecha Michalika.` | genitive | `PER 'Wojciecha Michalika'` |
30
+
31
+ ## Tricky surface forms
32
+
33
+ | Input | Why it's hard | Model |
34
+ |---|---|---|
35
+ | `Powódka Anna Nowak-Kowalska złożyła wniosek.` | hyphenated double surname → one span | `PER 'Anna Nowak-Kowalska'` |
36
+ | `Stawiła się ALEKSANDRA ŚWIDERSKA, pełnomocnik.` | ALL-CAPS header/signature style | `PER 'ALEKSANDRA ŚWIDERSKA'` |
37
+ | `Z poważaniem,` / `Iga Mełech` / `radca prawny` | short female first + rare surname in a low-cue signature | `PER 'Iga Mełech'` |
38
+
39
+ ## Multiple types + public/private distinction
40
+
41
+ ```
42
+ Jan Kowalski, zam. ul. Słoneczna 5 w Krakowie, PESEL 02070803628.
43
+ PER 'Jan Kowalski'
44
+ LOC 'ul. Słoneczna 5' ← private address (masked)
45
+ LOC_PUB 'Krakowie' ← public city (kept by default)
46
+ ID '02070803628' ← national id (PESEL)
47
+ ```
48
+
49
+ The model keeps the public place (`Krakowie`) visible while masking the private
50
+ street address (`ul. Słoneczna 5`) — so the legal context survives anonymisation.
examples/strong_cases.py ADDED
@@ -0,0 +1,81 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env python3
2
+ """Hard cases the model handles WELL — the flip side of KNOWN_LIMITATIONS.md.
3
+
4
+ All inputs are synthetic (fictional names). Outputs are the model's actual
5
+ predictions (quantized ONNX, recall-first PER threshold 0.2). Reproduce:
6
+ python examples/strong_cases.py
7
+ """
8
+ import json
9
+ from pathlib import Path
10
+
11
+ import numpy as np
12
+ import onnxruntime as ort
13
+ from transformers import AutoTokenizer
14
+
15
+ ROOT = Path(__file__).resolve().parent.parent
16
+ PER_THRESHOLD = 0.2
17
+
18
+ tok = AutoTokenizer.from_pretrained(str(ROOT))
19
+ cfg = json.load(open(ROOT / "config.json", encoding="utf-8"))
20
+ id2label = {int(k): v for k, v in cfg["id2label"].items()}
21
+ label2id = cfg["label2id"]
22
+ sess = ort.InferenceSession(str(ROOT / "onnx" / "model_quantized.onnx"))
23
+ in_names = {i.name for i in sess.get_inputs()}
24
+
25
+
26
+ def softmax(x):
27
+ e = np.exp(x - x.max(-1, keepdims=True)); return e / e.sum(-1, keepdims=True)
28
+
29
+
30
+ def predict(text):
31
+ enc = tok(text, return_offsets_mapping=True, return_tensors="np", truncation=True, max_length=512)
32
+ offs = enc["offset_mapping"][0]
33
+ feeds = {"input_ids": enc["input_ids"].astype(np.int64),
34
+ "attention_mask": enc["attention_mask"].astype(np.int64)}
35
+ if "token_type_ids" in in_names:
36
+ feeds["token_type_ids"] = np.zeros_like(enc["input_ids"], dtype=np.int64)
37
+ probs = softmax(sess.run(None, feeds)[0][0])
38
+ ids = probs.argmax(-1)
39
+ pb, pi = label2id["B-PER"], label2id["I-PER"]
40
+ spans, cur = [], None
41
+ for i, (s, e) in enumerate(offs):
42
+ if s == e:
43
+ if cur: spans.append(cur); cur = None
44
+ continue
45
+ lab = id2label[int(ids[i])]
46
+ if lab == "O" and probs[i, pb] + probs[i, pi] >= PER_THRESHOLD:
47
+ lab = "I-PER" if (cur and cur[2] == "PER") else "B-PER"
48
+ if lab == "O":
49
+ if cur: spans.append(cur); cur = None
50
+ continue
51
+ tag, et = lab.split("-", 1)
52
+ if tag == "B" or cur is None or cur[2] != et:
53
+ if cur: spans.append(cur)
54
+ cur = [int(s), int(e), et]
55
+ else:
56
+ cur[1] = int(e)
57
+ if cur: spans.append(cur)
58
+ return [{"type": t, "start": a, "end": b, "text": text[a:b]} for a, b, t in spans]
59
+
60
+
61
+ CASES = [
62
+ ("Pozwany Damian Zotadek wniósł odpowiedź na pozew.", "OCR: missing diacritics (Żołądek)"),
63
+ ("Z powództwa Jiga Bękowiki przeciwko spółce.", "OCR: heavily garbled name (Jaga Bętkowski)"),
64
+ ("Pełnomocnik: Ptek Gostoiski, adwokat.", "OCR garble in a signature line (Patryk Gostomski)"),
65
+ ("Pozew skierowano przeciwko Jerzemu Sekule.", "inflected (dative) Polish name"),
66
+ ("Sprawa dotyczy akt Wojciecha Michalika.", "inflected (genitive) Polish name"),
67
+ ("Powódka Anna Nowak-Kowalska złożyła wniosek.", "hyphenated double surname (single span)"),
68
+ ("Stawiła się ALEKSANDRA ŚWIDERSKA, pełnomocnik.", "ALL-CAPS name (header/signature style)"),
69
+ ("Z poważaniem,\nIga Mełech\nradca prawny", "short female first + rare surname in a signature"),
70
+ ("Jan Kowalski, zam. ul. Słoneczna 5 w Krakowie, PESEL 02070803628.",
71
+ "private address (LOC) vs public city (LOC_PUB) + national ID"),
72
+ ]
73
+
74
+ if __name__ == "__main__":
75
+ print(f"PER threshold = {PER_THRESHOLD}\n")
76
+ for text, note in CASES:
77
+ print(f"# {note}")
78
+ print(f" input : {text.replace(chr(10), ' / ')}")
79
+ for e in predict(text):
80
+ print(f" {e['type']:8} {e['text']!r}")
81
+ print()