i008 commited on
Commit
399294e
·
verified ·
1 Parent(s): 9d68e1b

HerBERT Polish legal NER (general): weights + ONNX + card + examples

Browse files
.gitattributes CHANGED
@@ -1,35 +1,3 @@
1
- *.7z filter=lfs diff=lfs merge=lfs -text
2
- *.arrow filter=lfs diff=lfs merge=lfs -text
3
- *.bin filter=lfs diff=lfs merge=lfs -text
4
- *.bz2 filter=lfs diff=lfs merge=lfs -text
5
- *.ckpt filter=lfs diff=lfs merge=lfs -text
6
- *.ftz filter=lfs diff=lfs merge=lfs -text
7
- *.gz filter=lfs diff=lfs merge=lfs -text
8
- *.h5 filter=lfs diff=lfs merge=lfs -text
9
- *.joblib filter=lfs diff=lfs merge=lfs -text
10
- *.lfs.* filter=lfs diff=lfs merge=lfs -text
11
- *.mlmodel filter=lfs diff=lfs merge=lfs -text
12
- *.model filter=lfs diff=lfs merge=lfs -text
13
- *.msgpack filter=lfs diff=lfs merge=lfs -text
14
- *.npy filter=lfs diff=lfs merge=lfs -text
15
- *.npz filter=lfs diff=lfs merge=lfs -text
16
- *.onnx filter=lfs diff=lfs merge=lfs -text
17
- *.ot filter=lfs diff=lfs merge=lfs -text
18
- *.parquet filter=lfs diff=lfs merge=lfs -text
19
- *.pb filter=lfs diff=lfs merge=lfs -text
20
- *.pickle filter=lfs diff=lfs merge=lfs -text
21
- *.pkl filter=lfs diff=lfs merge=lfs -text
22
- *.pt filter=lfs diff=lfs merge=lfs -text
23
- *.pth filter=lfs diff=lfs merge=lfs -text
24
- *.rar filter=lfs diff=lfs merge=lfs -text
25
  *.safetensors filter=lfs diff=lfs merge=lfs -text
26
- saved_model/**/* filter=lfs diff=lfs merge=lfs -text
27
- *.tar.* filter=lfs diff=lfs merge=lfs -text
28
- *.tar filter=lfs diff=lfs merge=lfs -text
29
- *.tflite filter=lfs diff=lfs merge=lfs -text
30
- *.tgz filter=lfs diff=lfs merge=lfs -text
31
- *.wasm filter=lfs diff=lfs merge=lfs -text
32
- *.xz filter=lfs diff=lfs merge=lfs -text
33
- *.zip filter=lfs diff=lfs merge=lfs -text
34
- *.zst filter=lfs diff=lfs merge=lfs -text
35
- *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  *.safetensors filter=lfs diff=lfs merge=lfs -text
2
+ *.onnx filter=lfs diff=lfs merge=lfs -text
3
+ *.bin filter=lfs diff=lfs merge=lfs -text
 
 
 
 
 
 
 
 
README.md ADDED
@@ -0,0 +1,210 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language:
3
+ - pl
4
+ license: cc-by-4.0
5
+ base_model: allegro/herbert-base-cased
6
+ base_model_relation: finetune
7
+ pipeline_tag: token-classification
8
+ library_name: transformers
9
+ metrics:
10
+ - f1
11
+ tags:
12
+ - token-classification
13
+ - named-entity-recognition
14
+ - ner
15
+ - pii
16
+ - pii-detection
17
+ - anonymization
18
+ - privacy
19
+ - gdpr
20
+ - polish
21
+ - legal
22
+ - legal-nlp
23
+ - herbert
24
+ - bert
25
+ - onnx
26
+ model-index:
27
+ - name: herbert-polish-legal-ner
28
+ results:
29
+ - task:
30
+ type: token-classification
31
+ name: Named Entity Recognition (PII)
32
+ dataset:
33
+ type: internal-polish-legal-eval
34
+ name: Internal Polish legal documents (identity-level eval)
35
+ metrics:
36
+ - type: f1
37
+ value: 0.94
38
+ name: Token-level F1 (held-out test)
39
+ widget:
40
+ - text: "Pozwany Jan Kowalski, zam. ul. Słoneczna 5 w Krakowie, PESEL 02070803628."
41
+ ---
42
+
43
+ # HerBERT Polish Legal NER — PII / Anonymization
44
+
45
+ A Polish token-classification (NER) model for **detecting personally identifiable
46
+ information (PII) in legal and administrative text**, fine-tuned from
47
+ [`allegro/herbert-base-cased`](https://huggingface.co/allegro/herbert-base-cased).
48
+
49
+ This is the **general-purpose** variant — the best overall accuracy on **clean,
50
+ digital** Polish legal documents.
51
+
52
+ > For **scanned / photographed (OCR'd)** documents, use the OCR-robust sibling
53
+ > **[`lexedit/herbert-polish-legal-ner-ocr`](https://huggingface.co/lexedit/herbert-polish-legal-ner-ocr)**,
54
+ > which cuts the person-name leak on scanned text by ~35–41 % (at a small precision
55
+ > cost on clean text).
56
+
57
+ > **Intended use is defensive:** flagging PII so it can be masked / anonymised
58
+ > before a document is shared or processed. It is **not** a guarantee of complete
59
+ > anonymisation — see *Limitations*.
60
+
61
+ **▶ Try it in your browser:** [lexedit.ai/lexedit-privacy](https://lexedit.ai/lexedit-privacy)
62
+ — an interactive, **fully client-side** Polish legal-document anonymisation demo
63
+ (the text never leaves your browser).
64
+
65
+ **📂 Example cases:** [`examples/STRONG_CASES.md`](examples/STRONG_CASES.md) (hard
66
+ inputs it handles) · [`examples/KNOWN_LIMITATIONS.md`](examples/KNOWN_LIMITATIONS.md)
67
+ (where it still fails).
68
+
69
+ ## Labels (29, BIO scheme)
70
+
71
+ `PER` (person) · `ORG` (organisation) · `LOC` (private address/location) ·
72
+ `LOC_PUB` (public place: city, country) · `DATE` · `MONEY` · `EMAIL` · `PHONE` ·
73
+ `ID` (national id / case / document number) · `IBAN` · `DIAGNOSIS` ·
74
+ `HEALTH_FACILITY` · `MEDICAL_ID` · `WATERMARK`, each as `B-…` / `I-…`, plus `O`.
75
+
76
+ The model distinguishes **private** locations (`LOC`, masked) from **public**
77
+ ones (`LOC_PUB`, usually kept), and treats `DATE` / `MONEY` as non-anonymised by
78
+ default.
79
+
80
+ ## Usage (ONNX, no PyTorch required)
81
+
82
+ ```bash
83
+ pip install onnxruntime transformers numpy
84
+ python examples/inference_onnx.py
85
+ ```
86
+
87
+ ```python
88
+ import numpy as np, onnxruntime as ort
89
+ from transformers import AutoTokenizer
90
+
91
+ tok = AutoTokenizer.from_pretrained("lexedit/herbert-polish-legal-ner")
92
+ sess = ort.InferenceSession("onnx/model_quantized.onnx")
93
+ enc = tok("Pozwany Jan Kowalski, PESEL 02070803628.",
94
+ return_offsets_mapping=True, return_tensors="np")
95
+ feeds = {"input_ids": enc["input_ids"].astype(np.int64),
96
+ "attention_mask": enc["attention_mask"].astype(np.int64)}
97
+ logits = sess.run(None, feeds)[0] # (1, seq, 29)
98
+ # argmax per token -> map ids via config.id2label -> group B-/I- with offset_mapping
99
+ ```
100
+
101
+ ## Usage (PyTorch / transformers pipeline)
102
+
103
+ ```python
104
+ from transformers import pipeline
105
+ ner = pipeline("token-classification",
106
+ model="lexedit/herbert-polish-legal-ner",
107
+ aggregation_strategy="first")
108
+ ner("Pozwany Jan Kowalski, zam. ul. Słoneczna 5 w Krakowie.")
109
+ ```
110
+
111
+ ## Runs in the browser (the intended setup)
112
+
113
+ This model is designed to run **entirely client-side**. The quantized ONNX
114
+ (~125 MB, int8) loads in the browser via
115
+ [onnxruntime-web](https://onnxruntime.ai/docs/tutorials/web/) or
116
+ [transformers.js](https://huggingface.co/docs/transformers.js) (WASM), so **the
117
+ document never leaves the user's device** — which is the whole point for sensitive
118
+ legal / medical text. It also runs anywhere ONNX Runtime does (Python, Node,
119
+ server, mobile). The demo at
120
+ [lexedit.ai/lexedit-privacy](https://lexedit.ai/lexedit-privacy) is exactly this.
121
+
122
+ ### Speed (rough)
123
+
124
+ Quantized ONNX, measured on a laptop CPU (Apple Silicon):
125
+
126
+ | input | 1 core | all cores |
127
+ |---|---|---|
128
+ | short sentence (~30 tokens) | ~18 ms | ~13 ms |
129
+ | full chunk (~500 tokens) | ~0.2 s | ~0.1 s |
130
+
131
+ ≈ **4–5 chunks/second single-threaded** natively. In the browser (WASM,
132
+ single-threaded) it is slower but practical: short text stays interactive, a 1–2
133
+ page document takes a few seconds, a large scanned document can take ~a minute or
134
+ two; cross-origin isolation (COOP/COEP → multi-threaded WASM) speeds it up.
135
+
136
+ ## Recommended production setup
137
+
138
+ This model is **recall-first** and is one layer of a pipeline, not the whole
139
+ solution. For best anonymisation, pair it with:
140
+
141
+ 1. **A recall-first threshold** for `PER` (flip a token to PER when the summed PER
142
+ probability ≥ ~0.2, even if it is not the arg-max).
143
+ 2. **A deterministic document-safety post-pass** — snap spans to whole words, merge
144
+ hyphenated surnames, and propagate a detected surname to its other inflected /
145
+ OCR-variant mentions across the document.
146
+ 3. **Checksum-validated regex** for structured PII (PESEL, NIP, REGON, IBAN, …).
147
+ 4. **Human review** for high-stakes use.
148
+
149
+ ## Evaluation
150
+
151
+ Identity-level **leak rate** = a person is "leaked" if *any* mention of them is
152
+ missed. Internal set of 50 real Polish legal documents (179 persons), recall-first
153
+ threshold 0.2.
154
+
155
+ | Document type | metric | this model | OCR-robust variant |
156
+ |---|---|---|---|
157
+ | **Clean / digital** | leak (threshold) | **10.1%** | 11.7% |
158
+ | **Clean / digital** | leak (+ post-pass) | **7.3%** | 7.3% |
159
+ | Scanned / OCR'd | leak (threshold) | 31.8% | **20.7%** |
160
+ | Scanned / OCR'd | leak (+ post-pass) | 25.7% | **15.1%** |
161
+
162
+ Token-level test F1 ≈ **0.94**.
163
+
164
+ **Takeaway:** this is the strongest variant on **clean digital** text (lowest leak,
165
+ best precision). On **scanned / OCR'd** text it is weaker — there the
166
+ [OCR-robust variant](https://huggingface.co/lexedit/herbert-polish-legal-ner-ocr)
167
+ wins. If you process both, **route by document type**.
168
+
169
+ ## Training data
170
+
171
+ Fine-tuned on Polish legal and administrative documents — a mix of document
172
+ templates, programmatically-generated labelled examples (valid checksum-correct
173
+ synthetic identifiers, rule-based Polish name declension), and real-world
174
+ legal-document samples.
175
+
176
+ No raw personal data is distributed with this model. Because this is a
177
+ **token-classification** model (it outputs a label per input token and cannot
178
+ generate text), the weights do not reproduce or expose any training document.
179
+
180
+ ## Limitations
181
+
182
+ > **Concrete, reproducible failure cases:** see
183
+ > [`examples/KNOWN_LIMITATIONS.md`](examples/KNOWN_LIMITATIONS.md) (synthetic inputs).
184
+ > Heavily OCR-garbled names are the main weakness — for scanned documents prefer the
185
+ > [OCR-robust variant](https://huggingface.co/lexedit/herbert-polish-legal-ner-ocr).
186
+
187
+ - **Polish only.**
188
+ - **Not a guarantee.** A residual leak rate remains (≈7–10 % identity-level on clean
189
+ text); always combine with the deterministic post-pass + checksum regex and human
190
+ review for high-stakes use.
191
+ - **Scanned / OCR'd text** is the weak spot of this variant (heavily garbled names
192
+ can be missed) — route those to the OCR-robust variant.
193
+ - Small evaluation set; numbers are indicative, not a benchmark.
194
+ - Not legal advice; not a substitute for a privacy/compliance review.
195
+
196
+ ## License — CC BY 4.0 (attribution required, commercial use allowed)
197
+
198
+ Released under **[Creative Commons Attribution 4.0](https://creativecommons.org/licenses/by/4.0/)**.
199
+ You may use, modify and redistribute this model — **including in commercial
200
+ products** — **provided you give appropriate credit**. Attribution is required in
201
+ any use, commercial or not; no other restrictions are added.
202
+
203
+ Suggested attribution:
204
+
205
+ > Polish legal NER / anonymisation model by **lexedit** (https://lexedit.ai),
206
+ > licensed CC BY 4.0, fine-tuned from HerBERT
207
+ > ([`allegro/herbert-base-cased`](https://huggingface.co/allegro/herbert-base-cased)).
208
+
209
+ This model is a derivative of HerBERT (Allegro), which is itself CC BY 4.0 — please
210
+ retain attribution to the base model as well.
config.json ADDED
@@ -0,0 +1,88 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "BertForTokenClassification"
4
+ ],
5
+ "attention_probs_dropout_prob": 0.1,
6
+ "bos_token_id": 0,
7
+ "classifier_dropout": null,
8
+ "dtype": "float32",
9
+ "hidden_act": "gelu",
10
+ "hidden_dropout_prob": 0.1,
11
+ "hidden_size": 768,
12
+ "id2label": {
13
+ "0": "B-DATE",
14
+ "1": "B-DIAGNOSIS",
15
+ "2": "B-EMAIL",
16
+ "3": "B-HEALTH_FACILITY",
17
+ "4": "B-IBAN",
18
+ "5": "B-ID",
19
+ "6": "B-LOC",
20
+ "7": "B-LOC_PUB",
21
+ "8": "B-MEDICAL_ID",
22
+ "9": "B-MONEY",
23
+ "10": "B-ORG",
24
+ "11": "B-PER",
25
+ "12": "B-PHONE",
26
+ "13": "B-WATERMARK",
27
+ "14": "I-DATE",
28
+ "15": "I-DIAGNOSIS",
29
+ "16": "I-EMAIL",
30
+ "17": "I-HEALTH_FACILITY",
31
+ "18": "I-IBAN",
32
+ "19": "I-ID",
33
+ "20": "I-LOC",
34
+ "21": "I-LOC_PUB",
35
+ "22": "I-MEDICAL_ID",
36
+ "23": "I-MONEY",
37
+ "24": "I-ORG",
38
+ "25": "I-PER",
39
+ "26": "I-PHONE",
40
+ "27": "I-WATERMARK",
41
+ "28": "O"
42
+ },
43
+ "initializer_range": 0.02,
44
+ "intermediate_size": 3072,
45
+ "label2id": {
46
+ "B-DATE": 0,
47
+ "B-DIAGNOSIS": 1,
48
+ "B-EMAIL": 2,
49
+ "B-HEALTH_FACILITY": 3,
50
+ "B-IBAN": 4,
51
+ "B-ID": 5,
52
+ "B-LOC": 6,
53
+ "B-LOC_PUB": 7,
54
+ "B-MEDICAL_ID": 8,
55
+ "B-MONEY": 9,
56
+ "B-ORG": 10,
57
+ "B-PER": 11,
58
+ "B-PHONE": 12,
59
+ "B-WATERMARK": 13,
60
+ "I-DATE": 14,
61
+ "I-DIAGNOSIS": 15,
62
+ "I-EMAIL": 16,
63
+ "I-HEALTH_FACILITY": 17,
64
+ "I-IBAN": 18,
65
+ "I-ID": 19,
66
+ "I-LOC": 20,
67
+ "I-LOC_PUB": 21,
68
+ "I-MEDICAL_ID": 22,
69
+ "I-MONEY": 23,
70
+ "I-ORG": 24,
71
+ "I-PER": 25,
72
+ "I-PHONE": 26,
73
+ "I-WATERMARK": 27,
74
+ "O": 28
75
+ },
76
+ "layer_norm_eps": 1e-12,
77
+ "max_position_embeddings": 514,
78
+ "model_type": "bert",
79
+ "num_attention_heads": 12,
80
+ "num_hidden_layers": 12,
81
+ "pad_token_id": 1,
82
+ "position_embedding_type": "absolute",
83
+ "tokenizer_class": "HerbertTokenizerFast",
84
+ "transformers_version": "4.57.6",
85
+ "type_vocab_size": 2,
86
+ "use_cache": true,
87
+ "vocab_size": 50000
88
+ }
examples/KNOWN_LIMITATIONS.md ADDED
@@ -0,0 +1,48 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Known limitations — real failure cases
2
+
3
+ This general-purpose variant is strongest on **clean digital** text. Its main
4
+ weakness is **heavily OCR-garbled names** (scanned / photographed documents).
5
+
6
+ > 👉 For scanned / OCR'd documents, use the OCR-robust sibling
7
+ > **[`lexedit/herbert-polish-legal-ner-ocr`](https://huggingface.co/lexedit/herbert-polish-legal-ner-ocr)**
8
+ > — it catches several of the cases below that this model misses.
9
+
10
+ All inputs are **synthetic** — fictional names, **no real or evaluation data**.
11
+ Outputs are the model's **actual** predictions (quantized ONNX, recall-first PER
12
+ threshold 0.2). Reproduce with [`known_limitations.py`](known_limitations.py).
13
+
14
+ ## Heavy OCR garble → name missed entirely
15
+
16
+ When OCR destroys most of a name's recognizable structure, this variant can fail to
17
+ flag it at all:
18
+
19
+ | Input | Behind the garble | Model `PER` |
20
+ |---|---|---|
21
+ | `Pozwana ote tobodaaka wniosła sprzeciw.` | *Bożena Łobodzińska* | — (nothing) |
22
+ | `Powód go Hea stawił się osobiście.` | *Igor Heliasz* | — (nothing) |
23
+ | `Decyzję doręczono: ist yi daziak.` | *Justyna Idziak* | — (nothing) |
24
+ | `Wniosek złożył wiar Żoją w dniu 5 maja.` | *Wiktor Żołądek* | — (nothing) |
25
+
26
+ *(The OCR-robust variant catches part of the last one — `['Żoją']` — instead of
27
+ missing it entirely.)*
28
+
29
+ ## Partial detection → residue leak
30
+
31
+ | Input | Behind the garble | Model `PER` | Problem |
32
+ |---|---|---|---|
33
+ | `WIESEAW KUE stawił się na rozprawie.` | *Wiesław Kuc* | `['WIE', 'AW KUE']` | split into two spans → the gap (`SE`) leaks |
34
+
35
+ ## Other known behaviours
36
+
37
+ - **Foreign / out-of-distribution names** (e.g. German, Dutch) are less reliable —
38
+ the model is tuned for Polish.
39
+ - **Boundaries** can occasionally over- or under-extend.
40
+
41
+ ## Mitigations (recommended)
42
+
43
+ 1. **Use the OCR-robust variant for scanned documents** (route by document type).
44
+ 2. **Document-safety post-pass** — propagate masking from any clean mention of a
45
+ surname to its other inflected / OCR-variant mentions in the same document.
46
+ 3. **Checksum regex** for structured PII (PESEL, NIP, IBAN, …).
47
+ 4. **Lower the threshold on noisy/scanned documents.**
48
+ 5. **Human review** for high-stakes anonymisation.
examples/STRONG_CASES.md ADDED
@@ -0,0 +1,52 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Where the model shines — hard cases handled well
2
+
3
+ Difficult inputs this general-purpose model gets **right**. All inputs are
4
+ synthetic (fictional names); outputs are the model's **actual** predictions
5
+ (quantized ONNX, recall-first PER threshold 0.2). Reproduce with
6
+ [`strong_cases.py`](strong_cases.py).
7
+
8
+ ## Polish inflection (declension)
9
+
10
+ Polish surnames change form by grammatical case; the model tags the full inflected
11
+ surface, not just the nominative:
12
+
13
+ | Input | Case | Model |
14
+ |---|---|---|
15
+ | `Pozew skierowano przeciwko Jerzemu Sekule.` | dative | `PER 'Jerzemu Sekule'` |
16
+ | `Sprawa dotyczy akt Wojciecha Michalika.` | genitive | `PER 'Wojciecha Michalika'` |
17
+
18
+ ## Tricky surface forms
19
+
20
+ | Input | Why it's hard | Model |
21
+ |---|---|---|
22
+ | `Powódka Anna Nowak-Kowalska złożyła wniosek.` | hyphenated double surname → one span | `PER 'Anna Nowak-Kowalska'` |
23
+ | `Stawiła się ALEKSANDRA ŚWIDERSKA, pełnomocnik.` | ALL-CAPS header/signature style | `PER 'ALEKSANDRA ŚWIDERSKA'` |
24
+ | `Z poważaniem,` / `Iga Mełech` / `radca prawny` | short female first + rare surname in a low-cue signature | `PER 'Iga Mełech'` |
25
+
26
+ ## Mild / moderate OCR noise
27
+
28
+ Light scan corruption (missing diacritics, a few confusable characters) is still
29
+ handled:
30
+
31
+ | Input | Behind the garble | Model |
32
+ |---|---|---|
33
+ | `Pozwany Damian Zotadek wniósł odpowiedź na pozew.` | *Damian Żołądek* (lost diacritics) | `PER 'Damian Zotadek'` |
34
+ | `Z powództwa Jiga Bękowiki przeciwko spółce.` | *Jaga Bętkowski* | `PER 'Jiga Bękowiki'` |
35
+ | `Pełnomocnik: Ptek Gostoiski, adwokat.` | *Patryk Gostomski* | `PER 'Ptek Gostoiski'` |
36
+
37
+ > **Heavy** OCR garble is this variant's weak spot — see
38
+ > [`KNOWN_LIMITATIONS.md`](KNOWN_LIMITATIONS.md). For scanned documents prefer the
39
+ > [OCR-robust variant](https://huggingface.co/lexedit/herbert-polish-legal-ner-ocr).
40
+
41
+ ## Multiple types + public/private distinction
42
+
43
+ ```
44
+ Jan Kowalski, zam. ul. Słoneczna 5 w Krakowie, PESEL 02070803628.
45
+ PER 'Jan Kowalski'
46
+ LOC 'ul. Słoneczna 5' ← private address (masked)
47
+ LOC_PUB 'Krakowie' ← public city (kept by default)
48
+ ID '02070803628' ← national id (PESEL)
49
+ ```
50
+
51
+ The model keeps the public place (`Krakowie`) visible while masking the private
52
+ street address — so the legal context survives anonymisation.
examples/inference_onnx.py ADDED
@@ -0,0 +1,76 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env python3
2
+ """Self-contained ONNX inference for the HerBERT Polish legal NER model.
3
+ No PyTorch needed. pip install onnxruntime transformers numpy
4
+
5
+ Run from the repo root: python examples/inference_onnx.py
6
+ """
7
+ import json
8
+ from pathlib import Path
9
+
10
+ import numpy as np
11
+ import onnxruntime as ort
12
+ from transformers import AutoTokenizer
13
+
14
+ ROOT = Path(__file__).resolve().parent.parent
15
+ PER_THRESHOLD = 0.2 # recall-first: flip a token to PER if summed PER prob >= this
16
+
17
+ tok = AutoTokenizer.from_pretrained(str(ROOT))
18
+ cfg = json.load(open(ROOT / "config.json", encoding="utf-8"))
19
+ id2label = {int(k): v for k, v in cfg["id2label"].items()}
20
+ label2id = cfg["label2id"]
21
+ sess = ort.InferenceSession(str(ROOT / "onnx" / "model_quantized.onnx"))
22
+ in_names = {i.name for i in sess.get_inputs()}
23
+
24
+
25
+ def softmax(x):
26
+ e = np.exp(x - x.max(-1, keepdims=True))
27
+ return e / e.sum(-1, keepdims=True)
28
+
29
+
30
+ def predict(text):
31
+ enc = tok(text, return_offsets_mapping=True, return_tensors="np",
32
+ truncation=True, max_length=512)
33
+ offsets = enc["offset_mapping"][0]
34
+ feeds = {"input_ids": enc["input_ids"].astype(np.int64),
35
+ "attention_mask": enc["attention_mask"].astype(np.int64)}
36
+ if "token_type_ids" in in_names:
37
+ feeds["token_type_ids"] = np.zeros_like(enc["input_ids"], dtype=np.int64)
38
+ probs = softmax(sess.run(None, feeds)[0][0]) # (seq, num_labels)
39
+ ids = probs.argmax(-1)
40
+ per_b, per_i = label2id["B-PER"], label2id["I-PER"]
41
+
42
+ spans, cur = [], None # cur = [start, end, type]
43
+ for i, (s, e) in enumerate(offsets):
44
+ if s == e: # special token
45
+ if cur:
46
+ spans.append(cur); cur = None
47
+ continue
48
+ label = id2label[int(ids[i])]
49
+ # recall-first override for persons
50
+ if label == "O" and probs[i, per_b] + probs[i, per_i] >= PER_THRESHOLD:
51
+ label = "I-PER" if (cur and cur[2] == "PER") else "B-PER"
52
+ if label == "O":
53
+ if cur:
54
+ spans.append(cur); cur = None
55
+ continue
56
+ tag, etype = label.split("-", 1)
57
+ if tag == "B" or cur is None or cur[2] != etype:
58
+ if cur:
59
+ spans.append(cur)
60
+ cur = [int(s), int(e), etype]
61
+ else:
62
+ cur[1] = int(e)
63
+ if cur:
64
+ spans.append(cur)
65
+ return [{"type": t, "start": a, "end": b, "text": text[a:b]} for a, b, t in spans]
66
+
67
+
68
+ if __name__ == "__main__":
69
+ samples = [
70
+ "Pozwany Jan Kowalski, zam. ul. Słoneczna 5 w Krakowie, PESEL 02070803628.",
71
+ "Powódka Anna Nowak-Kowalska, e-mail a.nowak@example.pl, tel. 501 234 567.",
72
+ ]
73
+ for t in samples:
74
+ print("\n" + t)
75
+ for ent in predict(t):
76
+ print(f" {ent['type']:8} [{ent['start']:>3}:{ent['end']:<3}] {ent['text']!r}")
examples/inference_pytorch.py ADDED
@@ -0,0 +1,32 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env python3
2
+ """PyTorch / transformers inference for the HerBERT Polish legal NER model.
3
+ pip install torch transformers
4
+
5
+ Run from the repo root: python examples/inference_pytorch.py
6
+ """
7
+ from pathlib import Path
8
+
9
+ from transformers import pipeline
10
+
11
+ ROOT = Path(__file__).resolve().parent.parent
12
+
13
+ ner = pipeline(
14
+ "token-classification",
15
+ model=str(ROOT),
16
+ tokenizer=str(ROOT),
17
+ aggregation_strategy="first", # group B-/I- subwords into whole entities
18
+ )
19
+
20
+ SAMPLES = [
21
+ "Pozwany Jan Kowalski, zam. ul. Słoneczna 5 w Krakowie, PESEL 02070803628.",
22
+ "Powódka Anna Nowak-Kowalska, e-mail a.nowak@example.pl, tel. 501 234 567.",
23
+ ]
24
+
25
+ if __name__ == "__main__":
26
+ for text in SAMPLES:
27
+ print("\n" + text)
28
+ for ent in ner(text):
29
+ print(f" {ent['entity_group']:8} [{ent['start']:>3}:{ent['end']:<3}] "
30
+ f"{ent['word']!r} ({ent['score']:.2f})")
31
+ # NOTE: for anonymisation, prefer a recall-first PER threshold (see
32
+ # examples/inference_onnx.py) over the arg-max grouping the pipeline uses.
examples/known_limitations.py ADDED
@@ -0,0 +1,81 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env python3
2
+ """Reproduce the model's KNOWN FAILURE CASES (see KNOWN_LIMITATIONS.md).
3
+
4
+ All inputs are SYNTHETIC — fictional names, no real personal data. The point is
5
+ to show, honestly and reproducibly, where the model still misses or mislabels
6
+ PERSON names so you can decide where extra safeguards (post-pass, review) are
7
+ needed. pip install onnxruntime transformers numpy
8
+
9
+ Run from the repo root: python examples/known_limitations.py
10
+ """
11
+ import json
12
+ from pathlib import Path
13
+
14
+ import numpy as np
15
+ import onnxruntime as ort
16
+ from transformers import AutoTokenizer
17
+
18
+ ROOT = Path(__file__).resolve().parent.parent
19
+ PER_THRESHOLD = 0.2
20
+
21
+ tok = AutoTokenizer.from_pretrained(str(ROOT))
22
+ cfg = json.load(open(ROOT / "config.json", encoding="utf-8"))
23
+ id2label = {int(k): v for k, v in cfg["id2label"].items()}
24
+ label2id = cfg["label2id"]
25
+ sess = ort.InferenceSession(str(ROOT / "onnx" / "model_quantized.onnx"))
26
+ in_names = {i.name for i in sess.get_inputs()}
27
+
28
+
29
+ def softmax(x):
30
+ e = np.exp(x - x.max(-1, keepdims=True)); return e / e.sum(-1, keepdims=True)
31
+
32
+
33
+ def predict_per(text):
34
+ enc = tok(text, return_offsets_mapping=True, return_tensors="np", truncation=True, max_length=512)
35
+ offs = enc["offset_mapping"][0]
36
+ feeds = {"input_ids": enc["input_ids"].astype(np.int64),
37
+ "attention_mask": enc["attention_mask"].astype(np.int64)}
38
+ if "token_type_ids" in in_names:
39
+ feeds["token_type_ids"] = np.zeros_like(enc["input_ids"], dtype=np.int64)
40
+ probs = softmax(sess.run(None, feeds)[0][0])
41
+ ids = probs.argmax(-1)
42
+ pb, pi = label2id["B-PER"], label2id["I-PER"]
43
+ spans, cur = [], None
44
+ for i, (s, e) in enumerate(offs):
45
+ if s == e:
46
+ if cur: spans.append(cur); cur = None
47
+ continue
48
+ lab = id2label[int(ids[i])]
49
+ if lab == "O" and probs[i, pb] + probs[i, pi] >= PER_THRESHOLD:
50
+ lab = "I-PER" if cur else "B-PER"
51
+ if not lab.endswith("PER"):
52
+ if cur: spans.append(cur); cur = None
53
+ continue
54
+ if lab == "B-PER" or cur is None:
55
+ if cur: spans.append(cur)
56
+ cur = [int(s), int(e)]
57
+ else:
58
+ cur[1] = int(e)
59
+ if cur: spans.append(cur)
60
+ return [text[a:b] for a, b in spans]
61
+
62
+
63
+ # (input, fictional clean name behind the garble, failure mode)
64
+ CASES = [
65
+ # 1) Heavy OCR garble -> missed entirely
66
+ ("Pozwana ote tobodaaka wniosła sprzeciw.", "Bożena Łobodzińska", "heavy OCR garble -> MISSED"),
67
+ ("Powód go Hea stawił się osobiście.", "Igor Heliasz", "heavy OCR garble -> MISSED"),
68
+ ("Decyzję doręczono: ist yi daziak.", "Justyna Idziak", "heavy OCR garble -> MISSED"),
69
+ # 2) Partial detection -> residue leak
70
+ ("Wniosek złożył wiar Żoją w dniu 5 maja.", "Wiktor Żołądek", "partial -> only part caught (residue)"),
71
+ ("WIESEAW KUE stawił się na rozprawie.", "Wiesław Kuc", "ALL-CAPS garble -> fragmented (gap leaks)"),
72
+ ]
73
+
74
+ if __name__ == "__main__":
75
+ print(f"PER threshold = {PER_THRESHOLD}\n")
76
+ for text, clean, mode in CASES:
77
+ per = predict_per(text)
78
+ print(f"# {mode} (fictional: {clean})")
79
+ print(f" input : {text}")
80
+ print(f" model : PER = {per if per else '— (nothing)'}")
81
+ print()
examples/strong_cases.py ADDED
@@ -0,0 +1,81 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env python3
2
+ """Hard cases the model handles WELL — the flip side of KNOWN_LIMITATIONS.md.
3
+
4
+ All inputs are synthetic (fictional names). Outputs are the model's actual
5
+ predictions (quantized ONNX, recall-first PER threshold 0.2). Reproduce:
6
+ python examples/strong_cases.py
7
+ """
8
+ import json
9
+ from pathlib import Path
10
+
11
+ import numpy as np
12
+ import onnxruntime as ort
13
+ from transformers import AutoTokenizer
14
+
15
+ ROOT = Path(__file__).resolve().parent.parent
16
+ PER_THRESHOLD = 0.2
17
+
18
+ tok = AutoTokenizer.from_pretrained(str(ROOT))
19
+ cfg = json.load(open(ROOT / "config.json", encoding="utf-8"))
20
+ id2label = {int(k): v for k, v in cfg["id2label"].items()}
21
+ label2id = cfg["label2id"]
22
+ sess = ort.InferenceSession(str(ROOT / "onnx" / "model_quantized.onnx"))
23
+ in_names = {i.name for i in sess.get_inputs()}
24
+
25
+
26
+ def softmax(x):
27
+ e = np.exp(x - x.max(-1, keepdims=True)); return e / e.sum(-1, keepdims=True)
28
+
29
+
30
+ def predict(text):
31
+ enc = tok(text, return_offsets_mapping=True, return_tensors="np", truncation=True, max_length=512)
32
+ offs = enc["offset_mapping"][0]
33
+ feeds = {"input_ids": enc["input_ids"].astype(np.int64),
34
+ "attention_mask": enc["attention_mask"].astype(np.int64)}
35
+ if "token_type_ids" in in_names:
36
+ feeds["token_type_ids"] = np.zeros_like(enc["input_ids"], dtype=np.int64)
37
+ probs = softmax(sess.run(None, feeds)[0][0])
38
+ ids = probs.argmax(-1)
39
+ pb, pi = label2id["B-PER"], label2id["I-PER"]
40
+ spans, cur = [], None
41
+ for i, (s, e) in enumerate(offs):
42
+ if s == e:
43
+ if cur: spans.append(cur); cur = None
44
+ continue
45
+ lab = id2label[int(ids[i])]
46
+ if lab == "O" and probs[i, pb] + probs[i, pi] >= PER_THRESHOLD:
47
+ lab = "I-PER" if (cur and cur[2] == "PER") else "B-PER"
48
+ if lab == "O":
49
+ if cur: spans.append(cur); cur = None
50
+ continue
51
+ tag, et = lab.split("-", 1)
52
+ if tag == "B" or cur is None or cur[2] != et:
53
+ if cur: spans.append(cur)
54
+ cur = [int(s), int(e), et]
55
+ else:
56
+ cur[1] = int(e)
57
+ if cur: spans.append(cur)
58
+ return [{"type": t, "start": a, "end": b, "text": text[a:b]} for a, b, t in spans]
59
+
60
+
61
+ CASES = [
62
+ ("Pozwany Damian Zotadek wniósł odpowiedź na pozew.", "OCR: missing diacritics (Żołądek)"),
63
+ ("Z powództwa Jiga Bękowiki przeciwko spółce.", "OCR: heavily garbled name (Jaga Bętkowski)"),
64
+ ("Pełnomocnik: Ptek Gostoiski, adwokat.", "OCR garble in a signature line (Patryk Gostomski)"),
65
+ ("Pozew skierowano przeciwko Jerzemu Sekule.", "inflected (dative) Polish name"),
66
+ ("Sprawa dotyczy akt Wojciecha Michalika.", "inflected (genitive) Polish name"),
67
+ ("Powódka Anna Nowak-Kowalska złożyła wniosek.", "hyphenated double surname (single span)"),
68
+ ("Stawiła się ALEKSANDRA ŚWIDERSKA, pełnomocnik.", "ALL-CAPS name (header/signature style)"),
69
+ ("Z poważaniem,\nIga Mełech\nradca prawny", "short female first + rare surname in a signature"),
70
+ ("Jan Kowalski, zam. ul. Słoneczna 5 w Krakowie, PESEL 02070803628.",
71
+ "private address (LOC) vs public city (LOC_PUB) + national ID"),
72
+ ]
73
+
74
+ if __name__ == "__main__":
75
+ print(f"PER threshold = {PER_THRESHOLD}\n")
76
+ for text, note in CASES:
77
+ print(f"# {note}")
78
+ print(f" input : {text.replace(chr(10), ' / ')}")
79
+ for e in predict(text):
80
+ print(f" {e['type']:8} {e['text']!r}")
81
+ print()
merges.txt ADDED
The diff for this file is too large to render. See raw diff
 
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:f19a470d0590f1620950b01338544d5acd494c2827a12fb87d9a99a38d4fa7bf
3
+ size 495521716
onnx/model_quantized.onnx ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:5ed91aabec29e07d2d5c7f370a5463283d4d88acbf44dd1ceedd0bca67330b50
3
+ size 125082875
requirements.txt ADDED
@@ -0,0 +1,6 @@
 
 
 
 
 
 
 
1
+ # ONNX inference (no PyTorch needed)
2
+ transformers>=4.40
3
+ onnxruntime>=1.17
4
+ numpy
5
+ # For the PyTorch / transformers-pipeline example instead, also:
6
+ # torch>=2.0
special_tokens_map.json ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "bos_token": "<s>",
3
+ "cls_token": "<s>",
4
+ "mask_token": "<mask>",
5
+ "pad_token": "<pad>",
6
+ "sep_token": "</s>",
7
+ "unk_token": "<unk>"
8
+ }
tokenizer.json ADDED
The diff for this file is too large to render. See raw diff
 
tokenizer_config.json ADDED
@@ -0,0 +1,58 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "added_tokens_decoder": {
3
+ "0": {
4
+ "content": "<s>",
5
+ "lstrip": false,
6
+ "normalized": false,
7
+ "rstrip": false,
8
+ "single_word": false,
9
+ "special": true
10
+ },
11
+ "1": {
12
+ "content": "<pad>",
13
+ "lstrip": false,
14
+ "normalized": false,
15
+ "rstrip": false,
16
+ "single_word": false,
17
+ "special": true
18
+ },
19
+ "2": {
20
+ "content": "</s>",
21
+ "lstrip": false,
22
+ "normalized": false,
23
+ "rstrip": false,
24
+ "single_word": false,
25
+ "special": true
26
+ },
27
+ "3": {
28
+ "content": "<unk>",
29
+ "lstrip": false,
30
+ "normalized": false,
31
+ "rstrip": false,
32
+ "single_word": false,
33
+ "special": true
34
+ },
35
+ "4": {
36
+ "content": "<mask>",
37
+ "lstrip": false,
38
+ "normalized": false,
39
+ "rstrip": false,
40
+ "single_word": false,
41
+ "special": true
42
+ }
43
+ },
44
+ "additional_special_tokens": [],
45
+ "bos_token": "<s>",
46
+ "clean_up_tokenization_spaces": false,
47
+ "cls_token": "<s>",
48
+ "do_lowercase_and_remove_accent": false,
49
+ "extra_special_tokens": {},
50
+ "id2lang": null,
51
+ "lang2id": null,
52
+ "mask_token": "<mask>",
53
+ "model_max_length": 512,
54
+ "pad_token": "<pad>",
55
+ "sep_token": "</s>",
56
+ "tokenizer_class": "HerbertTokenizer",
57
+ "unk_token": "<unk>"
58
+ }
vocab.json ADDED
The diff for this file is too large to render. See raw diff