File size: 8,638 Bytes
399294e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3b0f37e
 
 
 
 
399294e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
---
language:
- pl
license: cc-by-4.0
base_model: allegro/herbert-base-cased
base_model_relation: finetune
pipeline_tag: token-classification
library_name: transformers
metrics:
- f1
tags:
- token-classification
- named-entity-recognition
- ner
- pii
- pii-detection
- anonymization
- privacy
- gdpr
- polish
- legal
- legal-nlp
- herbert
- bert
- onnx
model-index:
- name: herbert-polish-legal-ner
  results:
  - task:
      type: token-classification
      name: Named Entity Recognition (PII)
    dataset:
      type: internal-polish-legal-eval
      name: Internal Polish legal documents (identity-level eval)
    metrics:
    - type: f1
      value: 0.94
      name: Token-level F1 (held-out test)
widget:
- text: "Pozwany Jan Kowalski, zam. ul. Słoneczna 5 w Krakowie, PESEL 02070803628."
---

# HerBERT Polish Legal NER — PII / Anonymization

A Polish token-classification (NER) model for **detecting personally identifiable
information (PII) in legal and administrative text**, fine-tuned from
[`allegro/herbert-base-cased`](https://huggingface.co/allegro/herbert-base-cased).

This is the **general-purpose** variant — the best overall accuracy on **clean,
digital** Polish legal documents.

> For **scanned / photographed (OCR'd)** documents, use the OCR-robust sibling
> **[`lexedit/herbert-polish-legal-ner-ocr`](https://huggingface.co/lexedit/herbert-polish-legal-ner-ocr)**,
> which cuts the person-name leak on scanned text by ~35–41 % (at a small precision
> cost on clean text).

> **Intended use is defensive:** flagging PII so it can be masked / anonymised
> before a document is shared or processed. It is **not** a guarantee of complete
> anonymisation — see *Limitations*.

**▶ Try it in your browser:** [lexedit.ai/lexedit-privacy](https://lexedit.ai/lexedit-privacy)
— an interactive, **fully client-side** Polish legal-document anonymisation demo
(the text never leaves your browser).

**📂 Example cases:** [`examples/STRONG_CASES.md`](examples/STRONG_CASES.md) (hard
inputs it handles) · [`examples/KNOWN_LIMITATIONS.md`](examples/KNOWN_LIMITATIONS.md)
(where it still fails).

**🛠 Use as a Claude Code skill:**
[`tuul-ai/lexedit-anonymizer-skill`](https://github.com/tuul-ai/lexedit-anonymizer-skill)
— a drop-in skill that runs this model **locally** to anonymise Polish PII
(reversible masking) right inside your Claude Code workflow.

## Labels (29, BIO scheme)

`PER` (person) · `ORG` (organisation) · `LOC` (private address/location) ·
`LOC_PUB` (public place: city, country) · `DATE` · `MONEY` · `EMAIL` · `PHONE` ·
`ID` (national id / case / document number) · `IBAN` · `DIAGNOSIS` ·
`HEALTH_FACILITY` · `MEDICAL_ID` · `WATERMARK`, each as `B-…` / `I-…`, plus `O`.

The model distinguishes **private** locations (`LOC`, masked) from **public**
ones (`LOC_PUB`, usually kept), and treats `DATE` / `MONEY` as non-anonymised by
default.

## Usage (ONNX, no PyTorch required)

```bash
pip install onnxruntime transformers numpy
python examples/inference_onnx.py
```

```python
import numpy as np, onnxruntime as ort
from transformers import AutoTokenizer

tok = AutoTokenizer.from_pretrained("lexedit/herbert-polish-legal-ner")
sess = ort.InferenceSession("onnx/model_quantized.onnx")
enc = tok("Pozwany Jan Kowalski, PESEL 02070803628.",
          return_offsets_mapping=True, return_tensors="np")
feeds = {"input_ids": enc["input_ids"].astype(np.int64),
         "attention_mask": enc["attention_mask"].astype(np.int64)}
logits = sess.run(None, feeds)[0]          # (1, seq, 29)
# argmax per token -> map ids via config.id2label -> group B-/I- with offset_mapping
```

## Usage (PyTorch / transformers pipeline)

```python
from transformers import pipeline
ner = pipeline("token-classification",
               model="lexedit/herbert-polish-legal-ner",
               aggregation_strategy="first")
ner("Pozwany Jan Kowalski, zam. ul. Słoneczna 5 w Krakowie.")
```

## Runs in the browser (the intended setup)

This model is designed to run **entirely client-side**. The quantized ONNX
(~125 MB, int8) loads in the browser via
[onnxruntime-web](https://onnxruntime.ai/docs/tutorials/web/) or
[transformers.js](https://huggingface.co/docs/transformers.js) (WASM), so **the
document never leaves the user's device** — which is the whole point for sensitive
legal / medical text. It also runs anywhere ONNX Runtime does (Python, Node,
server, mobile). The demo at
[lexedit.ai/lexedit-privacy](https://lexedit.ai/lexedit-privacy) is exactly this.

### Speed (rough)

Quantized ONNX, measured on a laptop CPU (Apple Silicon):

| input | 1 core | all cores |
|---|---|---|
| short sentence (~30 tokens) | ~18 ms | ~13 ms |
| full chunk (~500 tokens) | ~0.2 s | ~0.1 s |

≈ **4–5 chunks/second single-threaded** natively. In the browser (WASM,
single-threaded) it is slower but practical: short text stays interactive, a 1–2
page document takes a few seconds, a large scanned document can take ~a minute or
two; cross-origin isolation (COOP/COEP → multi-threaded WASM) speeds it up.

## Recommended production setup

This model is **recall-first** and is one layer of a pipeline, not the whole
solution. For best anonymisation, pair it with:

1. **A recall-first threshold** for `PER` (flip a token to PER when the summed PER
   probability ≥ ~0.2, even if it is not the arg-max).
2. **A deterministic document-safety post-pass** — snap spans to whole words, merge
   hyphenated surnames, and propagate a detected surname to its other inflected /
   OCR-variant mentions across the document.
3. **Checksum-validated regex** for structured PII (PESEL, NIP, REGON, IBAN, …).
4. **Human review** for high-stakes use.

## Evaluation

Identity-level **leak rate** = a person is "leaked" if *any* mention of them is
missed. Internal set of 50 real Polish legal documents (179 persons), recall-first
threshold 0.2.

| Document type | metric | this model | OCR-robust variant |
|---|---|---|---|
| **Clean / digital** | leak (threshold) | **10.1%** | 11.7% |
| **Clean / digital** | leak (+ post-pass) | **7.3%** | 7.3% |
| Scanned / OCR'd | leak (threshold) | 31.8% | **20.7%** |
| Scanned / OCR'd | leak (+ post-pass) | 25.7% | **15.1%** |

Token-level test F1 ≈ **0.94**.

**Takeaway:** this is the strongest variant on **clean digital** text (lowest leak,
best precision). On **scanned / OCR'd** text it is weaker — there the
[OCR-robust variant](https://huggingface.co/lexedit/herbert-polish-legal-ner-ocr)
wins. If you process both, **route by document type**.

## Training data

Fine-tuned on Polish legal and administrative documents — a mix of document
templates, programmatically-generated labelled examples (valid checksum-correct
synthetic identifiers, rule-based Polish name declension), and real-world
legal-document samples.

No raw personal data is distributed with this model. Because this is a
**token-classification** model (it outputs a label per input token and cannot
generate text), the weights do not reproduce or expose any training document.

## Limitations

> **Concrete, reproducible failure cases:** see
> [`examples/KNOWN_LIMITATIONS.md`](examples/KNOWN_LIMITATIONS.md) (synthetic inputs).
> Heavily OCR-garbled names are the main weakness — for scanned documents prefer the
> [OCR-robust variant](https://huggingface.co/lexedit/herbert-polish-legal-ner-ocr).

- **Polish only.**
- **Not a guarantee.** A residual leak rate remains (≈7–10 % identity-level on clean
  text); always combine with the deterministic post-pass + checksum regex and human
  review for high-stakes use.
- **Scanned / OCR'd text** is the weak spot of this variant (heavily garbled names
  can be missed) — route those to the OCR-robust variant.
- Small evaluation set; numbers are indicative, not a benchmark.
- Not legal advice; not a substitute for a privacy/compliance review.

## License — CC BY 4.0 (attribution required, commercial use allowed)

Released under **[Creative Commons Attribution 4.0](https://creativecommons.org/licenses/by/4.0/)**.
You may use, modify and redistribute this model — **including in commercial
products** — **provided you give appropriate credit**. Attribution is required in
any use, commercial or not; no other restrictions are added.

Suggested attribution:

> Polish legal NER / anonymisation model by **lexedit** (https://lexedit.ai),
> licensed CC BY 4.0, fine-tuned from HerBERT
> ([`allegro/herbert-base-cased`](https://huggingface.co/allegro/herbert-base-cased)).

This model is a derivative of HerBERT (Allegro), which is itself CC BY 4.0 — please
retain attribution to the base model as well.