Token Classification
Transformers
ONNX
Safetensors
Polish
bert
named-entity-recognition
ner
pii
pii-detection
anonymization
privacy
gdpr
polish
legal
legal-nlp
herbert
Eval Results (legacy)
Instructions to use lexedit/herbert-polish-legal-ner with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use lexedit/herbert-polish-legal-ner with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="lexedit/herbert-polish-legal-ner")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForTokenClassification tokenizer = AutoTokenizer.from_pretrained("lexedit/herbert-polish-legal-ner") model = AutoModelForTokenClassification.from_pretrained("lexedit/herbert-polish-legal-ner", device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 8,638 Bytes
399294e 3b0f37e 399294e | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 | ---
language:
- pl
license: cc-by-4.0
base_model: allegro/herbert-base-cased
base_model_relation: finetune
pipeline_tag: token-classification
library_name: transformers
metrics:
- f1
tags:
- token-classification
- named-entity-recognition
- ner
- pii
- pii-detection
- anonymization
- privacy
- gdpr
- polish
- legal
- legal-nlp
- herbert
- bert
- onnx
model-index:
- name: herbert-polish-legal-ner
results:
- task:
type: token-classification
name: Named Entity Recognition (PII)
dataset:
type: internal-polish-legal-eval
name: Internal Polish legal documents (identity-level eval)
metrics:
- type: f1
value: 0.94
name: Token-level F1 (held-out test)
widget:
- text: "Pozwany Jan Kowalski, zam. ul. Słoneczna 5 w Krakowie, PESEL 02070803628."
---
# HerBERT Polish Legal NER — PII / Anonymization
A Polish token-classification (NER) model for **detecting personally identifiable
information (PII) in legal and administrative text**, fine-tuned from
[`allegro/herbert-base-cased`](https://huggingface.co/allegro/herbert-base-cased).
This is the **general-purpose** variant — the best overall accuracy on **clean,
digital** Polish legal documents.
> For **scanned / photographed (OCR'd)** documents, use the OCR-robust sibling
> **[`lexedit/herbert-polish-legal-ner-ocr`](https://huggingface.co/lexedit/herbert-polish-legal-ner-ocr)**,
> which cuts the person-name leak on scanned text by ~35–41 % (at a small precision
> cost on clean text).
> **Intended use is defensive:** flagging PII so it can be masked / anonymised
> before a document is shared or processed. It is **not** a guarantee of complete
> anonymisation — see *Limitations*.
**▶ Try it in your browser:** [lexedit.ai/lexedit-privacy](https://lexedit.ai/lexedit-privacy)
— an interactive, **fully client-side** Polish legal-document anonymisation demo
(the text never leaves your browser).
**📂 Example cases:** [`examples/STRONG_CASES.md`](examples/STRONG_CASES.md) (hard
inputs it handles) · [`examples/KNOWN_LIMITATIONS.md`](examples/KNOWN_LIMITATIONS.md)
(where it still fails).
**🛠 Use as a Claude Code skill:**
[`tuul-ai/lexedit-anonymizer-skill`](https://github.com/tuul-ai/lexedit-anonymizer-skill)
— a drop-in skill that runs this model **locally** to anonymise Polish PII
(reversible masking) right inside your Claude Code workflow.
## Labels (29, BIO scheme)
`PER` (person) · `ORG` (organisation) · `LOC` (private address/location) ·
`LOC_PUB` (public place: city, country) · `DATE` · `MONEY` · `EMAIL` · `PHONE` ·
`ID` (national id / case / document number) · `IBAN` · `DIAGNOSIS` ·
`HEALTH_FACILITY` · `MEDICAL_ID` · `WATERMARK`, each as `B-…` / `I-…`, plus `O`.
The model distinguishes **private** locations (`LOC`, masked) from **public**
ones (`LOC_PUB`, usually kept), and treats `DATE` / `MONEY` as non-anonymised by
default.
## Usage (ONNX, no PyTorch required)
```bash
pip install onnxruntime transformers numpy
python examples/inference_onnx.py
```
```python
import numpy as np, onnxruntime as ort
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("lexedit/herbert-polish-legal-ner")
sess = ort.InferenceSession("onnx/model_quantized.onnx")
enc = tok("Pozwany Jan Kowalski, PESEL 02070803628.",
return_offsets_mapping=True, return_tensors="np")
feeds = {"input_ids": enc["input_ids"].astype(np.int64),
"attention_mask": enc["attention_mask"].astype(np.int64)}
logits = sess.run(None, feeds)[0] # (1, seq, 29)
# argmax per token -> map ids via config.id2label -> group B-/I- with offset_mapping
```
## Usage (PyTorch / transformers pipeline)
```python
from transformers import pipeline
ner = pipeline("token-classification",
model="lexedit/herbert-polish-legal-ner",
aggregation_strategy="first")
ner("Pozwany Jan Kowalski, zam. ul. Słoneczna 5 w Krakowie.")
```
## Runs in the browser (the intended setup)
This model is designed to run **entirely client-side**. The quantized ONNX
(~125 MB, int8) loads in the browser via
[onnxruntime-web](https://onnxruntime.ai/docs/tutorials/web/) or
[transformers.js](https://huggingface.co/docs/transformers.js) (WASM), so **the
document never leaves the user's device** — which is the whole point for sensitive
legal / medical text. It also runs anywhere ONNX Runtime does (Python, Node,
server, mobile). The demo at
[lexedit.ai/lexedit-privacy](https://lexedit.ai/lexedit-privacy) is exactly this.
### Speed (rough)
Quantized ONNX, measured on a laptop CPU (Apple Silicon):
| input | 1 core | all cores |
|---|---|---|
| short sentence (~30 tokens) | ~18 ms | ~13 ms |
| full chunk (~500 tokens) | ~0.2 s | ~0.1 s |
≈ **4–5 chunks/second single-threaded** natively. In the browser (WASM,
single-threaded) it is slower but practical: short text stays interactive, a 1–2
page document takes a few seconds, a large scanned document can take ~a minute or
two; cross-origin isolation (COOP/COEP → multi-threaded WASM) speeds it up.
## Recommended production setup
This model is **recall-first** and is one layer of a pipeline, not the whole
solution. For best anonymisation, pair it with:
1. **A recall-first threshold** for `PER` (flip a token to PER when the summed PER
probability ≥ ~0.2, even if it is not the arg-max).
2. **A deterministic document-safety post-pass** — snap spans to whole words, merge
hyphenated surnames, and propagate a detected surname to its other inflected /
OCR-variant mentions across the document.
3. **Checksum-validated regex** for structured PII (PESEL, NIP, REGON, IBAN, …).
4. **Human review** for high-stakes use.
## Evaluation
Identity-level **leak rate** = a person is "leaked" if *any* mention of them is
missed. Internal set of 50 real Polish legal documents (179 persons), recall-first
threshold 0.2.
| Document type | metric | this model | OCR-robust variant |
|---|---|---|---|
| **Clean / digital** | leak (threshold) | **10.1%** | 11.7% |
| **Clean / digital** | leak (+ post-pass) | **7.3%** | 7.3% |
| Scanned / OCR'd | leak (threshold) | 31.8% | **20.7%** |
| Scanned / OCR'd | leak (+ post-pass) | 25.7% | **15.1%** |
Token-level test F1 ≈ **0.94**.
**Takeaway:** this is the strongest variant on **clean digital** text (lowest leak,
best precision). On **scanned / OCR'd** text it is weaker — there the
[OCR-robust variant](https://huggingface.co/lexedit/herbert-polish-legal-ner-ocr)
wins. If you process both, **route by document type**.
## Training data
Fine-tuned on Polish legal and administrative documents — a mix of document
templates, programmatically-generated labelled examples (valid checksum-correct
synthetic identifiers, rule-based Polish name declension), and real-world
legal-document samples.
No raw personal data is distributed with this model. Because this is a
**token-classification** model (it outputs a label per input token and cannot
generate text), the weights do not reproduce or expose any training document.
## Limitations
> **Concrete, reproducible failure cases:** see
> [`examples/KNOWN_LIMITATIONS.md`](examples/KNOWN_LIMITATIONS.md) (synthetic inputs).
> Heavily OCR-garbled names are the main weakness — for scanned documents prefer the
> [OCR-robust variant](https://huggingface.co/lexedit/herbert-polish-legal-ner-ocr).
- **Polish only.**
- **Not a guarantee.** A residual leak rate remains (≈7–10 % identity-level on clean
text); always combine with the deterministic post-pass + checksum regex and human
review for high-stakes use.
- **Scanned / OCR'd text** is the weak spot of this variant (heavily garbled names
can be missed) — route those to the OCR-robust variant.
- Small evaluation set; numbers are indicative, not a benchmark.
- Not legal advice; not a substitute for a privacy/compliance review.
## License — CC BY 4.0 (attribution required, commercial use allowed)
Released under **[Creative Commons Attribution 4.0](https://creativecommons.org/licenses/by/4.0/)**.
You may use, modify and redistribute this model — **including in commercial
products** — **provided you give appropriate credit**. Attribution is required in
any use, commercial or not; no other restrictions are added.
Suggested attribution:
> Polish legal NER / anonymisation model by **lexedit** (https://lexedit.ai),
> licensed CC BY 4.0, fine-tuned from HerBERT
> ([`allegro/herbert-base-cased`](https://huggingface.co/allegro/herbert-base-cased)).
This model is a derivative of HerBERT (Allegro), which is itself CC BY 4.0 — please
retain attribution to the base model as well.
|