Token Classification
Transformers
ONNX
Safetensors
Polish
bert
named-entity-recognition
ner
pii
pii-detection
anonymization
privacy
gdpr
polish
legal
legal-nlp
ocr-robust
herbert
Eval Results (legacy)
Instructions to use lexedit/herbert-polish-legal-ner-ocr with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use lexedit/herbert-polish-legal-ner-ocr with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="lexedit/herbert-polish-legal-ner-ocr")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForTokenClassification tokenizer = AutoTokenizer.from_pretrained("lexedit/herbert-polish-legal-ner-ocr") model = AutoModelForTokenClassification.from_pretrained("lexedit/herbert-polish-legal-ner-ocr", device_map="auto") - Notebooks
- Google Colab
- Kaggle
|
Download examples/STRONG_CASES.md from lexedit/herbert-polish-legal-ner-ocr: direct link, hf CLI and curl.
- Browser
- Download file 2.31 kB
-
https://huggingface.co/lexedit/herbert-polish-legal-ner-ocr/resolve/main/examples/STRONG_CASES.md
- Command line
-
hf download hf://lexedit/herbert-polish-legal-ner-ocr/examples/STRONG_CASES.md
-
curl -L -o STRONG_CASES.md https://huggingface.co/lexedit/herbert-polish-legal-ner-ocr/resolve/main/examples/STRONG_CASES.md
2.31 kB
| # Where the model shines — hard cases handled well | |
| The flip side of [`KNOWN_LIMITATIONS.md`](KNOWN_LIMITATIONS.md): difficult inputs | |
| the model gets **right**. These are exactly the cases a plain dictionary / regex | |
| or a non-OCR-augmented model tends to miss. | |
| All inputs are synthetic (fictional names). Outputs are the model's **actual** | |
| predictions (quantized ONNX, recall-first PER threshold 0.2). Reproduce with | |
| [`strong_cases.py`](strong_cases.py). | |
| ## OCR-garbled names (the variant's headline strength) | |
| Names corrupted by optical character recognition — still recognised: | |
| | Input | Behind the garble | Model | | |
| |---|---|---| | |
| | `Pozwany Damian Zotadek wniósł odpowiedź na pozew.` | *Damian Żołądek* (lost diacritics) | `PER 'Damian Zotadek'` | | |
| | `Z powództwa Jiga Bękowiki przeciwko spółce.` | *Jaga Bętkowski* (heavy garble) | `PER 'Jiga Bękowiki'` | | |
| | `Pełnomocnik: Ptek Gostoiski, adwokat.` | *Patryk Gostomski* (signature line) | `PER 'Ptek Gostoiski'` | | |
| ## Polish inflection (declension) | |
| Polish surnames change form by grammatical case; the model tags the full inflected | |
| surface, not just the nominative: | |
| | Input | Case | Model | | |
| |---|---|---| | |
| | `Pozew skierowano przeciwko Jerzemu Sekule.` | dative | `PER 'Jerzemu Sekule'` | | |
| | `Sprawa dotyczy akt Wojciecha Michalika.` | genitive | `PER 'Wojciecha Michalika'` | | |
| ## Tricky surface forms | |
| | Input | Why it's hard | Model | | |
| |---|---|---| | |
| | `Powódka Anna Nowak-Kowalska złożyła wniosek.` | hyphenated double surname → one span | `PER 'Anna Nowak-Kowalska'` | | |
| | `Stawiła się ALEKSANDRA ŚWIDERSKA, pełnomocnik.` | ALL-CAPS header/signature style | `PER 'ALEKSANDRA ŚWIDERSKA'` | | |
| | `Z poważaniem,` / `Iga Mełech` / `radca prawny` | short female first + rare surname in a low-cue signature | `PER 'Iga Mełech'` | | |
| ## Multiple types + public/private distinction | |
| ``` | |
| Jan Kowalski, zam. ul. Słoneczna 5 w Krakowie, PESEL 02070803628. | |
| PER 'Jan Kowalski' | |
| LOC 'ul. Słoneczna 5' ← private address (masked) | |
| LOC_PUB 'Krakowie' ← public city (kept by default) | |
| ID '02070803628' ← national id (PESEL) | |
| ``` | |
| The model keeps the public place (`Krakowie`) visible while masking the private | |
| street address (`ul. Słoneczna 5`) — so the legal context survives anonymisation. | |