Token Classification
Transformers
ONNX
Safetensors
Polish
bert
named-entity-recognition
ner
pii
pii-detection
anonymization
privacy
gdpr
polish
legal
legal-nlp
herbert
Eval Results (legacy)
Instructions to use lexedit/herbert-polish-legal-ner with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use lexedit/herbert-polish-legal-ner with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="lexedit/herbert-polish-legal-ner")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForTokenClassification tokenizer = AutoTokenizer.from_pretrained("lexedit/herbert-polish-legal-ner") model = AutoModelForTokenClassification.from_pretrained("lexedit/herbert-polish-legal-ner", device_map="auto") - Notebooks
- Google Colab
- Kaggle
HerBERT Polish legal NER (general): weights + ONNX + card + examples
Browse files- .gitattributes +2 -34
- README.md +210 -0
- config.json +88 -0
- examples/KNOWN_LIMITATIONS.md +48 -0
- examples/STRONG_CASES.md +52 -0
- examples/inference_onnx.py +76 -0
- examples/inference_pytorch.py +32 -0
- examples/known_limitations.py +81 -0
- examples/strong_cases.py +81 -0
- merges.txt +0 -0
- model.safetensors +3 -0
- onnx/model_quantized.onnx +3 -0
- requirements.txt +6 -0
- special_tokens_map.json +8 -0
- tokenizer.json +0 -0
- tokenizer_config.json +58 -0
- vocab.json +0 -0
.gitattributes
CHANGED
|
@@ -1,35 +1,3 @@
|
|
| 1 |
-
*.7z filter=lfs diff=lfs merge=lfs -text
|
| 2 |
-
*.arrow filter=lfs diff=lfs merge=lfs -text
|
| 3 |
-
*.bin filter=lfs diff=lfs merge=lfs -text
|
| 4 |
-
*.bz2 filter=lfs diff=lfs merge=lfs -text
|
| 5 |
-
*.ckpt filter=lfs diff=lfs merge=lfs -text
|
| 6 |
-
*.ftz filter=lfs diff=lfs merge=lfs -text
|
| 7 |
-
*.gz filter=lfs diff=lfs merge=lfs -text
|
| 8 |
-
*.h5 filter=lfs diff=lfs merge=lfs -text
|
| 9 |
-
*.joblib filter=lfs diff=lfs merge=lfs -text
|
| 10 |
-
*.lfs.* filter=lfs diff=lfs merge=lfs -text
|
| 11 |
-
*.mlmodel filter=lfs diff=lfs merge=lfs -text
|
| 12 |
-
*.model filter=lfs diff=lfs merge=lfs -text
|
| 13 |
-
*.msgpack filter=lfs diff=lfs merge=lfs -text
|
| 14 |
-
*.npy filter=lfs diff=lfs merge=lfs -text
|
| 15 |
-
*.npz filter=lfs diff=lfs merge=lfs -text
|
| 16 |
-
*.onnx filter=lfs diff=lfs merge=lfs -text
|
| 17 |
-
*.ot filter=lfs diff=lfs merge=lfs -text
|
| 18 |
-
*.parquet filter=lfs diff=lfs merge=lfs -text
|
| 19 |
-
*.pb filter=lfs diff=lfs merge=lfs -text
|
| 20 |
-
*.pickle filter=lfs diff=lfs merge=lfs -text
|
| 21 |
-
*.pkl filter=lfs diff=lfs merge=lfs -text
|
| 22 |
-
*.pt filter=lfs diff=lfs merge=lfs -text
|
| 23 |
-
*.pth filter=lfs diff=lfs merge=lfs -text
|
| 24 |
-
*.rar filter=lfs diff=lfs merge=lfs -text
|
| 25 |
*.safetensors filter=lfs diff=lfs merge=lfs -text
|
| 26 |
-
|
| 27 |
-
*.
|
| 28 |
-
*.tar filter=lfs diff=lfs merge=lfs -text
|
| 29 |
-
*.tflite filter=lfs diff=lfs merge=lfs -text
|
| 30 |
-
*.tgz filter=lfs diff=lfs merge=lfs -text
|
| 31 |
-
*.wasm filter=lfs diff=lfs merge=lfs -text
|
| 32 |
-
*.xz filter=lfs diff=lfs merge=lfs -text
|
| 33 |
-
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
-
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
-
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
*.safetensors filter=lfs diff=lfs merge=lfs -text
|
| 2 |
+
*.onnx filter=lfs diff=lfs merge=lfs -text
|
| 3 |
+
*.bin filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
README.md
ADDED
|
@@ -0,0 +1,210 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
language:
|
| 3 |
+
- pl
|
| 4 |
+
license: cc-by-4.0
|
| 5 |
+
base_model: allegro/herbert-base-cased
|
| 6 |
+
base_model_relation: finetune
|
| 7 |
+
pipeline_tag: token-classification
|
| 8 |
+
library_name: transformers
|
| 9 |
+
metrics:
|
| 10 |
+
- f1
|
| 11 |
+
tags:
|
| 12 |
+
- token-classification
|
| 13 |
+
- named-entity-recognition
|
| 14 |
+
- ner
|
| 15 |
+
- pii
|
| 16 |
+
- pii-detection
|
| 17 |
+
- anonymization
|
| 18 |
+
- privacy
|
| 19 |
+
- gdpr
|
| 20 |
+
- polish
|
| 21 |
+
- legal
|
| 22 |
+
- legal-nlp
|
| 23 |
+
- herbert
|
| 24 |
+
- bert
|
| 25 |
+
- onnx
|
| 26 |
+
model-index:
|
| 27 |
+
- name: herbert-polish-legal-ner
|
| 28 |
+
results:
|
| 29 |
+
- task:
|
| 30 |
+
type: token-classification
|
| 31 |
+
name: Named Entity Recognition (PII)
|
| 32 |
+
dataset:
|
| 33 |
+
type: internal-polish-legal-eval
|
| 34 |
+
name: Internal Polish legal documents (identity-level eval)
|
| 35 |
+
metrics:
|
| 36 |
+
- type: f1
|
| 37 |
+
value: 0.94
|
| 38 |
+
name: Token-level F1 (held-out test)
|
| 39 |
+
widget:
|
| 40 |
+
- text: "Pozwany Jan Kowalski, zam. ul. Słoneczna 5 w Krakowie, PESEL 02070803628."
|
| 41 |
+
---
|
| 42 |
+
|
| 43 |
+
# HerBERT Polish Legal NER — PII / Anonymization
|
| 44 |
+
|
| 45 |
+
A Polish token-classification (NER) model for **detecting personally identifiable
|
| 46 |
+
information (PII) in legal and administrative text**, fine-tuned from
|
| 47 |
+
[`allegro/herbert-base-cased`](https://huggingface.co/allegro/herbert-base-cased).
|
| 48 |
+
|
| 49 |
+
This is the **general-purpose** variant — the best overall accuracy on **clean,
|
| 50 |
+
digital** Polish legal documents.
|
| 51 |
+
|
| 52 |
+
> For **scanned / photographed (OCR'd)** documents, use the OCR-robust sibling
|
| 53 |
+
> **[`lexedit/herbert-polish-legal-ner-ocr`](https://huggingface.co/lexedit/herbert-polish-legal-ner-ocr)**,
|
| 54 |
+
> which cuts the person-name leak on scanned text by ~35–41 % (at a small precision
|
| 55 |
+
> cost on clean text).
|
| 56 |
+
|
| 57 |
+
> **Intended use is defensive:** flagging PII so it can be masked / anonymised
|
| 58 |
+
> before a document is shared or processed. It is **not** a guarantee of complete
|
| 59 |
+
> anonymisation — see *Limitations*.
|
| 60 |
+
|
| 61 |
+
**▶ Try it in your browser:** [lexedit.ai/lexedit-privacy](https://lexedit.ai/lexedit-privacy)
|
| 62 |
+
— an interactive, **fully client-side** Polish legal-document anonymisation demo
|
| 63 |
+
(the text never leaves your browser).
|
| 64 |
+
|
| 65 |
+
**📂 Example cases:** [`examples/STRONG_CASES.md`](examples/STRONG_CASES.md) (hard
|
| 66 |
+
inputs it handles) · [`examples/KNOWN_LIMITATIONS.md`](examples/KNOWN_LIMITATIONS.md)
|
| 67 |
+
(where it still fails).
|
| 68 |
+
|
| 69 |
+
## Labels (29, BIO scheme)
|
| 70 |
+
|
| 71 |
+
`PER` (person) · `ORG` (organisation) · `LOC` (private address/location) ·
|
| 72 |
+
`LOC_PUB` (public place: city, country) · `DATE` · `MONEY` · `EMAIL` · `PHONE` ·
|
| 73 |
+
`ID` (national id / case / document number) · `IBAN` · `DIAGNOSIS` ·
|
| 74 |
+
`HEALTH_FACILITY` · `MEDICAL_ID` · `WATERMARK`, each as `B-…` / `I-…`, plus `O`.
|
| 75 |
+
|
| 76 |
+
The model distinguishes **private** locations (`LOC`, masked) from **public**
|
| 77 |
+
ones (`LOC_PUB`, usually kept), and treats `DATE` / `MONEY` as non-anonymised by
|
| 78 |
+
default.
|
| 79 |
+
|
| 80 |
+
## Usage (ONNX, no PyTorch required)
|
| 81 |
+
|
| 82 |
+
```bash
|
| 83 |
+
pip install onnxruntime transformers numpy
|
| 84 |
+
python examples/inference_onnx.py
|
| 85 |
+
```
|
| 86 |
+
|
| 87 |
+
```python
|
| 88 |
+
import numpy as np, onnxruntime as ort
|
| 89 |
+
from transformers import AutoTokenizer
|
| 90 |
+
|
| 91 |
+
tok = AutoTokenizer.from_pretrained("lexedit/herbert-polish-legal-ner")
|
| 92 |
+
sess = ort.InferenceSession("onnx/model_quantized.onnx")
|
| 93 |
+
enc = tok("Pozwany Jan Kowalski, PESEL 02070803628.",
|
| 94 |
+
return_offsets_mapping=True, return_tensors="np")
|
| 95 |
+
feeds = {"input_ids": enc["input_ids"].astype(np.int64),
|
| 96 |
+
"attention_mask": enc["attention_mask"].astype(np.int64)}
|
| 97 |
+
logits = sess.run(None, feeds)[0] # (1, seq, 29)
|
| 98 |
+
# argmax per token -> map ids via config.id2label -> group B-/I- with offset_mapping
|
| 99 |
+
```
|
| 100 |
+
|
| 101 |
+
## Usage (PyTorch / transformers pipeline)
|
| 102 |
+
|
| 103 |
+
```python
|
| 104 |
+
from transformers import pipeline
|
| 105 |
+
ner = pipeline("token-classification",
|
| 106 |
+
model="lexedit/herbert-polish-legal-ner",
|
| 107 |
+
aggregation_strategy="first")
|
| 108 |
+
ner("Pozwany Jan Kowalski, zam. ul. Słoneczna 5 w Krakowie.")
|
| 109 |
+
```
|
| 110 |
+
|
| 111 |
+
## Runs in the browser (the intended setup)
|
| 112 |
+
|
| 113 |
+
This model is designed to run **entirely client-side**. The quantized ONNX
|
| 114 |
+
(~125 MB, int8) loads in the browser via
|
| 115 |
+
[onnxruntime-web](https://onnxruntime.ai/docs/tutorials/web/) or
|
| 116 |
+
[transformers.js](https://huggingface.co/docs/transformers.js) (WASM), so **the
|
| 117 |
+
document never leaves the user's device** — which is the whole point for sensitive
|
| 118 |
+
legal / medical text. It also runs anywhere ONNX Runtime does (Python, Node,
|
| 119 |
+
server, mobile). The demo at
|
| 120 |
+
[lexedit.ai/lexedit-privacy](https://lexedit.ai/lexedit-privacy) is exactly this.
|
| 121 |
+
|
| 122 |
+
### Speed (rough)
|
| 123 |
+
|
| 124 |
+
Quantized ONNX, measured on a laptop CPU (Apple Silicon):
|
| 125 |
+
|
| 126 |
+
| input | 1 core | all cores |
|
| 127 |
+
|---|---|---|
|
| 128 |
+
| short sentence (~30 tokens) | ~18 ms | ~13 ms |
|
| 129 |
+
| full chunk (~500 tokens) | ~0.2 s | ~0.1 s |
|
| 130 |
+
|
| 131 |
+
≈ **4–5 chunks/second single-threaded** natively. In the browser (WASM,
|
| 132 |
+
single-threaded) it is slower but practical: short text stays interactive, a 1–2
|
| 133 |
+
page document takes a few seconds, a large scanned document can take ~a minute or
|
| 134 |
+
two; cross-origin isolation (COOP/COEP → multi-threaded WASM) speeds it up.
|
| 135 |
+
|
| 136 |
+
## Recommended production setup
|
| 137 |
+
|
| 138 |
+
This model is **recall-first** and is one layer of a pipeline, not the whole
|
| 139 |
+
solution. For best anonymisation, pair it with:
|
| 140 |
+
|
| 141 |
+
1. **A recall-first threshold** for `PER` (flip a token to PER when the summed PER
|
| 142 |
+
probability ≥ ~0.2, even if it is not the arg-max).
|
| 143 |
+
2. **A deterministic document-safety post-pass** — snap spans to whole words, merge
|
| 144 |
+
hyphenated surnames, and propagate a detected surname to its other inflected /
|
| 145 |
+
OCR-variant mentions across the document.
|
| 146 |
+
3. **Checksum-validated regex** for structured PII (PESEL, NIP, REGON, IBAN, …).
|
| 147 |
+
4. **Human review** for high-stakes use.
|
| 148 |
+
|
| 149 |
+
## Evaluation
|
| 150 |
+
|
| 151 |
+
Identity-level **leak rate** = a person is "leaked" if *any* mention of them is
|
| 152 |
+
missed. Internal set of 50 real Polish legal documents (179 persons), recall-first
|
| 153 |
+
threshold 0.2.
|
| 154 |
+
|
| 155 |
+
| Document type | metric | this model | OCR-robust variant |
|
| 156 |
+
|---|---|---|---|
|
| 157 |
+
| **Clean / digital** | leak (threshold) | **10.1%** | 11.7% |
|
| 158 |
+
| **Clean / digital** | leak (+ post-pass) | **7.3%** | 7.3% |
|
| 159 |
+
| Scanned / OCR'd | leak (threshold) | 31.8% | **20.7%** |
|
| 160 |
+
| Scanned / OCR'd | leak (+ post-pass) | 25.7% | **15.1%** |
|
| 161 |
+
|
| 162 |
+
Token-level test F1 ≈ **0.94**.
|
| 163 |
+
|
| 164 |
+
**Takeaway:** this is the strongest variant on **clean digital** text (lowest leak,
|
| 165 |
+
best precision). On **scanned / OCR'd** text it is weaker — there the
|
| 166 |
+
[OCR-robust variant](https://huggingface.co/lexedit/herbert-polish-legal-ner-ocr)
|
| 167 |
+
wins. If you process both, **route by document type**.
|
| 168 |
+
|
| 169 |
+
## Training data
|
| 170 |
+
|
| 171 |
+
Fine-tuned on Polish legal and administrative documents — a mix of document
|
| 172 |
+
templates, programmatically-generated labelled examples (valid checksum-correct
|
| 173 |
+
synthetic identifiers, rule-based Polish name declension), and real-world
|
| 174 |
+
legal-document samples.
|
| 175 |
+
|
| 176 |
+
No raw personal data is distributed with this model. Because this is a
|
| 177 |
+
**token-classification** model (it outputs a label per input token and cannot
|
| 178 |
+
generate text), the weights do not reproduce or expose any training document.
|
| 179 |
+
|
| 180 |
+
## Limitations
|
| 181 |
+
|
| 182 |
+
> **Concrete, reproducible failure cases:** see
|
| 183 |
+
> [`examples/KNOWN_LIMITATIONS.md`](examples/KNOWN_LIMITATIONS.md) (synthetic inputs).
|
| 184 |
+
> Heavily OCR-garbled names are the main weakness — for scanned documents prefer the
|
| 185 |
+
> [OCR-robust variant](https://huggingface.co/lexedit/herbert-polish-legal-ner-ocr).
|
| 186 |
+
|
| 187 |
+
- **Polish only.**
|
| 188 |
+
- **Not a guarantee.** A residual leak rate remains (≈7–10 % identity-level on clean
|
| 189 |
+
text); always combine with the deterministic post-pass + checksum regex and human
|
| 190 |
+
review for high-stakes use.
|
| 191 |
+
- **Scanned / OCR'd text** is the weak spot of this variant (heavily garbled names
|
| 192 |
+
can be missed) — route those to the OCR-robust variant.
|
| 193 |
+
- Small evaluation set; numbers are indicative, not a benchmark.
|
| 194 |
+
- Not legal advice; not a substitute for a privacy/compliance review.
|
| 195 |
+
|
| 196 |
+
## License — CC BY 4.0 (attribution required, commercial use allowed)
|
| 197 |
+
|
| 198 |
+
Released under **[Creative Commons Attribution 4.0](https://creativecommons.org/licenses/by/4.0/)**.
|
| 199 |
+
You may use, modify and redistribute this model — **including in commercial
|
| 200 |
+
products** — **provided you give appropriate credit**. Attribution is required in
|
| 201 |
+
any use, commercial or not; no other restrictions are added.
|
| 202 |
+
|
| 203 |
+
Suggested attribution:
|
| 204 |
+
|
| 205 |
+
> Polish legal NER / anonymisation model by **lexedit** (https://lexedit.ai),
|
| 206 |
+
> licensed CC BY 4.0, fine-tuned from HerBERT
|
| 207 |
+
> ([`allegro/herbert-base-cased`](https://huggingface.co/allegro/herbert-base-cased)).
|
| 208 |
+
|
| 209 |
+
This model is a derivative of HerBERT (Allegro), which is itself CC BY 4.0 — please
|
| 210 |
+
retain attribution to the base model as well.
|
config.json
ADDED
|
@@ -0,0 +1,88 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"architectures": [
|
| 3 |
+
"BertForTokenClassification"
|
| 4 |
+
],
|
| 5 |
+
"attention_probs_dropout_prob": 0.1,
|
| 6 |
+
"bos_token_id": 0,
|
| 7 |
+
"classifier_dropout": null,
|
| 8 |
+
"dtype": "float32",
|
| 9 |
+
"hidden_act": "gelu",
|
| 10 |
+
"hidden_dropout_prob": 0.1,
|
| 11 |
+
"hidden_size": 768,
|
| 12 |
+
"id2label": {
|
| 13 |
+
"0": "B-DATE",
|
| 14 |
+
"1": "B-DIAGNOSIS",
|
| 15 |
+
"2": "B-EMAIL",
|
| 16 |
+
"3": "B-HEALTH_FACILITY",
|
| 17 |
+
"4": "B-IBAN",
|
| 18 |
+
"5": "B-ID",
|
| 19 |
+
"6": "B-LOC",
|
| 20 |
+
"7": "B-LOC_PUB",
|
| 21 |
+
"8": "B-MEDICAL_ID",
|
| 22 |
+
"9": "B-MONEY",
|
| 23 |
+
"10": "B-ORG",
|
| 24 |
+
"11": "B-PER",
|
| 25 |
+
"12": "B-PHONE",
|
| 26 |
+
"13": "B-WATERMARK",
|
| 27 |
+
"14": "I-DATE",
|
| 28 |
+
"15": "I-DIAGNOSIS",
|
| 29 |
+
"16": "I-EMAIL",
|
| 30 |
+
"17": "I-HEALTH_FACILITY",
|
| 31 |
+
"18": "I-IBAN",
|
| 32 |
+
"19": "I-ID",
|
| 33 |
+
"20": "I-LOC",
|
| 34 |
+
"21": "I-LOC_PUB",
|
| 35 |
+
"22": "I-MEDICAL_ID",
|
| 36 |
+
"23": "I-MONEY",
|
| 37 |
+
"24": "I-ORG",
|
| 38 |
+
"25": "I-PER",
|
| 39 |
+
"26": "I-PHONE",
|
| 40 |
+
"27": "I-WATERMARK",
|
| 41 |
+
"28": "O"
|
| 42 |
+
},
|
| 43 |
+
"initializer_range": 0.02,
|
| 44 |
+
"intermediate_size": 3072,
|
| 45 |
+
"label2id": {
|
| 46 |
+
"B-DATE": 0,
|
| 47 |
+
"B-DIAGNOSIS": 1,
|
| 48 |
+
"B-EMAIL": 2,
|
| 49 |
+
"B-HEALTH_FACILITY": 3,
|
| 50 |
+
"B-IBAN": 4,
|
| 51 |
+
"B-ID": 5,
|
| 52 |
+
"B-LOC": 6,
|
| 53 |
+
"B-LOC_PUB": 7,
|
| 54 |
+
"B-MEDICAL_ID": 8,
|
| 55 |
+
"B-MONEY": 9,
|
| 56 |
+
"B-ORG": 10,
|
| 57 |
+
"B-PER": 11,
|
| 58 |
+
"B-PHONE": 12,
|
| 59 |
+
"B-WATERMARK": 13,
|
| 60 |
+
"I-DATE": 14,
|
| 61 |
+
"I-DIAGNOSIS": 15,
|
| 62 |
+
"I-EMAIL": 16,
|
| 63 |
+
"I-HEALTH_FACILITY": 17,
|
| 64 |
+
"I-IBAN": 18,
|
| 65 |
+
"I-ID": 19,
|
| 66 |
+
"I-LOC": 20,
|
| 67 |
+
"I-LOC_PUB": 21,
|
| 68 |
+
"I-MEDICAL_ID": 22,
|
| 69 |
+
"I-MONEY": 23,
|
| 70 |
+
"I-ORG": 24,
|
| 71 |
+
"I-PER": 25,
|
| 72 |
+
"I-PHONE": 26,
|
| 73 |
+
"I-WATERMARK": 27,
|
| 74 |
+
"O": 28
|
| 75 |
+
},
|
| 76 |
+
"layer_norm_eps": 1e-12,
|
| 77 |
+
"max_position_embeddings": 514,
|
| 78 |
+
"model_type": "bert",
|
| 79 |
+
"num_attention_heads": 12,
|
| 80 |
+
"num_hidden_layers": 12,
|
| 81 |
+
"pad_token_id": 1,
|
| 82 |
+
"position_embedding_type": "absolute",
|
| 83 |
+
"tokenizer_class": "HerbertTokenizerFast",
|
| 84 |
+
"transformers_version": "4.57.6",
|
| 85 |
+
"type_vocab_size": 2,
|
| 86 |
+
"use_cache": true,
|
| 87 |
+
"vocab_size": 50000
|
| 88 |
+
}
|
examples/KNOWN_LIMITATIONS.md
ADDED
|
@@ -0,0 +1,48 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Known limitations — real failure cases
|
| 2 |
+
|
| 3 |
+
This general-purpose variant is strongest on **clean digital** text. Its main
|
| 4 |
+
weakness is **heavily OCR-garbled names** (scanned / photographed documents).
|
| 5 |
+
|
| 6 |
+
> 👉 For scanned / OCR'd documents, use the OCR-robust sibling
|
| 7 |
+
> **[`lexedit/herbert-polish-legal-ner-ocr`](https://huggingface.co/lexedit/herbert-polish-legal-ner-ocr)**
|
| 8 |
+
> — it catches several of the cases below that this model misses.
|
| 9 |
+
|
| 10 |
+
All inputs are **synthetic** — fictional names, **no real or evaluation data**.
|
| 11 |
+
Outputs are the model's **actual** predictions (quantized ONNX, recall-first PER
|
| 12 |
+
threshold 0.2). Reproduce with [`known_limitations.py`](known_limitations.py).
|
| 13 |
+
|
| 14 |
+
## Heavy OCR garble → name missed entirely
|
| 15 |
+
|
| 16 |
+
When OCR destroys most of a name's recognizable structure, this variant can fail to
|
| 17 |
+
flag it at all:
|
| 18 |
+
|
| 19 |
+
| Input | Behind the garble | Model `PER` |
|
| 20 |
+
|---|---|---|
|
| 21 |
+
| `Pozwana ote tobodaaka wniosła sprzeciw.` | *Bożena Łobodzińska* | — (nothing) |
|
| 22 |
+
| `Powód go Hea stawił się osobiście.` | *Igor Heliasz* | — (nothing) |
|
| 23 |
+
| `Decyzję doręczono: ist yi daziak.` | *Justyna Idziak* | — (nothing) |
|
| 24 |
+
| `Wniosek złożył wiar Żoją w dniu 5 maja.` | *Wiktor Żołądek* | — (nothing) |
|
| 25 |
+
|
| 26 |
+
*(The OCR-robust variant catches part of the last one — `['Żoją']` — instead of
|
| 27 |
+
missing it entirely.)*
|
| 28 |
+
|
| 29 |
+
## Partial detection → residue leak
|
| 30 |
+
|
| 31 |
+
| Input | Behind the garble | Model `PER` | Problem |
|
| 32 |
+
|---|---|---|---|
|
| 33 |
+
| `WIESEAW KUE stawił się na rozprawie.` | *Wiesław Kuc* | `['WIE', 'AW KUE']` | split into two spans → the gap (`SE`) leaks |
|
| 34 |
+
|
| 35 |
+
## Other known behaviours
|
| 36 |
+
|
| 37 |
+
- **Foreign / out-of-distribution names** (e.g. German, Dutch) are less reliable —
|
| 38 |
+
the model is tuned for Polish.
|
| 39 |
+
- **Boundaries** can occasionally over- or under-extend.
|
| 40 |
+
|
| 41 |
+
## Mitigations (recommended)
|
| 42 |
+
|
| 43 |
+
1. **Use the OCR-robust variant for scanned documents** (route by document type).
|
| 44 |
+
2. **Document-safety post-pass** — propagate masking from any clean mention of a
|
| 45 |
+
surname to its other inflected / OCR-variant mentions in the same document.
|
| 46 |
+
3. **Checksum regex** for structured PII (PESEL, NIP, IBAN, …).
|
| 47 |
+
4. **Lower the threshold on noisy/scanned documents.**
|
| 48 |
+
5. **Human review** for high-stakes anonymisation.
|
examples/STRONG_CASES.md
ADDED
|
@@ -0,0 +1,52 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Where the model shines — hard cases handled well
|
| 2 |
+
|
| 3 |
+
Difficult inputs this general-purpose model gets **right**. All inputs are
|
| 4 |
+
synthetic (fictional names); outputs are the model's **actual** predictions
|
| 5 |
+
(quantized ONNX, recall-first PER threshold 0.2). Reproduce with
|
| 6 |
+
[`strong_cases.py`](strong_cases.py).
|
| 7 |
+
|
| 8 |
+
## Polish inflection (declension)
|
| 9 |
+
|
| 10 |
+
Polish surnames change form by grammatical case; the model tags the full inflected
|
| 11 |
+
surface, not just the nominative:
|
| 12 |
+
|
| 13 |
+
| Input | Case | Model |
|
| 14 |
+
|---|---|---|
|
| 15 |
+
| `Pozew skierowano przeciwko Jerzemu Sekule.` | dative | `PER 'Jerzemu Sekule'` |
|
| 16 |
+
| `Sprawa dotyczy akt Wojciecha Michalika.` | genitive | `PER 'Wojciecha Michalika'` |
|
| 17 |
+
|
| 18 |
+
## Tricky surface forms
|
| 19 |
+
|
| 20 |
+
| Input | Why it's hard | Model |
|
| 21 |
+
|---|---|---|
|
| 22 |
+
| `Powódka Anna Nowak-Kowalska złożyła wniosek.` | hyphenated double surname → one span | `PER 'Anna Nowak-Kowalska'` |
|
| 23 |
+
| `Stawiła się ALEKSANDRA ŚWIDERSKA, pełnomocnik.` | ALL-CAPS header/signature style | `PER 'ALEKSANDRA ŚWIDERSKA'` |
|
| 24 |
+
| `Z poważaniem,` / `Iga Mełech` / `radca prawny` | short female first + rare surname in a low-cue signature | `PER 'Iga Mełech'` |
|
| 25 |
+
|
| 26 |
+
## Mild / moderate OCR noise
|
| 27 |
+
|
| 28 |
+
Light scan corruption (missing diacritics, a few confusable characters) is still
|
| 29 |
+
handled:
|
| 30 |
+
|
| 31 |
+
| Input | Behind the garble | Model |
|
| 32 |
+
|---|---|---|
|
| 33 |
+
| `Pozwany Damian Zotadek wniósł odpowiedź na pozew.` | *Damian Żołądek* (lost diacritics) | `PER 'Damian Zotadek'` |
|
| 34 |
+
| `Z powództwa Jiga Bękowiki przeciwko spółce.` | *Jaga Bętkowski* | `PER 'Jiga Bękowiki'` |
|
| 35 |
+
| `Pełnomocnik: Ptek Gostoiski, adwokat.` | *Patryk Gostomski* | `PER 'Ptek Gostoiski'` |
|
| 36 |
+
|
| 37 |
+
> **Heavy** OCR garble is this variant's weak spot — see
|
| 38 |
+
> [`KNOWN_LIMITATIONS.md`](KNOWN_LIMITATIONS.md). For scanned documents prefer the
|
| 39 |
+
> [OCR-robust variant](https://huggingface.co/lexedit/herbert-polish-legal-ner-ocr).
|
| 40 |
+
|
| 41 |
+
## Multiple types + public/private distinction
|
| 42 |
+
|
| 43 |
+
```
|
| 44 |
+
Jan Kowalski, zam. ul. Słoneczna 5 w Krakowie, PESEL 02070803628.
|
| 45 |
+
PER 'Jan Kowalski'
|
| 46 |
+
LOC 'ul. Słoneczna 5' ← private address (masked)
|
| 47 |
+
LOC_PUB 'Krakowie' ← public city (kept by default)
|
| 48 |
+
ID '02070803628' ← national id (PESEL)
|
| 49 |
+
```
|
| 50 |
+
|
| 51 |
+
The model keeps the public place (`Krakowie`) visible while masking the private
|
| 52 |
+
street address — so the legal context survives anonymisation.
|
examples/inference_onnx.py
ADDED
|
@@ -0,0 +1,76 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
#!/usr/bin/env python3
|
| 2 |
+
"""Self-contained ONNX inference for the HerBERT Polish legal NER model.
|
| 3 |
+
No PyTorch needed. pip install onnxruntime transformers numpy
|
| 4 |
+
|
| 5 |
+
Run from the repo root: python examples/inference_onnx.py
|
| 6 |
+
"""
|
| 7 |
+
import json
|
| 8 |
+
from pathlib import Path
|
| 9 |
+
|
| 10 |
+
import numpy as np
|
| 11 |
+
import onnxruntime as ort
|
| 12 |
+
from transformers import AutoTokenizer
|
| 13 |
+
|
| 14 |
+
ROOT = Path(__file__).resolve().parent.parent
|
| 15 |
+
PER_THRESHOLD = 0.2 # recall-first: flip a token to PER if summed PER prob >= this
|
| 16 |
+
|
| 17 |
+
tok = AutoTokenizer.from_pretrained(str(ROOT))
|
| 18 |
+
cfg = json.load(open(ROOT / "config.json", encoding="utf-8"))
|
| 19 |
+
id2label = {int(k): v for k, v in cfg["id2label"].items()}
|
| 20 |
+
label2id = cfg["label2id"]
|
| 21 |
+
sess = ort.InferenceSession(str(ROOT / "onnx" / "model_quantized.onnx"))
|
| 22 |
+
in_names = {i.name for i in sess.get_inputs()}
|
| 23 |
+
|
| 24 |
+
|
| 25 |
+
def softmax(x):
|
| 26 |
+
e = np.exp(x - x.max(-1, keepdims=True))
|
| 27 |
+
return e / e.sum(-1, keepdims=True)
|
| 28 |
+
|
| 29 |
+
|
| 30 |
+
def predict(text):
|
| 31 |
+
enc = tok(text, return_offsets_mapping=True, return_tensors="np",
|
| 32 |
+
truncation=True, max_length=512)
|
| 33 |
+
offsets = enc["offset_mapping"][0]
|
| 34 |
+
feeds = {"input_ids": enc["input_ids"].astype(np.int64),
|
| 35 |
+
"attention_mask": enc["attention_mask"].astype(np.int64)}
|
| 36 |
+
if "token_type_ids" in in_names:
|
| 37 |
+
feeds["token_type_ids"] = np.zeros_like(enc["input_ids"], dtype=np.int64)
|
| 38 |
+
probs = softmax(sess.run(None, feeds)[0][0]) # (seq, num_labels)
|
| 39 |
+
ids = probs.argmax(-1)
|
| 40 |
+
per_b, per_i = label2id["B-PER"], label2id["I-PER"]
|
| 41 |
+
|
| 42 |
+
spans, cur = [], None # cur = [start, end, type]
|
| 43 |
+
for i, (s, e) in enumerate(offsets):
|
| 44 |
+
if s == e: # special token
|
| 45 |
+
if cur:
|
| 46 |
+
spans.append(cur); cur = None
|
| 47 |
+
continue
|
| 48 |
+
label = id2label[int(ids[i])]
|
| 49 |
+
# recall-first override for persons
|
| 50 |
+
if label == "O" and probs[i, per_b] + probs[i, per_i] >= PER_THRESHOLD:
|
| 51 |
+
label = "I-PER" if (cur and cur[2] == "PER") else "B-PER"
|
| 52 |
+
if label == "O":
|
| 53 |
+
if cur:
|
| 54 |
+
spans.append(cur); cur = None
|
| 55 |
+
continue
|
| 56 |
+
tag, etype = label.split("-", 1)
|
| 57 |
+
if tag == "B" or cur is None or cur[2] != etype:
|
| 58 |
+
if cur:
|
| 59 |
+
spans.append(cur)
|
| 60 |
+
cur = [int(s), int(e), etype]
|
| 61 |
+
else:
|
| 62 |
+
cur[1] = int(e)
|
| 63 |
+
if cur:
|
| 64 |
+
spans.append(cur)
|
| 65 |
+
return [{"type": t, "start": a, "end": b, "text": text[a:b]} for a, b, t in spans]
|
| 66 |
+
|
| 67 |
+
|
| 68 |
+
if __name__ == "__main__":
|
| 69 |
+
samples = [
|
| 70 |
+
"Pozwany Jan Kowalski, zam. ul. Słoneczna 5 w Krakowie, PESEL 02070803628.",
|
| 71 |
+
"Powódka Anna Nowak-Kowalska, e-mail a.nowak@example.pl, tel. 501 234 567.",
|
| 72 |
+
]
|
| 73 |
+
for t in samples:
|
| 74 |
+
print("\n" + t)
|
| 75 |
+
for ent in predict(t):
|
| 76 |
+
print(f" {ent['type']:8} [{ent['start']:>3}:{ent['end']:<3}] {ent['text']!r}")
|
examples/inference_pytorch.py
ADDED
|
@@ -0,0 +1,32 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
#!/usr/bin/env python3
|
| 2 |
+
"""PyTorch / transformers inference for the HerBERT Polish legal NER model.
|
| 3 |
+
pip install torch transformers
|
| 4 |
+
|
| 5 |
+
Run from the repo root: python examples/inference_pytorch.py
|
| 6 |
+
"""
|
| 7 |
+
from pathlib import Path
|
| 8 |
+
|
| 9 |
+
from transformers import pipeline
|
| 10 |
+
|
| 11 |
+
ROOT = Path(__file__).resolve().parent.parent
|
| 12 |
+
|
| 13 |
+
ner = pipeline(
|
| 14 |
+
"token-classification",
|
| 15 |
+
model=str(ROOT),
|
| 16 |
+
tokenizer=str(ROOT),
|
| 17 |
+
aggregation_strategy="first", # group B-/I- subwords into whole entities
|
| 18 |
+
)
|
| 19 |
+
|
| 20 |
+
SAMPLES = [
|
| 21 |
+
"Pozwany Jan Kowalski, zam. ul. Słoneczna 5 w Krakowie, PESEL 02070803628.",
|
| 22 |
+
"Powódka Anna Nowak-Kowalska, e-mail a.nowak@example.pl, tel. 501 234 567.",
|
| 23 |
+
]
|
| 24 |
+
|
| 25 |
+
if __name__ == "__main__":
|
| 26 |
+
for text in SAMPLES:
|
| 27 |
+
print("\n" + text)
|
| 28 |
+
for ent in ner(text):
|
| 29 |
+
print(f" {ent['entity_group']:8} [{ent['start']:>3}:{ent['end']:<3}] "
|
| 30 |
+
f"{ent['word']!r} ({ent['score']:.2f})")
|
| 31 |
+
# NOTE: for anonymisation, prefer a recall-first PER threshold (see
|
| 32 |
+
# examples/inference_onnx.py) over the arg-max grouping the pipeline uses.
|
examples/known_limitations.py
ADDED
|
@@ -0,0 +1,81 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
#!/usr/bin/env python3
|
| 2 |
+
"""Reproduce the model's KNOWN FAILURE CASES (see KNOWN_LIMITATIONS.md).
|
| 3 |
+
|
| 4 |
+
All inputs are SYNTHETIC — fictional names, no real personal data. The point is
|
| 5 |
+
to show, honestly and reproducibly, where the model still misses or mislabels
|
| 6 |
+
PERSON names so you can decide where extra safeguards (post-pass, review) are
|
| 7 |
+
needed. pip install onnxruntime transformers numpy
|
| 8 |
+
|
| 9 |
+
Run from the repo root: python examples/known_limitations.py
|
| 10 |
+
"""
|
| 11 |
+
import json
|
| 12 |
+
from pathlib import Path
|
| 13 |
+
|
| 14 |
+
import numpy as np
|
| 15 |
+
import onnxruntime as ort
|
| 16 |
+
from transformers import AutoTokenizer
|
| 17 |
+
|
| 18 |
+
ROOT = Path(__file__).resolve().parent.parent
|
| 19 |
+
PER_THRESHOLD = 0.2
|
| 20 |
+
|
| 21 |
+
tok = AutoTokenizer.from_pretrained(str(ROOT))
|
| 22 |
+
cfg = json.load(open(ROOT / "config.json", encoding="utf-8"))
|
| 23 |
+
id2label = {int(k): v for k, v in cfg["id2label"].items()}
|
| 24 |
+
label2id = cfg["label2id"]
|
| 25 |
+
sess = ort.InferenceSession(str(ROOT / "onnx" / "model_quantized.onnx"))
|
| 26 |
+
in_names = {i.name for i in sess.get_inputs()}
|
| 27 |
+
|
| 28 |
+
|
| 29 |
+
def softmax(x):
|
| 30 |
+
e = np.exp(x - x.max(-1, keepdims=True)); return e / e.sum(-1, keepdims=True)
|
| 31 |
+
|
| 32 |
+
|
| 33 |
+
def predict_per(text):
|
| 34 |
+
enc = tok(text, return_offsets_mapping=True, return_tensors="np", truncation=True, max_length=512)
|
| 35 |
+
offs = enc["offset_mapping"][0]
|
| 36 |
+
feeds = {"input_ids": enc["input_ids"].astype(np.int64),
|
| 37 |
+
"attention_mask": enc["attention_mask"].astype(np.int64)}
|
| 38 |
+
if "token_type_ids" in in_names:
|
| 39 |
+
feeds["token_type_ids"] = np.zeros_like(enc["input_ids"], dtype=np.int64)
|
| 40 |
+
probs = softmax(sess.run(None, feeds)[0][0])
|
| 41 |
+
ids = probs.argmax(-1)
|
| 42 |
+
pb, pi = label2id["B-PER"], label2id["I-PER"]
|
| 43 |
+
spans, cur = [], None
|
| 44 |
+
for i, (s, e) in enumerate(offs):
|
| 45 |
+
if s == e:
|
| 46 |
+
if cur: spans.append(cur); cur = None
|
| 47 |
+
continue
|
| 48 |
+
lab = id2label[int(ids[i])]
|
| 49 |
+
if lab == "O" and probs[i, pb] + probs[i, pi] >= PER_THRESHOLD:
|
| 50 |
+
lab = "I-PER" if cur else "B-PER"
|
| 51 |
+
if not lab.endswith("PER"):
|
| 52 |
+
if cur: spans.append(cur); cur = None
|
| 53 |
+
continue
|
| 54 |
+
if lab == "B-PER" or cur is None:
|
| 55 |
+
if cur: spans.append(cur)
|
| 56 |
+
cur = [int(s), int(e)]
|
| 57 |
+
else:
|
| 58 |
+
cur[1] = int(e)
|
| 59 |
+
if cur: spans.append(cur)
|
| 60 |
+
return [text[a:b] for a, b in spans]
|
| 61 |
+
|
| 62 |
+
|
| 63 |
+
# (input, fictional clean name behind the garble, failure mode)
|
| 64 |
+
CASES = [
|
| 65 |
+
# 1) Heavy OCR garble -> missed entirely
|
| 66 |
+
("Pozwana ote tobodaaka wniosła sprzeciw.", "Bożena Łobodzińska", "heavy OCR garble -> MISSED"),
|
| 67 |
+
("Powód go Hea stawił się osobiście.", "Igor Heliasz", "heavy OCR garble -> MISSED"),
|
| 68 |
+
("Decyzję doręczono: ist yi daziak.", "Justyna Idziak", "heavy OCR garble -> MISSED"),
|
| 69 |
+
# 2) Partial detection -> residue leak
|
| 70 |
+
("Wniosek złożył wiar Żoją w dniu 5 maja.", "Wiktor Żołądek", "partial -> only part caught (residue)"),
|
| 71 |
+
("WIESEAW KUE stawił się na rozprawie.", "Wiesław Kuc", "ALL-CAPS garble -> fragmented (gap leaks)"),
|
| 72 |
+
]
|
| 73 |
+
|
| 74 |
+
if __name__ == "__main__":
|
| 75 |
+
print(f"PER threshold = {PER_THRESHOLD}\n")
|
| 76 |
+
for text, clean, mode in CASES:
|
| 77 |
+
per = predict_per(text)
|
| 78 |
+
print(f"# {mode} (fictional: {clean})")
|
| 79 |
+
print(f" input : {text}")
|
| 80 |
+
print(f" model : PER = {per if per else '— (nothing)'}")
|
| 81 |
+
print()
|
examples/strong_cases.py
ADDED
|
@@ -0,0 +1,81 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
#!/usr/bin/env python3
|
| 2 |
+
"""Hard cases the model handles WELL — the flip side of KNOWN_LIMITATIONS.md.
|
| 3 |
+
|
| 4 |
+
All inputs are synthetic (fictional names). Outputs are the model's actual
|
| 5 |
+
predictions (quantized ONNX, recall-first PER threshold 0.2). Reproduce:
|
| 6 |
+
python examples/strong_cases.py
|
| 7 |
+
"""
|
| 8 |
+
import json
|
| 9 |
+
from pathlib import Path
|
| 10 |
+
|
| 11 |
+
import numpy as np
|
| 12 |
+
import onnxruntime as ort
|
| 13 |
+
from transformers import AutoTokenizer
|
| 14 |
+
|
| 15 |
+
ROOT = Path(__file__).resolve().parent.parent
|
| 16 |
+
PER_THRESHOLD = 0.2
|
| 17 |
+
|
| 18 |
+
tok = AutoTokenizer.from_pretrained(str(ROOT))
|
| 19 |
+
cfg = json.load(open(ROOT / "config.json", encoding="utf-8"))
|
| 20 |
+
id2label = {int(k): v for k, v in cfg["id2label"].items()}
|
| 21 |
+
label2id = cfg["label2id"]
|
| 22 |
+
sess = ort.InferenceSession(str(ROOT / "onnx" / "model_quantized.onnx"))
|
| 23 |
+
in_names = {i.name for i in sess.get_inputs()}
|
| 24 |
+
|
| 25 |
+
|
| 26 |
+
def softmax(x):
|
| 27 |
+
e = np.exp(x - x.max(-1, keepdims=True)); return e / e.sum(-1, keepdims=True)
|
| 28 |
+
|
| 29 |
+
|
| 30 |
+
def predict(text):
|
| 31 |
+
enc = tok(text, return_offsets_mapping=True, return_tensors="np", truncation=True, max_length=512)
|
| 32 |
+
offs = enc["offset_mapping"][0]
|
| 33 |
+
feeds = {"input_ids": enc["input_ids"].astype(np.int64),
|
| 34 |
+
"attention_mask": enc["attention_mask"].astype(np.int64)}
|
| 35 |
+
if "token_type_ids" in in_names:
|
| 36 |
+
feeds["token_type_ids"] = np.zeros_like(enc["input_ids"], dtype=np.int64)
|
| 37 |
+
probs = softmax(sess.run(None, feeds)[0][0])
|
| 38 |
+
ids = probs.argmax(-1)
|
| 39 |
+
pb, pi = label2id["B-PER"], label2id["I-PER"]
|
| 40 |
+
spans, cur = [], None
|
| 41 |
+
for i, (s, e) in enumerate(offs):
|
| 42 |
+
if s == e:
|
| 43 |
+
if cur: spans.append(cur); cur = None
|
| 44 |
+
continue
|
| 45 |
+
lab = id2label[int(ids[i])]
|
| 46 |
+
if lab == "O" and probs[i, pb] + probs[i, pi] >= PER_THRESHOLD:
|
| 47 |
+
lab = "I-PER" if (cur and cur[2] == "PER") else "B-PER"
|
| 48 |
+
if lab == "O":
|
| 49 |
+
if cur: spans.append(cur); cur = None
|
| 50 |
+
continue
|
| 51 |
+
tag, et = lab.split("-", 1)
|
| 52 |
+
if tag == "B" or cur is None or cur[2] != et:
|
| 53 |
+
if cur: spans.append(cur)
|
| 54 |
+
cur = [int(s), int(e), et]
|
| 55 |
+
else:
|
| 56 |
+
cur[1] = int(e)
|
| 57 |
+
if cur: spans.append(cur)
|
| 58 |
+
return [{"type": t, "start": a, "end": b, "text": text[a:b]} for a, b, t in spans]
|
| 59 |
+
|
| 60 |
+
|
| 61 |
+
CASES = [
|
| 62 |
+
("Pozwany Damian Zotadek wniósł odpowiedź na pozew.", "OCR: missing diacritics (Żołądek)"),
|
| 63 |
+
("Z powództwa Jiga Bękowiki przeciwko spółce.", "OCR: heavily garbled name (Jaga Bętkowski)"),
|
| 64 |
+
("Pełnomocnik: Ptek Gostoiski, adwokat.", "OCR garble in a signature line (Patryk Gostomski)"),
|
| 65 |
+
("Pozew skierowano przeciwko Jerzemu Sekule.", "inflected (dative) Polish name"),
|
| 66 |
+
("Sprawa dotyczy akt Wojciecha Michalika.", "inflected (genitive) Polish name"),
|
| 67 |
+
("Powódka Anna Nowak-Kowalska złożyła wniosek.", "hyphenated double surname (single span)"),
|
| 68 |
+
("Stawiła się ALEKSANDRA ŚWIDERSKA, pełnomocnik.", "ALL-CAPS name (header/signature style)"),
|
| 69 |
+
("Z poważaniem,\nIga Mełech\nradca prawny", "short female first + rare surname in a signature"),
|
| 70 |
+
("Jan Kowalski, zam. ul. Słoneczna 5 w Krakowie, PESEL 02070803628.",
|
| 71 |
+
"private address (LOC) vs public city (LOC_PUB) + national ID"),
|
| 72 |
+
]
|
| 73 |
+
|
| 74 |
+
if __name__ == "__main__":
|
| 75 |
+
print(f"PER threshold = {PER_THRESHOLD}\n")
|
| 76 |
+
for text, note in CASES:
|
| 77 |
+
print(f"# {note}")
|
| 78 |
+
print(f" input : {text.replace(chr(10), ' / ')}")
|
| 79 |
+
for e in predict(text):
|
| 80 |
+
print(f" {e['type']:8} {e['text']!r}")
|
| 81 |
+
print()
|
merges.txt
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
model.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:f19a470d0590f1620950b01338544d5acd494c2827a12fb87d9a99a38d4fa7bf
|
| 3 |
+
size 495521716
|
onnx/model_quantized.onnx
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:5ed91aabec29e07d2d5c7f370a5463283d4d88acbf44dd1ceedd0bca67330b50
|
| 3 |
+
size 125082875
|
requirements.txt
ADDED
|
@@ -0,0 +1,6 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# ONNX inference (no PyTorch needed)
|
| 2 |
+
transformers>=4.40
|
| 3 |
+
onnxruntime>=1.17
|
| 4 |
+
numpy
|
| 5 |
+
# For the PyTorch / transformers-pipeline example instead, also:
|
| 6 |
+
# torch>=2.0
|
special_tokens_map.json
ADDED
|
@@ -0,0 +1,8 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"bos_token": "<s>",
|
| 3 |
+
"cls_token": "<s>",
|
| 4 |
+
"mask_token": "<mask>",
|
| 5 |
+
"pad_token": "<pad>",
|
| 6 |
+
"sep_token": "</s>",
|
| 7 |
+
"unk_token": "<unk>"
|
| 8 |
+
}
|
tokenizer.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
tokenizer_config.json
ADDED
|
@@ -0,0 +1,58 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"added_tokens_decoder": {
|
| 3 |
+
"0": {
|
| 4 |
+
"content": "<s>",
|
| 5 |
+
"lstrip": false,
|
| 6 |
+
"normalized": false,
|
| 7 |
+
"rstrip": false,
|
| 8 |
+
"single_word": false,
|
| 9 |
+
"special": true
|
| 10 |
+
},
|
| 11 |
+
"1": {
|
| 12 |
+
"content": "<pad>",
|
| 13 |
+
"lstrip": false,
|
| 14 |
+
"normalized": false,
|
| 15 |
+
"rstrip": false,
|
| 16 |
+
"single_word": false,
|
| 17 |
+
"special": true
|
| 18 |
+
},
|
| 19 |
+
"2": {
|
| 20 |
+
"content": "</s>",
|
| 21 |
+
"lstrip": false,
|
| 22 |
+
"normalized": false,
|
| 23 |
+
"rstrip": false,
|
| 24 |
+
"single_word": false,
|
| 25 |
+
"special": true
|
| 26 |
+
},
|
| 27 |
+
"3": {
|
| 28 |
+
"content": "<unk>",
|
| 29 |
+
"lstrip": false,
|
| 30 |
+
"normalized": false,
|
| 31 |
+
"rstrip": false,
|
| 32 |
+
"single_word": false,
|
| 33 |
+
"special": true
|
| 34 |
+
},
|
| 35 |
+
"4": {
|
| 36 |
+
"content": "<mask>",
|
| 37 |
+
"lstrip": false,
|
| 38 |
+
"normalized": false,
|
| 39 |
+
"rstrip": false,
|
| 40 |
+
"single_word": false,
|
| 41 |
+
"special": true
|
| 42 |
+
}
|
| 43 |
+
},
|
| 44 |
+
"additional_special_tokens": [],
|
| 45 |
+
"bos_token": "<s>",
|
| 46 |
+
"clean_up_tokenization_spaces": false,
|
| 47 |
+
"cls_token": "<s>",
|
| 48 |
+
"do_lowercase_and_remove_accent": false,
|
| 49 |
+
"extra_special_tokens": {},
|
| 50 |
+
"id2lang": null,
|
| 51 |
+
"lang2id": null,
|
| 52 |
+
"mask_token": "<mask>",
|
| 53 |
+
"model_max_length": 512,
|
| 54 |
+
"pad_token": "<pad>",
|
| 55 |
+
"sep_token": "</s>",
|
| 56 |
+
"tokenizer_class": "HerbertTokenizer",
|
| 57 |
+
"unk_token": "<unk>"
|
| 58 |
+
}
|
vocab.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|