Instructions to use SlayerLab/NERGAL with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use SlayerLab/NERGAL with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="SlayerLab/NERGAL")# Load model directly from transformers import AutoTokenizer, AutoModelForTokenClassification tokenizer = AutoTokenizer.from_pretrained("SlayerLab/NERGAL") model = AutoModelForTokenClassification.from_pretrained("SlayerLab/NERGAL", device_map="auto") - Notebooks
- Google Colab
- Kaggle
1.2.0: phone policy v3 at mask time
Browse filesSame weights, threshold and API. Rules 1.1.3 (candidate 73c185f7, comments
reworded): one span per number of 7+ digits, standalone numbers under 7 digits
not masked. Model phone spans are cut by the same _phone_spans before the union
(nergal.model_keep). Card basis moves to gold restated to policy v3 plus
gold_check_v4 (315 spans). 841-dev, 1.1.2 -> 1.2.0 on that gold: whole 303/315
unchanged, rules FP 83 -> 24, union FP 333 -> 80. No whole value lost in any
benchmark cohort. Dynaword reviewed delta: phone whole 57 -> 65 of 70, 0 lost.
Full Dynaword: masked characters -10,766 (0.21%).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- CHANGELOG.md +23 -1
- README.md +18 -13
- hybrid.json +17 -13
- nergal.py +15 -7
- scrub_pii.py +113 -50
- test_nergal.py +25 -2
CHANGELOG.md
CHANGED
|
@@ -6,7 +6,29 @@ Semver for this island:
|
|
| 6 |
- **MINOR** — new capability, same API
|
| 7 |
- **PATCH** — rules or card fix, same weights and API
|
| 8 |
|
| 9 |
-
Accuracy is 841-dev, union at 0.95
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 10 |
|
| 11 |
## 1.1.2
|
| 12 |
|
|
|
|
| 6 |
- **MINOR** — new capability, same API
|
| 7 |
- **PATCH** — rules or card fix, same weights and API
|
| 8 |
|
| 9 |
+
Accuracy is 841-dev, union at 0.95. Through 1.1.2 the gold has 354 spans; from 1.2.0 it is restated to phone policy v3 (315 spans), and the two are not comparable. A version that changes those numbers must update `hybrid.json` `eval` and the tables below.
|
| 10 |
+
|
| 11 |
+
## 1.2.0
|
| 12 |
+
|
| 13 |
+
Same weights, threshold and API. Rules SHA `b238d5b8…`. The rules file is the frozen and scored candidate `73c185f7…` with two comments reworded; its syntax tree is identical. Phone policy v3 (labelling policy amended 2026-10-02) now applies at mask time, to the rules and to the model spans.
|
| 14 |
+
|
| 15 |
+
- **Phone policy v3:** each number of 7+ digits is its own span, and connectors between numbers (`/`, `,`, `lub`, a spaced dash) stay as text. A short part (an extension `wew. 101`, an alternative ending `/90`, a wrapped line) stays inside its number's span. A standalone number under 7 digits is not masked: emergency numbers (`112`, `997`), helplines and service numbers (`116 111`), short codes. Country and area codes count as digits, and so do keypad letters. A code before a slash or in parentheses belongs to the number after it (`032/2345678 — 032/2345679` is two numbers), and a `+` slash chain (`+420/55/123456`) is one.
|
| 16 |
+
- **Rules:** the service-number and lone-extension passes are removed. Detected phones are merged into runs and cut by `scrub_pii._phone_spans`. A written connector joins two detections, but a bare line break joins only two short detections (one wrapped number). A lone slash followed by at least 5 digits joins its two sides into one number (`022/123-456`).
|
| 17 |
+
- **Model spans:** at 0.95, each model `phone` span is cut by the same `_phone_spans` before the union: one span per full number, and short numbers are dropped. `pii` spans are unchanged. `nergal.model_keep(text, model_spans)` returns the spans that are kept. The weights did not change, so `predict()` still returns short numbers; only `scrub()` / `scrub_many()` drop them.
|
| 18 |
+
- **Gold:** 841-dev is restated to policy v3 mechanically by the policy's own rule (`policy_restatement_v2`: 55 passages change; checked on a seeded 41-passage human spot-check, which it matches except for one passage a later ruling overrode), and 5 passages are amended by a reviewed gold check (`gold_check_v4`). Gold spans: 354 → 315 (phone 169 → 130). Both are listed in `hybrid.json` `eval`.
|
| 19 |
+
|
| 20 |
+
841-dev (restated gold), 1.1.2 → 1.2.0: whole 303/315 and residual 9 are unchanged. Rules FP 83 → 24, union FP 333 → 80, passages with false characters 42 → 1. Char P 94.37% → 98.59%, char R 97.83%, exact-span F1 82.28% → 92.33%. Without the model-span cut, rules 1.1.3 alone give union FP 327: the published model masks short numbers. The restatement and the mask-time cut apply the same policy, so these gains measure agreement with it and are not an independent test.
|
| 21 |
+
|
| 22 |
+
On the benchmark halves of the human test sets (restated the same way), no whole value is lost in any cohort. Union FP: human_test_v2 random 51 → 39, email 136 → 127, registry 4 → 0, bare-ID 3 → 0, e-Delivery 1 → 0; mC4 negatives email 224 → 222; the other cohorts are unchanged. The human_test_v2 half is not independent: one of its rows set the line-break rule and another the 5-digit slash cut.
|
| 23 |
+
|
| 24 |
+
Dynaword rules delta, human-reviewed (40 passages around changed masks; 38 scored, 2 left uncertain and unscored): phone values whole 57 → 65 of 70, false phone characters 373 → 113, no value lost. Two of these passages set the `+` chain and spaced-dash rules, so this is development evidence. Full Dynaword rules delta against 1.1.2 (36 sources, 4,094,255 documents, counts only, unreviewed): rule spans 294,819 → 291,651, masked characters −10,766 (0.21%). In steps: the 1.1.3 v2 candidate changes 2,092 documents (4,661 spans removed, 1,153 added), and v3 then changes 293 (eurlex 287, wikivoyage 6; 335 removed, 675 added). Both steps read the same per-unit input hashes. The 1.1.2 entry's run covered 21 sources on an older checkout, and `samorzad_gov_pl` has changed upstream since, so document counts across entries are not comparable.
|
| 25 |
+
|
| 26 |
+
Known limits: two slashes between short parts (`12 / 345 / 678`) are left as text. The 5 reviewed Dynaword phone values that 1.2.0 does not wholly mask are also missed by 1.1.2. The model-span cut was not reviewed on new text: no new inference was run for this release.
|
| 27 |
+
|
| 28 |
+
| Version | Whole /315 | Residual | Rules FP | Union FP | Char P | Char R | What changed |
|
| 29 |
+
|---|---:|---:|---:|---:|---:|---:|---|
|
| 30 |
+
| 1.1.2, restated gold | 303 | 9 | 83 | 333 | 94.37% | 97.83% | Restated on phone policy v3 and `gold_check_v4` |
|
| 31 |
+
| 1.2.0 | 303 | 9 | 24 | 80 | 98.59% | 97.83% | Phone policy v3 at mask time, rules and model spans |
|
| 32 |
|
| 33 |
## 1.1.2
|
| 34 |
|
README.md
CHANGED
|
@@ -13,7 +13,7 @@ tags:
|
|
| 13 |
- hybrid
|
| 14 |
---
|
| 15 |
|
| 16 |
-
# NERGAL 1.
|
| 17 |
|
| 18 |
**Named Entity Recognition with Grounded Additive Labels**
|
| 19 |
|
|
@@ -21,10 +21,10 @@ SlayerLab hybrid PII cleaner for Polish. Not a chat model. Not a drop-in `pipeli
|
|
| 21 |
|
| 22 |
## TL;DR
|
| 23 |
|
| 24 |
-
Python rules do the identifiers they can prove. A transformer NER head adds phone and other PII spans the regex misses. The cleaner **unions** the two on the original text, then replaces hits with `[Telefon]` or `[PII]`.
|
| 25 |
|
| 26 |
-
- **Version:** `1.
|
| 27 |
-
- **Ground:** `scrub_pii` regex (SHA256 `
|
| 28 |
- **Additive labels:** XLM-RoBERTa-large token classifier, BIO tags `phone` / `pii`, threshold 0.95
|
| 29 |
- **Throughput:** about 80k chars/s on one RTX 4090 with `scrub_many` + `dtype="float16"` and 3 processes (1.0.3: 23k)
|
| 30 |
- **This snapshot:** seed `202609160`, **epoch 5** of a seven-epoch schedule
|
|
@@ -39,7 +39,7 @@ These are target categories, not a guarantee that every occurrence or format is
|
|
| 39 |
|
| 40 |
| Category | Values in scope | Replacement |
|
| 41 |
|---|---|---|
|
| 42 |
-
| Phone contacts | Phone, fax and SMS contact numbers, including foreign,
|
| 43 |
| Email | Email addresses, including recoverable broken or incomplete addresses | `[PII]` |
|
| 44 |
| Personal identifiers | PESEL, passport and identity-document numbers | `[PII]` |
|
| 45 |
| Organization identifiers | NIP/VAT, REGON, KRS, GEMI, LEI and equivalent foreign company registration numbers; labelled DUNS, BDO and RPWDL register-book numbers | `[PII]` |
|
|
@@ -62,13 +62,18 @@ NERGAL is not designed to remove:
|
|
| 62 |
|
| 63 |
These are intended exclusions; false positives can still mask some of this content.
|
| 64 |
|
| 65 |
-
### Known gaps in 1.
|
| 66 |
|
| 67 |
-
Unlabelled phones and identifiers (the rules take a bare phone only in the grouped national forms `601 234 567` and `22 123 45 67`; a plain `601234567` needs a label or the model), unusual formatting and damaged text can escape detection. **VINs and obfuscated emails** (such as `name (at) domain.pl`) are approved annotation targets, but that approval alone does not establish reliable support in the released 1.
|
| 68 |
|
| 69 |
## Versions
|
| 70 |
|
| 71 |
-
841-dev, union at 0.95
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 72 |
|
| 73 |
| Version | Whole /354 | Residual | Rules FP | Union FP | Char P | Char R | What changed |
|
| 74 |
|---|---:|---:|---:|---:|---:|---:|---|
|
|
@@ -79,7 +84,7 @@ Unlabelled phones and identifiers (the rules take a bare phone only in the group
|
|
| 79 |
| 1.1.0 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Batch API (`predict_many`, `scrub_many`), opt-in float16 |
|
| 80 |
| 1.1.1 | 326 | 23 | 98 | 123 | 97.94% | 96.50% | Grouped national phones without a label |
|
| 81 |
| 1.1.1, amended gold | 326 | 23 | 108 | 133 | 97.77% | 96.49% | Restated on `gold_check_v2` |
|
| 82 |
-
|
|
| 83 |
|
| 84 |
## Cue-less phone test
|
| 85 |
|
|
@@ -93,13 +98,13 @@ Unlabelled phones and identifiers (the rules take a bare phone only in the group
|
|
| 93 |
| Passages with false masks /150 | 3 | 4 |
|
| 94 |
| False characters | 83 | 95 |
|
| 95 |
|
| 96 |
-
1.1.1 lost no value 1.1.0 masked; its 12 new false characters are in one passage. The set is enriched by selection, so these numbers say nothing about how common such phones are, and it has a single reviewer. 1.1.2 changes one passage (an email match): false characters 95 → 78, phones unchanged.
|
| 97 |
|
| 98 |
## 841-dev
|
| 99 |
|
| 100 |
Tables use one development split: 841 passages, 215 with gold PII, **354 spans** (169 phone, 185 other PII). It is the `dev` side of a 4,500-passage labelled tranche (3,655 train / 841 dev). Sources match Dynaword (EUR-Lex, HPLT, Wikipedia, parliamentary and government text, plus smaller news/literary slices). Labels mix unchanged silver with human review.
|
| 101 |
|
| 102 |
-
One span, an address that ran on into a glued URL host, was trimmed by a reviewed post-hoc gold check (`gold_check_v2`); numbers from 1.1.2 on use the amended gold.
|
| 103 |
|
| 104 |
The files contain real identifiers, so they are not released with the weights.
|
| 105 |
|
|
@@ -139,7 +144,7 @@ Same 841-dev split, threshold **0.95**. **Naked** is the transformer alone. **
|
|
| 139 |
| XLM-R epoch 5 | naked | 298 | 42 | 67 | 98.78% | 89.36% |
|
| 140 |
| **NERGAL 1.1.2** | **∪ regex** | 326 | 23 | 80 | 98.65% | 96.49% |
|
| 141 |
|
| 142 |
-
Zero-shot [GLiNER 2.5](https://huggingface.co/fastino/gliner2.5-multi-v1) is not competitive here, especially on non-phone PII (7/185 whole vs 160 naked / 179 union). Fine-tuned historical GLiNER is the precise naked baseline (10 false characters) and still trails XLM-R on coverage. NERGAL is XLM-R epoch 5 plus the regex: 147/169 phone, 179/185 other PII. Exact-span precision 86.41%, recall 89.83%, F1 88.09%. The GLiNER rows use the original gold (10 characters of one span differ), and the two GLiNER ∪ regex rows use the 1.1.0 rules; the other rows use the amended gold.
|
| 143 |
|
| 144 |
Trained GLiNER, HerBERT-large, and XLM-R-large were also compared on this split (plot above). GLiNER’s 0.50 coverage lead is the precision collapse in that figure; no GLiNER or HerBERT operating point passed the content-preservation gate.
|
| 145 |
|
|
@@ -174,6 +179,6 @@ masked, counts = nergal.scrub(text)
|
|
| 174 |
|
| 175 |
For many texts, `nergal.scrub_many(texts)` (or `predict_many` for the raw model spans) batches windows across texts: about 2× the throughput of calling `scrub` in a loop on a CUDA GPU. `from_pretrained(..., dtype="float16")` casts the weights at load time on CUDA or MPS: another 1.9× on an RTX 4090, and 841-dev union numbers are unchanged, but scores are not bit-identical to float32 (0.95 spans can differ in rare cases). For a corpus, run 2–3 processes per GPU, because float16 inference is limited by CPU-side tokenization. Measurements are in `CHANGELOG.md`.
|
| 176 |
|
| 177 |
-
`hybrid.json` records version `1.
|
| 178 |
|
| 179 |
Base weights: [`FacebookAI/xlm-roberta-large`](https://huggingface.co/FacebookAI/xlm-roberta-large) revision `c23d21b0620b635a76227c604d44e43a9f0ee389` (MIT).
|
|
|
|
| 13 |
- hybrid
|
| 14 |
---
|
| 15 |
|
| 16 |
+
# NERGAL 1.2.0
|
| 17 |
|
| 18 |
**Named Entity Recognition with Grounded Additive Labels**
|
| 19 |
|
|
|
|
| 21 |
|
| 22 |
## TL;DR
|
| 23 |
|
| 24 |
+
Python rules do the identifiers they can prove. A transformer NER head adds phone and other PII spans the regex misses. The cleaner **unions** the two on the original text, then replaces hits with `[Telefon]` or `[PII]`. Phone spans from both follow phone policy v3: one span per number of 7+ digits, and shorter numbers are not masked.
|
| 25 |
|
| 26 |
+
- **Version:** `1.2.0` (`hybrid.json`, `CHANGELOG.md`)
|
| 27 |
+
- **Ground:** `scrub_pii` regex (SHA256 `b238d5b8…`)
|
| 28 |
- **Additive labels:** XLM-RoBERTa-large token classifier, BIO tags `phone` / `pii`, threshold 0.95
|
| 29 |
- **Throughput:** about 80k chars/s on one RTX 4090 with `scrub_many` + `dtype="float16"` and 3 processes (1.0.3: 23k)
|
| 30 |
- **This snapshot:** seed `202609160`, **epoch 5** of a seven-epoch schedule
|
|
|
|
| 39 |
|
| 40 |
| Category | Values in scope | Replacement |
|
| 41 |
|---|---|---|
|
| 42 |
+
| Phone contacts | Phone, fax and SMS contact numbers of 7+ digits (country and area codes and keypad letters count), including foreign and vanity numbers, one span per number; an extension or alternative ending stays inside its number. Emergency, helpline, service and other numbers under 7 digits are not masked on their own | `[Telefon]` |
|
| 43 |
| Email | Email addresses, including recoverable broken or incomplete addresses | `[PII]` |
|
| 44 |
| Personal identifiers | PESEL, passport and identity-document numbers | `[PII]` |
|
| 45 |
| Organization identifiers | NIP/VAT, REGON, KRS, GEMI, LEI and equivalent foreign company registration numbers; labelled DUNS, BDO and RPWDL register-book numbers | `[PII]` |
|
|
|
|
| 62 |
|
| 63 |
These are intended exclusions; false positives can still mask some of this content.
|
| 64 |
|
| 65 |
+
### Known gaps in 1.2.0
|
| 66 |
|
| 67 |
+
Unlabelled phones and identifiers (the rules take a bare phone only in the grouped national forms `601 234 567` and `22 123 45 67`; a plain `601234567` needs a label or the model), unusual formatting and damaged text can escape detection. **VINs and obfuscated emails** (such as `name (at) domain.pl`) are approved annotation targets, but that approval alone does not establish reliable support in the released 1.2.0 model. Do not rely on it to remove them consistently. An email with a lowercase word glued onto its domain (`jan@firma.plkontakt`) is masked together with that word, and text glued before an address can be masked with it. A phone written as short parts joined by two slashes (`12 / 345 / 678`) is left as text.
|
| 68 |
|
| 69 |
## Versions
|
| 70 |
|
| 71 |
+
841-dev, union at 0.95. Same weights and API is a patch; new capability is minor; API, threshold, or weight recipe is major. A release that changes these numbers updates `hybrid.json` `eval` and this table. From 1.1.2 on, one 841-dev span is amended by a reviewed gold check (`gold_check_v2`, 10 characters trimmed); 1.1.1 is restated on it, and earlier rows use the original gold. From 1.2.0 the gold is restated to phone policy v3 (315 spans, see 841-dev below); 1.1.2 is restated on it, and the two tables are not comparable.
|
| 72 |
+
|
| 73 |
+
| Version | Whole /315 | Residual | Rules FP | Union FP | Char P | Char R | What changed |
|
| 74 |
+
|---|---:|---:|---:|---:|---:|---:|---|
|
| 75 |
+
| 1.1.2, restated gold | 303 | 9 | 83 | 333 | 94.37% | 97.83% | Restated on phone policy v3 and `gold_check_v4` |
|
| 76 |
+
| **1.2.0** | **303** | **9** | **24** | **80** | **98.59%** | **97.83%** | Phone policy v3 at mask time, rules and model spans |
|
| 77 |
|
| 78 |
| Version | Whole /354 | Residual | Rules FP | Union FP | Char P | Char R | What changed |
|
| 79 |
|---|---:|---:|---:|---:|---:|---:|---|
|
|
|
|
| 84 |
| 1.1.0 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Batch API (`predict_many`, `scrub_many`), opt-in float16 |
|
| 85 |
| 1.1.1 | 326 | 23 | 98 | 123 | 97.94% | 96.50% | Grouped national phones without a label |
|
| 86 |
| 1.1.1, amended gold | 326 | 23 | 108 | 133 | 97.77% | 96.49% | Restated on `gold_check_v2` |
|
| 87 |
+
| 1.1.2 | 326 | 23 | 24 | 80 | 98.65% | 96.49% | Email boundary rules |
|
| 88 |
|
| 89 |
## Cue-less phone test
|
| 90 |
|
|
|
|
| 98 |
| Passages with false masks /150 | 3 | 4 |
|
| 99 |
| False characters | 83 | 95 |
|
| 100 |
|
| 101 |
+
1.1.1 lost no value 1.1.0 masked; its 12 new false characters are in one passage. The set is enriched by selection, so these numbers say nothing about how common such phones are, and it has a single reviewer. 1.1.2 changes one passage (an email match): false characters 95 → 78, phones unchanged. Half of this set is now a benchmark on gold restated to phone policy v3 (70 passages): 1.2.0 and 1.1.2 both wholly mask 62 of 83 values there, with 33 false characters each.
|
| 102 |
|
| 103 |
## 841-dev
|
| 104 |
|
| 105 |
Tables use one development split: 841 passages, 215 with gold PII, **354 spans** (169 phone, 185 other PII). It is the `dev` side of a 4,500-passage labelled tranche (3,655 train / 841 dev). Sources match Dynaword (EUR-Lex, HPLT, Wikipedia, parliamentary and government text, plus smaller news/literary slices). Labels mix unchanged silver with human review.
|
| 106 |
|
| 107 |
+
One span, an address that ran on into a glued URL host, was trimmed by a reviewed post-hoc gold check (`gold_check_v2`); numbers from 1.1.2 on use the amended gold. From 1.2.0 the gold is restated to phone policy v3 (55 passages; short and emergency numbers dropped, joined numbers split) and 5 more passages are amended by a reviewed check (`gold_check_v4`): **315 spans** (130 phone, 185 other PII). The tables below that are marked /354 use the earlier gold.
|
| 108 |
|
| 109 |
The files contain real identifiers, so they are not released with the weights.
|
| 110 |
|
|
|
|
| 144 |
| XLM-R epoch 5 | naked | 298 | 42 | 67 | 98.78% | 89.36% |
|
| 145 |
| **NERGAL 1.1.2** | **∪ regex** | 326 | 23 | 80 | 98.65% | 96.49% |
|
| 146 |
|
| 147 |
+
Zero-shot [GLiNER 2.5](https://huggingface.co/fastino/gliner2.5-multi-v1) is not competitive here, especially on non-phone PII (7/185 whole vs 160 naked / 179 union). Fine-tuned historical GLiNER is the precise naked baseline (10 false characters) and still trails XLM-R on coverage. NERGAL 1.1.2 is XLM-R epoch 5 plus the regex: 147/169 phone, 179/185 other PII. Exact-span precision 86.41%, recall 89.83%, F1 88.09%. 1.2.0 is scored only on the restated gold (/315): 124/130 phone, 179/185 other PII, union FP 80, exact-span precision 91.05%, recall 93.65%, F1 92.33%. The GLiNER rows use the original gold (10 characters of one span differ), and the two GLiNER ∪ regex rows use the 1.1.0 rules; the other rows use the amended gold.
|
| 148 |
|
| 149 |
Trained GLiNER, HerBERT-large, and XLM-R-large were also compared on this split (plot above). GLiNER’s 0.50 coverage lead is the precision collapse in that figure; no GLiNER or HerBERT operating point passed the content-preservation gate.
|
| 150 |
|
|
|
|
| 179 |
|
| 180 |
For many texts, `nergal.scrub_many(texts)` (or `predict_many` for the raw model spans) batches windows across texts: about 2× the throughput of calling `scrub` in a loop on a CUDA GPU. `from_pretrained(..., dtype="float16")` casts the weights at load time on CUDA or MPS: another 1.9× on an RTX 4090, and 841-dev union numbers are unchanged, but scores are not bit-identical to float32 (0.95 spans can differ in rare cases). For a corpus, run 2–3 processes per GPU, because float16 inference is limited by CPU-side tokenization. Measurements are in `CHANGELOG.md`.
|
| 181 |
|
| 182 |
+
`hybrid.json` records version `1.2.0`, threshold 0.95, gap ids `250002` / `250003`, the 841-dev `eval` block (with its `gold_amendments` and `gold_restatements`), and two weight hashes: `model_safetensors_sha256` for the published file and `source_checkpoint_sha256` for the training checkpoint it was packed from. `test_nergal.py` is synthetic (no corpus text). From this snapshot: `python -m unittest test_nergal`.
|
| 183 |
|
| 184 |
Base weights: [`FacebookAI/xlm-roberta-large`](https://huggingface.co/FacebookAI/xlm-roberta-large) revision `c23d21b0620b635a76227c604d44e43a9f0ee389` (MIT).
|
hybrid.json
CHANGED
|
@@ -1,29 +1,33 @@
|
|
| 1 |
{
|
| 2 |
"full_name": "Named Entity Recognition with Grounded Additive Labels",
|
| 3 |
"hub_id": "SlayerLab/NERGAL",
|
| 4 |
-
"version": "1.
|
| 5 |
"mode": "rules_union",
|
| 6 |
"epoch": 5,
|
| 7 |
"seed": 202609160,
|
| 8 |
"threshold": 0.95,
|
| 9 |
-
"rules_sha256": "
|
| 10 |
"eval": {
|
| 11 |
"split": "841-dev",
|
| 12 |
"gold_amendments": [
|
| 13 |
-
"gold_check_v2"
|
|
|
|
| 14 |
],
|
| 15 |
-
"
|
| 16 |
-
|
| 17 |
-
|
|
|
|
|
|
|
|
|
|
| 18 |
"union_fp": 80,
|
| 19 |
"rules_fp": 24,
|
| 20 |
-
"character_precision": 0.
|
| 21 |
-
"character_recall": 0.
|
| 22 |
-
"exact_precision": 0.
|
| 23 |
-
"exact_recall": 0.
|
| 24 |
-
"exact_f1": 0.
|
| 25 |
-
"phone_whole":
|
| 26 |
-
"phone_gold":
|
| 27 |
"pii_whole": 179,
|
| 28 |
"pii_gold": 185
|
| 29 |
},
|
|
|
|
| 1 |
{
|
| 2 |
"full_name": "Named Entity Recognition with Grounded Additive Labels",
|
| 3 |
"hub_id": "SlayerLab/NERGAL",
|
| 4 |
+
"version": "1.2.0",
|
| 5 |
"mode": "rules_union",
|
| 6 |
"epoch": 5,
|
| 7 |
"seed": 202609160,
|
| 8 |
"threshold": 0.95,
|
| 9 |
+
"rules_sha256": "b238d5b88aa3f3d55a24bb051ec93f9179dfb14b2c650441c8e0acb225d81d59",
|
| 10 |
"eval": {
|
| 11 |
"split": "841-dev",
|
| 12 |
"gold_amendments": [
|
| 13 |
+
"gold_check_v2",
|
| 14 |
+
"gold_check_v4"
|
| 15 |
],
|
| 16 |
+
"gold_restatements": [
|
| 17 |
+
"policy_restatement_v2"
|
| 18 |
+
],
|
| 19 |
+
"gold_entities": 315,
|
| 20 |
+
"whole_entities": 303,
|
| 21 |
+
"residual_passages": 9,
|
| 22 |
"union_fp": 80,
|
| 23 |
"rules_fp": 24,
|
| 24 |
+
"character_precision": 0.9859,
|
| 25 |
+
"character_recall": 0.9783,
|
| 26 |
+
"exact_precision": 0.9105,
|
| 27 |
+
"exact_recall": 0.9365,
|
| 28 |
+
"exact_f1": 0.9233,
|
| 29 |
+
"phone_whole": 124,
|
| 30 |
+
"phone_gold": 130,
|
| 31 |
"pii_whole": 179,
|
| 32 |
"pii_gold": 185
|
| 33 |
},
|
nergal.py
CHANGED
|
@@ -17,13 +17,13 @@ import scrub_pii
|
|
| 17 |
from scrub_pii import PHONE_TAG, PII_TAG
|
| 18 |
|
| 19 |
HUB_ID = 'SlayerLab/NERGAL'
|
| 20 |
-
VERSION = '1.
|
| 21 |
GAPS = ['[PII_SPACE]', '[PII_BREAK]']
|
| 22 |
GAP_IDS = [250002, 250003]
|
| 23 |
BIO_LABELS = ['O', 'B-phone', 'I-phone', 'B-pii', 'I-pii']
|
| 24 |
LABELS = ['phone', 'pii']
|
| 25 |
THRESHOLD = 0.95
|
| 26 |
-
RULES_SHA = '
|
| 27 |
DTYPES = ('float32', 'float16')
|
| 28 |
|
| 29 |
|
|
@@ -210,12 +210,20 @@ def apply_union(text, spans):
|
|
| 210 |
return ''.join(out), chars, n_phone, n_pii
|
| 211 |
|
| 212 |
|
| 213 |
-
def
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 214 |
rule_keys = {(s['start'], s['end'], s['label']) for s in rule_spans}
|
| 215 |
-
|
| 216 |
-
extra = sum(1 for s in
|
| 217 |
_, rules_chars, _, _ = apply_union(text, rule_spans)
|
| 218 |
-
masked, union_chars, n_phone, n_pii = apply_union(text, list(rule_spans) +
|
| 219 |
return masked, {
|
| 220 |
'phone': n_phone,
|
| 221 |
'pii': n_pii,
|
|
@@ -336,7 +344,7 @@ class Nergal:
|
|
| 336 |
|
| 337 |
def scrub_many(self, texts):
|
| 338 |
texts = list(texts)
|
| 339 |
-
return [scrub_spans(text, self.rule_spans(text), model, threshold=self.threshold)
|
| 340 |
for text, model in zip(texts, self.predict_many(texts), strict=True)]
|
| 341 |
|
| 342 |
|
|
|
|
| 17 |
from scrub_pii import PHONE_TAG, PII_TAG
|
| 18 |
|
| 19 |
HUB_ID = 'SlayerLab/NERGAL'
|
| 20 |
+
VERSION = '1.2.0'
|
| 21 |
GAPS = ['[PII_SPACE]', '[PII_BREAK]']
|
| 22 |
GAP_IDS = [250002, 250003]
|
| 23 |
BIO_LABELS = ['O', 'B-phone', 'I-phone', 'B-pii', 'I-pii']
|
| 24 |
LABELS = ['phone', 'pii']
|
| 25 |
THRESHOLD = 0.95
|
| 26 |
+
RULES_SHA = 'b238d5b88aa3f3d55a24bb051ec93f9179dfb14b2c650441c8e0acb225d81d59'
|
| 27 |
DTYPES = ('float32', 'float16')
|
| 28 |
|
| 29 |
|
|
|
|
| 210 |
return ''.join(out), chars, n_phone, n_pii
|
| 211 |
|
| 212 |
|
| 213 |
+
def model_keep(text, model_spans, *, threshold=THRESHOLD, module=scrub_pii):
|
| 214 |
+
"""Decoded model spans, each phone span cut by the rules' phone policy: one span per number of 7+ digits; short
|
| 215 |
+
and emergency numbers alone are dropped. Other labels pass unchanged."""
|
| 216 |
+
return [dict(s, start=a, end=b) for s in decode(model_spans, threshold)
|
| 217 |
+
for a, b in (module._phone_spans(text, s['start'], s['end']) if s['label'] == 'phone'
|
| 218 |
+
else [(s['start'], s['end'])])]
|
| 219 |
+
|
| 220 |
+
|
| 221 |
+
def scrub_spans(text, rule_spans, model_spans, *, threshold=THRESHOLD, module=scrub_pii):
|
| 222 |
rule_keys = {(s['start'], s['end'], s['label']) for s in rule_spans}
|
| 223 |
+
keep = model_keep(text, model_spans, threshold=threshold, module=module)
|
| 224 |
+
extra = sum(1 for s in keep if (s['start'], s['end'], s['label']) not in rule_keys)
|
| 225 |
_, rules_chars, _, _ = apply_union(text, rule_spans)
|
| 226 |
+
masked, union_chars, n_phone, n_pii = apply_union(text, list(rule_spans) + keep)
|
| 227 |
return masked, {
|
| 228 |
'phone': n_phone,
|
| 229 |
'pii': n_pii,
|
|
|
|
| 344 |
|
| 345 |
def scrub_many(self, texts):
|
| 346 |
texts = list(texts)
|
| 347 |
+
return [scrub_spans(text, self.rule_spans(text), model, threshold=self.threshold, module=self._scrub)
|
| 348 |
for text, model in zip(texts, self.predict_many(texts), strict=True)]
|
| 349 |
|
| 350 |
|
scrub_pii.py
CHANGED
|
@@ -26,10 +26,14 @@ formatted phone fields. Labelled full-number ranges retain shared prefixes.
|
|
| 26 |
Phones require a nearby contact cue, a Polish +48/0048 prefix, or explicit
|
| 27 |
international country/trunk notation such as +CC (0). With a cue, the EUR-Lex
|
| 28 |
"(32-2) 299 11 11" country-area form counts as international. Strong labels also
|
| 29 |
-
admit one-digit country codes and wider hyphenated area codes.
|
| 30 |
-
|
| 31 |
-
|
| 32 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 33 |
Flattened tables can glue labels to both neighbours ("Mödlingtel.: … 38112faks:");
|
| 34 |
such glued labels count as labels and end the preceding number.
|
| 35 |
|
|
@@ -181,14 +185,6 @@ _PHONE_BREAK_RE = re.compile(
|
|
| 181 |
r"|\d{1,2}[.)][ \t]+\d{1,2}[./]\d{1,2}[./](?:19|20)\d{2})")
|
| 182 |
_SECTION_MARKER_RE = re.compile(r"\d{1,2}[.)](?!\d)")
|
| 183 |
_PHONE_EXTENSION_RE =re.compile(r"(?:(?:[ \t]+,?[ \t]*|,[ \t]*)wew(?:n(?:ętrzny)?)?\.?[ \t]*\d{1,5}|[ \t]+do[ \t]+\d{1,3}(?=[ \t]*(?:[.;,](?!\d)|\r?\n|\Z)))(?!\w|[ \t]*\d)", re.I)
|
| 184 |
-
_SERVICE_PHONE_RE = re.compile(r"(?<![\w+])[1-9]\d{2}(?![\w\d]|[ \t./()-]*\d)")
|
| 185 |
-
_SERVICE_PHONE_LABEL_RE = re.compile(
|
| 186 |
-
r"(?:\b(?:tel(?:efon\w*)?|phone)[.: \t]{0,8}(?:alarmow\w*[.: \t]{0,8})?|"
|
| 187 |
-
r"\b(?:numer[ \t]+)?bezpłatny[.: \t]{0,8}|"
|
| 188 |
-
r"\b(?:bezpłatny[ \t]+)?numer[ \t]+alarmowy[.: \t]{0,8}|"
|
| 189 |
-
r"(?:^|\n)[ \t]*(?:Straż pożarna|Policja|Pogotowie|Ratunek|Lekarz pogotowia ratunkowego)"
|
| 190 |
-
r"(?:[ \t]+\([^()\n]{1,60}\))?[ \t]*:[ \t]*)"
|
| 191 |
-
r"(?:[1-9]\d{2}[ \t]*(?:[,;]|i|lub|oraz)[ \t]*)*\Z", re.I)
|
| 192 |
_OTHER_NUMBER_LABEL_RE = re.compile(r"\b(?:NIP|REGON|PESEL|KRS|ISBN|kod)\b[^\d\n]{0,20}\Z", re.I)
|
| 193 |
# Grouped national phones need no cue: web contact blocks write them bare. Plain 9-digit strings stay cue-gated,
|
| 194 |
# since many unlabelled ones are not phones.
|
|
@@ -653,9 +649,105 @@ def _replace_checked(text: str, pattern: re.Pattern, tag: str, ok,
|
|
| 653 |
return pattern.sub(_sub, text), n
|
| 654 |
|
| 655 |
|
| 656 |
-
|
| 657 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 658 |
out = []
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 659 |
pos = 0
|
| 660 |
last_phone_end = None
|
| 661 |
last_phone_complete = False
|
|
@@ -663,11 +755,10 @@ def _replace_phones(text: str, *, mask=_tag) -> tuple[str, int]:
|
|
| 663 |
while True:
|
| 664 |
m = _PHONE_RE.search(text, pos)
|
| 665 |
if not m:
|
| 666 |
-
out.append(text[pos:])
|
| 667 |
break
|
| 668 |
raw = m[0]
|
| 669 |
end = m.end()
|
| 670 |
-
|
| 671 |
line_end = text.find('\n', m.start(), end)
|
| 672 |
if line_end >= 0 and re.fullmatch(r'\d{1,3}', text[m.start():line_end]):
|
| 673 |
next_end = text.find('\n', line_end+1, end)
|
|
@@ -675,7 +766,6 @@ def _replace_phones(text: str, *, mask=_tag) -> tuple[str, int]:
|
|
| 675 |
if (_directory_context(text, line_end+1, next_end)
|
| 676 |
or _grouped_phone(text, line_end+1, text[line_end+1:next_end])):
|
| 677 |
# A flattened table's room cell or an address's house number is not a phone country prefix.
|
| 678 |
-
out.append(text[pos:line_end+1])
|
| 679 |
pos = line_end+1
|
| 680 |
continue
|
| 681 |
immediate_phone_label = bool(_SHORT_PHONE_LABEL_RE.search(text[max(0, m.start()-100):m.start()]))
|
|
@@ -717,7 +807,6 @@ def _replace_phones(text: str, *, mask=_tag) -> tuple[str, int]:
|
|
| 717 |
if novel_country_area_ok and (cued or _grouped_phone(text, m.start(), raw)):
|
| 718 |
if (_is_amount(text, m.start(), end)
|
| 719 |
and re.fullmatch(r'\d{1,3}(?:\.[ \t]*\d{3})+', raw)):
|
| 720 |
-
out.extend((text[pos:m.start()], raw))
|
| 721 |
pos = end
|
| 722 |
continue
|
| 723 |
# Stop before a new line or opening hours, but only after a complete
|
|
@@ -737,10 +826,8 @@ def _replace_phones(text: str, *, mask=_tag) -> tuple[str, int]:
|
|
| 737 |
if _phone_ok(prefix, short and section) and not _is_amount(text, m.start(), gap.start()):
|
| 738 |
end = m.start() + len(prefix)
|
| 739 |
raw = prefix
|
| 740 |
-
replacement = raw
|
| 741 |
break
|
| 742 |
if _is_amount(text, m.start(), end):
|
| 743 |
-
out.extend((text[pos:m.start()], raw))
|
| 744 |
pos = end
|
| 745 |
continue
|
| 746 |
# A suffix range repeats the final extension digits, not a second
|
|
@@ -765,8 +852,8 @@ def _replace_phones(text: str, *, mask=_tag) -> tuple[str, int]:
|
|
| 765 |
# Parentheses enclosing prose are punctuation, not phone syntax.
|
| 766 |
if raw.startswith('(+') and raw.count('(') > raw.count(')') and text[end-1:end] != ')':
|
| 767 |
start += 1
|
| 768 |
-
|
| 769 |
-
|
| 770 |
# Split only pairs (<=22 digits); longer lists need label
|
| 771 |
# context to distinguish them from accounts. Never redact ID tails.
|
| 772 |
elif 18 <= len(_digits(raw)) <= 22 and "." not in raw:
|
|
@@ -775,36 +862,15 @@ def _replace_phones(text: str, *, mask=_tag) -> tuple[str, int]:
|
|
| 775 |
gaps = sorted(re.finditer(r"[ \t]*/[ \t]*|[ \t\n]+", raw), key=lambda g: '/' not in g[0])
|
| 776 |
for gap in gaps:
|
| 777 |
if _phone_ok(raw[:gap.start()]) and _phone_ok(raw[gap.end():], '/' in gap[0]):
|
| 778 |
-
|
| 779 |
-
|
| 780 |
-
n += 2
|
| 781 |
break
|
| 782 |
-
|
| 783 |
-
if replacement != raw:
|
| 784 |
last_phone_end = end
|
| 785 |
last_phone_complete = _phone_ok(raw)
|
| 786 |
last_phone_cued = bool(cued)
|
| 787 |
pos = end
|
| 788 |
-
return
|
| 789 |
-
|
| 790 |
-
|
| 791 |
-
def _replace_extensions(text: str, *, mask=_tag) -> tuple[str, int]:
|
| 792 |
-
count = 0
|
| 793 |
-
|
| 794 |
-
def replace(m):
|
| 795 |
-
nonlocal count
|
| 796 |
-
before = text[max(0, m.start()-500):m.start()]
|
| 797 |
-
# Only a telephone followed by extension/name entries establishes scope.
|
| 798 |
-
heading = re.search(r'\btel(?:efon)?[.: \t]*\[Telefon\]'
|
| 799 |
-
r'(?P<entries>(?:\s*wewn?\.?[ \t]+\d{1,5}[ \t]+[^\d\n]+)*\s*)\Z', before, re.I)
|
| 800 |
-
if not heading or _is_amount(text, m.start('number'), m.end('number')):
|
| 801 |
-
return m[0]
|
| 802 |
-
count += 1
|
| 803 |
-
return m['label'] + mask(m.start('number'), m.end('number'), PHONE_TAG)
|
| 804 |
-
|
| 805 |
-
output = re.sub(r'(?P<label>(?:^|\n)[ \t]*wewn?\.?[ \t]+)(?P<number>\d{1,5})(?!\w)',
|
| 806 |
-
replace, text, flags=re.I)
|
| 807 |
-
return output, count
|
| 808 |
|
| 809 |
|
| 810 |
def scrub_pii(text: str, *, spans: list | None = None) -> tuple[str, dict[str, int]]:
|
|
@@ -880,9 +946,6 @@ def scrub_pii(text: str, *, spans: list | None = None) -> tuple[str, dict[str, i
|
|
| 880 |
counts["email"] += apply(_replace_checked, _EMAIL_FRAGMENT_RE, PII_TAG,
|
| 881 |
lambda _: True, _EMAIL_LABEL_RE)
|
| 882 |
counts["phone"] = apply(_replace_phones)
|
| 883 |
-
counts["phone"] += apply(_replace_checked, _SERVICE_PHONE_RE, PHONE_TAG,
|
| 884 |
-
lambda _: True, _SERVICE_PHONE_LABEL_RE)
|
| 885 |
-
counts['phone'] += apply(_replace_extensions)
|
| 886 |
# Contact context takes precedence over coincidental PESEL checksums in
|
| 887 |
# foreign phone numbers; explicitly labelled identifiers were handled first.
|
| 888 |
bare_pesel = apply(_replace_checked, _PESEL_RE, PII_TAG, _pesel_ok)
|
|
|
|
| 26 |
Phones require a nearby contact cue, a Polish +48/0048 prefix, or explicit
|
| 27 |
international country/trunk notation such as +CC (0). With a cue, the EUR-Lex
|
| 28 |
"(32-2) 299 11 11" country-area form counts as international. Strong labels also
|
| 29 |
+
admit one-digit country codes and wider hyphenated area codes. Phone policy v3: each
|
| 30 |
+
number of at least 7 digits (keypad letters count) is its own [Telefon], connectors
|
| 31 |
+
between numbers stay unmasked, and a shorter part (an extension, "/90") stays inside
|
| 32 |
+
its number's span. A phone match with no such number (a short service or emergency
|
| 33 |
+
number, a lone extension) is left as text; a line break alone never splits a
|
| 34 |
+
number, nor does one slash before 5+ digits. Unlabelled domestic numbers
|
| 35 |
+
are left for audit because table cells have the same shapes, except the grouped
|
| 36 |
+
national forms: mobile 3-3-3 and landline 2-3-2-2, one separator kind.
|
| 37 |
Flattened tables can glue labels to both neighbours ("Mödlingtel.: … 38112faks:");
|
| 38 |
such glued labels count as labels and end the preceding number.
|
| 39 |
|
|
|
|
| 185 |
r"|\d{1,2}[.)][ \t]+\d{1,2}[./]\d{1,2}[./](?:19|20)\d{2})")
|
| 186 |
_SECTION_MARKER_RE = re.compile(r"\d{1,2}[.)](?!\d)")
|
| 187 |
_PHONE_EXTENSION_RE =re.compile(r"(?:(?:[ \t]+,?[ \t]*|,[ \t]*)wew(?:n(?:ętrzny)?)?\.?[ \t]*\d{1,5}|[ \t]+do[ \t]+\d{1,3}(?=[ \t]*(?:[.;,](?!\d)|\r?\n|\Z)))(?!\w|[ \t]*\d)", re.I)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 188 |
_OTHER_NUMBER_LABEL_RE = re.compile(r"\b(?:NIP|REGON|PESEL|KRS|ISBN|kod)\b[^\d\n]{0,20}\Z", re.I)
|
| 189 |
# Grouped national phones need no cue: web contact blocks write them bare. Plain 9-digit strings stay cue-gated,
|
| 190 |
# since many unlabelled ones are not phones.
|
|
|
|
| 649 |
return pattern.sub(_sub, text), n
|
| 650 |
|
| 651 |
|
| 652 |
+
# Phone policy v3 (labelling policy amended 2026-10-02): the mask-time copy of the rule that restated the evaluation
|
| 653 |
+
# gold. Each number of 7+ digits is its own span; shorter and emergency numbers are not masked alone.
|
| 654 |
+
_PHONE_CONNECTOR_RE = re.compile(r'\s*(?:[,;/&\n]|\b(?:lub|albo|i|oraz|and)\b)\s*', re.I)
|
| 655 |
+
_KEYPAD_RE = re.compile(r'(?<=[\d-])[A-Z]+')
|
| 656 |
+
_PHONE_CODE_RE = re.compile(r'\(?\+?\d{2,4}\)?')
|
| 657 |
+
_PHONE_DASH_RE = re.compile(r'\s+[-–—]\s+')
|
| 658 |
+
_FULL_PHONE = 7
|
| 659 |
+
|
| 660 |
+
|
| 661 |
+
def _phone_size(part: str) -> int:
|
| 662 |
+
return len(_digits(part)) + sum(map(len, _KEYPAD_RE.findall(part)))
|
| 663 |
+
|
| 664 |
+
|
| 665 |
+
def _phone_parts(text: str, start: int, end: int) -> list[tuple[int, int]]:
|
| 666 |
+
"""The parts of one phone match between connectors. A spaced dash also separates once a full number precedes it
|
| 667 |
+
("ddd/ddddddd – ddd/ddddddd"), not inside one ("601 – 234 – 567"). A part led by "+" takes the slash-joined parts
|
| 668 |
+
after it until it is full ("+ddd/dd/dddddd"): a dialling chain, not short numbers."""
|
| 669 |
+
bounds = [start, *(start + i for c in _PHONE_CONNECTOR_RE.finditer(text[start:end]) for i in c.span()), end]
|
| 670 |
+
ps = []
|
| 671 |
+
for a, b in zip(bounds[::2], bounds[1::2]):
|
| 672 |
+
if not text[a:b].strip():
|
| 673 |
+
continue
|
| 674 |
+
for d in _PHONE_DASH_RE.finditer(text, a, b):
|
| 675 |
+
if _phone_size(text[a:d.start()]) >= _FULL_PHONE and _phone_size(text[d.end():b]):
|
| 676 |
+
ps.append((a, d.start()))
|
| 677 |
+
a = d.end()
|
| 678 |
+
ps.append((a, b))
|
| 679 |
out = []
|
| 680 |
+
for a, b in ps:
|
| 681 |
+
if out and text[out[-1][1]:a].strip() == '/' and re.match(r'\(?\+', text[out[-1][0]:out[-1][1]].strip()) and (
|
| 682 |
+
_phone_size(text[out[-1][0]:out[-1][1]]) < _FULL_PHONE):
|
| 683 |
+
out[-1] = (out[-1][0], b)
|
| 684 |
+
else:
|
| 685 |
+
out.append((a, b))
|
| 686 |
+
return out
|
| 687 |
+
|
| 688 |
+
|
| 689 |
+
def _phone_spans(text: str, start: int, end: int) -> list[tuple[int, int]]:
|
| 690 |
+
"""The policy v3 spans of one phone match: none, the match itself, or one per full number."""
|
| 691 |
+
ps = _phone_parts(text, start, end)
|
| 692 |
+
full = [i for i, (a, b) in enumerate(ps) if _phone_size(text[a:b]) >= _FULL_PHONE]
|
| 693 |
+
if not full:
|
| 694 |
+
# A line break alone never separates entries, so a number wrapped into short lines keeps its mask;
|
| 695 |
+
# and one slash between two short parts is a code separator ("(+48)1234/567890"), not two numbers (user,
|
| 696 |
+
# 2026-10-03), when 5+ digits follow it: a subscriber part, not a year ("123/2019", "2019/2020"). Two
|
| 697 |
+
# slashes ("12 / 345 / 678") still drop unless a "+" leads them (_phone_parts). Each stretch between
|
| 698 |
+
# other written connectors is judged alone, so a list of two such numbers keeps both.
|
| 699 |
+
gaps = [text[b:a].strip() for (_, b), (a, _) in zip(ps, ps[1:])]
|
| 700 |
+
out, i = [], 0
|
| 701 |
+
for j in [*(k + 1 for k, g in enumerate(gaps) if g not in ('', '/')), len(ps)]:
|
| 702 |
+
whole = text[ps[i][0]:ps[j - 1][1]]
|
| 703 |
+
if gaps[i:j - 1].count('/') <= 1 and _phone_size(whole) >= _FULL_PHONE and _phone_size(
|
| 704 |
+
whole.rpartition('/')[2]) >= 5:
|
| 705 |
+
out.append((ps[i][0], ps[j - 1][1]))
|
| 706 |
+
i = j
|
| 707 |
+
return [(start, end)] if out == [(ps[0][0], ps[-1][1])] else out
|
| 708 |
+
if len(full) == 1:
|
| 709 |
+
return [(start, end)]
|
| 710 |
+
|
| 711 |
+
def code_prefix(i): # "(22) / 601 234 567": an area code before a slash belongs to the next number
|
| 712 |
+
part, before = text[ps[i][0]:ps[i][1]].strip(), text[ps[i - 1][1]:ps[i][0]] if i else ''
|
| 713 |
+
return bool(_PHONE_CODE_RE.fullmatch(part)) and '/' not in before and (
|
| 714 |
+
part.startswith('(') or text[ps[i][1]:ps[i + 1][0]].strip() == '/')
|
| 715 |
+
|
| 716 |
+
cuts = [0, *(i - code_prefix(i - 1) for i in full[1:]), len(ps)]
|
| 717 |
+
out = []
|
| 718 |
+
for i, j in zip(cuts, cuts[1:]):
|
| 719 |
+
a = start if i == 0 else ps[i][0] + re.search(r'[+(\d]', text[ps[i][0]:ps[i][1]]).start()
|
| 720 |
+
# the last part that ends like a number: a word between connectors ("601 234 567, fax, 602 345 678", which a
|
| 721 |
+
# model span can hold) ends nothing. ps[i] is full or a code, so one always does.
|
| 722 |
+
b = end if j == len(ps) else max(ps[k][0] + m.start() + 1 for k in range(i, j)
|
| 723 |
+
if (m := re.search(r'[\dA-Z)][^\dA-Z)]*$', text[ps[k][0]:ps[k][1]])))
|
| 724 |
+
out.append((a, b))
|
| 725 |
+
return out
|
| 726 |
+
|
| 727 |
+
|
| 728 |
+
def _mask_phone_runs(text: str, found: list[tuple[int, int]], mask) -> tuple[str, int]:
|
| 729 |
+
"""Mask detected phones as policy v3 spans. Detections joined by one written connector form one run, as a joined
|
| 730 |
+
gold span did, so a short continuation stays inside its full number's span. A bare line break joins only two
|
| 731 |
+
short detections, the fragments of one wrapped number; a short line after a full number is not its part."""
|
| 732 |
+
runs = []
|
| 733 |
+
for a, b in found:
|
| 734 |
+
gap = text[runs[-1][1]:a] if runs else ''
|
| 735 |
+
wrapped = runs and not gap.strip() and max(
|
| 736 |
+
_phone_size(text[runs[-1][0]:runs[-1][1]]), _phone_size(text[a:b])) < _FULL_PHONE
|
| 737 |
+
if runs and _PHONE_CONNECTOR_RE.fullmatch(gap) and (gap.strip() or wrapped):
|
| 738 |
+
runs[-1][1] = b
|
| 739 |
+
else:
|
| 740 |
+
runs.append([a, b])
|
| 741 |
+
out, pos, n = [], 0, 0
|
| 742 |
+
for a, b in runs:
|
| 743 |
+
for start, end in _phone_spans(text, a, b):
|
| 744 |
+
out += [text[pos:start], mask(start, end, PHONE_TAG)]
|
| 745 |
+
pos, n = end, n + 1
|
| 746 |
+
return ''.join(out) + text[pos:], n
|
| 747 |
+
|
| 748 |
+
|
| 749 |
+
def _replace_phones(text: str, *, mask=_tag) -> tuple[str, int]:
|
| 750 |
+
found = [] # (start, end) of each detected phone, masked by policy v3 at the end
|
| 751 |
pos = 0
|
| 752 |
last_phone_end = None
|
| 753 |
last_phone_complete = False
|
|
|
|
| 755 |
while True:
|
| 756 |
m = _PHONE_RE.search(text, pos)
|
| 757 |
if not m:
|
|
|
|
| 758 |
break
|
| 759 |
raw = m[0]
|
| 760 |
end = m.end()
|
| 761 |
+
hit = False
|
| 762 |
line_end = text.find('\n', m.start(), end)
|
| 763 |
if line_end >= 0 and re.fullmatch(r'\d{1,3}', text[m.start():line_end]):
|
| 764 |
next_end = text.find('\n', line_end+1, end)
|
|
|
|
| 766 |
if (_directory_context(text, line_end+1, next_end)
|
| 767 |
or _grouped_phone(text, line_end+1, text[line_end+1:next_end])):
|
| 768 |
# A flattened table's room cell or an address's house number is not a phone country prefix.
|
|
|
|
| 769 |
pos = line_end+1
|
| 770 |
continue
|
| 771 |
immediate_phone_label = bool(_SHORT_PHONE_LABEL_RE.search(text[max(0, m.start()-100):m.start()]))
|
|
|
|
| 807 |
if novel_country_area_ok and (cued or _grouped_phone(text, m.start(), raw)):
|
| 808 |
if (_is_amount(text, m.start(), end)
|
| 809 |
and re.fullmatch(r'\d{1,3}(?:\.[ \t]*\d{3})+', raw)):
|
|
|
|
| 810 |
pos = end
|
| 811 |
continue
|
| 812 |
# Stop before a new line or opening hours, but only after a complete
|
|
|
|
| 826 |
if _phone_ok(prefix, short and section) and not _is_amount(text, m.start(), gap.start()):
|
| 827 |
end = m.start() + len(prefix)
|
| 828 |
raw = prefix
|
|
|
|
| 829 |
break
|
| 830 |
if _is_amount(text, m.start(), end):
|
|
|
|
| 831 |
pos = end
|
| 832 |
continue
|
| 833 |
# A suffix range repeats the final extension digits, not a second
|
|
|
|
| 852 |
# Parentheses enclosing prose are punctuation, not phone syntax.
|
| 853 |
if raw.startswith('(+') and raw.count('(') > raw.count(')') and text[end-1:end] != ')':
|
| 854 |
start += 1
|
| 855 |
+
found.append((start, end))
|
| 856 |
+
hit = True
|
| 857 |
# Split only pairs (<=22 digits); longer lists need label
|
| 858 |
# context to distinguish them from accounts. Never redact ID tails.
|
| 859 |
elif 18 <= len(_digits(raw)) <= 22 and "." not in raw:
|
|
|
|
| 862 |
gaps = sorted(re.finditer(r"[ \t]*/[ \t]*|[ \t\n]+", raw), key=lambda g: '/' not in g[0])
|
| 863 |
for gap in gaps:
|
| 864 |
if _phone_ok(raw[:gap.start()]) and _phone_ok(raw[gap.end():], '/' in gap[0]):
|
| 865 |
+
found += [(m.start(), m.start()+gap.start()), (m.start()+gap.end(), end)]
|
| 866 |
+
hit = True
|
|
|
|
| 867 |
break
|
| 868 |
+
if hit:
|
|
|
|
| 869 |
last_phone_end = end
|
| 870 |
last_phone_complete = _phone_ok(raw)
|
| 871 |
last_phone_cued = bool(cued)
|
| 872 |
pos = end
|
| 873 |
+
return _mask_phone_runs(text, found, mask)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 874 |
|
| 875 |
|
| 876 |
def scrub_pii(text: str, *, spans: list | None = None) -> tuple[str, dict[str, int]]:
|
|
|
|
| 946 |
counts["email"] += apply(_replace_checked, _EMAIL_FRAGMENT_RE, PII_TAG,
|
| 947 |
lambda _: True, _EMAIL_LABEL_RE)
|
| 948 |
counts["phone"] = apply(_replace_phones)
|
|
|
|
|
|
|
|
|
|
| 949 |
# Contact context takes precedence over coincidental PESEL checksums in
|
| 950 |
# foreign phone numbers; explicitly labelled identifiers were handled first.
|
| 951 |
bare_pesel = apply(_replace_checked, _PESEL_RE, PII_TAG, _pesel_ok)
|
test_nergal.py
CHANGED
|
@@ -5,7 +5,7 @@ import unittest
|
|
| 5 |
from pathlib import Path
|
| 6 |
|
| 7 |
HERE = Path(__file__).resolve().parent
|
| 8 |
-
RULES_SHA = '
|
| 9 |
|
| 10 |
|
| 11 |
class NergalTests(unittest.TestCase):
|
|
@@ -13,10 +13,11 @@ class NergalTests(unittest.TestCase):
|
|
| 13 |
from nergal import GAP_IDS, GAPS, HUB_ID, RULES_SHA as PINNED, THRESHOLD, VERSION
|
| 14 |
card = json.loads((HERE / 'hybrid.json').read_text())
|
| 15 |
self.assertEqual(HUB_ID, 'SlayerLab/NERGAL')
|
| 16 |
-
self.assertEqual(VERSION, '1.
|
| 17 |
self.assertEqual(card['version'], VERSION)
|
| 18 |
self.assertEqual(card['eval']['union_fp'], 80)
|
| 19 |
self.assertEqual(card['eval']['rules_fp'], 24)
|
|
|
|
| 20 |
self.assertEqual(GAPS, card['gaps'])
|
| 21 |
self.assertEqual(GAP_IDS, card['gap_ids'])
|
| 22 |
self.assertEqual(THRESHOLD, card['threshold'])
|
|
@@ -105,6 +106,28 @@ class NergalTests(unittest.TestCase):
|
|
| 105 |
self.assertNotIn('000000000', masked)
|
| 106 |
self.assertNotIn('extra', masked)
|
| 107 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 108 |
|
| 109 |
if __name__ == '__main__':
|
| 110 |
unittest.main()
|
|
|
|
| 5 |
from pathlib import Path
|
| 6 |
|
| 7 |
HERE = Path(__file__).resolve().parent
|
| 8 |
+
RULES_SHA = 'b238d5b88aa3f3d55a24bb051ec93f9179dfb14b2c650441c8e0acb225d81d59'
|
| 9 |
|
| 10 |
|
| 11 |
class NergalTests(unittest.TestCase):
|
|
|
|
| 13 |
from nergal import GAP_IDS, GAPS, HUB_ID, RULES_SHA as PINNED, THRESHOLD, VERSION
|
| 14 |
card = json.loads((HERE / 'hybrid.json').read_text())
|
| 15 |
self.assertEqual(HUB_ID, 'SlayerLab/NERGAL')
|
| 16 |
+
self.assertEqual(VERSION, '1.2.0')
|
| 17 |
self.assertEqual(card['version'], VERSION)
|
| 18 |
self.assertEqual(card['eval']['union_fp'], 80)
|
| 19 |
self.assertEqual(card['eval']['rules_fp'], 24)
|
| 20 |
+
self.assertEqual((card['eval']['whole_entities'], card['eval']['gold_entities']), (303, 315)) # restated gold
|
| 21 |
self.assertEqual(GAPS, card['gaps'])
|
| 22 |
self.assertEqual(GAP_IDS, card['gap_ids'])
|
| 23 |
self.assertEqual(THRESHOLD, card['threshold'])
|
|
|
|
| 106 |
self.assertNotIn('000000000', masked)
|
| 107 |
self.assertNotIn('extra', masked)
|
| 108 |
|
| 109 |
+
def test_model_phone_spans_follow_the_phone_policy(self):
|
| 110 |
+
from nergal import model_keep, scrub_spans
|
| 111 |
+
span = lambda text, part, label='phone', score=0.99: dict(
|
| 112 |
+
start=text.index(part), end=text.index(part) + len(part), label=label, score=score)
|
| 113 |
+
for text, part, kept in (('tel. 112', '112', []), # emergency number
|
| 114 |
+
('tel. 51 23 45', '51 23 45', []), # under 7 digits
|
| 115 |
+
('tel. 601 234 567/602 345 678', '601 234 567/602 345 678',
|
| 116 |
+
['601 234 567', '602 345 678']), # one span per number
|
| 117 |
+
('tel. 601 234 567, fax, 602 345 678', '601 234 567, fax, 602 345 678',
|
| 118 |
+
['601 234 567', '602 345 678']), # a word between parts
|
| 119 |
+
('tel. 22 123 45 67 wew. 101', '22 123 45 67 wew. 101',
|
| 120 |
+
['22 123 45 67 wew. 101']), # extension stays inside
|
| 121 |
+
('Jan Kowalski, 112', 'Jan Kowalski', ['Jan Kowalski'])): # other labels unchanged
|
| 122 |
+
label = 'pii' if part[0].isalpha() else 'phone'
|
| 123 |
+
with self.subTest(text=text):
|
| 124 |
+
keep = model_keep(text, [span(text, part, label)])
|
| 125 |
+
self.assertEqual([text[s['start']:s['end']] for s in keep], kept)
|
| 126 |
+
self.assertTrue(all(s['label'] == label and s['score'] == 0.99 for s in keep))
|
| 127 |
+
text = 'tel. 112'
|
| 128 |
+
self.assertEqual(scrub_spans(text, [], [span(text, '112')])[0], text)
|
| 129 |
+
self.assertEqual(model_keep(text, [span(text, '112', score=0.9)], threshold=0.95), [])
|
| 130 |
+
|
| 131 |
|
| 132 |
if __name__ == '__main__':
|
| 133 |
unittest.main()
|