ppuzio Claude Opus 5.5 commited on
Commit
bbd1d02
·
1 Parent(s): e0caa8b

1.2.0: phone policy v3 at mask time

Browse files

Same weights, threshold and API. Rules 1.1.3 (candidate 73c185f7, comments
reworded): one span per number of 7+ digits, standalone numbers under 7 digits
not masked. Model phone spans are cut by the same _phone_spans before the union
(nergal.model_keep). Card basis moves to gold restated to policy v3 plus
gold_check_v4 (315 spans). 841-dev, 1.1.2 -> 1.2.0 on that gold: whole 303/315
unchanged, rules FP 83 -> 24, union FP 333 -> 80. No whole value lost in any
benchmark cohort. Dynaword reviewed delta: phone whole 57 -> 65 of 70, 0 lost.
Full Dynaword: masked characters -10,766 (0.21%).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

Files changed (6) hide show
  1. CHANGELOG.md +23 -1
  2. README.md +18 -13
  3. hybrid.json +17 -13
  4. nergal.py +15 -7
  5. scrub_pii.py +113 -50
  6. test_nergal.py +25 -2
CHANGELOG.md CHANGED
@@ -6,7 +6,29 @@ Semver for this island:
6
  - **MINOR** — new capability, same API
7
  - **PATCH** — rules or card fix, same weights and API
8
 
9
- Accuracy is 841-dev, union at 0.95, 354 gold spans. A version that changes those numbers must update `hybrid.json` `eval` and the tables below.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
10
 
11
  ## 1.1.2
12
 
 
6
  - **MINOR** — new capability, same API
7
  - **PATCH** — rules or card fix, same weights and API
8
 
9
+ Accuracy is 841-dev, union at 0.95. Through 1.1.2 the gold has 354 spans; from 1.2.0 it is restated to phone policy v3 (315 spans), and the two are not comparable. A version that changes those numbers must update `hybrid.json` `eval` and the tables below.
10
+
11
+ ## 1.2.0
12
+
13
+ Same weights, threshold and API. Rules SHA `b238d5b8…`. The rules file is the frozen and scored candidate `73c185f7…` with two comments reworded; its syntax tree is identical. Phone policy v3 (labelling policy amended 2026-10-02) now applies at mask time, to the rules and to the model spans.
14
+
15
+ - **Phone policy v3:** each number of 7+ digits is its own span, and connectors between numbers (`/`, `,`, `lub`, a spaced dash) stay as text. A short part (an extension `wew. 101`, an alternative ending `/90`, a wrapped line) stays inside its number's span. A standalone number under 7 digits is not masked: emergency numbers (`112`, `997`), helplines and service numbers (`116 111`), short codes. Country and area codes count as digits, and so do keypad letters. A code before a slash or in parentheses belongs to the number after it (`032/2345678 — 032/2345679` is two numbers), and a `+` slash chain (`+420/55/123456`) is one.
16
+ - **Rules:** the service-number and lone-extension passes are removed. Detected phones are merged into runs and cut by `scrub_pii._phone_spans`. A written connector joins two detections, but a bare line break joins only two short detections (one wrapped number). A lone slash followed by at least 5 digits joins its two sides into one number (`022/123-456`).
17
+ - **Model spans:** at 0.95, each model `phone` span is cut by the same `_phone_spans` before the union: one span per full number, and short numbers are dropped. `pii` spans are unchanged. `nergal.model_keep(text, model_spans)` returns the spans that are kept. The weights did not change, so `predict()` still returns short numbers; only `scrub()` / `scrub_many()` drop them.
18
+ - **Gold:** 841-dev is restated to policy v3 mechanically by the policy's own rule (`policy_restatement_v2`: 55 passages change; checked on a seeded 41-passage human spot-check, which it matches except for one passage a later ruling overrode), and 5 passages are amended by a reviewed gold check (`gold_check_v4`). Gold spans: 354 → 315 (phone 169 → 130). Both are listed in `hybrid.json` `eval`.
19
+
20
+ 841-dev (restated gold), 1.1.2 → 1.2.0: whole 303/315 and residual 9 are unchanged. Rules FP 83 → 24, union FP 333 → 80, passages with false characters 42 → 1. Char P 94.37% → 98.59%, char R 97.83%, exact-span F1 82.28% → 92.33%. Without the model-span cut, rules 1.1.3 alone give union FP 327: the published model masks short numbers. The restatement and the mask-time cut apply the same policy, so these gains measure agreement with it and are not an independent test.
21
+
22
+ On the benchmark halves of the human test sets (restated the same way), no whole value is lost in any cohort. Union FP: human_test_v2 random 51 → 39, email 136 → 127, registry 4 → 0, bare-ID 3 → 0, e-Delivery 1 → 0; mC4 negatives email 224 → 222; the other cohorts are unchanged. The human_test_v2 half is not independent: one of its rows set the line-break rule and another the 5-digit slash cut.
23
+
24
+ Dynaword rules delta, human-reviewed (40 passages around changed masks; 38 scored, 2 left uncertain and unscored): phone values whole 57 → 65 of 70, false phone characters 373 → 113, no value lost. Two of these passages set the `+` chain and spaced-dash rules, so this is development evidence. Full Dynaword rules delta against 1.1.2 (36 sources, 4,094,255 documents, counts only, unreviewed): rule spans 294,819 → 291,651, masked characters −10,766 (0.21%). In steps: the 1.1.3 v2 candidate changes 2,092 documents (4,661 spans removed, 1,153 added), and v3 then changes 293 (eurlex 287, wikivoyage 6; 335 removed, 675 added). Both steps read the same per-unit input hashes. The 1.1.2 entry's run covered 21 sources on an older checkout, and `samorzad_gov_pl` has changed upstream since, so document counts across entries are not comparable.
25
+
26
+ Known limits: two slashes between short parts (`12 / 345 / 678`) are left as text. The 5 reviewed Dynaword phone values that 1.2.0 does not wholly mask are also missed by 1.1.2. The model-span cut was not reviewed on new text: no new inference was run for this release.
27
+
28
+ | Version | Whole /315 | Residual | Rules FP | Union FP | Char P | Char R | What changed |
29
+ |---|---:|---:|---:|---:|---:|---:|---|
30
+ | 1.1.2, restated gold | 303 | 9 | 83 | 333 | 94.37% | 97.83% | Restated on phone policy v3 and `gold_check_v4` |
31
+ | 1.2.0 | 303 | 9 | 24 | 80 | 98.59% | 97.83% | Phone policy v3 at mask time, rules and model spans |
32
 
33
  ## 1.1.2
34
 
README.md CHANGED
@@ -13,7 +13,7 @@ tags:
13
  - hybrid
14
  ---
15
 
16
- # NERGAL 1.1.2
17
 
18
  **Named Entity Recognition with Grounded Additive Labels**
19
 
@@ -21,10 +21,10 @@ SlayerLab hybrid PII cleaner for Polish. Not a chat model. Not a drop-in `pipeli
21
 
22
  ## TL;DR
23
 
24
- Python rules do the identifiers they can prove. A transformer NER head adds phone and other PII spans the regex misses. The cleaner **unions** the two on the original text, then replaces hits with `[Telefon]` or `[PII]`.
25
 
26
- - **Version:** `1.1.2` (`hybrid.json`, `CHANGELOG.md`)
27
- - **Ground:** `scrub_pii` regex (SHA256 `08faef84…`)
28
  - **Additive labels:** XLM-RoBERTa-large token classifier, BIO tags `phone` / `pii`, threshold 0.95
29
  - **Throughput:** about 80k chars/s on one RTX 4090 with `scrub_many` + `dtype="float16"` and 3 processes (1.0.3: 23k)
30
  - **This snapshot:** seed `202609160`, **epoch 5** of a seven-epoch schedule
@@ -39,7 +39,7 @@ These are target categories, not a guarantee that every occurrence or format is
39
 
40
  | Category | Values in scope | Replacement |
41
  |---|---|---|
42
- | Phone contacts | Phone, fax and SMS contact numbers, including foreign, emergency, short/service and vanity numbers; extensions and number alternatives | `[Telefon]` |
43
  | Email | Email addresses, including recoverable broken or incomplete addresses | `[PII]` |
44
  | Personal identifiers | PESEL, passport and identity-document numbers | `[PII]` |
45
  | Organization identifiers | NIP/VAT, REGON, KRS, GEMI, LEI and equivalent foreign company registration numbers; labelled DUNS, BDO and RPWDL register-book numbers | `[PII]` |
@@ -62,13 +62,18 @@ NERGAL is not designed to remove:
62
 
63
  These are intended exclusions; false positives can still mask some of this content.
64
 
65
- ### Known gaps in 1.1.2
66
 
67
- Unlabelled phones and identifiers (the rules take a bare phone only in the grouped national forms `601 234 567` and `22 123 45 67`; a plain `601234567` needs a label or the model), unusual formatting and damaged text can escape detection. **VINs and obfuscated emails** (such as `name (at) domain.pl`) are approved annotation targets, but that approval alone does not establish reliable support in the released 1.1.2 model. Do not rely on it to remove them consistently. An email with a lowercase word glued onto its domain (`jan@firma.plkontakt`) is masked together with that word, and text glued before an address can be masked with it.
68
 
69
  ## Versions
70
 
71
- 841-dev, union at 0.95, 354 gold spans. Same weights and API is a patch; new capability is minor; API, threshold, or weight recipe is major. A release that changes these numbers updates `hybrid.json` `eval` and this table. From 1.1.2 on, one 841-dev span is amended by a reviewed gold check (`gold_check_v2`, 10 characters trimmed); 1.1.1 is restated on it, and earlier rows use the original gold.
 
 
 
 
 
72
 
73
  | Version | Whole /354 | Residual | Rules FP | Union FP | Char P | Char R | What changed |
74
  |---|---:|---:|---:|---:|---:|---:|---|
@@ -79,7 +84,7 @@ Unlabelled phones and identifiers (the rules take a bare phone only in the group
79
  | 1.1.0 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Batch API (`predict_many`, `scrub_many`), opt-in float16 |
80
  | 1.1.1 | 326 | 23 | 98 | 123 | 97.94% | 96.50% | Grouped national phones without a label |
81
  | 1.1.1, amended gold | 326 | 23 | 108 | 133 | 97.77% | 96.49% | Restated on `gold_check_v2` |
82
- | **1.1.2** | **326** | **23** | **24** | **80** | **98.65%** | **96.49%** | Email boundary rules |
83
 
84
  ## Cue-less phone test
85
 
@@ -93,13 +98,13 @@ Unlabelled phones and identifiers (the rules take a bare phone only in the group
93
  | Passages with false masks /150 | 3 | 4 |
94
  | False characters | 83 | 95 |
95
 
96
- 1.1.1 lost no value 1.1.0 masked; its 12 new false characters are in one passage. The set is enriched by selection, so these numbers say nothing about how common such phones are, and it has a single reviewer. 1.1.2 changes one passage (an email match): false characters 95 → 78, phones unchanged.
97
 
98
  ## 841-dev
99
 
100
  Tables use one development split: 841 passages, 215 with gold PII, **354 spans** (169 phone, 185 other PII). It is the `dev` side of a 4,500-passage labelled tranche (3,655 train / 841 dev). Sources match Dynaword (EUR-Lex, HPLT, Wikipedia, parliamentary and government text, plus smaller news/literary slices). Labels mix unchanged silver with human review.
101
 
102
- One span, an address that ran on into a glued URL host, was trimmed by a reviewed post-hoc gold check (`gold_check_v2`); numbers from 1.1.2 on use the amended gold.
103
 
104
  The files contain real identifiers, so they are not released with the weights.
105
 
@@ -139,7 +144,7 @@ Same 841-dev split, threshold **0.95**. **Naked** is the transformer alone. **
139
  | XLM-R epoch 5 | naked | 298 | 42 | 67 | 98.78% | 89.36% |
140
  | **NERGAL 1.1.2** | **∪ regex** | 326 | 23 | 80 | 98.65% | 96.49% |
141
 
142
- Zero-shot [GLiNER 2.5](https://huggingface.co/fastino/gliner2.5-multi-v1) is not competitive here, especially on non-phone PII (7/185 whole vs 160 naked / 179 union). Fine-tuned historical GLiNER is the precise naked baseline (10 false characters) and still trails XLM-R on coverage. NERGAL is XLM-R epoch 5 plus the regex: 147/169 phone, 179/185 other PII. Exact-span precision 86.41%, recall 89.83%, F1 88.09%. The GLiNER rows use the original gold (10 characters of one span differ), and the two GLiNER ∪ regex rows use the 1.1.0 rules; the other rows use the amended gold.
143
 
144
  Trained GLiNER, HerBERT-large, and XLM-R-large were also compared on this split (plot above). GLiNER’s 0.50 coverage lead is the precision collapse in that figure; no GLiNER or HerBERT operating point passed the content-preservation gate.
145
 
@@ -174,6 +179,6 @@ masked, counts = nergal.scrub(text)
174
 
175
  For many texts, `nergal.scrub_many(texts)` (or `predict_many` for the raw model spans) batches windows across texts: about 2× the throughput of calling `scrub` in a loop on a CUDA GPU. `from_pretrained(..., dtype="float16")` casts the weights at load time on CUDA or MPS: another 1.9× on an RTX 4090, and 841-dev union numbers are unchanged, but scores are not bit-identical to float32 (0.95 spans can differ in rare cases). For a corpus, run 2–3 processes per GPU, because float16 inference is limited by CPU-side tokenization. Measurements are in `CHANGELOG.md`.
176
 
177
- `hybrid.json` records version `1.1.2`, threshold 0.95, gap ids `250002` / `250003`, the 841-dev `eval` block (with its `gold_amendments`), and two weight hashes: `model_safetensors_sha256` for the published file and `source_checkpoint_sha256` for the training checkpoint it was packed from. `test_nergal.py` is synthetic (no corpus text). From this snapshot: `python -m unittest test_nergal`.
178
 
179
  Base weights: [`FacebookAI/xlm-roberta-large`](https://huggingface.co/FacebookAI/xlm-roberta-large) revision `c23d21b0620b635a76227c604d44e43a9f0ee389` (MIT).
 
13
  - hybrid
14
  ---
15
 
16
+ # NERGAL 1.2.0
17
 
18
  **Named Entity Recognition with Grounded Additive Labels**
19
 
 
21
 
22
  ## TL;DR
23
 
24
+ Python rules do the identifiers they can prove. A transformer NER head adds phone and other PII spans the regex misses. The cleaner **unions** the two on the original text, then replaces hits with `[Telefon]` or `[PII]`. Phone spans from both follow phone policy v3: one span per number of 7+ digits, and shorter numbers are not masked.
25
 
26
+ - **Version:** `1.2.0` (`hybrid.json`, `CHANGELOG.md`)
27
+ - **Ground:** `scrub_pii` regex (SHA256 `b238d5b8…`)
28
  - **Additive labels:** XLM-RoBERTa-large token classifier, BIO tags `phone` / `pii`, threshold 0.95
29
  - **Throughput:** about 80k chars/s on one RTX 4090 with `scrub_many` + `dtype="float16"` and 3 processes (1.0.3: 23k)
30
  - **This snapshot:** seed `202609160`, **epoch 5** of a seven-epoch schedule
 
39
 
40
  | Category | Values in scope | Replacement |
41
  |---|---|---|
42
+ | Phone contacts | Phone, fax and SMS contact numbers of 7+ digits (country and area codes and keypad letters count), including foreign and vanity numbers, one span per number; an extension or alternative ending stays inside its number. Emergency, helpline, service and other numbers under 7 digits are not masked on their own | `[Telefon]` |
43
  | Email | Email addresses, including recoverable broken or incomplete addresses | `[PII]` |
44
  | Personal identifiers | PESEL, passport and identity-document numbers | `[PII]` |
45
  | Organization identifiers | NIP/VAT, REGON, KRS, GEMI, LEI and equivalent foreign company registration numbers; labelled DUNS, BDO and RPWDL register-book numbers | `[PII]` |
 
62
 
63
  These are intended exclusions; false positives can still mask some of this content.
64
 
65
+ ### Known gaps in 1.2.0
66
 
67
+ Unlabelled phones and identifiers (the rules take a bare phone only in the grouped national forms `601 234 567` and `22 123 45 67`; a plain `601234567` needs a label or the model), unusual formatting and damaged text can escape detection. **VINs and obfuscated emails** (such as `name (at) domain.pl`) are approved annotation targets, but that approval alone does not establish reliable support in the released 1.2.0 model. Do not rely on it to remove them consistently. An email with a lowercase word glued onto its domain (`jan@firma.plkontakt`) is masked together with that word, and text glued before an address can be masked with it. A phone written as short parts joined by two slashes (`12 / 345 / 678`) is left as text.
68
 
69
  ## Versions
70
 
71
+ 841-dev, union at 0.95. Same weights and API is a patch; new capability is minor; API, threshold, or weight recipe is major. A release that changes these numbers updates `hybrid.json` `eval` and this table. From 1.1.2 on, one 841-dev span is amended by a reviewed gold check (`gold_check_v2`, 10 characters trimmed); 1.1.1 is restated on it, and earlier rows use the original gold. From 1.2.0 the gold is restated to phone policy v3 (315 spans, see 841-dev below); 1.1.2 is restated on it, and the two tables are not comparable.
72
+
73
+ | Version | Whole /315 | Residual | Rules FP | Union FP | Char P | Char R | What changed |
74
+ |---|---:|---:|---:|---:|---:|---:|---|
75
+ | 1.1.2, restated gold | 303 | 9 | 83 | 333 | 94.37% | 97.83% | Restated on phone policy v3 and `gold_check_v4` |
76
+ | **1.2.0** | **303** | **9** | **24** | **80** | **98.59%** | **97.83%** | Phone policy v3 at mask time, rules and model spans |
77
 
78
  | Version | Whole /354 | Residual | Rules FP | Union FP | Char P | Char R | What changed |
79
  |---|---:|---:|---:|---:|---:|---:|---|
 
84
  | 1.1.0 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Batch API (`predict_many`, `scrub_many`), opt-in float16 |
85
  | 1.1.1 | 326 | 23 | 98 | 123 | 97.94% | 96.50% | Grouped national phones without a label |
86
  | 1.1.1, amended gold | 326 | 23 | 108 | 133 | 97.77% | 96.49% | Restated on `gold_check_v2` |
87
+ | 1.1.2 | 326 | 23 | 24 | 80 | 98.65% | 96.49% | Email boundary rules |
88
 
89
  ## Cue-less phone test
90
 
 
98
  | Passages with false masks /150 | 3 | 4 |
99
  | False characters | 83 | 95 |
100
 
101
+ 1.1.1 lost no value 1.1.0 masked; its 12 new false characters are in one passage. The set is enriched by selection, so these numbers say nothing about how common such phones are, and it has a single reviewer. 1.1.2 changes one passage (an email match): false characters 95 → 78, phones unchanged. Half of this set is now a benchmark on gold restated to phone policy v3 (70 passages): 1.2.0 and 1.1.2 both wholly mask 62 of 83 values there, with 33 false characters each.
102
 
103
  ## 841-dev
104
 
105
  Tables use one development split: 841 passages, 215 with gold PII, **354 spans** (169 phone, 185 other PII). It is the `dev` side of a 4,500-passage labelled tranche (3,655 train / 841 dev). Sources match Dynaword (EUR-Lex, HPLT, Wikipedia, parliamentary and government text, plus smaller news/literary slices). Labels mix unchanged silver with human review.
106
 
107
+ One span, an address that ran on into a glued URL host, was trimmed by a reviewed post-hoc gold check (`gold_check_v2`); numbers from 1.1.2 on use the amended gold. From 1.2.0 the gold is restated to phone policy v3 (55 passages; short and emergency numbers dropped, joined numbers split) and 5 more passages are amended by a reviewed check (`gold_check_v4`): **315 spans** (130 phone, 185 other PII). The tables below that are marked /354 use the earlier gold.
108
 
109
  The files contain real identifiers, so they are not released with the weights.
110
 
 
144
  | XLM-R epoch 5 | naked | 298 | 42 | 67 | 98.78% | 89.36% |
145
  | **NERGAL 1.1.2** | **∪ regex** | 326 | 23 | 80 | 98.65% | 96.49% |
146
 
147
+ Zero-shot [GLiNER 2.5](https://huggingface.co/fastino/gliner2.5-multi-v1) is not competitive here, especially on non-phone PII (7/185 whole vs 160 naked / 179 union). Fine-tuned historical GLiNER is the precise naked baseline (10 false characters) and still trails XLM-R on coverage. NERGAL 1.1.2 is XLM-R epoch 5 plus the regex: 147/169 phone, 179/185 other PII. Exact-span precision 86.41%, recall 89.83%, F1 88.09%. 1.2.0 is scored only on the restated gold (/315): 124/130 phone, 179/185 other PII, union FP 80, exact-span precision 91.05%, recall 93.65%, F1 92.33%. The GLiNER rows use the original gold (10 characters of one span differ), and the two GLiNER ∪ regex rows use the 1.1.0 rules; the other rows use the amended gold.
148
 
149
  Trained GLiNER, HerBERT-large, and XLM-R-large were also compared on this split (plot above). GLiNER’s 0.50 coverage lead is the precision collapse in that figure; no GLiNER or HerBERT operating point passed the content-preservation gate.
150
 
 
179
 
180
  For many texts, `nergal.scrub_many(texts)` (or `predict_many` for the raw model spans) batches windows across texts: about 2× the throughput of calling `scrub` in a loop on a CUDA GPU. `from_pretrained(..., dtype="float16")` casts the weights at load time on CUDA or MPS: another 1.9× on an RTX 4090, and 841-dev union numbers are unchanged, but scores are not bit-identical to float32 (0.95 spans can differ in rare cases). For a corpus, run 2–3 processes per GPU, because float16 inference is limited by CPU-side tokenization. Measurements are in `CHANGELOG.md`.
181
 
182
+ `hybrid.json` records version `1.2.0`, threshold 0.95, gap ids `250002` / `250003`, the 841-dev `eval` block (with its `gold_amendments` and `gold_restatements`), and two weight hashes: `model_safetensors_sha256` for the published file and `source_checkpoint_sha256` for the training checkpoint it was packed from. `test_nergal.py` is synthetic (no corpus text). From this snapshot: `python -m unittest test_nergal`.
183
 
184
  Base weights: [`FacebookAI/xlm-roberta-large`](https://huggingface.co/FacebookAI/xlm-roberta-large) revision `c23d21b0620b635a76227c604d44e43a9f0ee389` (MIT).
hybrid.json CHANGED
@@ -1,29 +1,33 @@
1
  {
2
  "full_name": "Named Entity Recognition with Grounded Additive Labels",
3
  "hub_id": "SlayerLab/NERGAL",
4
- "version": "1.1.2",
5
  "mode": "rules_union",
6
  "epoch": 5,
7
  "seed": 202609160,
8
  "threshold": 0.95,
9
- "rules_sha256": "08faef844c850bcd438c904d0b3f898df47c8bde8dd827d36ebffd39dc1594fb",
10
  "eval": {
11
  "split": "841-dev",
12
  "gold_amendments": [
13
- "gold_check_v2"
 
14
  ],
15
- "gold_entities": 354,
16
- "whole_entities": 326,
17
- "residual_passages": 23,
 
 
 
18
  "union_fp": 80,
19
  "rules_fp": 24,
20
- "character_precision": 0.9865,
21
- "character_recall": 0.9649,
22
- "exact_precision": 0.8641,
23
- "exact_recall": 0.8983,
24
- "exact_f1": 0.8809,
25
- "phone_whole": 147,
26
- "phone_gold": 169,
27
  "pii_whole": 179,
28
  "pii_gold": 185
29
  },
 
1
  {
2
  "full_name": "Named Entity Recognition with Grounded Additive Labels",
3
  "hub_id": "SlayerLab/NERGAL",
4
+ "version": "1.2.0",
5
  "mode": "rules_union",
6
  "epoch": 5,
7
  "seed": 202609160,
8
  "threshold": 0.95,
9
+ "rules_sha256": "b238d5b88aa3f3d55a24bb051ec93f9179dfb14b2c650441c8e0acb225d81d59",
10
  "eval": {
11
  "split": "841-dev",
12
  "gold_amendments": [
13
+ "gold_check_v2",
14
+ "gold_check_v4"
15
  ],
16
+ "gold_restatements": [
17
+ "policy_restatement_v2"
18
+ ],
19
+ "gold_entities": 315,
20
+ "whole_entities": 303,
21
+ "residual_passages": 9,
22
  "union_fp": 80,
23
  "rules_fp": 24,
24
+ "character_precision": 0.9859,
25
+ "character_recall": 0.9783,
26
+ "exact_precision": 0.9105,
27
+ "exact_recall": 0.9365,
28
+ "exact_f1": 0.9233,
29
+ "phone_whole": 124,
30
+ "phone_gold": 130,
31
  "pii_whole": 179,
32
  "pii_gold": 185
33
  },
nergal.py CHANGED
@@ -17,13 +17,13 @@ import scrub_pii
17
  from scrub_pii import PHONE_TAG, PII_TAG
18
 
19
  HUB_ID = 'SlayerLab/NERGAL'
20
- VERSION = '1.1.2'
21
  GAPS = ['[PII_SPACE]', '[PII_BREAK]']
22
  GAP_IDS = [250002, 250003]
23
  BIO_LABELS = ['O', 'B-phone', 'I-phone', 'B-pii', 'I-pii']
24
  LABELS = ['phone', 'pii']
25
  THRESHOLD = 0.95
26
- RULES_SHA = '08faef844c850bcd438c904d0b3f898df47c8bde8dd827d36ebffd39dc1594fb'
27
  DTYPES = ('float32', 'float16')
28
 
29
 
@@ -210,12 +210,20 @@ def apply_union(text, spans):
210
  return ''.join(out), chars, n_phone, n_pii
211
 
212
 
213
- def scrub_spans(text, rule_spans, model_spans, *, threshold=THRESHOLD):
 
 
 
 
 
 
 
 
214
  rule_keys = {(s['start'], s['end'], s['label']) for s in rule_spans}
215
- model_keep = decode(model_spans, threshold)
216
- extra = sum(1 for s in model_keep if (s['start'], s['end'], s['label']) not in rule_keys)
217
  _, rules_chars, _, _ = apply_union(text, rule_spans)
218
- masked, union_chars, n_phone, n_pii = apply_union(text, list(rule_spans) + model_keep)
219
  return masked, {
220
  'phone': n_phone,
221
  'pii': n_pii,
@@ -336,7 +344,7 @@ class Nergal:
336
 
337
  def scrub_many(self, texts):
338
  texts = list(texts)
339
- return [scrub_spans(text, self.rule_spans(text), model, threshold=self.threshold)
340
  for text, model in zip(texts, self.predict_many(texts), strict=True)]
341
 
342
 
 
17
  from scrub_pii import PHONE_TAG, PII_TAG
18
 
19
  HUB_ID = 'SlayerLab/NERGAL'
20
+ VERSION = '1.2.0'
21
  GAPS = ['[PII_SPACE]', '[PII_BREAK]']
22
  GAP_IDS = [250002, 250003]
23
  BIO_LABELS = ['O', 'B-phone', 'I-phone', 'B-pii', 'I-pii']
24
  LABELS = ['phone', 'pii']
25
  THRESHOLD = 0.95
26
+ RULES_SHA = 'b238d5b88aa3f3d55a24bb051ec93f9179dfb14b2c650441c8e0acb225d81d59'
27
  DTYPES = ('float32', 'float16')
28
 
29
 
 
210
  return ''.join(out), chars, n_phone, n_pii
211
 
212
 
213
+ def model_keep(text, model_spans, *, threshold=THRESHOLD, module=scrub_pii):
214
+ """Decoded model spans, each phone span cut by the rules' phone policy: one span per number of 7+ digits; short
215
+ and emergency numbers alone are dropped. Other labels pass unchanged."""
216
+ return [dict(s, start=a, end=b) for s in decode(model_spans, threshold)
217
+ for a, b in (module._phone_spans(text, s['start'], s['end']) if s['label'] == 'phone'
218
+ else [(s['start'], s['end'])])]
219
+
220
+
221
+ def scrub_spans(text, rule_spans, model_spans, *, threshold=THRESHOLD, module=scrub_pii):
222
  rule_keys = {(s['start'], s['end'], s['label']) for s in rule_spans}
223
+ keep = model_keep(text, model_spans, threshold=threshold, module=module)
224
+ extra = sum(1 for s in keep if (s['start'], s['end'], s['label']) not in rule_keys)
225
  _, rules_chars, _, _ = apply_union(text, rule_spans)
226
+ masked, union_chars, n_phone, n_pii = apply_union(text, list(rule_spans) + keep)
227
  return masked, {
228
  'phone': n_phone,
229
  'pii': n_pii,
 
344
 
345
  def scrub_many(self, texts):
346
  texts = list(texts)
347
+ return [scrub_spans(text, self.rule_spans(text), model, threshold=self.threshold, module=self._scrub)
348
  for text, model in zip(texts, self.predict_many(texts), strict=True)]
349
 
350
 
scrub_pii.py CHANGED
@@ -26,10 +26,14 @@ formatted phone fields. Labelled full-number ranges retain shared prefixes.
26
  Phones require a nearby contact cue, a Polish +48/0048 prefix, or explicit
27
  international country/trunk notation such as +CC (0). With a cue, the EUR-Lex
28
  "(32-2) 299 11 11" country-area form counts as international. Strong labels also
29
- admit one-digit country codes and wider hyphenated area codes. Short service numbers
30
- need strong labels; 116xxx numbers also accept nearby telephone prose. Unlabelled
31
- domestic numbers are left for audit because table cells have the same shapes, except
32
- the grouped national forms: mobile 3-3-3 and landline 2-3-2-2, one separator kind.
 
 
 
 
33
  Flattened tables can glue labels to both neighbours ("Mödlingtel.: … 38112faks:");
34
  such glued labels count as labels and end the preceding number.
35
 
@@ -181,14 +185,6 @@ _PHONE_BREAK_RE = re.compile(
181
  r"|\d{1,2}[.)][ \t]+\d{1,2}[./]\d{1,2}[./](?:19|20)\d{2})")
182
  _SECTION_MARKER_RE = re.compile(r"\d{1,2}[.)](?!\d)")
183
  _PHONE_EXTENSION_RE =re.compile(r"(?:(?:[ \t]+,?[ \t]*|,[ \t]*)wew(?:n(?:ętrzny)?)?\.?[ \t]*\d{1,5}|[ \t]+do[ \t]+\d{1,3}(?=[ \t]*(?:[.;,](?!\d)|\r?\n|\Z)))(?!\w|[ \t]*\d)", re.I)
184
- _SERVICE_PHONE_RE = re.compile(r"(?<![\w+])[1-9]\d{2}(?![\w\d]|[ \t./()-]*\d)")
185
- _SERVICE_PHONE_LABEL_RE = re.compile(
186
- r"(?:\b(?:tel(?:efon\w*)?|phone)[.: \t]{0,8}(?:alarmow\w*[.: \t]{0,8})?|"
187
- r"\b(?:numer[ \t]+)?bezpłatny[.: \t]{0,8}|"
188
- r"\b(?:bezpłatny[ \t]+)?numer[ \t]+alarmowy[.: \t]{0,8}|"
189
- r"(?:^|\n)[ \t]*(?:Straż pożarna|Policja|Pogotowie|Ratunek|Lekarz pogotowia ratunkowego)"
190
- r"(?:[ \t]+\([^()\n]{1,60}\))?[ \t]*:[ \t]*)"
191
- r"(?:[1-9]\d{2}[ \t]*(?:[,;]|i|lub|oraz)[ \t]*)*\Z", re.I)
192
  _OTHER_NUMBER_LABEL_RE = re.compile(r"\b(?:NIP|REGON|PESEL|KRS|ISBN|kod)\b[^\d\n]{0,20}\Z", re.I)
193
  # Grouped national phones need no cue: web contact blocks write them bare. Plain 9-digit strings stay cue-gated,
194
  # since many unlabelled ones are not phones.
@@ -653,9 +649,105 @@ def _replace_checked(text: str, pattern: re.Pattern, tag: str, ok,
653
  return pattern.sub(_sub, text), n
654
 
655
 
656
- def _replace_phones(text: str, *, mask=_tag) -> tuple[str, int]:
657
- n = 0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
658
  out = []
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
659
  pos = 0
660
  last_phone_end = None
661
  last_phone_complete = False
@@ -663,11 +755,10 @@ def _replace_phones(text: str, *, mask=_tag) -> tuple[str, int]:
663
  while True:
664
  m = _PHONE_RE.search(text, pos)
665
  if not m:
666
- out.append(text[pos:])
667
  break
668
  raw = m[0]
669
  end = m.end()
670
- replacement = raw
671
  line_end = text.find('\n', m.start(), end)
672
  if line_end >= 0 and re.fullmatch(r'\d{1,3}', text[m.start():line_end]):
673
  next_end = text.find('\n', line_end+1, end)
@@ -675,7 +766,6 @@ def _replace_phones(text: str, *, mask=_tag) -> tuple[str, int]:
675
  if (_directory_context(text, line_end+1, next_end)
676
  or _grouped_phone(text, line_end+1, text[line_end+1:next_end])):
677
  # A flattened table's room cell or an address's house number is not a phone country prefix.
678
- out.append(text[pos:line_end+1])
679
  pos = line_end+1
680
  continue
681
  immediate_phone_label = bool(_SHORT_PHONE_LABEL_RE.search(text[max(0, m.start()-100):m.start()]))
@@ -717,7 +807,6 @@ def _replace_phones(text: str, *, mask=_tag) -> tuple[str, int]:
717
  if novel_country_area_ok and (cued or _grouped_phone(text, m.start(), raw)):
718
  if (_is_amount(text, m.start(), end)
719
  and re.fullmatch(r'\d{1,3}(?:\.[ \t]*\d{3})+', raw)):
720
- out.extend((text[pos:m.start()], raw))
721
  pos = end
722
  continue
723
  # Stop before a new line or opening hours, but only after a complete
@@ -737,10 +826,8 @@ def _replace_phones(text: str, *, mask=_tag) -> tuple[str, int]:
737
  if _phone_ok(prefix, short and section) and not _is_amount(text, m.start(), gap.start()):
738
  end = m.start() + len(prefix)
739
  raw = prefix
740
- replacement = raw
741
  break
742
  if _is_amount(text, m.start(), end):
743
- out.extend((text[pos:m.start()], raw))
744
  pos = end
745
  continue
746
  # A suffix range repeats the final extension digits, not a second
@@ -765,8 +852,8 @@ def _replace_phones(text: str, *, mask=_tag) -> tuple[str, int]:
765
  # Parentheses enclosing prose are punctuation, not phone syntax.
766
  if raw.startswith('(+') and raw.count('(') > raw.count(')') and text[end-1:end] != ')':
767
  start += 1
768
- replacement = text[m.start():start] + mask(start, end, PHONE_TAG)
769
- n += 1
770
  # Split only pairs (<=22 digits); longer lists need label
771
  # context to distinguish them from accounts. Never redact ID tails.
772
  elif 18 <= len(_digits(raw)) <= 22 and "." not in raw:
@@ -775,36 +862,15 @@ def _replace_phones(text: str, *, mask=_tag) -> tuple[str, int]:
775
  gaps = sorted(re.finditer(r"[ \t]*/[ \t]*|[ \t\n]+", raw), key=lambda g: '/' not in g[0])
776
  for gap in gaps:
777
  if _phone_ok(raw[:gap.start()]) and _phone_ok(raw[gap.end():], '/' in gap[0]):
778
- replacement = (mask(m.start(), m.start()+gap.start(), PHONE_TAG)
779
- + gap[0] + mask(m.start()+gap.end(), end, PHONE_TAG))
780
- n += 2
781
  break
782
- out.extend((text[pos:m.start()], replacement))
783
- if replacement != raw:
784
  last_phone_end = end
785
  last_phone_complete = _phone_ok(raw)
786
  last_phone_cued = bool(cued)
787
  pos = end
788
- return "".join(out), n
789
-
790
-
791
- def _replace_extensions(text: str, *, mask=_tag) -> tuple[str, int]:
792
- count = 0
793
-
794
- def replace(m):
795
- nonlocal count
796
- before = text[max(0, m.start()-500):m.start()]
797
- # Only a telephone followed by extension/name entries establishes scope.
798
- heading = re.search(r'\btel(?:efon)?[.: \t]*\[Telefon\]'
799
- r'(?P<entries>(?:\s*wewn?\.?[ \t]+\d{1,5}[ \t]+[^\d\n]+)*\s*)\Z', before, re.I)
800
- if not heading or _is_amount(text, m.start('number'), m.end('number')):
801
- return m[0]
802
- count += 1
803
- return m['label'] + mask(m.start('number'), m.end('number'), PHONE_TAG)
804
-
805
- output = re.sub(r'(?P<label>(?:^|\n)[ \t]*wewn?\.?[ \t]+)(?P<number>\d{1,5})(?!\w)',
806
- replace, text, flags=re.I)
807
- return output, count
808
 
809
 
810
  def scrub_pii(text: str, *, spans: list | None = None) -> tuple[str, dict[str, int]]:
@@ -880,9 +946,6 @@ def scrub_pii(text: str, *, spans: list | None = None) -> tuple[str, dict[str, i
880
  counts["email"] += apply(_replace_checked, _EMAIL_FRAGMENT_RE, PII_TAG,
881
  lambda _: True, _EMAIL_LABEL_RE)
882
  counts["phone"] = apply(_replace_phones)
883
- counts["phone"] += apply(_replace_checked, _SERVICE_PHONE_RE, PHONE_TAG,
884
- lambda _: True, _SERVICE_PHONE_LABEL_RE)
885
- counts['phone'] += apply(_replace_extensions)
886
  # Contact context takes precedence over coincidental PESEL checksums in
887
  # foreign phone numbers; explicitly labelled identifiers were handled first.
888
  bare_pesel = apply(_replace_checked, _PESEL_RE, PII_TAG, _pesel_ok)
 
26
  Phones require a nearby contact cue, a Polish +48/0048 prefix, or explicit
27
  international country/trunk notation such as +CC (0). With a cue, the EUR-Lex
28
  "(32-2) 299 11 11" country-area form counts as international. Strong labels also
29
+ admit one-digit country codes and wider hyphenated area codes. Phone policy v3: each
30
+ number of at least 7 digits (keypad letters count) is its own [Telefon], connectors
31
+ between numbers stay unmasked, and a shorter part (an extension, "/90") stays inside
32
+ its number's span. A phone match with no such number (a short service or emergency
33
+ number, a lone extension) is left as text; a line break alone never splits a
34
+ number, nor does one slash before 5+ digits. Unlabelled domestic numbers
35
+ are left for audit because table cells have the same shapes, except the grouped
36
+ national forms: mobile 3-3-3 and landline 2-3-2-2, one separator kind.
37
  Flattened tables can glue labels to both neighbours ("Mödlingtel.: … 38112faks:");
38
  such glued labels count as labels and end the preceding number.
39
 
 
185
  r"|\d{1,2}[.)][ \t]+\d{1,2}[./]\d{1,2}[./](?:19|20)\d{2})")
186
  _SECTION_MARKER_RE = re.compile(r"\d{1,2}[.)](?!\d)")
187
  _PHONE_EXTENSION_RE =re.compile(r"(?:(?:[ \t]+,?[ \t]*|,[ \t]*)wew(?:n(?:ętrzny)?)?\.?[ \t]*\d{1,5}|[ \t]+do[ \t]+\d{1,3}(?=[ \t]*(?:[.;,](?!\d)|\r?\n|\Z)))(?!\w|[ \t]*\d)", re.I)
 
 
 
 
 
 
 
 
188
  _OTHER_NUMBER_LABEL_RE = re.compile(r"\b(?:NIP|REGON|PESEL|KRS|ISBN|kod)\b[^\d\n]{0,20}\Z", re.I)
189
  # Grouped national phones need no cue: web contact blocks write them bare. Plain 9-digit strings stay cue-gated,
190
  # since many unlabelled ones are not phones.
 
649
  return pattern.sub(_sub, text), n
650
 
651
 
652
+ # Phone policy v3 (labelling policy amended 2026-10-02): the mask-time copy of the rule that restated the evaluation
653
+ # gold. Each number of 7+ digits is its own span; shorter and emergency numbers are not masked alone.
654
+ _PHONE_CONNECTOR_RE = re.compile(r'\s*(?:[,;/&\n]|\b(?:lub|albo|i|oraz|and)\b)\s*', re.I)
655
+ _KEYPAD_RE = re.compile(r'(?<=[\d-])[A-Z]+')
656
+ _PHONE_CODE_RE = re.compile(r'\(?\+?\d{2,4}\)?')
657
+ _PHONE_DASH_RE = re.compile(r'\s+[-–—]\s+')
658
+ _FULL_PHONE = 7
659
+
660
+
661
+ def _phone_size(part: str) -> int:
662
+ return len(_digits(part)) + sum(map(len, _KEYPAD_RE.findall(part)))
663
+
664
+
665
+ def _phone_parts(text: str, start: int, end: int) -> list[tuple[int, int]]:
666
+ """The parts of one phone match between connectors. A spaced dash also separates once a full number precedes it
667
+ ("ddd/ddddddd – ddd/ddddddd"), not inside one ("601 – 234 – 567"). A part led by "+" takes the slash-joined parts
668
+ after it until it is full ("+ddd/dd/dddddd"): a dialling chain, not short numbers."""
669
+ bounds = [start, *(start + i for c in _PHONE_CONNECTOR_RE.finditer(text[start:end]) for i in c.span()), end]
670
+ ps = []
671
+ for a, b in zip(bounds[::2], bounds[1::2]):
672
+ if not text[a:b].strip():
673
+ continue
674
+ for d in _PHONE_DASH_RE.finditer(text, a, b):
675
+ if _phone_size(text[a:d.start()]) >= _FULL_PHONE and _phone_size(text[d.end():b]):
676
+ ps.append((a, d.start()))
677
+ a = d.end()
678
+ ps.append((a, b))
679
  out = []
680
+ for a, b in ps:
681
+ if out and text[out[-1][1]:a].strip() == '/' and re.match(r'\(?\+', text[out[-1][0]:out[-1][1]].strip()) and (
682
+ _phone_size(text[out[-1][0]:out[-1][1]]) < _FULL_PHONE):
683
+ out[-1] = (out[-1][0], b)
684
+ else:
685
+ out.append((a, b))
686
+ return out
687
+
688
+
689
+ def _phone_spans(text: str, start: int, end: int) -> list[tuple[int, int]]:
690
+ """The policy v3 spans of one phone match: none, the match itself, or one per full number."""
691
+ ps = _phone_parts(text, start, end)
692
+ full = [i for i, (a, b) in enumerate(ps) if _phone_size(text[a:b]) >= _FULL_PHONE]
693
+ if not full:
694
+ # A line break alone never separates entries, so a number wrapped into short lines keeps its mask;
695
+ # and one slash between two short parts is a code separator ("(+48)1234/567890"), not two numbers (user,
696
+ # 2026-10-03), when 5+ digits follow it: a subscriber part, not a year ("123/2019", "2019/2020"). Two
697
+ # slashes ("12 / 345 / 678") still drop unless a "+" leads them (_phone_parts). Each stretch between
698
+ # other written connectors is judged alone, so a list of two such numbers keeps both.
699
+ gaps = [text[b:a].strip() for (_, b), (a, _) in zip(ps, ps[1:])]
700
+ out, i = [], 0
701
+ for j in [*(k + 1 for k, g in enumerate(gaps) if g not in ('', '/')), len(ps)]:
702
+ whole = text[ps[i][0]:ps[j - 1][1]]
703
+ if gaps[i:j - 1].count('/') <= 1 and _phone_size(whole) >= _FULL_PHONE and _phone_size(
704
+ whole.rpartition('/')[2]) >= 5:
705
+ out.append((ps[i][0], ps[j - 1][1]))
706
+ i = j
707
+ return [(start, end)] if out == [(ps[0][0], ps[-1][1])] else out
708
+ if len(full) == 1:
709
+ return [(start, end)]
710
+
711
+ def code_prefix(i): # "(22) / 601 234 567": an area code before a slash belongs to the next number
712
+ part, before = text[ps[i][0]:ps[i][1]].strip(), text[ps[i - 1][1]:ps[i][0]] if i else ''
713
+ return bool(_PHONE_CODE_RE.fullmatch(part)) and '/' not in before and (
714
+ part.startswith('(') or text[ps[i][1]:ps[i + 1][0]].strip() == '/')
715
+
716
+ cuts = [0, *(i - code_prefix(i - 1) for i in full[1:]), len(ps)]
717
+ out = []
718
+ for i, j in zip(cuts, cuts[1:]):
719
+ a = start if i == 0 else ps[i][0] + re.search(r'[+(\d]', text[ps[i][0]:ps[i][1]]).start()
720
+ # the last part that ends like a number: a word between connectors ("601 234 567, fax, 602 345 678", which a
721
+ # model span can hold) ends nothing. ps[i] is full or a code, so one always does.
722
+ b = end if j == len(ps) else max(ps[k][0] + m.start() + 1 for k in range(i, j)
723
+ if (m := re.search(r'[\dA-Z)][^\dA-Z)]*$', text[ps[k][0]:ps[k][1]])))
724
+ out.append((a, b))
725
+ return out
726
+
727
+
728
+ def _mask_phone_runs(text: str, found: list[tuple[int, int]], mask) -> tuple[str, int]:
729
+ """Mask detected phones as policy v3 spans. Detections joined by one written connector form one run, as a joined
730
+ gold span did, so a short continuation stays inside its full number's span. A bare line break joins only two
731
+ short detections, the fragments of one wrapped number; a short line after a full number is not its part."""
732
+ runs = []
733
+ for a, b in found:
734
+ gap = text[runs[-1][1]:a] if runs else ''
735
+ wrapped = runs and not gap.strip() and max(
736
+ _phone_size(text[runs[-1][0]:runs[-1][1]]), _phone_size(text[a:b])) < _FULL_PHONE
737
+ if runs and _PHONE_CONNECTOR_RE.fullmatch(gap) and (gap.strip() or wrapped):
738
+ runs[-1][1] = b
739
+ else:
740
+ runs.append([a, b])
741
+ out, pos, n = [], 0, 0
742
+ for a, b in runs:
743
+ for start, end in _phone_spans(text, a, b):
744
+ out += [text[pos:start], mask(start, end, PHONE_TAG)]
745
+ pos, n = end, n + 1
746
+ return ''.join(out) + text[pos:], n
747
+
748
+
749
+ def _replace_phones(text: str, *, mask=_tag) -> tuple[str, int]:
750
+ found = [] # (start, end) of each detected phone, masked by policy v3 at the end
751
  pos = 0
752
  last_phone_end = None
753
  last_phone_complete = False
 
755
  while True:
756
  m = _PHONE_RE.search(text, pos)
757
  if not m:
 
758
  break
759
  raw = m[0]
760
  end = m.end()
761
+ hit = False
762
  line_end = text.find('\n', m.start(), end)
763
  if line_end >= 0 and re.fullmatch(r'\d{1,3}', text[m.start():line_end]):
764
  next_end = text.find('\n', line_end+1, end)
 
766
  if (_directory_context(text, line_end+1, next_end)
767
  or _grouped_phone(text, line_end+1, text[line_end+1:next_end])):
768
  # A flattened table's room cell or an address's house number is not a phone country prefix.
 
769
  pos = line_end+1
770
  continue
771
  immediate_phone_label = bool(_SHORT_PHONE_LABEL_RE.search(text[max(0, m.start()-100):m.start()]))
 
807
  if novel_country_area_ok and (cued or _grouped_phone(text, m.start(), raw)):
808
  if (_is_amount(text, m.start(), end)
809
  and re.fullmatch(r'\d{1,3}(?:\.[ \t]*\d{3})+', raw)):
 
810
  pos = end
811
  continue
812
  # Stop before a new line or opening hours, but only after a complete
 
826
  if _phone_ok(prefix, short and section) and not _is_amount(text, m.start(), gap.start()):
827
  end = m.start() + len(prefix)
828
  raw = prefix
 
829
  break
830
  if _is_amount(text, m.start(), end):
 
831
  pos = end
832
  continue
833
  # A suffix range repeats the final extension digits, not a second
 
852
  # Parentheses enclosing prose are punctuation, not phone syntax.
853
  if raw.startswith('(+') and raw.count('(') > raw.count(')') and text[end-1:end] != ')':
854
  start += 1
855
+ found.append((start, end))
856
+ hit = True
857
  # Split only pairs (<=22 digits); longer lists need label
858
  # context to distinguish them from accounts. Never redact ID tails.
859
  elif 18 <= len(_digits(raw)) <= 22 and "." not in raw:
 
862
  gaps = sorted(re.finditer(r"[ \t]*/[ \t]*|[ \t\n]+", raw), key=lambda g: '/' not in g[0])
863
  for gap in gaps:
864
  if _phone_ok(raw[:gap.start()]) and _phone_ok(raw[gap.end():], '/' in gap[0]):
865
+ found += [(m.start(), m.start()+gap.start()), (m.start()+gap.end(), end)]
866
+ hit = True
 
867
  break
868
+ if hit:
 
869
  last_phone_end = end
870
  last_phone_complete = _phone_ok(raw)
871
  last_phone_cued = bool(cued)
872
  pos = end
873
+ return _mask_phone_runs(text, found, mask)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
874
 
875
 
876
  def scrub_pii(text: str, *, spans: list | None = None) -> tuple[str, dict[str, int]]:
 
946
  counts["email"] += apply(_replace_checked, _EMAIL_FRAGMENT_RE, PII_TAG,
947
  lambda _: True, _EMAIL_LABEL_RE)
948
  counts["phone"] = apply(_replace_phones)
 
 
 
949
  # Contact context takes precedence over coincidental PESEL checksums in
950
  # foreign phone numbers; explicitly labelled identifiers were handled first.
951
  bare_pesel = apply(_replace_checked, _PESEL_RE, PII_TAG, _pesel_ok)
test_nergal.py CHANGED
@@ -5,7 +5,7 @@ import unittest
5
  from pathlib import Path
6
 
7
  HERE = Path(__file__).resolve().parent
8
- RULES_SHA = '08faef844c850bcd438c904d0b3f898df47c8bde8dd827d36ebffd39dc1594fb'
9
 
10
 
11
  class NergalTests(unittest.TestCase):
@@ -13,10 +13,11 @@ class NergalTests(unittest.TestCase):
13
  from nergal import GAP_IDS, GAPS, HUB_ID, RULES_SHA as PINNED, THRESHOLD, VERSION
14
  card = json.loads((HERE / 'hybrid.json').read_text())
15
  self.assertEqual(HUB_ID, 'SlayerLab/NERGAL')
16
- self.assertEqual(VERSION, '1.1.2')
17
  self.assertEqual(card['version'], VERSION)
18
  self.assertEqual(card['eval']['union_fp'], 80)
19
  self.assertEqual(card['eval']['rules_fp'], 24)
 
20
  self.assertEqual(GAPS, card['gaps'])
21
  self.assertEqual(GAP_IDS, card['gap_ids'])
22
  self.assertEqual(THRESHOLD, card['threshold'])
@@ -105,6 +106,28 @@ class NergalTests(unittest.TestCase):
105
  self.assertNotIn('000000000', masked)
106
  self.assertNotIn('extra', masked)
107
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
108
 
109
  if __name__ == '__main__':
110
  unittest.main()
 
5
  from pathlib import Path
6
 
7
  HERE = Path(__file__).resolve().parent
8
+ RULES_SHA = 'b238d5b88aa3f3d55a24bb051ec93f9179dfb14b2c650441c8e0acb225d81d59'
9
 
10
 
11
  class NergalTests(unittest.TestCase):
 
13
  from nergal import GAP_IDS, GAPS, HUB_ID, RULES_SHA as PINNED, THRESHOLD, VERSION
14
  card = json.loads((HERE / 'hybrid.json').read_text())
15
  self.assertEqual(HUB_ID, 'SlayerLab/NERGAL')
16
+ self.assertEqual(VERSION, '1.2.0')
17
  self.assertEqual(card['version'], VERSION)
18
  self.assertEqual(card['eval']['union_fp'], 80)
19
  self.assertEqual(card['eval']['rules_fp'], 24)
20
+ self.assertEqual((card['eval']['whole_entities'], card['eval']['gold_entities']), (303, 315)) # restated gold
21
  self.assertEqual(GAPS, card['gaps'])
22
  self.assertEqual(GAP_IDS, card['gap_ids'])
23
  self.assertEqual(THRESHOLD, card['threshold'])
 
106
  self.assertNotIn('000000000', masked)
107
  self.assertNotIn('extra', masked)
108
 
109
+ def test_model_phone_spans_follow_the_phone_policy(self):
110
+ from nergal import model_keep, scrub_spans
111
+ span = lambda text, part, label='phone', score=0.99: dict(
112
+ start=text.index(part), end=text.index(part) + len(part), label=label, score=score)
113
+ for text, part, kept in (('tel. 112', '112', []), # emergency number
114
+ ('tel. 51 23 45', '51 23 45', []), # under 7 digits
115
+ ('tel. 601 234 567/602 345 678', '601 234 567/602 345 678',
116
+ ['601 234 567', '602 345 678']), # one span per number
117
+ ('tel. 601 234 567, fax, 602 345 678', '601 234 567, fax, 602 345 678',
118
+ ['601 234 567', '602 345 678']), # a word between parts
119
+ ('tel. 22 123 45 67 wew. 101', '22 123 45 67 wew. 101',
120
+ ['22 123 45 67 wew. 101']), # extension stays inside
121
+ ('Jan Kowalski, 112', 'Jan Kowalski', ['Jan Kowalski'])): # other labels unchanged
122
+ label = 'pii' if part[0].isalpha() else 'phone'
123
+ with self.subTest(text=text):
124
+ keep = model_keep(text, [span(text, part, label)])
125
+ self.assertEqual([text[s['start']:s['end']] for s in keep], kept)
126
+ self.assertTrue(all(s['label'] == label and s['score'] == 0.99 for s in keep))
127
+ text = 'tel. 112'
128
+ self.assertEqual(scrub_spans(text, [], [span(text, '112')])[0], text)
129
+ self.assertEqual(model_keep(text, [span(text, '112', score=0.9)], threshold=0.95), [])
130
+
131
 
132
  if __name__ == '__main__':
133
  unittest.main()