# NERGAL versions Semver for this island: - **MAJOR** — API, threshold, or weight recipe changes - **MINOR** — new capability, same API - **PATCH** — rules or card fix, same weights and API Accuracy is 841-dev, union at 0.95. Through 1.1.2 the gold has 354 spans; from 1.2.0 it is restated to phone policy v3 (315 spans), and the two are not comparable. A version that changes those numbers must update `hybrid.json` `eval` and the tables below. ## 1.2.0 Same weights, threshold and API. Rules SHA `b238d5b8…`. The rules file is the frozen and scored candidate `73c185f7…` with two comments reworded; its syntax tree is identical. Phone policy v3 (labelling policy amended 2026-10-02) now applies at mask time, to the rules and to the model spans. - **Phone policy v3:** each number of 7+ digits is its own span, and connectors between numbers (`/`, `,`, `lub`, a spaced dash) stay as text. A short part (an extension `wew. 101`, an alternative ending `/90`, a wrapped line) stays inside its number's span. A standalone number under 7 digits is not masked: emergency numbers (`112`, `997`), helplines and service numbers (`116 111`), short codes. Country and area codes count as digits, and so do keypad letters. A code before a slash or in parentheses belongs to the number after it (`032/2345678 — 032/2345679` is two numbers), and a `+` slash chain (`+420/55/123456`) is one. - **Rules:** the service-number and lone-extension passes are removed. Detected phones are merged into runs and cut by `scrub_pii._phone_spans`. A written connector joins two detections, but a bare line break joins only two short detections (one wrapped number). A lone slash followed by at least 5 digits joins its two sides into one number (`022/123-456`). - **Model spans:** at 0.95, each model `phone` span is cut by the same `_phone_spans` before the union: one span per full number, and short numbers are dropped. `pii` spans are unchanged. `nergal.model_keep(text, model_spans)` returns the spans that are kept. The weights did not change, so `predict()` still returns short numbers; only `scrub()` / `scrub_many()` drop them. - **Gold:** 841-dev is restated to policy v3 mechanically by the policy's own rule (`policy_restatement_v2`: 55 passages change; checked on a seeded 41-passage human spot-check, which it matches except for one passage a later ruling overrode), and 5 passages are amended by a reviewed gold check (`gold_check_v4`). Gold spans: 354 → 315 (phone 169 → 130). Both are listed in `hybrid.json` `eval`. 841-dev (restated gold), 1.1.2 → 1.2.0: whole 303/315 and residual 9 are unchanged. Rules FP 83 → 24, union FP 333 → 80, passages with false characters 42 → 1. Char P 94.37% → 98.59%, char R 97.83%, exact-span F1 82.28% → 92.33%. Without the model-span cut, rules 1.1.3 alone give union FP 327: the published model masks short numbers. The restatement and the mask-time cut apply the same policy, so these gains measure agreement with it and are not an independent test. On the benchmark halves of the human test sets (restated the same way), no whole value is lost in any cohort. Union FP: human_test_v2 random 51 → 39, email 136 → 127, registry 4 → 0, bare-ID 3 → 0, e-Delivery 1 → 0; mC4 negatives email 224 → 222; the other cohorts are unchanged. The human_test_v2 half is not independent: one of its rows set the line-break rule and another the 5-digit slash cut. Dynaword rules delta, human-reviewed (40 passages around changed masks; 38 scored, 2 left uncertain and unscored): phone values whole 57 → 65 of 70, false phone characters 373 → 113, no value lost. Two of these passages set the `+` chain and spaced-dash rules, so this is development evidence. Full Dynaword rules delta against 1.1.2 (36 sources, 4,094,255 documents, counts only, unreviewed): rule spans 294,819 → 291,651, masked characters −10,766 (0.21%). In steps: the 1.1.3 v2 candidate changes 2,092 documents (4,661 spans removed, 1,153 added), and v3 then changes 293 (eurlex 287, wikivoyage 6; 335 removed, 675 added). Both steps read the same per-unit input hashes. The 1.1.2 entry's run covered 21 sources on an older checkout, and `samorzad_gov_pl` has changed upstream since, so document counts across entries are not comparable. Known limits: two slashes between short parts (`12 / 345 / 678`) are left as text. The 5 reviewed Dynaword phone values that 1.2.0 does not wholly mask are also missed by 1.1.2. The model-span cut was not reviewed on new text: no new inference was run for this release. | Version | Whole /315 | Residual | Rules FP | Union FP | Char P | Char R | What changed | |---|---:|---:|---:|---:|---:|---:|---| | 1.1.2, restated gold | 303 | 9 | 83 | 333 | 94.37% | 97.83% | Restated on phone policy v3 and `gold_check_v4` | | 1.2.0 | 303 | 9 | 24 | 80 | 98.59% | 97.83% | Phone policy v3 at mask time, rules and model spans | ## 1.1.2 Same weights, threshold and API. Rules SHA `08faef84…`. The rules file is the frozen and scored candidate `08faef84…`. - **Rules (email only):** an address match ends where glued text begins, before a `www.` host glued onto the domain (`jan@firma.plwww.…`) or at a capital glued onto a lowercase domain ending (`jan@firma.plKontakt`). A cut is kept only when what remains is a complete address. A match right after another `@` is dropped when a space comes before the match's own `@` (a list of @-mentions). Without the space it keeps its mask: a handle (`@jan@firma.social`), a label glued on with `@` or glued addresses can hold a real one. - **Gold:** one 841-dev email span ran on into a glued URL host, unlike the other seven addresses in its passage and the labelling policy. It was reviewed and trimmed by 10 characters (`gold_check_v2`, `hybrid.json` `eval.gold_amendments`). Numbers from 1.1.2 on use the amended gold; earlier rows use the original. On the amended gold, 1.1.1 has rules FP 108, union FP 133, char P 97.77% and char R 96.49%. 841-dev (amended gold): whole 326/354 and residual 23 unchanged; union FP 133 → 80, rules FP 108 → 24; exact-span and per-label numbers unchanged. All changes are in one passage. Its 24 remaining rules FP are capitalised words glued before a local part, which the 1.0.1 prefix trim cuts only partly. Web development set (41,204 passages): 13 passages change, 0 characters added, and a review ruled all dropped runs out of scope. Human test v2 email benchmark half: rules FP 89 → 65, union FP 162 → 138; not independent, since that half surfaced the cases. Cue-less phone test (150 passages): one passage changes, false characters 95 → 78, phones unchanged. Full Dynaword (2,797,400 documents): 165 documents change, almost all EUR-Lex. There are 553 cuts: 521 at a capital and 32 at a glued `www.`. No cut frees an `@`, and 26 cuts free a glued `Tel`/`Fax` label, whose phone number the rules now mask. No mask is dropped. A first candidate also dropped a match after an `@` inside a word (`a@b@firma.pl`). On full Dynaword that lost real addresses in glued lists, so it was removed. Known limits: a lowercase word glued onto the domain (`jan@firma.plkontakt`) stays masked with the address, because the TLD list does not separate the two; text glued before the local part (a `www.` host or a word) is not cut. | Version | Whole /354 | Residual | Rules FP | Union FP | Char P | Char R | What changed | |---|---:|---:|---:|---:|---:|---:|---| | 1.0.0 | 323 | 25 | 133 | 133 | 97.76% | 95.95% | First Hub snapshot | | 1.0.1 | 323 | 25 | 98 | 123 | 97.93% | 95.95% | Prefix-only glued-email trim | | 1.0.2 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Labelled country-area phone fix | | 1.0.3 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Label-note, e-Delivery and registry rules; placeholder and card fixes | | 1.1.0 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Batch API (`predict_many`, `scrub_many`), opt-in float16 | | 1.1.1 | 326 | 23 | 98 | 123 | 97.94% | 96.50% | Grouped national phones without a label | | 1.1.1, amended gold | 326 | 23 | 108 | 133 | 97.77% | 96.49% | Restated on `gold_check_v2` | | 1.1.2 | 326 | 23 | 24 | 80 | 98.65% | 96.49% | Email boundary rules | ## 1.1.1 Same weights, threshold and API. Rules SHA `ad5c51f7…`. The rules file is the frozen and scored candidate `7412261c…` with comment tags removed; its syntax tree is identical. - **Rules:** Polish phones in the grouped national forms are masked without a contact label: mobile `601 234 567` and landline `22 123 45 67` (area code optionally in parentheses), optional `+48`, one separator kind (space or hyphen) throughout. Not taken: plain 9-digit strings (still label-gated), amounts, round counts (`500 000 000`), decimal figures (`601 234 567,89`), and numbers after another identifier's label (NIP, REGON, codes). After a bare grouped phone, a list goes on only with complete Polish numbers. 841-dev: 326/354 whole (1.1.0: 324), residual 23 (24), 0 new false characters; phone 147/169 (145). A frozen blind set of 150 web passages selected for unlabelled phones: phones wholly masked 66/111 (1.1.0: 37), all values 127/178 (98), 0 lost, 1 passage with new false characters. Dynaword (60,000 documents): 79 documents change, 247 phone spans added, 0 removed; in a reviewed sample 1 of 23 sampled documents' added spans was not a phone. | Version | Whole /354 | Residual | Rules FP | Union FP | Char P | Char R | What changed | |---|---:|---:|---:|---:|---:|---:|---| | 1.0.0 | 323 | 25 | 133 | 133 | 97.76% | 95.95% | First Hub snapshot | | 1.0.1 | 323 | 25 | 98 | 123 | 97.93% | 95.95% | Prefix-only glued-email trim | | 1.0.2 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Labelled country-area phone fix | | 1.0.3 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Label-note, e-Delivery and registry rules; placeholder and card fixes | | 1.1.0 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Batch API (`predict_many`, `scrub_many`), opt-in float16 | | 1.1.1 | 326 | 23 | 98 | 123 | 97.94% | 96.50% | Grouped national phones without a label | ## 1.1.0 Same weights, rules, threshold and default outputs; `hybrid.json` `eval` is unchanged. New batch API and opt-in fp16. - **`Nergal.predict_many(texts)` / `scrub_many(texts)`:** windows from up to 64 texts are sorted by token length and packed into batches of at most 32,768 padded tokens and 128 rows. `predict` / `scrub` are now the one-text case of these. - **`dtype='float16'`** on `Nergal(...)` / `from_pretrained(...)` (CLI `--dtype`): casts the float32 weights at load time. Needs CUDA or MPS. The default stays `float32`, and `model.safetensors` is unchanged. - **Faster window sizing:** `Encoding.count` adds up cached unit pieces instead of re-tokenizing inside the window search. It yields the same windows, because `encode()` still rejects any unit whose pieces change with context. - **Default device:** `from_pretrained` now tries CUDA, then MPS, then CPU. 1.0.x used CPU even on CUDA machines. Throughput on one RTX 4090 (13.88M chars of FineWeb-2, kchar/s): 1.0.3-style per-document batches 11.0 (23.0 with 3 processes); `predict_many` float32 21.3 (28.1 with 2); `predict_many` float16 39.4 (79.6 with 3). The float16 path is limited by CPU-side tokenization, so run 2–3 processes per GPU. Equivalence: on the 1,685 labelled dev rows, float32 `predict_many` gives 0 span changes at 0.95 against the cached model spans of the published weights (841-dev max score change 2.6e-5). Float16 adds 2 spans on gold (1 on 841-dev, already covered by the rules) and removes none; 841-dev union numbers are identical. On 6,182 FineWeb-2 documents (1,236 with MPS and 4,946 with CUDA float32 references), CUDA float16 gave 0 span changes (max score change 0.012). | Version | Whole /354 | Residual | Rules FP | Union FP | Char P | Char R | What changed | |---|---:|---:|---:|---:|---:|---:|---| | 1.0.0 | 323 | 25 | 133 | 133 | 97.76% | 95.95% | First Hub snapshot | | 1.0.1 | 323 | 25 | 98 | 123 | 97.93% | 95.95% | Prefix-only glued-email trim | | 1.0.2 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Labelled country-area phone fix | | 1.0.3 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Label-note, e-Delivery and registry rules; placeholder and card fixes | | 1.1.0 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Batch API (`predict_many`, `scrub_many`), opt-in float16 | ## 1.0.3 Same weights, threshold, and API. Rules SHA `f32d5c54…`. - **Rules:** `NIP`/`REGON` labels with footnote marks or a short gloss (`NIP*:`, `NIP (Wykonawcy):`); e-Delivery addresses broken across a line; `DUNS`, `BDO` and `RPWDL` numbers right after their own label, at a fixed length. - **Wrapper:** existing `[PII]`/`[Telefon]` placeholders no longer switch the rules off. Before this, one placeholder anywhere in the text dropped every rule span for the whole document. - **Card:** `weights_sha256` was the hash of the source training checkpoint, not of `model.safetensors`. It is now `source_checkpoint_sha256`, and `model_safetensors_sha256` holds the published file's hash. Readers of `weights_sha256` must switch keys. 841-dev is unchanged: 0 rule spans change, and no 841-dev passage contains a placeholder. On simulated pre-masked input (gold spans replaced by their own placeholder), the rules now cover 1,812 of 2,606 remaining gold spans (1.0.2: 0), with 0 new false characters versus the same rules on unmasked text. Full Dynaword (2.8M documents): 27 documents change, 42 registry spans added, 0 removed. Known limit: an identifier with a placeholder inside it or right before it (`NIP [PII] …`, `8503[PII]…`) can still be missed by both layers, and a phone that follows an already-masked phone in a list can be missed by the rules. | Version | Whole /354 | Residual | Rules FP | Union FP | Char P | Char R | What changed | |---|---:|---:|---:|---:|---:|---:|---| | 1.0.0 | 323 | 25 | 133 | 133 | 97.76% | 95.95% | First Hub snapshot | | 1.0.1 | 323 | 25 | 98 | 123 | 97.93% | 95.95% | Prefix-only glued-email trim | | 1.0.2 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Labelled country-area phone fix | | 1.0.3 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Label-note, e-Delivery and registry rules; placeholder and card fixes | ## 1.0.2 Fix the public wrapper’s tokenizer batch shape so local model inference works. The packed tokenizer regression and an end-to-end local inference check cover this path. Fix immediately labelled hyphenated country-area phones with one-digit country codes or wider area codes. Numeric continuations, slash lists, weak contact cues, and nearby title prose still abstain. Same weights, threshold, and API. 841-dev: rules recover 14 more whole spans; the union recovers one (323 → 324), reducing residual passages 25 → 24. Rules FP remain 98; union FP remain 123. Lost gold, new false characters, and new clean-passage damage are all zero. Rules SHA `3016ae5b…`. Exact-span scores are recomputed from deduplicated raw rule/model span triples: precision 86.34%, recall 89.27%, F1 87.78%. The 1.0.1 card's precision/F1 were stale after the prefix trim changed exact rule/model duplicate counts; its character metrics were correct. | Version | Whole /354 | Residual | Rules FP | Union FP | Char P | Char R | What changed | |---|---:|---:|---:|---:|---:|---:|---| | 1.0.0 | 323 | 25 | 133 | 133 | 97.76% | 95.95% | First Hub snapshot | | 1.0.1 | 323 | 25 | 98 | 123 | 97.93% | 95.95% | Prefix-only glued-email trim | | 1.0.2 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Labelled country-area phone fix | ## 1.0.1 Prefix-only glued-email trim. Cluster-gated title-case 5–11 letter prefixes are dropped from the redaction when the remainder is already a lowercase-local email. Same epoch-5 weights. Lost gold 0. New clean-passage damage 0. | Version | Whole /354 | Residual | Rules FP | Union FP | Char P | Char R | What changed | |---|---:|---:|---:|---:|---:|---:|---| | 1.0.0 | 323 | 25 | 133 | 133 | 97.76% | 95.95% | First Hub snapshot | | 1.0.1 | 323 | 25 | 98 | 123 | 97.93% | 95.95% | Prefix-only glued-email trim | Rules SHA `4dcc441c…`. Seven emails trimmed, 35 prefix characters dropped; the model still covers 25 of those positions. Suffix glue is unchanged. ## 1.0.0 First Hub snapshot. Seed `202609160`, epoch 5, rules `547c0428…`. Union FP 133 was the glued-email regex floor. This seed added no new false characters versus the historical incumbent.