NERGAL / CHANGELOG.md
ppuzio's picture
Claude Opus 5.5
1.1.2: email boundary rules
e0caa8b
|
Raw History Blame Contribute Delete
11.8 kB
# NERGAL versions
Semver for this island:
- **MAJOR** — API, threshold, or weight recipe changes
- **MINOR** — new capability, same API
- **PATCH** — rules or card fix, same weights and API
Accuracy is 841-dev, union at 0.95, 354 gold spans. A version that changes those numbers must update `hybrid.json` `eval` and the tables below.
## 1.1.2
Same weights, threshold and API. Rules SHA `08faef84…`. The rules file is the frozen and scored candidate `08faef84…`.
- **Rules (email only):** an address match ends where glued text begins, before a `www.` host glued onto the domain (`jan@firma.plwww.…`) or at a capital glued onto a lowercase domain ending (`jan@firma.plKontakt`). A cut is kept only when what remains is a complete address. A match right after another `@` is dropped when a space comes before the match's own `@` (a list of @-mentions). Without the space it keeps its mask: a handle (`@jan@firma.social`), a label glued on with `@` or glued addresses can hold a real one.
- **Gold:** one 841-dev email span ran on into a glued URL host, unlike the other seven addresses in its passage and the labelling policy. It was reviewed and trimmed by 10 characters (`gold_check_v2`, `hybrid.json` `eval.gold_amendments`). Numbers from 1.1.2 on use the amended gold; earlier rows use the original. On the amended gold, 1.1.1 has rules FP 108, union FP 133, char P 97.77% and char R 96.49%.
841-dev (amended gold): whole 326/354 and residual 23 unchanged; union FP 133 → 80, rules FP 108 → 24; exact-span and per-label numbers unchanged. All changes are in one passage. Its 24 remaining rules FP are capitalised words glued before a local part, which the 1.0.1 prefix trim cuts only partly. Web development set (41,204 passages): 13 passages change, 0 characters added, and a review ruled all dropped runs out of scope. Human test v2 email benchmark half: rules FP 89 → 65, union FP 162 → 138; not independent, since that half surfaced the cases. Cue-less phone test (150 passages): one passage changes, false characters 95 → 78, phones unchanged. Full Dynaword (2,797,400 documents): 165 documents change, almost all EUR-Lex. There are 553 cuts: 521 at a capital and 32 at a glued `www.`. No cut frees an `@`, and 26 cuts free a glued `Tel`/`Fax` label, whose phone number the rules now mask. No mask is dropped. A first candidate also dropped a match after an `@` inside a word (`a@b@firma.pl`). On full Dynaword that lost real addresses in glued lists, so it was removed.
Known limits: a lowercase word glued onto the domain (`jan@firma.plkontakt`) stays masked with the address, because the TLD list does not separate the two; text glued before the local part (a `www.` host or a word) is not cut.
| Version | Whole /354 | Residual | Rules FP | Union FP | Char P | Char R | What changed |
|---|---:|---:|---:|---:|---:|---:|---|
| 1.0.0 | 323 | 25 | 133 | 133 | 97.76% | 95.95% | First Hub snapshot |
| 1.0.1 | 323 | 25 | 98 | 123 | 97.93% | 95.95% | Prefix-only glued-email trim |
| 1.0.2 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Labelled country-area phone fix |
| 1.0.3 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Label-note, e-Delivery and registry rules; placeholder and card fixes |
| 1.1.0 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Batch API (`predict_many`, `scrub_many`), opt-in float16 |
| 1.1.1 | 326 | 23 | 98 | 123 | 97.94% | 96.50% | Grouped national phones without a label |
| 1.1.1, amended gold | 326 | 23 | 108 | 133 | 97.77% | 96.49% | Restated on `gold_check_v2` |
| 1.1.2 | 326 | 23 | 24 | 80 | 98.65% | 96.49% | Email boundary rules |
## 1.1.1
Same weights, threshold and API. Rules SHA `ad5c51f7…`. The rules file is the frozen and scored candidate `7412261c…` with comment tags removed; its syntax tree is identical.
- **Rules:** Polish phones in the grouped national forms are masked without a contact label: mobile `601 234 567` and landline `22 123 45 67` (area code optionally in parentheses), optional `+48`, one separator kind (space or hyphen) throughout. Not taken: plain 9-digit strings (still label-gated), amounts, round counts (`500 000 000`), decimal figures (`601 234 567,89`), and numbers after another identifier's label (NIP, REGON, codes). After a bare grouped phone, a list goes on only with complete Polish numbers.
841-dev: 326/354 whole (1.1.0: 324), residual 23 (24), 0 new false characters; phone 147/169 (145). A frozen blind set of 150 web passages selected for unlabelled phones: phones wholly masked 66/111 (1.1.0: 37), all values 127/178 (98), 0 lost, 1 passage with new false characters. Dynaword (60,000 documents): 79 documents change, 247 phone spans added, 0 removed; in a reviewed sample 1 of 23 sampled documents' added spans was not a phone.
| Version | Whole /354 | Residual | Rules FP | Union FP | Char P | Char R | What changed |
|---|---:|---:|---:|---:|---:|---:|---|
| 1.0.0 | 323 | 25 | 133 | 133 | 97.76% | 95.95% | First Hub snapshot |
| 1.0.1 | 323 | 25 | 98 | 123 | 97.93% | 95.95% | Prefix-only glued-email trim |
| 1.0.2 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Labelled country-area phone fix |
| 1.0.3 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Label-note, e-Delivery and registry rules; placeholder and card fixes |
| 1.1.0 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Batch API (`predict_many`, `scrub_many`), opt-in float16 |
| 1.1.1 | 326 | 23 | 98 | 123 | 97.94% | 96.50% | Grouped national phones without a label |
## 1.1.0
Same weights, rules, threshold and default outputs; `hybrid.json` `eval` is unchanged. New batch API and opt-in fp16.
- **`Nergal.predict_many(texts)` / `scrub_many(texts)`:** windows from up to 64 texts are sorted by token length and packed into batches of at most 32,768 padded tokens and 128 rows. `predict` / `scrub` are now the one-text case of these.
- **`dtype='float16'`** on `Nergal(...)` / `from_pretrained(...)` (CLI `--dtype`): casts the float32 weights at load time. Needs CUDA or MPS. The default stays `float32`, and `model.safetensors` is unchanged.
- **Faster window sizing:** `Encoding.count` adds up cached unit pieces instead of re-tokenizing inside the window search. It yields the same windows, because `encode()` still rejects any unit whose pieces change with context.
- **Default device:** `from_pretrained` now tries CUDA, then MPS, then CPU. 1.0.x used CPU even on CUDA machines.
Throughput on one RTX 4090 (13.88M chars of FineWeb-2, kchar/s): 1.0.3-style per-document batches 11.0 (23.0 with 3 processes); `predict_many` float32 21.3 (28.1 with 2); `predict_many` float16 39.4 (79.6 with 3). The float16 path is limited by CPU-side tokenization, so run 2–3 processes per GPU.
Equivalence: on the 1,685 labelled dev rows, float32 `predict_many` gives 0 span changes at 0.95 against the cached model spans of the published weights (841-dev max score change 2.6e-5). Float16 adds 2 spans on gold (1 on 841-dev, already covered by the rules) and removes none; 841-dev union numbers are identical. On 6,182 FineWeb-2 documents (1,236 with MPS and 4,946 with CUDA float32 references), CUDA float16 gave 0 span changes (max score change 0.012).
| Version | Whole /354 | Residual | Rules FP | Union FP | Char P | Char R | What changed |
|---|---:|---:|---:|---:|---:|---:|---|
| 1.0.0 | 323 | 25 | 133 | 133 | 97.76% | 95.95% | First Hub snapshot |
| 1.0.1 | 323 | 25 | 98 | 123 | 97.93% | 95.95% | Prefix-only glued-email trim |
| 1.0.2 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Labelled country-area phone fix |
| 1.0.3 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Label-note, e-Delivery and registry rules; placeholder and card fixes |
| 1.1.0 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Batch API (`predict_many`, `scrub_many`), opt-in float16 |
## 1.0.3
Same weights, threshold, and API. Rules SHA `f32d5c54…`.
- **Rules:** `NIP`/`REGON` labels with footnote marks or a short gloss (`NIP*:`, `NIP (Wykonawcy):`); e-Delivery addresses broken across a line; `DUNS`, `BDO` and `RPWDL` numbers right after their own label, at a fixed length.
- **Wrapper:** existing `[PII]`/`[Telefon]` placeholders no longer switch the rules off. Before this, one placeholder anywhere in the text dropped every rule span for the whole document.
- **Card:** `weights_sha256` was the hash of the source training checkpoint, not of `model.safetensors`. It is now `source_checkpoint_sha256`, and `model_safetensors_sha256` holds the published file's hash. Readers of `weights_sha256` must switch keys.
841-dev is unchanged: 0 rule spans change, and no 841-dev passage contains a placeholder. On simulated pre-masked input (gold spans replaced by their own placeholder), the rules now cover 1,812 of 2,606 remaining gold spans (1.0.2: 0), with 0 new false characters versus the same rules on unmasked text. Full Dynaword (2.8M documents): 27 documents change, 42 registry spans added, 0 removed.
Known limit: an identifier with a placeholder inside it or right before it (`NIP [PII] …`, `8503[PII]…`) can still be missed by both layers, and a phone that follows an already-masked phone in a list can be missed by the rules.
| Version | Whole /354 | Residual | Rules FP | Union FP | Char P | Char R | What changed |
|---|---:|---:|---:|---:|---:|---:|---|
| 1.0.0 | 323 | 25 | 133 | 133 | 97.76% | 95.95% | First Hub snapshot |
| 1.0.1 | 323 | 25 | 98 | 123 | 97.93% | 95.95% | Prefix-only glued-email trim |
| 1.0.2 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Labelled country-area phone fix |
| 1.0.3 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Label-note, e-Delivery and registry rules; placeholder and card fixes |
## 1.0.2
Fix the public wrapper’s tokenizer batch shape so local model inference works. The packed tokenizer regression and an end-to-end local inference check cover this path.
Fix immediately labelled hyphenated country-area phones with one-digit country codes or wider area codes. Numeric continuations, slash lists, weak contact cues, and nearby title prose still abstain. Same weights, threshold, and API.
841-dev: rules recover 14 more whole spans; the union recovers one (323 → 324), reducing residual passages 25 → 24. Rules FP remain 98; union FP remain 123. Lost gold, new false characters, and new clean-passage damage are all zero. Rules SHA `3016ae5b…`.
Exact-span scores are recomputed from deduplicated raw rule/model span triples: precision 86.34%, recall 89.27%, F1 87.78%. The 1.0.1 card's precision/F1 were stale after the prefix trim changed exact rule/model duplicate counts; its character metrics were correct.
| Version | Whole /354 | Residual | Rules FP | Union FP | Char P | Char R | What changed |
|---|---:|---:|---:|---:|---:|---:|---|
| 1.0.0 | 323 | 25 | 133 | 133 | 97.76% | 95.95% | First Hub snapshot |
| 1.0.1 | 323 | 25 | 98 | 123 | 97.93% | 95.95% | Prefix-only glued-email trim |
| 1.0.2 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Labelled country-area phone fix |
## 1.0.1
Prefix-only glued-email trim. Cluster-gated title-case 5–11 letter prefixes are dropped from the redaction when the remainder is already a lowercase-local email. Same epoch-5 weights. Lost gold 0. New clean-passage damage 0.
| Version | Whole /354 | Residual | Rules FP | Union FP | Char P | Char R | What changed |
|---|---:|---:|---:|---:|---:|---:|---|
| 1.0.0 | 323 | 25 | 133 | 133 | 97.76% | 95.95% | First Hub snapshot |
| 1.0.1 | 323 | 25 | 98 | 123 | 97.93% | 95.95% | Prefix-only glued-email trim |
Rules SHA `4dcc441c…`. Seven emails trimmed, 35 prefix characters dropped; the model still covers 25 of those positions. Suffix glue is unchanged.
## 1.0.0
First Hub snapshot. Seed `202609160`, epoch 5, rules `547c0428…`. Union FP 133 was the glued-email regex floor. This seed added no new false characters versus the historical incumbent.