# NERGAL versions Semver for this island: - **MAJOR** — API, threshold, or weight recipe changes - **MINOR** — new capability, same API - **PATCH** — rules or card fix, same weights and API Accuracy is 841-dev, union at 0.95, 354 gold spans. A version that changes those numbers must update `hybrid.json` `eval` and the tables below. ## 1.1.0 Same weights, rules, threshold and default outputs; `hybrid.json` `eval` is unchanged. New batch API and opt-in fp16. - **`Nergal.predict_many(texts)` / `scrub_many(texts)`:** windows from up to 64 texts are sorted by token length and packed into batches of at most 32,768 padded tokens and 128 rows. `predict` / `scrub` are now the one-text case of these. - **`dtype='float16'`** on `Nergal(...)` / `from_pretrained(...)` (CLI `--dtype`): casts the float32 weights at load time. Needs CUDA or MPS. The default stays `float32`, and `model.safetensors` is unchanged. - **Faster window sizing:** `Encoding.count` adds up cached unit pieces instead of re-tokenizing inside the window search. It yields the same windows, because `encode()` still rejects any unit whose pieces change with context. - **Default device:** `from_pretrained` now tries CUDA, then MPS, then CPU. 1.0.x used CPU even on CUDA machines. Throughput on one RTX 4090 (13.88M chars of FineWeb-2, kchar/s): 1.0.3-style per-document batches 11.0 (23.0 with 3 processes); `predict_many` float32 21.3 (28.1 with 2); `predict_many` float16 39.4 (79.6 with 3). The float16 path is limited by CPU-side tokenization, so run 2–3 processes per GPU. Equivalence: on the 1,685 labelled dev rows, float32 `predict_many` gives 0 span changes at 0.95 against the cached model spans of the published weights (841-dev max score change 2.6e-5). Float16 adds 2 spans on gold (1 on 841-dev, already covered by the rules) and removes none; 841-dev union numbers are identical. On 6,182 FineWeb-2 documents (1,236 with MPS and 4,946 with CUDA float32 references), CUDA float16 gave 0 span changes (max score change 0.012). | Version | Whole /354 | Residual | Rules FP | Union FP | Char P | Char R | What changed | |---|---:|---:|---:|---:|---:|---:|---| | 1.0.0 | 323 | 25 | 133 | 133 | 97.76% | 95.95% | First Hub snapshot | | 1.0.1 | 323 | 25 | 98 | 123 | 97.93% | 95.95% | Prefix-only glued-email trim | | 1.0.2 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Labelled country-area phone fix | | 1.0.3 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Label-note, e-Delivery and registry rules; placeholder and card fixes | | 1.1.0 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Batch API (`predict_many`, `scrub_many`), opt-in float16 | ## 1.0.3 Same weights, threshold, and API. Rules SHA `f32d5c54…`. - **Rules:** `NIP`/`REGON` labels with footnote marks or a short gloss (`NIP*:`, `NIP (Wykonawcy):`); e-Delivery addresses broken across a line; `DUNS`, `BDO` and `RPWDL` numbers right after their own label, at a fixed length. - **Wrapper:** existing `[PII]`/`[Telefon]` placeholders no longer switch the rules off. Before this, one placeholder anywhere in the text dropped every rule span for the whole document. - **Card:** `weights_sha256` was the hash of the source training checkpoint, not of `model.safetensors`. It is now `source_checkpoint_sha256`, and `model_safetensors_sha256` holds the published file's hash. Readers of `weights_sha256` must switch keys. 841-dev is unchanged: 0 rule spans change, and no 841-dev passage contains a placeholder. On simulated pre-masked input (gold spans replaced by their own placeholder), the rules now cover 1,812 of 2,606 remaining gold spans (1.0.2: 0), with 0 new false characters versus the same rules on unmasked text. Full Dynaword (2.8M documents): 27 documents change, 42 registry spans added, 0 removed. Known limit: an identifier with a placeholder inside it or right before it (`NIP [PII] …`, `8503[PII]…`) can still be missed by both layers, and a phone that follows an already-masked phone in a list can be missed by the rules. | Version | Whole /354 | Residual | Rules FP | Union FP | Char P | Char R | What changed | |---|---:|---:|---:|---:|---:|---:|---| | 1.0.0 | 323 | 25 | 133 | 133 | 97.76% | 95.95% | First Hub snapshot | | 1.0.1 | 323 | 25 | 98 | 123 | 97.93% | 95.95% | Prefix-only glued-email trim | | 1.0.2 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Labelled country-area phone fix | | 1.0.3 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Label-note, e-Delivery and registry rules; placeholder and card fixes | ## 1.0.2 Fix the public wrapper’s tokenizer batch shape so local model inference works. The packed tokenizer regression and an end-to-end local inference check cover this path. Fix immediately labelled hyphenated country-area phones with one-digit country codes or wider area codes. Numeric continuations, slash lists, weak contact cues, and nearby title prose still abstain. Same weights, threshold, and API. 841-dev: rules recover 14 more whole spans; the union recovers one (323 → 324), reducing residual passages 25 → 24. Rules FP remain 98; union FP remain 123. Lost gold, new false characters, and new clean-passage damage are all zero. Rules SHA `3016ae5b…`. Exact-span scores are recomputed from deduplicated raw rule/model span triples: precision 86.34%, recall 89.27%, F1 87.78%. The 1.0.1 card's precision/F1 were stale after the prefix trim changed exact rule/model duplicate counts; its character metrics were correct. | Version | Whole /354 | Residual | Rules FP | Union FP | Char P | Char R | What changed | |---|---:|---:|---:|---:|---:|---:|---| | 1.0.0 | 323 | 25 | 133 | 133 | 97.76% | 95.95% | First Hub snapshot | | 1.0.1 | 323 | 25 | 98 | 123 | 97.93% | 95.95% | Prefix-only glued-email trim | | 1.0.2 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Labelled country-area phone fix | ## 1.0.1 Prefix-only glued-email trim. Cluster-gated title-case 5–11 letter prefixes are dropped from the redaction when the remainder is already a lowercase-local email. Same epoch-5 weights. Lost gold 0. New clean-passage damage 0. | Version | Whole /354 | Residual | Rules FP | Union FP | Char P | Char R | What changed | |---|---:|---:|---:|---:|---:|---:|---| | 1.0.0 | 323 | 25 | 133 | 133 | 97.76% | 95.95% | First Hub snapshot | | 1.0.1 | 323 | 25 | 98 | 123 | 97.93% | 95.95% | Prefix-only glued-email trim | Rules SHA `4dcc441c…`. Seven emails trimmed, 35 prefix characters dropped; the model still covers 25 of those positions. Suffix glue is unchanged. ## 1.0.0 First Hub snapshot. Seed `202609160`, epoch 5, rules `547c0428…`. Union FP 133 was the glued-email regex floor. This seed added no new false characters versus the historical incumbent.