NERGAL / CHANGELOG.md
ppuzio's picture
1.1.0: batch API (predict_many, scrub_many), opt-in float16 (#2)
78cc85a
|
Raw
History Blame Contribute Delete
6.67 kB

NERGAL versions

Semver for this island:

  • MAJOR — API, threshold, or weight recipe changes
  • MINOR — new capability, same API
  • PATCH — rules or card fix, same weights and API

Accuracy is 841-dev, union at 0.95, 354 gold spans. A version that changes those numbers must update hybrid.json eval and the tables below.

1.1.0

Same weights, rules, threshold and default outputs; hybrid.json eval is unchanged. New batch API and opt-in fp16.

  • Nergal.predict_many(texts) / scrub_many(texts): windows from up to 64 texts are sorted by token length and packed into batches of at most 32,768 padded tokens and 128 rows. predict / scrub are now the one-text case of these.
  • dtype='float16' on Nergal(...) / from_pretrained(...) (CLI --dtype): casts the float32 weights at load time. Needs CUDA or MPS. The default stays float32, and model.safetensors is unchanged.
  • Faster window sizing: Encoding.count adds up cached unit pieces instead of re-tokenizing inside the window search. It yields the same windows, because encode() still rejects any unit whose pieces change with context.
  • Default device: from_pretrained now tries CUDA, then MPS, then CPU. 1.0.x used CPU even on CUDA machines.

Throughput on one RTX 4090 (13.88M chars of FineWeb-2, kchar/s): 1.0.3-style per-document batches 11.0 (23.0 with 3 processes); predict_many float32 21.3 (28.1 with 2); predict_many float16 39.4 (79.6 with 3). The float16 path is limited by CPU-side tokenization, so run 2–3 processes per GPU.

Equivalence: on the 1,685 labelled dev rows, float32 predict_many gives 0 span changes at 0.95 against the cached model spans of the published weights (841-dev max score change 2.6e-5). Float16 adds 2 spans on gold (1 on 841-dev, already covered by the rules) and removes none; 841-dev union numbers are identical. On 6,182 FineWeb-2 documents (1,236 with MPS and 4,946 with CUDA float32 references), CUDA float16 gave 0 span changes (max score change 0.012).

Version Whole /354 Residual Rules FP Union FP Char P Char R What changed
1.0.0 323 25 133 133 97.76% 95.95% First Hub snapshot
1.0.1 323 25 98 123 97.93% 95.95% Prefix-only glued-email trim
1.0.2 324 24 98 123 97.93% 96.12% Labelled country-area phone fix
1.0.3 324 24 98 123 97.93% 96.12% Label-note, e-Delivery and registry rules; placeholder and card fixes
1.1.0 324 24 98 123 97.93% 96.12% Batch API (predict_many, scrub_many), opt-in float16

1.0.3

Same weights, threshold, and API. Rules SHA f32d5c54….

  • Rules: NIP/REGON labels with footnote marks or a short gloss (NIP*:, NIP (Wykonawcy):); e-Delivery addresses broken across a line; DUNS, BDO and RPWDL numbers right after their own label, at a fixed length.
  • Wrapper: existing [PII]/[Telefon] placeholders no longer switch the rules off. Before this, one placeholder anywhere in the text dropped every rule span for the whole document.
  • Card: weights_sha256 was the hash of the source training checkpoint, not of model.safetensors. It is now source_checkpoint_sha256, and model_safetensors_sha256 holds the published file's hash. Readers of weights_sha256 must switch keys.

841-dev is unchanged: 0 rule spans change, and no 841-dev passage contains a placeholder. On simulated pre-masked input (gold spans replaced by their own placeholder), the rules now cover 1,812 of 2,606 remaining gold spans (1.0.2: 0), with 0 new false characters versus the same rules on unmasked text. Full Dynaword (2.8M documents): 27 documents change, 42 registry spans added, 0 removed.

Known limit: an identifier with a placeholder inside it or right before it (NIP [PII] …, 8503[PII]…) can still be missed by both layers, and a phone that follows an already-masked phone in a list can be missed by the rules.

Version Whole /354 Residual Rules FP Union FP Char P Char R What changed
1.0.0 323 25 133 133 97.76% 95.95% First Hub snapshot
1.0.1 323 25 98 123 97.93% 95.95% Prefix-only glued-email trim
1.0.2 324 24 98 123 97.93% 96.12% Labelled country-area phone fix
1.0.3 324 24 98 123 97.93% 96.12% Label-note, e-Delivery and registry rules; placeholder and card fixes

1.0.2

Fix the public wrapper’s tokenizer batch shape so local model inference works. The packed tokenizer regression and an end-to-end local inference check cover this path.

Fix immediately labelled hyphenated country-area phones with one-digit country codes or wider area codes. Numeric continuations, slash lists, weak contact cues, and nearby title prose still abstain. Same weights, threshold, and API.

841-dev: rules recover 14 more whole spans; the union recovers one (323 → 324), reducing residual passages 25 → 24. Rules FP remain 98; union FP remain 123. Lost gold, new false characters, and new clean-passage damage are all zero. Rules SHA 3016ae5b….

Exact-span scores are recomputed from deduplicated raw rule/model span triples: precision 86.34%, recall 89.27%, F1 87.78%. The 1.0.1 card's precision/F1 were stale after the prefix trim changed exact rule/model duplicate counts; its character metrics were correct.

Version Whole /354 Residual Rules FP Union FP Char P Char R What changed
1.0.0 323 25 133 133 97.76% 95.95% First Hub snapshot
1.0.1 323 25 98 123 97.93% 95.95% Prefix-only glued-email trim
1.0.2 324 24 98 123 97.93% 96.12% Labelled country-area phone fix

1.0.1

Prefix-only glued-email trim. Cluster-gated title-case 5–11 letter prefixes are dropped from the redaction when the remainder is already a lowercase-local email. Same epoch-5 weights. Lost gold 0. New clean-passage damage 0.

Version Whole /354 Residual Rules FP Union FP Char P Char R What changed
1.0.0 323 25 133 133 97.76% 95.95% First Hub snapshot
1.0.1 323 25 98 123 97.93% 95.95% Prefix-only glued-email trim

Rules SHA 4dcc441c…. Seven emails trimmed, 35 prefix characters dropped; the model still covers 25 of those positions. Suffix glue is unchanged.

1.0.0

First Hub snapshot. Seed 202609160, epoch 5, rules 547c0428…. Union FP 133 was the glued-email regex floor. This seed added no new false characters versus the historical incumbent.