Instructions to use SlayerLab/NERGAL with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use SlayerLab/NERGAL with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="SlayerLab/NERGAL")# Load model directly from transformers import AutoTokenizer, AutoModelForTokenClassification tokenizer = AutoTokenizer.from_pretrained("SlayerLab/NERGAL") model = AutoModelForTokenClassification.from_pretrained("SlayerLab/NERGAL", device_map="auto") - Notebooks
- Google Colab
- Kaggle
|
Download CHANGELOG.md from SlayerLab/NERGAL: direct link, hf CLI and curl.
- Browser
- Download file 11.8 kB
-
https://huggingface.co/SlayerLab/NERGAL/resolve/main/CHANGELOG.md
- Command line
-
hf download hf://SlayerLab/NERGAL/CHANGELOG.md
-
curl -L -o CHANGELOG.md https://huggingface.co/SlayerLab/NERGAL/resolve/main/CHANGELOG.md
11.8 kB
| # NERGAL versions | |
| Semver for this island: | |
| - **MAJOR** — API, threshold, or weight recipe changes | |
| - **MINOR** — new capability, same API | |
| - **PATCH** — rules or card fix, same weights and API | |
| Accuracy is 841-dev, union at 0.95, 354 gold spans. A version that changes those numbers must update `hybrid.json` `eval` and the tables below. | |
| ## 1.1.2 | |
| Same weights, threshold and API. Rules SHA `08faef84…`. The rules file is the frozen and scored candidate `08faef84…`. | |
| - **Rules (email only):** an address match ends where glued text begins, before a `www.` host glued onto the domain (`jan@firma.plwww.…`) or at a capital glued onto a lowercase domain ending (`jan@firma.plKontakt`). A cut is kept only when what remains is a complete address. A match right after another `@` is dropped when a space comes before the match's own `@` (a list of @-mentions). Without the space it keeps its mask: a handle (`@jan@firma.social`), a label glued on with `@` or glued addresses can hold a real one. | |
| - **Gold:** one 841-dev email span ran on into a glued URL host, unlike the other seven addresses in its passage and the labelling policy. It was reviewed and trimmed by 10 characters (`gold_check_v2`, `hybrid.json` `eval.gold_amendments`). Numbers from 1.1.2 on use the amended gold; earlier rows use the original. On the amended gold, 1.1.1 has rules FP 108, union FP 133, char P 97.77% and char R 96.49%. | |
| 841-dev (amended gold): whole 326/354 and residual 23 unchanged; union FP 133 → 80, rules FP 108 → 24; exact-span and per-label numbers unchanged. All changes are in one passage. Its 24 remaining rules FP are capitalised words glued before a local part, which the 1.0.1 prefix trim cuts only partly. Web development set (41,204 passages): 13 passages change, 0 characters added, and a review ruled all dropped runs out of scope. Human test v2 email benchmark half: rules FP 89 → 65, union FP 162 → 138; not independent, since that half surfaced the cases. Cue-less phone test (150 passages): one passage changes, false characters 95 → 78, phones unchanged. Full Dynaword (2,797,400 documents): 165 documents change, almost all EUR-Lex. There are 553 cuts: 521 at a capital and 32 at a glued `www.`. No cut frees an `@`, and 26 cuts free a glued `Tel`/`Fax` label, whose phone number the rules now mask. No mask is dropped. A first candidate also dropped a match after an `@` inside a word (`a@b@firma.pl`). On full Dynaword that lost real addresses in glued lists, so it was removed. | |
| Known limits: a lowercase word glued onto the domain (`jan@firma.plkontakt`) stays masked with the address, because the TLD list does not separate the two; text glued before the local part (a `www.` host or a word) is not cut. | |
| | Version | Whole /354 | Residual | Rules FP | Union FP | Char P | Char R | What changed | | |
| |---|---:|---:|---:|---:|---:|---:|---| | |
| | 1.0.0 | 323 | 25 | 133 | 133 | 97.76% | 95.95% | First Hub snapshot | | |
| | 1.0.1 | 323 | 25 | 98 | 123 | 97.93% | 95.95% | Prefix-only glued-email trim | | |
| | 1.0.2 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Labelled country-area phone fix | | |
| | 1.0.3 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Label-note, e-Delivery and registry rules; placeholder and card fixes | | |
| | 1.1.0 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Batch API (`predict_many`, `scrub_many`), opt-in float16 | | |
| | 1.1.1 | 326 | 23 | 98 | 123 | 97.94% | 96.50% | Grouped national phones without a label | | |
| | 1.1.1, amended gold | 326 | 23 | 108 | 133 | 97.77% | 96.49% | Restated on `gold_check_v2` | | |
| | 1.1.2 | 326 | 23 | 24 | 80 | 98.65% | 96.49% | Email boundary rules | | |
| ## 1.1.1 | |
| Same weights, threshold and API. Rules SHA `ad5c51f7…`. The rules file is the frozen and scored candidate `7412261c…` with comment tags removed; its syntax tree is identical. | |
| - **Rules:** Polish phones in the grouped national forms are masked without a contact label: mobile `601 234 567` and landline `22 123 45 67` (area code optionally in parentheses), optional `+48`, one separator kind (space or hyphen) throughout. Not taken: plain 9-digit strings (still label-gated), amounts, round counts (`500 000 000`), decimal figures (`601 234 567,89`), and numbers after another identifier's label (NIP, REGON, codes). After a bare grouped phone, a list goes on only with complete Polish numbers. | |
| 841-dev: 326/354 whole (1.1.0: 324), residual 23 (24), 0 new false characters; phone 147/169 (145). A frozen blind set of 150 web passages selected for unlabelled phones: phones wholly masked 66/111 (1.1.0: 37), all values 127/178 (98), 0 lost, 1 passage with new false characters. Dynaword (60,000 documents): 79 documents change, 247 phone spans added, 0 removed; in a reviewed sample 1 of 23 sampled documents' added spans was not a phone. | |
| | Version | Whole /354 | Residual | Rules FP | Union FP | Char P | Char R | What changed | | |
| |---|---:|---:|---:|---:|---:|---:|---| | |
| | 1.0.0 | 323 | 25 | 133 | 133 | 97.76% | 95.95% | First Hub snapshot | | |
| | 1.0.1 | 323 | 25 | 98 | 123 | 97.93% | 95.95% | Prefix-only glued-email trim | | |
| | 1.0.2 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Labelled country-area phone fix | | |
| | 1.0.3 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Label-note, e-Delivery and registry rules; placeholder and card fixes | | |
| | 1.1.0 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Batch API (`predict_many`, `scrub_many`), opt-in float16 | | |
| | 1.1.1 | 326 | 23 | 98 | 123 | 97.94% | 96.50% | Grouped national phones without a label | | |
| ## 1.1.0 | |
| Same weights, rules, threshold and default outputs; `hybrid.json` `eval` is unchanged. New batch API and opt-in fp16. | |
| - **`Nergal.predict_many(texts)` / `scrub_many(texts)`:** windows from up to 64 texts are sorted by token length and packed into batches of at most 32,768 padded tokens and 128 rows. `predict` / `scrub` are now the one-text case of these. | |
| - **`dtype='float16'`** on `Nergal(...)` / `from_pretrained(...)` (CLI `--dtype`): casts the float32 weights at load time. Needs CUDA or MPS. The default stays `float32`, and `model.safetensors` is unchanged. | |
| - **Faster window sizing:** `Encoding.count` adds up cached unit pieces instead of re-tokenizing inside the window search. It yields the same windows, because `encode()` still rejects any unit whose pieces change with context. | |
| - **Default device:** `from_pretrained` now tries CUDA, then MPS, then CPU. 1.0.x used CPU even on CUDA machines. | |
| Throughput on one RTX 4090 (13.88M chars of FineWeb-2, kchar/s): 1.0.3-style per-document batches 11.0 (23.0 with 3 processes); `predict_many` float32 21.3 (28.1 with 2); `predict_many` float16 39.4 (79.6 with 3). The float16 path is limited by CPU-side tokenization, so run 2–3 processes per GPU. | |
| Equivalence: on the 1,685 labelled dev rows, float32 `predict_many` gives 0 span changes at 0.95 against the cached model spans of the published weights (841-dev max score change 2.6e-5). Float16 adds 2 spans on gold (1 on 841-dev, already covered by the rules) and removes none; 841-dev union numbers are identical. On 6,182 FineWeb-2 documents (1,236 with MPS and 4,946 with CUDA float32 references), CUDA float16 gave 0 span changes (max score change 0.012). | |
| | Version | Whole /354 | Residual | Rules FP | Union FP | Char P | Char R | What changed | | |
| |---|---:|---:|---:|---:|---:|---:|---| | |
| | 1.0.0 | 323 | 25 | 133 | 133 | 97.76% | 95.95% | First Hub snapshot | | |
| | 1.0.1 | 323 | 25 | 98 | 123 | 97.93% | 95.95% | Prefix-only glued-email trim | | |
| | 1.0.2 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Labelled country-area phone fix | | |
| | 1.0.3 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Label-note, e-Delivery and registry rules; placeholder and card fixes | | |
| | 1.1.0 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Batch API (`predict_many`, `scrub_many`), opt-in float16 | | |
| ## 1.0.3 | |
| Same weights, threshold, and API. Rules SHA `f32d5c54…`. | |
| - **Rules:** `NIP`/`REGON` labels with footnote marks or a short gloss (`NIP*:`, `NIP (Wykonawcy):`); e-Delivery addresses broken across a line; `DUNS`, `BDO` and `RPWDL` numbers right after their own label, at a fixed length. | |
| - **Wrapper:** existing `[PII]`/`[Telefon]` placeholders no longer switch the rules off. Before this, one placeholder anywhere in the text dropped every rule span for the whole document. | |
| - **Card:** `weights_sha256` was the hash of the source training checkpoint, not of `model.safetensors`. It is now `source_checkpoint_sha256`, and `model_safetensors_sha256` holds the published file's hash. Readers of `weights_sha256` must switch keys. | |
| 841-dev is unchanged: 0 rule spans change, and no 841-dev passage contains a placeholder. On simulated pre-masked input (gold spans replaced by their own placeholder), the rules now cover 1,812 of 2,606 remaining gold spans (1.0.2: 0), with 0 new false characters versus the same rules on unmasked text. Full Dynaword (2.8M documents): 27 documents change, 42 registry spans added, 0 removed. | |
| Known limit: an identifier with a placeholder inside it or right before it (`NIP [PII] …`, `8503[PII]…`) can still be missed by both layers, and a phone that follows an already-masked phone in a list can be missed by the rules. | |
| | Version | Whole /354 | Residual | Rules FP | Union FP | Char P | Char R | What changed | | |
| |---|---:|---:|---:|---:|---:|---:|---| | |
| | 1.0.0 | 323 | 25 | 133 | 133 | 97.76% | 95.95% | First Hub snapshot | | |
| | 1.0.1 | 323 | 25 | 98 | 123 | 97.93% | 95.95% | Prefix-only glued-email trim | | |
| | 1.0.2 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Labelled country-area phone fix | | |
| | 1.0.3 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Label-note, e-Delivery and registry rules; placeholder and card fixes | | |
| ## 1.0.2 | |
| Fix the public wrapper’s tokenizer batch shape so local model inference works. The packed tokenizer regression and an end-to-end local inference check cover this path. | |
| Fix immediately labelled hyphenated country-area phones with one-digit country codes or wider area codes. Numeric continuations, slash lists, weak contact cues, and nearby title prose still abstain. Same weights, threshold, and API. | |
| 841-dev: rules recover 14 more whole spans; the union recovers one (323 → 324), reducing residual passages 25 → 24. Rules FP remain 98; union FP remain 123. Lost gold, new false characters, and new clean-passage damage are all zero. Rules SHA `3016ae5b…`. | |
| Exact-span scores are recomputed from deduplicated raw rule/model span triples: precision 86.34%, recall 89.27%, F1 87.78%. The 1.0.1 card's precision/F1 were stale after the prefix trim changed exact rule/model duplicate counts; its character metrics were correct. | |
| | Version | Whole /354 | Residual | Rules FP | Union FP | Char P | Char R | What changed | | |
| |---|---:|---:|---:|---:|---:|---:|---| | |
| | 1.0.0 | 323 | 25 | 133 | 133 | 97.76% | 95.95% | First Hub snapshot | | |
| | 1.0.1 | 323 | 25 | 98 | 123 | 97.93% | 95.95% | Prefix-only glued-email trim | | |
| | 1.0.2 | 324 | 24 | 98 | 123 | 97.93% | 96.12% | Labelled country-area phone fix | | |
| ## 1.0.1 | |
| Prefix-only glued-email trim. Cluster-gated title-case 5–11 letter prefixes are dropped from the redaction when the remainder is already a lowercase-local email. Same epoch-5 weights. Lost gold 0. New clean-passage damage 0. | |
| | Version | Whole /354 | Residual | Rules FP | Union FP | Char P | Char R | What changed | | |
| |---|---:|---:|---:|---:|---:|---:|---| | |
| | 1.0.0 | 323 | 25 | 133 | 133 | 97.76% | 95.95% | First Hub snapshot | | |
| | 1.0.1 | 323 | 25 | 98 | 123 | 97.93% | 95.95% | Prefix-only glued-email trim | | |
| Rules SHA `4dcc441c…`. Seven emails trimmed, 35 prefix characters dropped; the model still covers 25 of those positions. Suffix glue is unchanged. | |
| ## 1.0.0 | |
| First Hub snapshot. Seed `202609160`, epoch 5, rules `547c0428…`. Union FP 133 was the glued-email regex floor. This seed added no new false characters versus the historical incumbent. | |