--- language: - en - la - de - he - hbo - smp - sam - fr - fi - el - is - syc - syr - sv - ar - chu license: apache-2.0 library_name: regnetx-det tags: - computer-vision - object-detection - baseline-detection - document-analysis - historical-documents - handwriting - page-xml - alto-xml - polyline - regnetx - pytorch - safetensors pipeline_tag: object-detection base_model: - timm/regnetx_080.tv2_in1k - johnlockejrr/dfine-det-large-baseline-stage0 base_model_relation: finetune datasets: - custom metrics: - f1 model-index: - name: regnetx-8gf-polyline-baseline-stage0 results: - task: type: object-detection name: Text-line baseline detection (polyline) dataset: name: Stage-0 multiscript holdout (5% val split) type: custom split: validation metrics: - type: f1 name: cbad_f1_max (conf sweep) value: 0.929 - type: f1 name: cbad_f1 @ conf=0.4 value: 0.929 - type: precision name: Precision @ conf=0.4 (best epoch) value: 0.930 - type: recall name: Recall @ conf=0.4 (best epoch) value: 0.929 - type: f1 name: cbad_f1 @ conf=0.1 (logged, not recommended) value: 0.751 --- # regnetx-8gf-polyline-baseline-stage0 **Stage-0 multiscript pretrained weights** for [`regnetx-det`](https://github.com/johnlockejrr/regnetx-det) — **RegNetX-8GF** + D-FINE **HybridEncoder** + **PolylineTransformer**, predicting **text-line baselines as polylines** (B-spline control points), not boxes. > Hub repo: [`johnlockejrr/regnetx-8gf-polyline-baseline-stage0`](https://huggingface.co/johnlockejrr/regnetx-8gf-polyline-baseline-stage0) > Primary file: `best_cbad_f1.safetensors` (~208 MB) > **Paper (Zenodo preprint):** [doi:10.5281/zenodo.22059333](https://doi.org/10.5281/zenodo.22059333) > **Recommended operating confidence: `0.40`** (CLI `--conf 0.4`). `0.50` is a close alternative. This checkpoint is a **generic baseline pretrain** for downstream fine-tunes (Hebrew/Samaritan, ICDAR cBAD Stage-1, other PAGE/ALTO corpora). It improves on the D-FINE Stage-0 teacher on the **same multiscript Arrow validation split**. Sibling checkpoints: - D-FINE Stage-0 teacher: [`johnlockejrr/dfine-det-large-baseline-stage0`](https://huggingface.co/johnlockejrr/dfine-det-large-baseline-stage0) (`cbad_f1_max` ≈ **0.893**) - Hebrew/Samaritan domain fine-tune (older warm-start): [`johnlockejrr/regnetx-8gf-polyline-hebrew-samaritan`](https://huggingface.co/johnlockejrr/regnetx-8gf-polyline-hebrew-samaritan) (`cbad_f1_max` ≈ **0.942** on internal val) --- ## Model summary | | | |---|---| | **Architecture** | `RegNetXPolylineModel` — timm RegNetX-8GF → HybridEncoder (C3/C4/C5) → PolylineTransformer | | **Backbone** | `regnetx_080.tv2_in1k` (ImageNet-1K V2) | | **Neck** | D-FINE HybridEncoder (`hidden_dim=256`, strides 8/16/32) | | **Parameters** | ~54.1 M | | **Queries** | 300 | | **Geometry** | \(K=8\) cubic B-spline control points per line + height | | **Canvas** | \(1280\times1280\) letterbox | | **Init** | ImageNet backbone + **shape-matched load** from D-FINE Stage-0 (`enc≈543`, `dec≈187`; `input_proj.*.conv.weight` stay RegNetX-specific) | | **Task** | Document **baseline / text-line** detection → PAGE or ALTO export | Detection is **set prediction** of scored polylines. Polygonal line environments (BLLA) are applied only at PAGE/ALTO serialize time. --- ## Intended use **Use for** - Starting point for **Stage-1 / domain fine-tuning** on a target script or corpus - Zero-shot probing of baselines on heterogeneous manuscripts (expect domain gap) - Research comparing RegNetX + HybridEncoder vs D-FINE / ConvNeXt polyline siblings **Not for** - Claiming official **ICDAR cBAD 2019 test** numbers from these Stage-0 val scores - Production OCR **transcription** (lines only; no text recognition) - Layout *region* detection (paragraphs, tables, illustrations) - Treating `cbad_f1` @ conf **0.1** as the quality number (see operating point below) --- ## Training data (Stage-0) Same multiscript PAGE-XML / ALTO baseline Arrow as [`dfine-det` Stage-0](https://huggingface.co/johnlockejrr/dfine-det-large-baseline-stage0): HTR-United–style open corpora + private packs, compiled with adaptive simplify → uniform cubic B-spline \(K=8\). | Split | Pages | |------:|------:| | Train | **46 511** | | Val (holdout) | **1 923** (~5%, seed 42) | Corpus mix matches the D-FINE Stage-0 card (29 corpora, multilingual historical scripts). See the [dfine-det Stage-0 model card](https://huggingface.co/johnlockejrr/dfine-det-large-baseline-stage0) for the full corpus table. **Important:** ICDAR cBAD pages appear **inside** this pretrain mixture. Stage-0 validation F1 is an **internal multiscript holdout**, **not** a fair official cBAD 2019 test score. For apples-to-apples cBAD bake-offs, use a separate **Stage-1-cBAD** fine-tune and the official test split. --- ## Training recipe | Hyperparameter | Value | |----------------|-------| | Init | `regnetx_080.tv2_in1k` + D-FINE Stage-0 encoder/decoder transfer | | Optimizer | AdamW; base LR \(1\times10^{-4}\), backbone \(0.2\times\) | | Schedule | Linear warmup 2 epochs + cosine | | Precision | `bf16-mixed` | | Effective batch | 32 (micro-batch 32 × accum 1 on MI300X) | | Epochs (cap) | 80; early-stop on `cbad_f1_max` (patience 12, min 15) | | Augment | Mild photometric + ±2° rotation (OOB lines dropped) | | Match distance | 20 px on the 1280 canvas | | Monitor | `cbad_f1_max` over conf ∈ {0.1…0.6} with **`fast_conf_sweep`** | | Seed | 42 | Config reference: `configs/baseline_pretrain_stage0_droplet.yaml` in the **regnetx-det** repo. Hardware: **AMD Instinct MI300X** on [AMD Developer Cloud](https://devcloud.amd.com/) (ROCm 7.14, PyTorch HIP). --- ## Evaluation (Stage-0 val) Metric: cBAD-style line F1 — Hungarian matching of densified polylines with mean bidirectional Chamfer cost; TP if cost ≤ 20 px. ### Released checkpoint | Metric | Value | Notes | |--------|------:|-------| | **`cbad_f1_max`** | **≈ 0.929** | Best monitored snapshot (`best_0.9293.safetensors`) | | Best epoch | **32** | From bundled TensorBoard logs | | Operating conf | **≈ 0.40** | Sweep winner at best epoch | | Precision @ 0.4 | **≈ 0.930** | Same epoch | | Recall @ 0.4 | **≈ 0.929** | Same epoch | | Mean Chamfer @ best | **≈ 4.0 px** | Same epoch | | Weights file | `best_cbad_f1.safetensors` | Same tensor data as `best_0.9293.safetensors` | Compared to the D-FINE Stage-0 teacher on the **same val split**: **+0.036** absolute `cbad_f1_max` (0.893 → 0.929). Fixed conf **0.1** F1 (`cbad_f1` on the progress bar) is substantially lower than `cbad_f1_max` — always **sweep** (or use ~**0.4** for this checkpoint) before comparing systems. --- ## How to use ### Install ```bash git clone cd regnetx-det python3 -m venv .venv && source .venv/bin/activate pip install -e ".[dev]" ``` Use this project's `.venv` (not a sibling detector repo). ### Download weights ```bash huggingface-cli download johnlockejrr/regnetx-8gf-polyline-baseline-stage0 \ best_cbad_f1.safetensors \ --local-dir PRETRAIN/regnetx_stage0 ``` ### Inference (PAGE XML) — preferred path ```bash regnetx-det infer-page \ -m PRETRAIN/regnetx_stage0/best_cbad_f1.safetensors \ -d cuda:0 \ --format page \ --conf 0.4 \ --text-direction horizontal-rl \ page.jpg ``` Single-image export with extra NMS/polygon knobs: ```bash regnetx-det infer \ -w PRETRAIN/regnetx_stage0/best_cbad_f1.safetensors \ --image page.jpg \ -o page.xml \ --format page \ --conf 0.4 \ --reading-order rtl ``` ### Fine-tune (Stage-1 / domain) ```yaml train: load: PRETRAIN/regnetx_stage0/best_cbad_f1.safetensors backbone: regnetx_080.tv2_in1k pretrained_backbone: false lr: 3.0e-5 backbone_lr_mult: 0.2 monitor: cbad_f1_max ``` ```bash regnetx-det --config configs/hebrew_samaritan_regnetx_stage0_finetune_balanced.yaml -d cuda:0 train ``` Hebrew/Samaritan and cBAD Stage-1 YAML templates ship in the **regnetx-det** repo. --- ## Files in this release | File | Upload to Hub? | Description | |------|:------------:|-------------| | `best_cbad_f1.safetensors` | **Yes** | **Recommended** Stage-0 weights | | `README.md` | **Yes** | This model card | --- ## Limitations - Val F1 is **not** official cBAD 2019 test performance. - Pretrain includes cBAD pages → do not treat Stage-0 zero-shot cBAD scores as an external bake-off. - \(K=8\) control points underfit strongly curved / torn lines. - EMA is used during training; released safetensors are the monitored best snapshot. - Some listed corpora may have missing sidecars at compile time; Arrow compile skips unresolved pages. --- ## Citation & credits ### Paper (prefer this) Preprint: [regnetx-det: Set-Prediction Polyline Baselines for Historical Documents with HybridEncoder Transfer](https://doi.org/10.5281/zenodo.22059333) (Zenodo, Aug 2026). ```bibtex @misc{bors2026regnetxdet, author = {Bors, Teodor}, title = {regnetx-det: Set-Prediction Polyline Baselines for Historical Documents with HybridEncoder Transfer}, year = {2026}, month = aug, publisher = {Zenodo}, version = {v1}, doi = {10.5281/zenodo.22059333}, url = {https://doi.org/10.5281/zenodo.22059333} } ``` ### D-FINE (HybridEncoder / decoder init) ```bibtex @inproceedings{peng2025dfine, title = {D-FINE: Redefine Regression Task in DETRs as Fine-grained Distribution Refinement}, author = {Peng, Yansong and Li, Hebei and Wu, Peixi and Zhang, Yueyi and Sun, Xiaoyan and Wu, Feng}, booktitle = {The Thirteenth International Conference on Learning Representations}, year = {2025}, url = {https://arxiv.org/abs/2410.13842} } ``` - Code: [github.com/Peterande/D-FINE](https://github.com/Peterande/D-FINE) - Stage-0 teacher: [johnlockejrr/dfine-det-large-baseline-stage0](https://huggingface.co/johnlockejrr/dfine-det-large-baseline-stage0) ### This Stage-0 checkpoint (weights) ```bibtex @software{regnetx_det_stage0_8gf, title = {regnetx-det Stage-0: RegNetX-8GF Multiscript Polyline Baseline Pretrain}, author = {Bors, Teodor}, year = {2026}, note = {RegNetX-8GF + HybridEncoder + PolylineTransformer; see Zenodo preprint doi:10.5281/zenodo.22059333}, url = {https://huggingface.co/johnlockejrr/regnetx-8gf-polyline-baseline-stage0} } ``` ## Acknowledgements Compute for Stage-0 pretraining was provided via complimentary credits on the [AMD Developer Cloud](https://devcloud.amd.com/) (AMD Instinct™ MI300X) through the [AMD AI Developer Program](https://www.amd.com/en/developer/ai-dev-program.html). ### Additional notices - Vendored D-FINE modules: see `THIRD_PARTY_NOTICES.md` in the repo (Apache-2.0). - PAGE polygonalization helpers derive from Kraken/BLLA (Apache-2.0) — used at export, not in the pretrain loss. - Training corpora: respect each dataset's original license. Private packs are **not** redistributed with this card; only model weights trained with them are released where license allows. --- ## License **Apache License 2.0** for the **regnetx-det** code and these weights. Downstream users must comply with licenses of any datasets used in further fine-tuning.