johnlockejrr's picture
Upload README.md with huggingface_hub
678d12f verified
|
Raw
History Blame Contribute Delete
11.5 kB
---
language:
- en
- la
- de
- he
- hbo
- smp
- sam
- fr
- fi
- el
- is
- syc
- syr
- sv
- ar
- chu
license: apache-2.0
library_name: regnetx-det
tags:
- computer-vision
- object-detection
- baseline-detection
- document-analysis
- historical-documents
- handwriting
- page-xml
- alto-xml
- polyline
- regnetx
- pytorch
- safetensors
pipeline_tag: object-detection
base_model:
- timm/regnetx_080.tv2_in1k
- johnlockejrr/dfine-det-large-baseline-stage0
base_model_relation: finetune
datasets:
- custom
metrics:
- f1
model-index:
- name: regnetx-8gf-polyline-baseline-stage0
results:
- task:
type: object-detection
name: Text-line baseline detection (polyline)
dataset:
name: Stage-0 multiscript holdout (5% val split)
type: custom
split: validation
metrics:
- type: f1
name: cbad_f1_max (conf sweep)
value: 0.929
- type: f1
name: cbad_f1 @ conf=0.4
value: 0.929
- type: precision
name: Precision @ conf=0.4 (best epoch)
value: 0.930
- type: recall
name: Recall @ conf=0.4 (best epoch)
value: 0.929
- type: f1
name: cbad_f1 @ conf=0.1 (logged, not recommended)
value: 0.751
---
# regnetx-8gf-polyline-baseline-stage0
**Stage-0 multiscript pretrained weights** for [`regnetx-det`](https://github.com/johnlockejrr/regnetx-det) — **RegNetX-8GF** + D-FINE **HybridEncoder** + **PolylineTransformer**, predicting **text-line baselines as polylines** (B-spline control points), not boxes.
> Hub repo: [`johnlockejrr/regnetx-8gf-polyline-baseline-stage0`](https://huggingface.co/johnlockejrr/regnetx-8gf-polyline-baseline-stage0)
> Primary file: `best_cbad_f1.safetensors` (~208 MB)
> **Paper (Zenodo preprint):** [doi:10.5281/zenodo.22059333](https://doi.org/10.5281/zenodo.22059333)
> **Recommended operating confidence: `0.40`** (CLI `--conf 0.4`). `0.50` is a close alternative.
This checkpoint is a **generic baseline pretrain** for downstream fine-tunes (Hebrew/Samaritan, ICDAR cBAD Stage-1, other PAGE/ALTO corpora). It improves on the D-FINE Stage-0 teacher on the **same multiscript Arrow validation split**.
Sibling checkpoints:
- D-FINE Stage-0 teacher: [`johnlockejrr/dfine-det-large-baseline-stage0`](https://huggingface.co/johnlockejrr/dfine-det-large-baseline-stage0) (`cbad_f1_max`**0.893**)
- Hebrew/Samaritan domain fine-tune (older warm-start): [`johnlockejrr/regnetx-8gf-polyline-hebrew-samaritan`](https://huggingface.co/johnlockejrr/regnetx-8gf-polyline-hebrew-samaritan) (`cbad_f1_max`**0.942** on internal val)
---
## Model summary
| | |
|---|---|
| **Architecture** | `RegNetXPolylineModel` — timm RegNetX-8GF → HybridEncoder (C3/C4/C5) → PolylineTransformer |
| **Backbone** | `regnetx_080.tv2_in1k` (ImageNet-1K V2) |
| **Neck** | D-FINE HybridEncoder (`hidden_dim=256`, strides 8/16/32) |
| **Parameters** | ~54.1 M |
| **Queries** | 300 |
| **Geometry** | \(K=8\) cubic B-spline control points per line + height |
| **Canvas** | \(1280\times1280\) letterbox |
| **Init** | ImageNet backbone + **shape-matched load** from D-FINE Stage-0 (`enc≈543`, `dec≈187`; `input_proj.*.conv.weight` stay RegNetX-specific) |
| **Task** | Document **baseline / text-line** detection → PAGE or ALTO export |
Detection is **set prediction** of scored polylines. Polygonal line environments (BLLA) are applied only at PAGE/ALTO serialize time.
---
## Intended use
**Use for**
- Starting point for **Stage-1 / domain fine-tuning** on a target script or corpus
- Zero-shot probing of baselines on heterogeneous manuscripts (expect domain gap)
- Research comparing RegNetX + HybridEncoder vs D-FINE / ConvNeXt polyline siblings
**Not for**
- Claiming official **ICDAR cBAD 2019 test** numbers from these Stage-0 val scores
- Production OCR **transcription** (lines only; no text recognition)
- Layout *region* detection (paragraphs, tables, illustrations)
- Treating `cbad_f1` @ conf **0.1** as the quality number (see operating point below)
---
## Training data (Stage-0)
Same multiscript PAGE-XML / ALTO baseline Arrow as [`dfine-det` Stage-0](https://huggingface.co/johnlockejrr/dfine-det-large-baseline-stage0): HTR-United–style open corpora + private packs, compiled with adaptive simplify → uniform cubic B-spline \(K=8\).
| Split | Pages |
|------:|------:|
| Train | **46 511** |
| Val (holdout) | **1 923** (~5%, seed 42) |
Corpus mix matches the D-FINE Stage-0 card (29 corpora, multilingual historical scripts). See the [dfine-det Stage-0 model card](https://huggingface.co/johnlockejrr/dfine-det-large-baseline-stage0) for the full corpus table.
**Important:** ICDAR cBAD pages appear **inside** this pretrain mixture. Stage-0 validation F1 is an **internal multiscript holdout**, **not** a fair official cBAD 2019 test score. For apples-to-apples cBAD bake-offs, use a separate **Stage-1-cBAD** fine-tune and the official test split.
---
## Training recipe
| Hyperparameter | Value |
|----------------|-------|
| Init | `regnetx_080.tv2_in1k` + D-FINE Stage-0 encoder/decoder transfer |
| Optimizer | AdamW; base LR \(1\times10^{-4}\), backbone \(0.2\times\) |
| Schedule | Linear warmup 2 epochs + cosine |
| Precision | `bf16-mixed` |
| Effective batch | 32 (micro-batch 32 × accum 1 on MI300X) |
| Epochs (cap) | 80; early-stop on `cbad_f1_max` (patience 12, min 15) |
| Augment | Mild photometric + ±2° rotation (OOB lines dropped) |
| Match distance | 20 px on the 1280 canvas |
| Monitor | `cbad_f1_max` over conf ∈ {0.1…0.6} with **`fast_conf_sweep`** |
| Seed | 42 |
Config reference: `configs/baseline_pretrain_stage0_droplet.yaml` in the **regnetx-det** repo.
Hardware: **AMD Instinct MI300X** on [AMD Developer Cloud](https://devcloud.amd.com/) (ROCm 7.14, PyTorch HIP).
---
## Evaluation (Stage-0 val)
Metric: cBAD-style line F1 — Hungarian matching of densified polylines with mean bidirectional Chamfer cost; TP if cost ≤ 20 px.
### Released checkpoint
| Metric | Value | Notes |
|--------|------:|-------|
| **`cbad_f1_max`** | **≈ 0.929** | Best monitored snapshot (`best_0.9293.safetensors`) |
| Best epoch | **32** | From bundled TensorBoard logs |
| Operating conf | **≈ 0.40** | Sweep winner at best epoch |
| Precision @ 0.4 | **≈ 0.930** | Same epoch |
| Recall @ 0.4 | **≈ 0.929** | Same epoch |
| Mean Chamfer @ best | **≈ 4.0 px** | Same epoch |
| Weights file | `best_cbad_f1.safetensors` | Same tensor data as `best_0.9293.safetensors` |
Compared to the D-FINE Stage-0 teacher on the **same val split**: **+0.036** absolute `cbad_f1_max` (0.893 → 0.929).
Fixed conf **0.1** F1 (`cbad_f1` on the progress bar) is substantially lower than `cbad_f1_max` — always **sweep** (or use ~**0.4** for this checkpoint) before comparing systems.
---
## How to use
### Install
```bash
git clone <regnetx-det-repo>
cd regnetx-det
python3 -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
```
Use this project's `.venv` (not a sibling detector repo).
### Download weights
```bash
huggingface-cli download johnlockejrr/regnetx-8gf-polyline-baseline-stage0 \
best_cbad_f1.safetensors \
--local-dir PRETRAIN/regnetx_stage0
```
### Inference (PAGE XML) — preferred path
```bash
regnetx-det infer-page \
-m PRETRAIN/regnetx_stage0/best_cbad_f1.safetensors \
-d cuda:0 \
--format page \
--conf 0.4 \
--text-direction horizontal-rl \
page.jpg
```
Single-image export with extra NMS/polygon knobs:
```bash
regnetx-det infer \
-w PRETRAIN/regnetx_stage0/best_cbad_f1.safetensors \
--image page.jpg \
-o page.xml \
--format page \
--conf 0.4 \
--reading-order rtl
```
### Fine-tune (Stage-1 / domain)
```yaml
train:
load: PRETRAIN/regnetx_stage0/best_cbad_f1.safetensors
backbone: regnetx_080.tv2_in1k
pretrained_backbone: false
lr: 3.0e-5
backbone_lr_mult: 0.2
monitor: cbad_f1_max
```
```bash
regnetx-det --config configs/hebrew_samaritan_regnetx_stage0_finetune_balanced.yaml -d cuda:0 train
```
Hebrew/Samaritan and cBAD Stage-1 YAML templates ship in the **regnetx-det** repo.
---
## Files in this release
| File | Upload to Hub? | Description |
|------|:------------:|-------------|
| `best_cbad_f1.safetensors` | **Yes** | **Recommended** Stage-0 weights |
| `README.md` | **Yes** | This model card |
---
## Limitations
- Val F1 is **not** official cBAD 2019 test performance.
- Pretrain includes cBAD pages → do not treat Stage-0 zero-shot cBAD scores as an external bake-off.
- \(K=8\) control points underfit strongly curved / torn lines.
- EMA is used during training; released safetensors are the monitored best snapshot.
- Some listed corpora may have missing sidecars at compile time; Arrow compile skips unresolved pages.
---
## Citation & credits
### Paper (prefer this)
Preprint: [regnetx-det: Set-Prediction Polyline Baselines for Historical Documents with HybridEncoder Transfer](https://doi.org/10.5281/zenodo.22059333) (Zenodo, Aug 2026).
```bibtex
@misc{bors2026regnetxdet,
author = {Bors, Teodor},
title = {regnetx-det: Set-Prediction Polyline Baselines for Historical Documents with HybridEncoder Transfer},
year = {2026},
month = aug,
publisher = {Zenodo},
version = {v1},
doi = {10.5281/zenodo.22059333},
url = {https://doi.org/10.5281/zenodo.22059333}
}
```
### D-FINE (HybridEncoder / decoder init)
```bibtex
@inproceedings{peng2025dfine,
title = {D-FINE: Redefine Regression Task in DETRs as Fine-grained Distribution Refinement},
author = {Peng, Yansong and Li, Hebei and Wu, Peixi and Zhang, Yueyi and Sun, Xiaoyan and Wu, Feng},
booktitle = {The Thirteenth International Conference on Learning Representations},
year = {2025},
url = {https://arxiv.org/abs/2410.13842}
}
```
- Code: [github.com/Peterande/D-FINE](https://github.com/Peterande/D-FINE)
- Stage-0 teacher: [johnlockejrr/dfine-det-large-baseline-stage0](https://huggingface.co/johnlockejrr/dfine-det-large-baseline-stage0)
### This Stage-0 checkpoint (weights)
```bibtex
@software{regnetx_det_stage0_8gf,
title = {regnetx-det Stage-0: RegNetX-8GF Multiscript Polyline Baseline Pretrain},
author = {Bors, Teodor},
year = {2026},
note = {RegNetX-8GF + HybridEncoder + PolylineTransformer; see Zenodo preprint doi:10.5281/zenodo.22059333},
url = {https://huggingface.co/johnlockejrr/regnetx-8gf-polyline-baseline-stage0}
}
```
## Acknowledgements
Compute for Stage-0 pretraining was provided via complimentary credits on the
[AMD Developer Cloud](https://devcloud.amd.com/) (AMD Instinct™ MI300X) through the
[AMD AI Developer Program](https://www.amd.com/en/developer/ai-dev-program.html).
### Additional notices
- Vendored D-FINE modules: see `THIRD_PARTY_NOTICES.md` in the repo (Apache-2.0).
- PAGE polygonalization helpers derive from Kraken/BLLA (Apache-2.0) — used at export, not in the pretrain loss.
- Training corpora: respect each dataset's original license. Private packs are **not** redistributed with this card; only model weights trained with them are released where license allows.
---
## License
**Apache License 2.0** for the **regnetx-det** code and these weights. Downstream users must comply with licenses of any datasets used in further fine-tuning.