TiBLA-RFDETR / README.md
Eroux's picture
Update eval numbers to post-crop Hidden Trespass at validation-selected operating point
74d2991 verified
|
Raw
History Blame Contribute Delete
4.72 kB
metadata
license: apache-2.0
tags:
  - object-detection
  - document-layout-analysis
  - tibetan
  - rf-detr
  - tibla
pipeline_tag: object-detection
datasets:
  - BDRC/TiBLAD

TiBLA-RFDETR

Permissive (Apache-2.0) alternative in TiBLA (Tibetan Book Layout Analysis) — a lighter, PyTorch-native RF-DETR-L detector for the page layout of modern Tibetan books (headers, text area, footers, footnotes).

Task

A 4-class detector — header, text-area, footer, footnote — kept as four classes at training time. Evaluation folds them into a 3-class canonical scheme: header+footer are combined into one header-footer class (matched individually, merged losslessly afterwards), text-area is merged to a single page/column envelope as a post-processing step (two boxes only on genuine two-column pages), and footnote is left as-is. All numbers below are in that canonical space, on the leak-free TiBLAD v4 833-page test set, unified scorer (pycocotools bbox mAP 0.50:0.05:0.95; F1 by greedy IoU≥0.5 at the best-mean-F1 operating point).

Inference

# pip install rfdetr
from rfdetr import RFDETRLarge

model = RFDETRLarge.from_checkpoint("rfdetr_tibetan_book_layout.pth")
det = model.predict("page.jpg", threshold=0.47, shape=(1024, 1024))
# checkpoint class ids are offset by 1 (id 0 = background):
#   1 header, 2 text-area, 3 footnote, 4 footer

A ready-made infer.py (batch, YOLO-format output, per-class thresholds) is included in this repo. Recommended global operating confidence: 0.47 (the validation-selected best-mean-F1 point); the bundled infer.py also ships per-class max-F1 thresholds (header 0.46, text-area 0.32, footnote 0.26, footer 0.52).

Evaluation (TiBLAD v4, 833-page test)

metric TiBLA-RTDETR TiBLA-PP-DocLayout-L TiBLA-RFDETR
license AGPL-3.0 Apache-2.0 Apache-2.0
base model RT-DETR-l (Ultralytics) PP-DocLayout-L (PaddleOCR, RT-DETR-L) RF-DETR-L (Roboflow)
mean F1 (canonical 3-class) 0.952 0.955 0.921
  header-footer F1 0.954 0.953 0.947
  text-area F1 0.999 0.998 0.996
  footnote F1 0.902 0.914 0.821
mean AP@0.50 0.974 0.959 0.925
mean AP@[0.50:0.95] 0.786 0.781 0.667
shared-class mAP@[.50:.95] (DocLayNet-aligned) 0.650 0.641 0.604
Hidden Trespass — header/footer 0.009 0.004 0.021
Hidden Trespass — footnote 0.043 0.037 0.178
COTe (Trespass) 0.975 (0.001) 0.978 (0.000) 0.974 (0.002)
operating confidence 0.64 0.61 0.47

"operating confidence" is the single global best-mean-F1 confidence, selected on the leak-free validation split and frozen for test (no test-set tuning). COCO AP rows are threshold-free (all detections above the fixed 0.05 floor).

Hidden Trespass = peripheral (header/footer/footnote) ground-truth area that survives in the actual OCR body crop C = E \ P, where E is the predicted text-area envelope and P is the union of the predicted peripheral boxes the pipeline subtracts; area-based, micro-averaged over the test set. Lower is better (less peripheral text bled into the OCR region). Formal definition in the paper.

Which checkpoint to pick

checkpoint license mean F1 shared mAP footnote HT
TiBLA-RTDETR (primary) AGPL-3.0 0.952 0.650 0.043
TiBLA-PP-DocLayout-L Apache-2.0 0.955 0.641 0.037
TiBLA-RFDETR Apache-2.0 0.921 0.604 0.178

RT-DETR-l leads on mAP, shared-class mAP and the 5-seed mean F1 (0.961 ± 0.009), but its weights are AGPL-3.0 (Ultralytics). If you need a permissive license, PP-DocLayout-L is an Apache-2.0 match (statistically on par on F1); RF-DETR is a lighter PyTorch-native Apache-2.0 option.

Citation

@misc{tibla2026,
  title        = {TiBLA: Tibetan Book Layout Analysis},
  author       = {Buddhist Digital Resource Center (BDRC)},
  year         = {2026},
  howpublished = {\url{https://github.com/buda-base/tibla}},
  note         = {arXiv link forthcoming}
}