File size: 4,717 Bytes
3efb1b0 fbce730 3efb1b0 fbce730 3efb1b0 fbce730 3efb1b0 fbce730 3efb1b0 fbce730 3efb1b0 fbce730 3efb1b0 fbce730 3efb1b0 fbce730 74d2991 fbce730 3efb1b0 fbce730 74d2991 fbce730 74d2991 fbce730 74d2991 fbce730 74d2991 fbce730 74d2991 fbce730 74d2991 fbce730 74d2991 fbce730 74d2991 fbce730 3efb1b0 fbce730 3efb1b0 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 | ---
license: apache-2.0
tags:
- object-detection
- document-layout-analysis
- tibetan
- rf-detr
- tibla
pipeline_tag: object-detection
datasets:
- BDRC/TiBLAD
---
# TiBLA-RFDETR
**Permissive (Apache-2.0) alternative in TiBLA (Tibetan Book Layout Analysis)** —
a lighter, PyTorch-native RF-DETR-L detector for the page layout of modern
Tibetan books (headers, text area, footers, footnotes).
- **Base model / provenance:** [RF-DETR-L](https://github.com/roboflow/rf-detr)
(Roboflow, DINOv2 backbone), fine-tuned on the leak-free **v4** `tam2col` split
of TiBLAD.
- **License:** Apache-2.0.
- **Dataset:** [BDRC/TiBLAD](https://huggingface.co/datasets/BDRC/TiBLAD)
- **Paper:** [buda-base/papers](https://github.com/buda-base/papers) (`papers/2026-tibetan-book-layout`) — *arXiv link forthcoming*
- **Code:** [github.com/buda-base/tibla](https://github.com/buda-base/tibla)
## Task
A **4-class** detector — `header`, `text-area`, `footer`, `footnote` — kept as
four classes at training time. Evaluation folds them into a **3-class canonical
scheme**: `header`+`footer` are combined into one `header-footer` class (matched
individually, merged losslessly afterwards), `text-area` is merged to a single
page/column envelope as a post-processing step (two boxes only on genuine
two-column pages), and `footnote` is left as-is. All numbers below are in that
canonical space, on the leak-free TiBLAD **v4** 833-page test set, unified scorer
(pycocotools bbox mAP 0.50:0.05:0.95; F1 by greedy IoU≥0.5 at the best-mean-F1
operating point).
## Inference
```python
# pip install rfdetr
from rfdetr import RFDETRLarge
model = RFDETRLarge.from_checkpoint("rfdetr_tibetan_book_layout.pth")
det = model.predict("page.jpg", threshold=0.47, shape=(1024, 1024))
# checkpoint class ids are offset by 1 (id 0 = background):
# 1 header, 2 text-area, 3 footnote, 4 footer
```
A ready-made `infer.py` (batch, YOLO-format output, per-class thresholds) is
included in this repo. Recommended global operating confidence: **0.47** (the
validation-selected best-mean-F1 point); the bundled `infer.py` also ships
per-class max-F1 thresholds (`header` 0.46, `text-area` 0.32, `footnote` 0.26,
`footer` 0.52).
## Evaluation (TiBLAD v4, 833-page test)
| metric | TiBLA-RTDETR | TiBLA-PP-DocLayout-L | **TiBLA-RFDETR** |
|---|---|---|---|
| license | AGPL-3.0 | Apache-2.0 | **Apache-2.0** |
| base model | RT-DETR-l (Ultralytics) | PP-DocLayout-L (PaddleOCR, RT-DETR-L) | **RF-DETR-L (Roboflow)** |
| mean F1 (canonical 3-class) | 0.952 | 0.955 | **0.921** |
| header-footer F1 | 0.954 | 0.953 | **0.947** |
| text-area F1 | 0.999 | 0.998 | **0.996** |
| footnote F1 | 0.902 | 0.914 | **0.821** |
| mean AP@0.50 | 0.974 | 0.959 | **0.925** |
| mean AP@[0.50:0.95] | 0.786 | 0.781 | **0.667** |
| shared-class mAP@[.50:.95] (DocLayNet-aligned) | 0.650 | 0.641 | **0.604** |
| Hidden Trespass — header/footer | 0.009 | 0.004 | **0.021** |
| Hidden Trespass — footnote | 0.043 | 0.037 | **0.178** |
| COTe (Trespass) | 0.975 (0.001) | 0.978 (0.000) | **0.974 (0.002)** |
| operating confidence | 0.64 | 0.61 | **0.47** |
*"operating confidence" is the single global best-mean-F1 confidence, selected on
the leak-free validation split and frozen for test (no test-set tuning). COCO AP
rows are threshold-free (all detections above the fixed 0.05 floor).*
**Hidden Trespass** = peripheral (header/footer/footnote) ground-truth **area**
that survives in the *actual OCR body crop* `C = E \ P`, where `E` is the predicted
`text-area` envelope and `P` is the union of the predicted peripheral boxes the
pipeline subtracts; area-based, micro-averaged over the test set. Lower is better
(less peripheral text bled into the OCR region). Formal definition in the
[paper](https://github.com/buda-base/papers).
## Which checkpoint to pick
| checkpoint | license | mean F1 | shared mAP | footnote HT |
|---|---|---|---|---|
| TiBLA-RTDETR (primary) | AGPL-3.0 | 0.952 | 0.650 | 0.043 |
| TiBLA-PP-DocLayout-L | Apache-2.0 | 0.955 | 0.641 | 0.037 |
| **TiBLA-RFDETR** | Apache-2.0 | 0.921 | 0.604 | 0.178 |
RT-DETR-l leads on mAP, shared-class mAP and the 5-seed mean F1 (0.961 ± 0.009),
but its weights are AGPL-3.0 (Ultralytics). If you need a permissive license,
PP-DocLayout-L is an Apache-2.0 match (statistically on par on F1); RF-DETR is a
lighter PyTorch-native Apache-2.0 option.
## Citation
```bibtex
@misc{tibla2026,
title = {TiBLA: Tibetan Book Layout Analysis},
author = {Buddhist Digital Resource Center (BDRC)},
year = {2026},
howpublished = {\url{https://github.com/buda-base/tibla}},
note = {arXiv link forthcoming}
}
```
|