File size: 4,717 Bytes
3efb1b0
 
 
fbce730
 
 
 
 
 
3efb1b0
fbce730
3efb1b0
 
fbce730
3efb1b0
fbce730
 
 
3efb1b0
fbce730
 
 
 
 
 
 
 
 
3efb1b0
fbce730
 
 
 
 
 
 
 
 
3efb1b0
fbce730
3efb1b0
 
fbce730
 
 
 
74d2991
fbce730
 
3efb1b0
 
fbce730
74d2991
 
 
 
fbce730
 
 
 
 
 
 
74d2991
 
 
 
fbce730
 
 
74d2991
 
fbce730
74d2991
fbce730
74d2991
 
 
fbce730
74d2991
 
 
 
 
 
fbce730
 
 
 
 
74d2991
 
 
fbce730
74d2991
 
 
fbce730
3efb1b0
 
 
 
fbce730
 
 
 
 
 
3efb1b0
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
---
license: apache-2.0
tags:
  - object-detection
  - document-layout-analysis
  - tibetan
  - rf-detr
  - tibla
pipeline_tag: object-detection
datasets:
  - BDRC/TiBLAD
---

# TiBLA-RFDETR

**Permissive (Apache-2.0) alternative in TiBLA (Tibetan Book Layout Analysis)** —
a lighter, PyTorch-native RF-DETR-L detector for the page layout of modern
Tibetan books (headers, text area, footers, footnotes).

- **Base model / provenance:** [RF-DETR-L](https://github.com/roboflow/rf-detr)
  (Roboflow, DINOv2 backbone), fine-tuned on the leak-free **v4** `tam2col` split
  of TiBLAD.
- **License:** Apache-2.0.
- **Dataset:** [BDRC/TiBLAD](https://huggingface.co/datasets/BDRC/TiBLAD)
- **Paper:** [buda-base/papers](https://github.com/buda-base/papers) (`papers/2026-tibetan-book-layout`) — *arXiv link forthcoming*
- **Code:** [github.com/buda-base/tibla](https://github.com/buda-base/tibla)

## Task

A **4-class** detector — `header`, `text-area`, `footer`, `footnote` — kept as
four classes at training time. Evaluation folds them into a **3-class canonical
scheme**: `header`+`footer` are combined into one `header-footer` class (matched
individually, merged losslessly afterwards), `text-area` is merged to a single
page/column envelope as a post-processing step (two boxes only on genuine
two-column pages), and `footnote` is left as-is. All numbers below are in that
canonical space, on the leak-free TiBLAD **v4** 833-page test set, unified scorer
(pycocotools bbox mAP 0.50:0.05:0.95; F1 by greedy IoU≥0.5 at the best-mean-F1
operating point).

## Inference

```python
# pip install rfdetr
from rfdetr import RFDETRLarge

model = RFDETRLarge.from_checkpoint("rfdetr_tibetan_book_layout.pth")
det = model.predict("page.jpg", threshold=0.47, shape=(1024, 1024))
# checkpoint class ids are offset by 1 (id 0 = background):
#   1 header, 2 text-area, 3 footnote, 4 footer
```

A ready-made `infer.py` (batch, YOLO-format output, per-class thresholds) is
included in this repo. Recommended global operating confidence: **0.47** (the
validation-selected best-mean-F1 point); the bundled `infer.py` also ships
per-class max-F1 thresholds (`header` 0.46, `text-area` 0.32, `footnote` 0.26,
`footer` 0.52).

## Evaluation (TiBLAD v4, 833-page test)

| metric | TiBLA-RTDETR | TiBLA-PP-DocLayout-L | **TiBLA-RFDETR** |
|---|---|---|---|
| license | AGPL-3.0 | Apache-2.0 | **Apache-2.0** |
| base model | RT-DETR-l (Ultralytics) | PP-DocLayout-L (PaddleOCR, RT-DETR-L) | **RF-DETR-L (Roboflow)** |
| mean F1 (canonical 3-class) | 0.952 | 0.955 | **0.921** |
|   header-footer F1 | 0.954 | 0.953 | **0.947** |
|   text-area F1 | 0.999 | 0.998 | **0.996** |
|   footnote F1 | 0.902 | 0.914 | **0.821** |
| mean AP@0.50 | 0.974 | 0.959 | **0.925** |
| mean AP@[0.50:0.95] | 0.786 | 0.781 | **0.667** |
| shared-class mAP@[.50:.95] (DocLayNet-aligned) | 0.650 | 0.641 | **0.604** |
| Hidden Trespass — header/footer | 0.009 | 0.004 | **0.021** |
| Hidden Trespass — footnote | 0.043 | 0.037 | **0.178** |
| COTe (Trespass) | 0.975 (0.001) | 0.978 (0.000) | **0.974 (0.002)** |
| operating confidence | 0.64 | 0.61 | **0.47** |

*"operating confidence" is the single global best-mean-F1 confidence, selected on
the leak-free validation split and frozen for test (no test-set tuning). COCO AP
rows are threshold-free (all detections above the fixed 0.05 floor).*

**Hidden Trespass** = peripheral (header/footer/footnote) ground-truth **area**
that survives in the *actual OCR body crop* `C = E \ P`, where `E` is the predicted
`text-area` envelope and `P` is the union of the predicted peripheral boxes the
pipeline subtracts; area-based, micro-averaged over the test set. Lower is better
(less peripheral text bled into the OCR region). Formal definition in the
[paper](https://github.com/buda-base/papers).

## Which checkpoint to pick

| checkpoint | license | mean F1 | shared mAP | footnote HT |
|---|---|---|---|---|
| TiBLA-RTDETR (primary) | AGPL-3.0 | 0.952 | 0.650 | 0.043 |
| TiBLA-PP-DocLayout-L | Apache-2.0 | 0.955 | 0.641 | 0.037 |
| **TiBLA-RFDETR** | Apache-2.0 | 0.921 | 0.604 | 0.178 |

RT-DETR-l leads on mAP, shared-class mAP and the 5-seed mean F1 (0.961 ± 0.009),
but its weights are AGPL-3.0 (Ultralytics). If you need a permissive license,
PP-DocLayout-L is an Apache-2.0 match (statistically on par on F1); RF-DETR is a
lighter PyTorch-native Apache-2.0 option.

## Citation

```bibtex
@misc{tibla2026,
  title        = {TiBLA: Tibetan Book Layout Analysis},
  author       = {Buddhist Digital Resource Center (BDRC)},
  year         = {2026},
  howpublished = {\url{https://github.com/buda-base/tibla}},
  note         = {arXiv link forthcoming}
}
```