donghyun95 commited on
Commit
dfe0dc7
·
verified ·
1 Parent(s): b718043

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +290 -0
README.md CHANGED
@@ -1,2 +1,292 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
 
2
 
 
1
+ ---
2
+ license: mit
3
+ language:
4
+ - zh
5
+ - ko
6
+ pipeline_tag: image-to-text
7
+ library_name: pytorch
8
+ tags:
9
+ - ocr
10
+ - document-ai
11
+ - computer-vision
12
+ - hanja
13
+ - takbon
14
+ - rubbing
15
+ - resnet
16
+ - hrcenternet
17
+ - google-vision-ocr
18
+ ---
19
+
20
+ # EpiText Hanja OCR (Takbon OCR)
21
+
22
+ <p align="center">
23
+ 🔧 <a href="#google-vision-api-setup-required">Setup</a> &nbsp;|&nbsp;
24
+ ▶️ <a href="#running-the-ocr">Run</a> &nbsp;|&nbsp;
25
+ 🖼️ <a href="#preprocessing-and-intermediate-outputs">Examples</a> &nbsp;|&nbsp;
26
+ 📦 <a href="#final-outputs">Outputs</a>
27
+ </p>
28
+
29
+ **Pipeline:** Input → Preprocess (gray for Swin, binary for OCR) → OCR (auto) → JSON + BBox
30
+
31
+
32
+ This repository provides a **damage-aware OCR pipeline specialized for Hanja rubbing (탁본) images**.
33
+ The system integrates **Google Vision OCR** with **custom deep learning models** to robustly recognize characters under severe degradation commonly found in stone inscriptions and epigraphic materials.
34
+
35
+ ---
36
+
37
+ ## Table of Contents
38
+ - [Overview](#overview)
39
+ - [Requirements](#requirements)
40
+ - [Google Vision API Setup](#google-vision-api-setup-required)
41
+ - [Running the OCR](#running-the-ocr)
42
+ - [Preprocessing and Intermediate Outputs](#preprocessing-and-intermediate-outputs)
43
+ - [Final Outputs](#final-outputs)
44
+ - [Why Specialized for Takbon](#why-this-ocr-is-specialized-for-rubbing-takbon-images)
45
+ - [License](#license)
46
+ - [Citation](#citation)
47
+
48
+ ---
49
+
50
+ ## Overview
51
+
52
+ Hanja rubbing images differ significantly from modern scanned documents.
53
+ They often exhibit erosion, ink bleeding, uneven backgrounds, and partially or fully missing characters.
54
+
55
+ To address these challenges, this project combines:
56
+
57
+ - Custom OCR models optimized for degraded inscription images
58
+ - Explicit modeling of character damage
59
+ - Layout-aware processing for vertical writing
60
+ - Auxiliary use of Google Vision OCR for complementary recognition
61
+
62
+ > ⚠️ **Google Vision OCR is not redistributed.**
63
+ > Users must provide their own Google Cloud API credentials.
64
+
65
+ ---
66
+
67
+ ## Features
68
+
69
+ - OCR ensemble: Google Vision OCR + custom OCR models
70
+ - Damage-aware character tokens: `[MASK1]`, `[MASK2]`
71
+ - Column-wise output for vertically written inscriptions
72
+ - Structured JSON OCR output
73
+ - Bounding box visualization for inspection
74
+ - Fully automated preprocessing → OCR pipeline
75
+
76
+ ---
77
+
78
+ ## Repository Structure
79
+
80
+ ```text
81
+ EpiText-Hanja-OCR/
82
+ ├─ assets/ # README example images
83
+ ├─ dong_ocr.py # Main execution script
84
+ ├─ ai_modules/ # OCR engine, preprocessing, model definitions
85
+ ├─ weights/ # Model weights (and user-provided API key)
86
+ ├─ requirements.txt
87
+ └─ README.md
88
+ ```
89
+
90
+ ---
91
+
92
+ ## Requirements
93
+
94
+ - Python 3.9+
95
+ - PyTorch
96
+ - Google Cloud Vision API credentials
97
+
98
+ Install dependencies:
99
+
100
+ ```bash
101
+ pip install -r requirements.txt
102
+ ```
103
+
104
+ ---
105
+
106
+ ## Google Vision API Setup (Required)
107
+
108
+ This project requires a **Google Vision API service account JSON file**.
109
+
110
+ ### Step 1. Create Google Cloud credentials
111
+
112
+ 1. Go to **Google Cloud Console**
113
+ 2. Create or select a project
114
+ 3. Enable **Cloud Vision API**
115
+ 4. Create a **Service Account**
116
+ 5. Generate and download a **JSON key file**
117
+
118
+ ---
119
+
120
+ ### Step 2. Place the JSON file in the `weights/` directory
121
+
122
+ ```text
123
+ weights/
124
+ ├─ best.pth
125
+ ├─ best_5000.pt
126
+ └─ google_key.json
127
+ ```
128
+
129
+ ⚠️ **Do NOT upload this JSON file to GitHub or Hugging Face.**
130
+ It must remain local to your machine.
131
+
132
+ ---
133
+
134
+ ### Step 3. Set environment variables
135
+
136
+ #### Linux / macOS
137
+
138
+ ```bash
139
+ export OCR_WEIGHTS_BASE_PATH=./weights
140
+ export GOOGLE_CREDENTIALS_JSON=google_key.json
141
+ ```
142
+
143
+ #### Windows (PowerShell)
144
+
145
+ ```powershell
146
+ $env:OCR_WEIGHTS_BASE_PATH=".\weights"
147
+ $env:GOOGLE_CREDENTIALS_JSON="google_key.json"
148
+ ```
149
+
150
+ ---
151
+
152
+ ## Running the OCR
153
+
154
+ ```bash
155
+ python dong_ocr.py path/to/image.jpg
156
+ ```
157
+
158
+ Example:
159
+
160
+ ```bash
161
+ python dong_ocr.py assets/input.jpg
162
+ ```
163
+
164
+ ---
165
+
166
+ ## Preprocessing and Intermediate Outputs
167
+
168
+ Before OCR inference, the input image is automatically preprocessed to generate
169
+ task-specific intermediate representations.
170
+
171
+ ### Preprocessing Examples
172
+
173
+ <table>
174
+ <tr>
175
+ <td align="center">
176
+ <strong>Input Image<br>(Rubbing / Takbon)</strong><br>
177
+ <img src="assets/input.jpg" width="250">
178
+ </td>
179
+ <td align="center">
180
+ <strong>Grayscale Image<br>(Swin Input)</strong><br>
181
+ <img src="assets/gray.jpg" width="250">
182
+ </td>
183
+ <td align="center">
184
+ <strong>Binarized Image<br>(OCR Input)</strong><br>
185
+ <img src="assets/binary.png" width="250">
186
+ </td>
187
+ </tr>
188
+ </table>
189
+
190
+ ---
191
+
192
+ ## Automatic OCR Pipeline Integration
193
+
194
+ The binarized (black-and-white) image is **automatically forwarded to the OCR engine**.
195
+
196
+ - No manual image selection is required
197
+ - OCR always consumes the internally generated binary image
198
+ - Preprocessing and OCR are fully coupled for reproducibility
199
+
200
+ > **Input image → preprocessing → binary image → OCR (automatic)**
201
+
202
+ ---
203
+
204
+ ## Final Outputs
205
+
206
+ - `*_gray.jpg`
207
+ Grayscale image used for Swin Transformer–based processing
208
+
209
+ - `*_binary.jpg`
210
+ Binarized image automatically used for OCR inference
211
+
212
+ - `*_ocr_result.json`
213
+ Structured OCR results including:
214
+ - bounding boxes
215
+ - recognized text
216
+ - damage type (`TEXT`, `MASK1`, `MASK2`)
217
+
218
+ - `*_bbox.jpg`
219
+ Visualization image with colored bounding boxes
220
+
221
+ ### Bounding Box Visualization Example
222
+
223
+ - **Green**: Google Vision OCR
224
+ - **Purple**: Custom OCR (HRCenterNet-based)
225
+ - **Blue**: `[MASK1]` (fully missing characters)
226
+ - **Red**: `[MASK2]` (partially damaged characters)
227
+
228
+ <img src="assets/bbox.jpg" width="50%">
229
+
230
+ ---
231
+
232
+ ## Why This OCR Is Specialized for Rubbing (Takbon) Images
233
+
234
+ - Ink bleeding and stone texture noise
235
+ - Partial or complete stroke erosion
236
+ - Non-uniform contrast
237
+ - Vertically arranged, tightly packed characters
238
+
239
+ ### Design Choices
240
+
241
+ #### 1. Dual Image Representation
242
+ - Grayscale for detection
243
+ - Binarized for OCR
244
+
245
+ #### 2. Damage-Aware Modeling
246
+ - `[MASK1]`: fully missing
247
+ - `[MASK2]`: partially damaged
248
+
249
+ #### 3. Layout Preservation
250
+ - Column-wise processing
251
+ - Correct reading order reconstruction
252
+
253
+ #### 4. Auxiliary Google Vision OCR
254
+ - Used as a complementary OCR engine
255
+ - Requires user-provided credentials
256
+
257
+ ---
258
+
259
+ ## Model Architecture
260
+
261
+ - Text Detection: HRCenterNet-based detector
262
+ - Text Recognition: ResNet-based recognizer
263
+ - Auxiliary OCR: Google Vision OCR
264
+
265
+ ---
266
+
267
+ ## Limitations
268
+
269
+ - Requires external Google Vision API credentials
270
+ - Performance may degrade under extreme blur
271
+ - Not intended as an end-to-end HF inference widget
272
+
273
+ ---
274
+
275
+ ## License
276
+
277
+ MIT License
278
+
279
+ ---
280
+
281
+ ## Citation
282
+
283
+ ```bibtex
284
+ @misc{epitext_hanja_ocr_2025,
285
+ title = {EpiText Hanja OCR: Damage-Aware OCR for Rubbing Images},
286
+ author = {donghyun95},
287
+ year = {2025},
288
+ howpublished = {Hugging Face Model Repository}
289
+ }
290
+ ```
291
 
292