dhchoi commited on
Commit
576c621
·
verified ·
1 Parent(s): d9af022

Fold the LoRA adapter in as lora/, one repository for both forms

Browse files
README.md CHANGED
@@ -16,6 +16,7 @@ tags:
16
  - hanja
17
  - classical-chinese
18
  - custom_code
 
19
  metrics:
20
  - cer
21
  ---
@@ -27,14 +28,17 @@ metrics:
27
  Bandit OCR reads a page of a Korean historical document and returns one region per printed column,
28
  in reading order, with the characters of that column. It is a LoRA fine-tune of
29
  [dots.mocr](https://huggingface.co/dots-studio/dots.mocr), a 1.7B-parameter document vision-language model,
30
- trained on 229,356 annotated pages of Joseon-dynasty woodblock prints and manuscripts. The
31
- adapter is merged into the weights, so it loads and serves exactly like the base.
32
 
33
  On the held-out test split it reads a page at 0.0195 character error rate with a column
34
  F1 of 0.9856, against 0.1436 for the AI Hub ResNet pipeline and
35
  0.2398 for NDLkotenOCR on the same pages.
36
 
37
- - LoRA adapter: [dhchoi/bandit-ocr-lora](https://huggingface.co/dhchoi/bandit-ocr-lora)
 
 
 
 
38
  - Desktop application and project site: https://bandit.dhchoi.net
39
  - Paper: JADH 2026, September 2026
40
 
@@ -45,8 +49,9 @@ F1 of 0.9856, against 0.1436 for the AI Hub ResNet pipeline and
45
  | Base model | [dots-studio/dots.mocr](https://huggingface.co/dots-studio/dots.mocr) at revision `e539fbb52280393adc081b289ec597430a0f9031` |
46
  | Parameters | 1.7B (1.2B language decoder + 0.4B vision tower) |
47
  | Adaptation | LoRA, rank 64, alpha 128, dropout 0.05, on the language projections and the vision tower |
48
- | Weights | the adapter folded in at scale 0.75; a plain bfloat16 checkpoint, no PEFT needed |
49
- | Files | 2 safetensors shards, about 6.1 GB |
 
50
  | Checkpoint | t3_full step 56000 of the full training run |
51
  | Trained on | 229,356 annotated pages of Korean historical documents |
52
  | Input | one page image, 3,136 to 11,289,600 pixels after `smart_resize` |
@@ -168,6 +173,40 @@ Send the image as a data URL with the prompt above, `temperature` 0, `max_tokens
168
  decode settings of the section below: `"stop_token_ids": [151673]` and
169
  `"logit_bias": {"151643": -100, "151672": -100}`.
170
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
171
  ## Decoding: stop on one token, suppress two
172
 
173
  This matters more than any other setting here. The training target ends with `<|endofassistant|>`
@@ -175,7 +214,9 @@ This matters more than any other setting here. The training target ends with `<|
175
  `<|assistant|>` (151672) as stop ids. The fine-tune sometimes emits `<|endoftext|>` in the middle of
176
  a column at a hard glyph, and a page that stops there is truncated or fails to parse.
177
 
178
- - `generation_config.json` in this repository already sets `eos_token_id` to 151673 alone.
 
 
179
  - Suppress 151643 and 151672 during decoding as well (`suppress_tokens` in `generate`, a -100
180
  `logit_bias` through an OpenAI-compatible server).
181
 
@@ -314,12 +355,14 @@ back to eager attention, which is quadratic in the number of patches; set
314
 
315
  Built with dots.mocr.
316
 
317
- This model is a fine-tune of [dots.mocr](https://huggingface.co/dots-studio/dots.mocr) by Xingyin
318
- Information Technology (Shanghai) Co., Ltd. (rednote / Xiaohongshu), used under the dots.mocr
319
- LICENSE AGREEMENT. A copy of that agreement and the base `NOTICE` ship in this repository (clauses
320
- 7.1 and 7.2), the notices in them are unchanged, and this card displays the attribution statement
321
- clause 7.3 requires. Nothing here implies endorsement by the licensor, and no name in this project
322
- is a licensor trade name (clause 4.4).
 
 
323
 
324
  The training data is licensed separately and is not redistributed; see the section above.
325
 
 
16
  - hanja
17
  - classical-chinese
18
  - custom_code
19
+ - lora
20
  metrics:
21
  - cer
22
  ---
 
28
  Bandit OCR reads a page of a Korean historical document and returns one region per printed column,
29
  in reading order, with the characters of that column. It is a LoRA fine-tune of
30
  [dots.mocr](https://huggingface.co/dots-studio/dots.mocr), a 1.7B-parameter document vision-language model,
31
+ trained on 229,356 annotated pages of Joseon-dynasty woodblock prints and manuscripts.
 
32
 
33
  On the held-out test split it reads a page at 0.0195 character error rate with a column
34
  F1 of 0.9856, against 0.1436 for the AI Hub ResNet pipeline and
35
  0.2398 for NDLkotenOCR on the same pages.
36
 
37
+ The repository holds both forms of the same fine-tune. At the root are the merged weights, the
38
+ adapter already folded in at scale 0.75, which load and serve exactly like the base and need
39
+ no PEFT. Under `lora/` is the LoRA adapter itself, for composing onto the base yourself,
40
+ changing the scale or training further. Take the root unless you know you want the adapter.
41
+
42
  - Desktop application and project site: https://bandit.dhchoi.net
43
  - Paper: JADH 2026, September 2026
44
 
 
49
  | Base model | [dots-studio/dots.mocr](https://huggingface.co/dots-studio/dots.mocr) at revision `e539fbb52280393adc081b289ec597430a0f9031` |
50
  | Parameters | 1.7B (1.2B language decoder + 0.4B vision tower) |
51
  | Adaptation | LoRA, rank 64, alpha 128, dropout 0.05, on the language projections and the vision tower |
52
+ | Weights | root: the adapter folded in at scale 0.75, a plain bfloat16 checkpoint, no PEFT needed |
53
+ | Adapter | `lora/`: the same LoRA unmerged, 580 MB, to apply at scale 0.75 yourself |
54
+ | Files | 2 safetensors shards at the root, about 6.1 GB, plus the adapter |
55
  | Checkpoint | t3_full step 56000 of the full training run |
56
  | Trained on | 229,356 annotated pages of Korean historical documents |
57
  | Input | one page image, 3,136 to 11,289,600 pixels after `smart_resize` |
 
173
  decode settings of the section below: `"stop_token_ids": [151673]` and
174
  `"logit_bias": {"151643": -100, "151672": -100}`.
175
 
176
+ ### The LoRA adapter
177
+
178
+ `lora/` holds the same fine-tune unmerged, for composing onto the base yourself, changing
179
+ the scale or training further. The delta is scaled by 0.75 before it is applied: that scale is
180
+ part of the released model, not a training artefact, and every number below was measured with it.
181
+
182
+ ```python
183
+ import torch
184
+ from peft import PeftModel
185
+ from transformers import AutoModelForCausalLM
186
+
187
+ base = AutoModelForCausalLM.from_pretrained(
188
+ "dots-studio/dots.mocr", revision="e539fbb52280393adc081b289ec597430a0f9031",
189
+ trust_remote_code=True, torch_dtype=torch.bfloat16, device_map="auto",
190
+ )
191
+ model = PeftModel.from_pretrained(base, "dhchoi/bandit-ocr", subfolder="lora")
192
+
193
+ for module in model.modules(): # every LoRA layer scales its delta by scaling[name]
194
+ scaling = getattr(module, "scaling", None)
195
+ if isinstance(scaling, dict):
196
+ for name in scaling:
197
+ scaling[name] *= 0.75
198
+
199
+ model = model.merge_and_unload()
200
+ ```
201
+
202
+ That reproduces the weights at the root of this repository. Prompt, decode and rescale boxes the
203
+ same way, and set the stop id yourself, because the composed model inherits the base's generation
204
+ config.
205
+
206
+ Applied at its trained scale of 1.0 the adapter is a different, measurably worse model: on the same
207
+ 540-page validation set an earlier checkpoint of this run read at 0.0656 page CER at scale 1.0
208
+ against 0.0376 at scale 0.75, both under the stop rule of the day.
209
+
210
  ## Decoding: stop on one token, suppress two
211
 
212
  This matters more than any other setting here. The training target ends with `<|endofassistant|>`
 
214
  `<|assistant|>` (151672) as stop ids. The fine-tune sometimes emits `<|endoftext|>` in the middle of
215
  a column at a hard glyph, and a page that stops there is truncated or fails to parse.
216
 
217
+ - `generation_config.json` at the root of this repository already sets `eos_token_id` to 151673
218
+ alone. If you compose the adapter onto the base yourself, set it there: the base's config stops
219
+ on all three.
220
  - Suppress 151643 and 151672 during decoding as well (`suppress_tokens` in `generate`, a -100
221
  `logit_bias` through an OpenAI-compatible server).
222
 
 
355
 
356
  Built with dots.mocr.
357
 
358
+ This model is a fine-tune of [dots.mocr](https://huggingface.co/dots-studio/dots.mocr), copyright
359
+ Xingyin Information Technology (Shanghai) Co., Ltd, used under the dots.mocr LICENSE AGREEMENT. A
360
+ copy of that agreement and the base `NOTICE` ship in this repository (clauses 7.1 and 7.2), the
361
+ notices in them are unchanged, and this card displays the attribution statement clause 7.3
362
+ requires. Nothing here implies endorsement by the licensor, and no name in this project is a
363
+ licensor trade name (clause 4.4). The weights are released on the terms the base is released
364
+ under; where the agreement and the metadata tag differ, the agreement is the document that
365
+ ships.
366
 
367
  The training data is licensed separately and is not redistributed; see the section above.
368
 
lora/adapter_config.json ADDED
@@ -0,0 +1,47 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "alpha_pattern": {},
3
+ "auto_mapping": null,
4
+ "base_model_name_or_path": "/home/dhchoi/bandit/models/dots_mocr",
5
+ "bias": "none",
6
+ "corda_config": null,
7
+ "eva_config": null,
8
+ "exclude_modules": null,
9
+ "fan_in_fan_out": false,
10
+ "inference_mode": true,
11
+ "init_lora_weights": true,
12
+ "layer_replication": null,
13
+ "layers_pattern": null,
14
+ "layers_to_transform": null,
15
+ "loftq_config": {},
16
+ "lora_alpha": 128,
17
+ "lora_bias": false,
18
+ "lora_dropout": 0.05,
19
+ "megatron_config": null,
20
+ "megatron_core": "megatron.core",
21
+ "modules_to_save": null,
22
+ "peft_type": "LORA",
23
+ "qalora_group_size": 16,
24
+ "r": 64,
25
+ "rank_pattern": {},
26
+ "revision": null,
27
+ "target_modules": [
28
+ "attn.qkv",
29
+ "mlp.fc1",
30
+ "attn.proj",
31
+ "down_proj",
32
+ "mlp.fc3",
33
+ "up_proj",
34
+ "q_proj",
35
+ "o_proj",
36
+ "gate_proj",
37
+ "mlp.fc2",
38
+ "k_proj",
39
+ "v_proj"
40
+ ],
41
+ "target_parameters": null,
42
+ "task_type": "CAUSAL_LM",
43
+ "trainable_token_indices": null,
44
+ "use_dora": false,
45
+ "use_qalora": false,
46
+ "use_rslora": false
47
+ }
lora/adapter_model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:9210f5f8b2a5ea50e5af6d83682cead973a3480817532bf490a9b9db0e3cb526
3
+ size 580430776