HoqueMahmudul commited on
Commit
ee7ccd9
·
verified ·
1 Parent(s): a04049b

Document DoRA inference overhead vs QLoRA across scales

Browse files
Files changed (1) hide show
  1. README.md +19 -0
README.md CHANGED
@@ -229,6 +229,25 @@ Hot-swapping is safe for these adapters, in either load order.
229
  > (`peft` 0.18.1, `transformers` 5.3.0, `torch` 2.11.0+cu128, `bitsandbytes` 0.49.2)
230
 
231
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
232
  ## Continuing fine-tuning from this adapter
233
 
234
  You can keep training these weights on your own data. Two options:
 
229
  > (`peft` 0.18.1, `transformers` 5.3.0, `torch` 2.11.0+cu128, `bitsandbytes` 0.49.2)
230
 
231
 
232
+ ## Inference cost of DoRA
233
+
234
+ DoRA recomputes weight column norms on **every forward pass through every linear
235
+ adapter module**, which makes it slower at inference than the equivalent QLoRA
236
+ adapter. At 72B — 8 target modules across 80 decoder layers — this overhead
237
+ dominates inference latency.
238
+
239
+ | Scale | Hardware | QLoRA | QDoRA | Overhead |
240
+ |---|---|---|---|---|
241
+ | 0.5B | A100 | 2.49 s/image | 4.73 s/image | **1.9x** |
242
+ | 7B | H200 | 1.63 s/image | 5.63 s/image | **3.4x** |
243
+ | 72B | A100 x8 | 12.5 s/image | 71.6 s/image | **5.7x** |
244
+
245
+ The penalty grows with decoder depth and the number of adapted modules.
246
+
247
+ **If inference latency matters to you, prefer the QLoRA adapter at the same
248
+ scale.** The two were trained with an identical recipe and differ only in the
249
+ adapter formulation.
250
+
251
  ## Continuing fine-tuning from this adapter
252
 
253
  You can keep training these weights on your own data. Two options: