HoqueMahmudul commited on
Commit
6facdbc
·
verified ·
1 Parent(s): b436505

Clarify latency comparability: base quantization must match, not just GPU

Browse files
Files changed (1) hide show
  1. README.md +6 -3
README.md CHANGED
@@ -186,9 +186,12 @@ device** holding a complete model — the model is replicated, not sharded.
186
  |---|---|---|---|---|---|
187
  | this variant (7B) | NVIDIA H200 (gpuH200x8) | ~16 GB | 4-bit | **23.70 GB** | **1.63 s/image** |
188
 
189
- > **Different variants in this series were measured on different GPUs** (0.5B on
190
- > A100, 7B and 72B on H200), so do not compare latency across scales. Comparisons
191
- > *within* a scale are valid — those pairs ran on the same partition.
 
 
 
192
 
193
  > **Why peak memory exceeds the weight size.** LLaVA-OneVision uses anyres tiling:
194
  > a large image expands into thousands of visual tokens (1024x768 -> ~5,100 tokens;
 
186
  |---|---|---|---|---|---|
187
  | this variant (7B) | NVIDIA H200 (gpuH200x8) | ~16 GB | 4-bit | **23.70 GB** | **1.63 s/image** |
188
 
189
+ > **Comparing these numbers across variants requires care.** Scales were measured
190
+ > on different GPUs (0.5B on A100, 7B and 72B on H200), so **latency is not
191
+ > comparable across scales**. Within a scale it is comparable only when the *base
192
+ > loading* also matches: QLoRA vs QDoRA is a fair comparison (same GPU, both
193
+ > 4-bit), but LoRA vs QLoRA is **not** — those differ in base quantization as well
194
+ > as PEFT method, so the gap conflates the two.
195
 
196
  > **Why peak memory exceeds the weight size.** LLaVA-OneVision uses anyres tiling:
197
  > a large image expands into thousands of visual tokens (1024x768 -> ~5,100 tokens;