Image-Text-to-Text
PEFT
Safetensors
English
lora
qlora
4-bit precision
adapter
medical
radiology
image-captioning
image-to-text
llava-onevision
qwen2
vision-language
imageclef
rocov2
conversational
Instructions to use HoqueMahmudul/llava-onevision-7b-qlora-radiology-image-caption with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use HoqueMahmudul/llava-onevision-7b-qlora-radiology-image-caption with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("llava-hf/llava-onevision-qwen2-7b-ov-hf") model = PeftModel.from_pretrained(base_model, "HoqueMahmudul/llava-onevision-7b-qlora-radiology-image-caption") - Notebooks
- Google Colab
- Kaggle
Clarify latency comparability: base quantization must match, not just GPU
Browse files
README.md
CHANGED
|
@@ -186,9 +186,12 @@ device** holding a complete model — the model is replicated, not sharded.
|
|
| 186 |
|---|---|---|---|---|---|
|
| 187 |
| this variant (7B) | NVIDIA H200 (gpuH200x8) | ~16 GB | 4-bit | **23.70 GB** | **1.63 s/image** |
|
| 188 |
|
| 189 |
-
> **
|
| 190 |
-
> A100, 7B and 72B on H200), so
|
| 191 |
-
> *
|
|
|
|
|
|
|
|
|
|
| 192 |
|
| 193 |
> **Why peak memory exceeds the weight size.** LLaVA-OneVision uses anyres tiling:
|
| 194 |
> a large image expands into thousands of visual tokens (1024x768 -> ~5,100 tokens;
|
|
|
|
| 186 |
|---|---|---|---|---|---|
|
| 187 |
| this variant (7B) | NVIDIA H200 (gpuH200x8) | ~16 GB | 4-bit | **23.70 GB** | **1.63 s/image** |
|
| 188 |
|
| 189 |
+
> **Comparing these numbers across variants requires care.** Scales were measured
|
| 190 |
+
> on different GPUs (0.5B on A100, 7B and 72B on H200), so **latency is not
|
| 191 |
+
> comparable across scales**. Within a scale it is comparable only when the *base
|
| 192 |
+
> loading* also matches: QLoRA vs QDoRA is a fair comparison (same GPU, both
|
| 193 |
+
> 4-bit), but LoRA vs QLoRA is **not** — those differ in base quantization as well
|
| 194 |
+
> as PEFT method, so the gap conflates the two.
|
| 195 |
|
| 196 |
> **Why peak memory exceeds the weight size.** LLaVA-OneVision uses anyres tiling:
|
| 197 |
> a large image expands into thousands of visual tokens (1024x768 -> ~5,100 tokens;
|