Instructions to use BonanDing/UniMVU with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use BonanDing/UniMVU with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
File size: 5,625 Bytes
4ee4d25 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 | ---
license: other
library_name: peft
base_model:
- lmms-lab/llava-onevision-qwen2-0.5b-ov
- lmms-lab/llava-onevision-qwen2-7b-ov
pipeline_tag: image-text-to-text
tags:
- multimodal
- video
- audio
- 3d
- peft
- lora
- safetensors
- llava-onevision
- qwen2
language:
- en
---
# UniMVU - LoRA Adapters for LLaVA-OneVision Qwen2
Open-source UniMVU release checkpoints for instruction-aware multimodal video understanding. This release covers audio-video QA, 3D QA, and unified multi-task adapters built on top of `lmms-lab/llava-onevision-qwen2-0.5b-ov` and `lmms-lab/llava-onevision-qwen2-7b-ov`.
Unlike plain LoRA releases, UniMVU checkpoints also include `non_lora_trainables.bin` for the extra modality-gating modules. Use the UniMVU loader instead of a PEFT-only `PeftModel.from_pretrained(...)` workflow.
[Paper PDF](./UniMVU_CVPR_2026__Camera_Ready_.pdf)
## Highlights
- Instruction-aware gating across video, audio, depth, and long-video evidence.
- Single-task adapters for AVQA, AVSD, Music-AVQA, ScanQA, and SQA3D.
- Unified multi-task adapters for the mixed-training UniMVU release.
- Gains of up to +13.5 CIDEr on AVSD over the reproduced PAVE baseline, as reported in the paper.
## Release Contents
| Folder | Scale | Type | Task(s) | Base model | Published size |
| --- | --- | --- | --- | --- | --- |
| `unimvu_0.5B_avqa` | 0.5B | Single-task | AVQA | `lmms-lab/llava-onevision-qwen2-0.5b-ov` | 96.4 MB |
| `unimvu_0.5B_avsd` | 0.5B | Single-task | AVSD | `lmms-lab/llava-onevision-qwen2-0.5b-ov` | 96.4 MB |
| `unimvu_0.5B_music_avqa` | 0.5B | Single-task | Music-AVQA | `lmms-lab/llava-onevision-qwen2-0.5b-ov` | 96.4 MB |
| `unimvu_0.5B_scanqa` | 0.5B | Single-task | ScanQA | `lmms-lab/llava-onevision-qwen2-0.5b-ov` | 96.4 MB |
| `unimvu_0.5B_sqa3d` | 0.5B | Single-task | SQA3D | `lmms-lab/llava-onevision-qwen2-0.5b-ov` | 96.4 MB |
| `unimvu_7B_avsd` | 7B | Single-task | AVSD | `lmms-lab/llava-onevision-qwen2-7b-ov` | 715.9 MB |
| `unimvu_7B_music_avqa` | 7B | Single-task | Music-AVQA | `lmms-lab/llava-onevision-qwen2-7b-ov` | 715.9 MB |
| `unimvu_7B_scanqa` | 7B | Single-task | ScanQA | `lmms-lab/llava-onevision-qwen2-7b-ov` | 1.04 GB |
| `unimvu_7B_sqa3d` | 7B | Single-task | SQA3D | `lmms-lab/llava-onevision-qwen2-7b-ov` | 1.04 GB |
| `unimvu_uni_0.5B` | 0.5B | Unified | Mixed multi-task release | `lmms-lab/llava-onevision-qwen2-0.5b-ov` | 103.7 MB |
| `unimvu_uni_7B` | 7B | Unified | Mixed multi-task release | `lmms-lab/llava-onevision-qwen2-7b-ov` | 745.3 MB |
The default upload manifest publishes only the final release files:
- `adapter_config.json`
- `adapter_model.safetensors`
- `config.json`
- `non_lora_trainables.bin`
Intermediate `checkpoint-*` folders inside `unimvu_uni_0.5B` are training snapshots and are excluded from the default Hugging Face upload.
## Requirements
Use these adapters with the open-source UniMVU codebase and its dependencies:
```bash
pip install -r requirements.txt
pip install huggingface_hub peft
```
If you only need one adapter, prefer `snapshot_download(...)` so you do not fetch the entire release repo.
## Quick Start
The example below downloads one subfolder from this repo and loads it through UniMVU's own evaluation loader, which merges the LoRA adapter and then restores `non_lora_trainables.bin`.
```python
import os
from huggingface_hub import snapshot_download
from unified_eval import load_trained_model_for_eval
REPO_ID = "BonanDing/UniMVU"
SUBFOLDER = "unimvu_uni_7B"
local_root = snapshot_download(
repo_id=REPO_ID,
allow_patterns=[f"{SUBFOLDER}/*"],
)
model_path = os.path.join(local_root, SUBFOLDER)
tokenizer, model, image_processor, context_len = load_trained_model_for_eval(
model_path=model_path,
model_base="lmms-lab/llava-onevision-qwen2-7b-ov",
model_arg_name="VideoFeatModelArgumentsUniMVU_Uni_7B",
model_type="unimvu_uni",
device="cuda",
)
model.eval()
```
## Loader Mapping
| Release family | `model_type` | `model_arg_name` | `model_base` |
| --- | --- | --- | --- |
| Single-task 0.5B adapters | `unimvu` | `VideoFeatModelArgumentsUniMVU` | `lmms-lab/llava-onevision-qwen2-0.5b-ov` |
| Single-task 7B adapters | `unimvu` | `VideoFeatModelArgumentsUniMVU_7B` | `lmms-lab/llava-onevision-qwen2-7b-ov` |
| Unified 0.5B adapter | `unimvu_uni` | `VideoFeatModelArgumentsUniMVU_Uni` | `lmms-lab/llava-onevision-qwen2-0.5b-ov` |
| Unified 7B adapter | `unimvu_uni` | `VideoFeatModelArgumentsUniMVU_Uni_7B` | `lmms-lab/llava-onevision-qwen2-7b-ov` |
## Evaluation Entry Points
- Use `unified_eval.py` for AVQA, AVSD, Music-AVQA, ScanQA, and SQA3D.
- Use `lmms_eval_start.py` for MVBench-style evaluation in the UniMVU codebase.
## License
The released adapters depend on third-party base models and should be used in compliance with the licenses of:
- `lmms-lab/llava-onevision-qwen2-0.5b-ov`
- `lmms-lab/llava-onevision-qwen2-7b-ov`
Please also follow the usage terms of the downstream datasets and features used in evaluation.
## Citation
If you use UniMVU in your work, please cite:
```bibtex
@inproceedings{ding2026unimvu,
title={Not All Modalities Are Equal: Instruction-Aware Gating for Multimodal Videos},
author={Ding, Bonan and Nawaz, Umair and Khan, Ufaq and Shaker, Abdelrahman M. and Khan, Muhammad Haris and Cao, Jiale and Xie, Jin and Khan, Fahad Shahbaz},
booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
year={2026}
}
```
## Acknowledgements
UniMVU builds on the open-source multimodal ecosystem around LLaVA-style training utilities, LMMS-Eval, PEFT, and Transformers.
|