--- pipeline_tag: image-text-to-text license: apache-2.0 base_model: Qwen/Qwen2.5-VL-7B-Instruct library_name: zeromodels language: - en tags: - keras - zeromodels - qwen2_5_vl - qwen2.5-vl - multimodal - vision - image-text-to-text - pytorch - jax - tf --- *See [our collection](https://huggingface.co/collections/zeromodels/qwen25-vl-6a8eae30c89eafe78d5dd8fa) for all Qwen2.5-VL sizes.* # Run Qwen2.5-VL with Keras 3: JAX, PyTorch, or TensorFlow [![GitHub](https://img.shields.io/badge/GitHub-ZeroModels-181717?logo=github)](https://github.com/IMvision12/ZeroModels) [![Docs](https://img.shields.io/badge/Docs-Qwen2.5--VL-1f6feb)](https://imvision12.github.io/ZeroModels/qwen2_5_vl/) [![HuggingFace](https://img.shields.io/badge/HuggingFace-Qwen2.5--VL-ffd21e?logo=huggingface&logoColor=black)](https://huggingface.co/collections/zeromodels/qwen25-vl-6a8eae30c89eafe78d5dd8fa) # zeromodels/qwen2.5-vl-7b-instruct Pure-**Keras 3** conversion of [`Qwen/Qwen2.5-VL-7B-Instruct`](https://huggingface.co/Qwen/Qwen2.5-VL-7B-Instruct) for [zeromodels](https://github.com/IMvision12/ZeroModels). One implementation runs unmodified on **TensorFlow / Torch / JAX**. This is the **7B** variant, served here as **image + text -> text** via `Qwen2_5VLProcessor`; weights are stored in **bfloat16**. For model details, license, and usage terms, see the upstream [model card](https://huggingface.co/Qwen/Qwen2.5-VL-7B-Instruct). Paper: [Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution (arXiv:2409.12191)](https://arxiv.org/abs/2409.12191) · [HF Papers](https://huggingface.co/papers/2409.12191) Paper: [YaRN: Efficient Context Window Extension of Large Language Models (arXiv:2309.00071)](https://arxiv.org/abs/2309.00071) · [HF Papers](https://huggingface.co/papers/2309.00071) Paper: [Qwen-VL: A Frontier Large Vision-Language Model with Versatile Abilities (arXiv:2308.12966)](https://arxiv.org/abs/2308.12966) · [HF Papers](https://huggingface.co/papers/2308.12966) ## ✨ Quick start ### Text-only ```python import os os.environ["KERAS_BACKEND"] = "torch" # or "jax" / "tensorflow" from zeromodels.models.qwen2_5_vl import Qwen2_5VLTextGenerate, Qwen2_5VLProcessor model = Qwen2_5VLTextGenerate.from_weights("zeromodels/qwen2.5-vl-7b-instruct") processor = Qwen2_5VLProcessor.from_weights("zeromodels/qwen2.5-vl-7b-instruct") inputs = processor(conversation=[ {"role": "user", "content": [{"type": "text", "text": "Hello, who are you?"}]} ]) outputs = model.generate(**inputs, max_new_tokens=64) print(processor.decode(outputs[0])) ``` ### Image + text ```python import os os.environ["KERAS_BACKEND"] = "torch" # or "jax" / "tensorflow" from PIL import Image from zeromodels.models.qwen2_5_vl import Qwen2_5VLConditionalGenerate, Qwen2_5VLProcessor model = Qwen2_5VLConditionalGenerate.from_weights("zeromodels/qwen2.5-vl-7b-instruct") processor = Qwen2_5VLProcessor.from_weights("zeromodels/qwen2.5-vl-7b-instruct") inputs = processor(conversation=[ {"role": "user", "content": [ {"type": "image", "image": Image.open("photo.jpg")}, {"type": "text", "text": "Describe this image in one sentence."}, ]} ]) outputs = model.generate(**inputs, max_new_tokens=64) print(processor.decode(outputs[0])) ``` Load any Qwen2.5-VL variant the same way with `from_weights("zeromodels/")`: | Variant | Hub | | --- | --- | | `qwen2.5-vl-3b-instruct` | [zeromodels/qwen2.5-vl-3b-instruct](https://huggingface.co/zeromodels/qwen2.5-vl-3b-instruct) | | `qwen2.5-vl-7b-instruct` | [zeromodels/qwen2.5-vl-7b-instruct](https://huggingface.co/zeromodels/qwen2.5-vl-7b-instruct) | | `qwen2.5-vl-32b-instruct` | [zeromodels/qwen2.5-vl-32b-instruct](https://huggingface.co/zeromodels/qwen2.5-vl-32b-instruct) | | `qwen2.5-vl-72b-instruct` | [zeromodels/qwen2.5-vl-72b-instruct](https://huggingface.co/zeromodels/qwen2.5-vl-72b-instruct) | ## Special Thanks A huge thank you to the Qwen team at Alibaba for creating and releasing these models. License: Apache 2.0.