Visual Question Answering
Transformers
Safetensors
minicpmv
feature-extraction
custom_code
4-bit precision
bitsandbytes
Instructions to use openbmb/MiniCPM-Llama3-V-2_5-int4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use openbmb/MiniCPM-Llama3-V-2_5-int4 with Transformers:
# Use a pipeline as a high-level helper # Warning: Pipeline type "visual-question-answering" is no longer supported in transformers v5. # You must load the model directly (see below) or downgrade to v4.x with: # pip install "transformers<5.0.0" from transformers import pipeline pipe = pipeline("visual-question-answering", model="openbmb/MiniCPM-Llama3-V-2_5-int4", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("openbmb/MiniCPM-Llama3-V-2_5-int4", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
|
Download README.md from openbmb/MiniCPM-Llama3-V-2_5-int4: direct link, hf CLI and curl.
- Browser
- Download file 1.57 kB
-
https://huggingface.co/openbmb/MiniCPM-Llama3-V-2_5-int4/resolve/ea02fa1b3b4a5ffd7646d16feccd76f391d98ab5/README.md
- Command line
-
hf download hf://openbmb/MiniCPM-Llama3-V-2_5-int4@ea02fa1b3b4a5ffd7646d16feccd76f391d98ab5/README.md
-
curl -L -o README.md https://huggingface.co/openbmb/MiniCPM-Llama3-V-2_5-int4/resolve/ea02fa1b3b4a5ffd7646d16feccd76f391d98ab5/README.md
1.57 kB
| pipeline_tag: visual-question-answering | |
| ## MiniCPM-Llama3-V 2.5 int4 | |
| This is the int4 quantized version of [MiniCPM-Llama3-V 2.5](https://huggingface.co/openbmb/MiniCPM-Llama3-V-2_5). | |
| Running with int4 version would use lower GPU mermory (about 9GB). | |
| ## Usage | |
| Inference using Huggingface transformers on NVIDIA GPUs. Requirements tested on python 3.10: | |
| ``` | |
| Pillow==10.1.0 | |
| torch==2.1.2 | |
| torchvision==0.16.2 | |
| transformers==4.40.0 | |
| sentencepiece==0.1.99 | |
| accelerate==0.30.1 | |
| bitsandbytes==0.43.1 | |
| ``` | |
| ```python | |
| # test.py | |
| import torch | |
| from PIL import Image | |
| from transformers import AutoModel, AutoTokenizer | |
| model = AutoModel.from_pretrained('openbmb/MiniCPM-Llama3-V-2_5-int4', trust_remote_code=True) | |
| tokenizer = AutoTokenizer.from_pretrained('openbmb/MiniCPM-Llama3-V-2_5-int4', trust_remote_code=True) | |
| model.eval() | |
| image = Image.open('xx.jpg').convert('RGB') | |
| question = 'What is in the image?' | |
| msgs = [{'role': 'user', 'content': question}] | |
| res = model.chat( | |
| image=image, | |
| msgs=msgs, | |
| tokenizer=tokenizer, | |
| sampling=True, # if sampling=False, beam_search will be used by default | |
| temperature=0.7, | |
| # system_prompt='' # pass system_prompt if needed | |
| ) | |
| print(res) | |
| ## if you want to use streaming, please make sure sampling=True and stream=True | |
| ## the model.chat will return a generator | |
| res = model.chat( | |
| image=image, | |
| msgs=msgs, | |
| tokenizer=tokenizer, | |
| sampling=True, | |
| temperature=0.7, | |
| stream=True | |
| ) | |
| generated_text = "" | |
| for new_text in res: | |
| generated_text += new_text | |
| print(new_text, flush=True, end='') | |
| ``` | |