Image-to-Text
Transformers
Safetensors
English
blip-2
text-generation
video-to-text
video-captioning
image-captioning
visual-question-answering
Instructions to use kpyu/eilev-blip2-flan-t5-xl with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use kpyu/eilev-blip2-flan-t5-xl with Transformers:
# Use a pipeline as a high-level helper # Warning: Pipeline type "image-to-text" is no longer supported in transformers v5. # You must load the model directly (see below) or downgrade to v4.x with: # 'pip install "transformers<5.0.0' from transformers import pipeline pipe = pipeline("image-to-text", model="kpyu/eilev-blip2-flan-t5-xl")# Load model directly from transformers import AutoProcessor, AutoModelForSeq2SeqLM processor = AutoProcessor.from_pretrained("kpyu/eilev-blip2-flan-t5-xl") model = AutoModelForSeq2SeqLM.from_pretrained("kpyu/eilev-blip2-flan-t5-xl", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Update README.md
Browse files
README.md
CHANGED
|
@@ -16,7 +16,7 @@ tags:
|
|
| 16 |
|
| 17 |

|
| 18 |
|
| 19 |
-
[Salesforce/blip2-flan-t5-xl](https://huggingface.co/Salesforce/blip2-flan-t5-xl) trained using [
|
| 20 |
|
| 21 |
## Model Details
|
| 22 |
|
|
|
|
| 16 |
|
| 17 |

|
| 18 |
|
| 19 |
+
[Salesforce/blip2-flan-t5-xl](https://huggingface.co/Salesforce/blip2-flan-t5-xl) trained using [EILeV](https://github.com/yukw777/EILEV), a novel training method that can elicit in-context learning in vision-language models (VLMs) for videos without requiring massive, naturalistic video datasets.
|
| 20 |
|
| 21 |
## Model Details
|
| 22 |
|