Image-Text-to-Text
Transformers
Safetensors
qwen3_vl
vision-language
image-classification
calibrated-probabilities
structured-outputs
typed-questions
jev-inspired
qwen3-vl
conversational
Instructions to use MeerDevelopment/Qevi-2B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use MeerDevelopment/Qevi-2B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="MeerDevelopment/Qevi-2B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("MeerDevelopment/Qevi-2B") model = AutoModelForMultimodalLM.from_pretrained("MeerDevelopment/Qevi-2B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use MeerDevelopment/Qevi-2B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "MeerDevelopment/Qevi-2B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "MeerDevelopment/Qevi-2B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/MeerDevelopment/Qevi-2B
- SGLang
How to use MeerDevelopment/Qevi-2B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "MeerDevelopment/Qevi-2B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "MeerDevelopment/Qevi-2B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "MeerDevelopment/Qevi-2B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "MeerDevelopment/Qevi-2B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use MeerDevelopment/Qevi-2B with Docker Model Runner:
docker model run hf.co/MeerDevelopment/Qevi-2B
Credit TypeSafe Jev as the inspiration; add structured-output tags
Browse files
README.md
CHANGED
|
@@ -7,6 +7,9 @@ tags:
|
|
| 7 |
- vision-language
|
| 8 |
- image-classification
|
| 9 |
- calibrated-probabilities
|
|
|
|
|
|
|
|
|
|
| 10 |
- qwen3-vl
|
| 11 |
---
|
| 12 |
|
|
@@ -330,6 +333,30 @@ per question rather than another full forward pass each.
|
|
| 330 |
The three files have no dependency on this repo's layout. Copy `qevi/` into your own project and
|
| 331 |
`import qevi` directly; only `torch`, `transformers` and `pillow` are required.
|
| 332 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 333 |
## Citation
|
| 334 |
|
| 335 |
```bibtex
|
|
|
|
| 7 |
- vision-language
|
| 8 |
- image-classification
|
| 9 |
- calibrated-probabilities
|
| 10 |
+
- structured-outputs
|
| 11 |
+
- typed-questions
|
| 12 |
+
- jev-inspired
|
| 13 |
- qwen3-vl
|
| 14 |
---
|
| 15 |
|
|
|
|
| 333 |
The three files have no dependency on this repo's layout. Copy `qevi/` into your own project and
|
| 334 |
`import qevi` directly; only `torch`, `transformers` and `pillow` are required.
|
| 335 |
|
| 336 |
+
## Inspiration and prior art
|
| 337 |
+
|
| 338 |
+
Qevi is an **independent, unaffiliated** implementation, for images, of an idea published by
|
| 339 |
+
[TypeSafe](https://typesafe.ai). Their "System One" model **Jev** answers typed questions with
|
| 340 |
+
type-safe structured values and calibrated probabilities rather than generating text: as they put
|
| 341 |
+
it, *"possible outputs and structure are defined in advance, the model never makes type errors,
|
| 342 |
+
all answers are accompanied with calibrated probabilities and confidence scores."*
|
| 343 |
+
|
| 344 |
+
Jev is a text model. Qevi asks the same question of images: if the answer is known in advance to
|
| 345 |
+
be one of a small fixed set, why make a vision-language model write a sentence to say it?
|
| 346 |
+
|
| 347 |
+
The name reflects the debt. This project began as **JEVI**, short for *Jev for images*, and was
|
| 348 |
+
later renamed Qevi (Qwen + Jevi) once it settled on a Qwen3-VL trunk.
|
| 349 |
+
|
| 350 |
+
**What is and is not shared.** The idea of typed, calibrated, non-generative outputs comes from
|
| 351 |
+
TypeSafe's public writing. Everything here is otherwise independent: no code, weights, data or
|
| 352 |
+
training method from TypeSafe is used, and this model is not endorsed by or affiliated with them.
|
| 353 |
+
In particular Qevi does **not** implement their RLCD training method: it is trained with ordinary
|
| 354 |
+
cross-entropy over candidate-answer logits with label smoothing, which is a proper scoring rule
|
| 355 |
+
and is what produces the calibration improvements reported above.
|
| 356 |
+
|
| 357 |
+
TypeSafe's published notes on Jev's failure modes (literal reading of questions, unreliable
|
| 358 |
+
counting) informed how this corpus was built and what this model does not claim to do.
|
| 359 |
+
|
| 360 |
## Citation
|
| 361 |
|
| 362 |
```bibtex
|