Image-Text-to-Text
Transformers
qwen3_5
embedding
multimodal
quantized
fp8
video-text-to-text
conversational
custom_code
Instructions to use Weidows/WeMM-Embedding-2B-FP8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Weidows/WeMM-Embedding-2B-FP8 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="Weidows/WeMM-Embedding-2B-FP8", trust_remote_code=True) messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Weidows/WeMM-Embedding-2B-FP8", trust_remote_code=True) model = AutoModelForMultimodalLM.from_pretrained("Weidows/WeMM-Embedding-2B-FP8", trust_remote_code=True, device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Weidows/WeMM-Embedding-2B-FP8 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Weidows/WeMM-Embedding-2B-FP8" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Weidows/WeMM-Embedding-2B-FP8", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Weidows/WeMM-Embedding-2B-FP8
- SGLang
How to use Weidows/WeMM-Embedding-2B-FP8 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Weidows/WeMM-Embedding-2B-FP8" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Weidows/WeMM-Embedding-2B-FP8", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Weidows/WeMM-Embedding-2B-FP8" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Weidows/WeMM-Embedding-2B-FP8", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use Weidows/WeMM-Embedding-2B-FP8 with Docker Model Runner:
docker model run hf.co/Weidows/WeMM-Embedding-2B-FP8
Download modeling_st_wemm.py from Weidows/WeMM-Embedding-2B-FP8: direct link, hf CLI and curl.
- Browser
- Download file 3.18 kB
-
https://huggingface.co/Weidows/WeMM-Embedding-2B-FP8/resolve/main/modeling_st_wemm.py
- Command line
-
hf download hf://Weidows/WeMM-Embedding-2B-FP8/modeling_st_wemm.py
-
curl -L -o modeling_st_wemm.py https://huggingface.co/Weidows/WeMM-Embedding-2B-FP8/resolve/main/modeling_st_wemm.py
3.18 kB
| """Sentence Transformers module for WeMM-Embedding. | |
| Reproduces the `transformers` usage from the model card inside a Sentence Transformers | |
| pipeline: vision inputs are prepared with `qwen_vl_utils.process_vision_info` and the | |
| embedding is read from `WeMMEmbedding.embedding`, which pools the `<embedding>` position and | |
| L2-normalizes. Everything else (batching, prompts, truncation, `encode_query` / | |
| `encode_document`, similarity) comes from the stock `Transformer` module. | |
| """ | |
| from __future__ import annotations | |
| import inspect | |
| from typing import Any | |
| from sentence_transformers.models import Transformer | |
| class WeMMTransformer(Transformer): | |
| """`Transformer` that prepares images and videos the way the model card's snippet does.""" | |
| def __init__(self, model_name_or_path: str, **kwargs: Any) -> None: | |
| super().__init__(model_name_or_path, **kwargs) | |
| vision_config = getattr(self.config, "vision_config", None) | |
| self.image_patch_size = int(getattr(vision_config, "patch_size", 16)) | |
| # `embedding` hands its **kwargs to the inner model, so filtering on its own signature | |
| # would drop `pixel_values`. Filter on the inner model's parameters instead, plus the | |
| # processor's input names for anything the model only accepts as **kwargs. | |
| inner_model = getattr(self.model, "model", self.model) | |
| signature = set(inspect.signature(inner_model.forward).parameters) | |
| signature |= set(getattr(self.processor, "model_input_names", ())) | |
| for modality_params in self.modality_config.values(): | |
| method_name = modality_params["method"] | |
| if method_name != "forward": | |
| self._method_signature_cache.setdefault(method_name, signature) | |
| def _apply_chat_template( | |
| self, | |
| messages: list[list[dict[str, Any]]], | |
| modality_kwargs: dict[str, dict[str, Any]], | |
| common_kwargs: dict[str, Any], | |
| chat_template_kwargs: dict[str, Any], | |
| ) -> dict[str, Any]: | |
| """Render the chat template and prepare images / videos exactly as the model card does.""" | |
| from qwen_vl_utils import process_vision_info | |
| chat_template_kwargs = {"add_generation_prompt": False, **chat_template_kwargs} | |
| texts = [ | |
| self.processor.apply_chat_template(conversation, tokenize=False, **chat_template_kwargs) | |
| for conversation in messages | |
| ] | |
| images, videos, video_kwargs = process_vision_info( | |
| [list(conversation) for conversation in messages], | |
| image_patch_size=self.image_patch_size, | |
| return_video_kwargs=True, | |
| return_video_metadata=True, | |
| ) | |
| if videos is not None: | |
| videos, video_metadata = (list(part) for part in zip(*videos)) | |
| video_kwargs = {**video_kwargs, "video_metadata": video_metadata} | |
| return self.processor( | |
| text=texts, | |
| images=images, | |
| videos=videos, | |
| text_kwargs=modality_kwargs["text"], | |
| images_kwargs=modality_kwargs["image"], | |
| videos_kwargs={**modality_kwargs["video"], **video_kwargs}, | |
| common_kwargs=common_kwargs, | |
| ) | |