--- license: openmdw-1.0 library_name: transformers pipeline_tag: feature-extraction tags: - vision-language - multimodal - image-text-retrieval - hyperbolic-embeddings ---
**hyper3-clip-v1** is a hyperbolic vision–language model for visual search. It brings images and text into a shared embedding space for broad-to-specific retrieval and compositional matching. - **Broad-to-specific retrieval:** Search for a broad concept such as “animal” to find different types of animals, then narrow your search with a specific description. - **Compositional matching:** Find images by combinations of objects, attributes, and relationships—for example, “a red car beside a tree.” This helps distinguish similar scenes where a color, object, or relationship changes. ## Results | Model | [SugarCrepe](https://github.com/RAIVNLab/sugar-crepe) macro accuracy | |---|---:| | **hyper3-clip-v1** | **79.54%** | | Jina CLIP v2 | 75.02% | | OpenAI CLIP ViT-B/16 | 73.06% | Image retrieval — mAP (%) on [evaluation subsets](https://hyper3labs.com/) of each dataset. | Dataset | hyper3-clip-v1 | OpenAI CLIP ViT-B/32 | |---|---:|---:| | Amazon Berkeley Objects — product types | **58.2** | 55.2 | | DeepFashion In-Shop — same item | **63.5** | 35.2 | | COCO — object categories | **55.4** | 53.2 | ## Quick start Complete the short access form on this page, then install and sign in. ```bash pip install "torch>=2.2" "transformers>=4.49,<5" "timm>=1.0" \ "safetensors>=0.4" "Pillow>=10" "huggingface_hub>=0.34" hf auth login ``` For SentenceTransformers, also run `pip install "sentence-transformers>=5.5.1,<6"`. This example ranks three descriptions of a sofa. The example image downloads automatically. ```python from PIL import Image from huggingface_hub import hf_hub_download from transformers import AutoModel model_id = "hyper3labs/hyper3-clip-v1" model = AutoModel.from_pretrained(model_id, trust_remote_code=True).eval() image = Image.open(hf_hub_download(model_id, "examples/grey-velvet-sofa.jpg")).convert("RGB") descriptions = ["a grey velvet sofa", "a blue velvet sofa", "a wooden chair"] scores = model.score(image, descriptions) for score, text in sorted(zip(scores.tolist(), descriptions), reverse=True): print(f"{score:.3f} {text}") ``` The default is Lorentz scoring. Scores are negative; higher means a closer match. To try your own, use `image = Image.open("your-image.jpg").convert("RGB")` and change the descriptions — broad words such as “furniture” work alongside specific ones. ## Scoring Use `scoring="lorentz"`, `"cosine"`, or `"cone"` with `model.score(image, descriptions, scoring=...)`. | Scoring | What it measures | Use | |---|---|---| | **Lorentz** (default) | Closeness in hyperbolic space; gives the same ranking as hyperbolic distance | General image–text retrieval | | **Cosine** | Similarity of embedding directions | Standard vector similarity search | | **Cone** | How well an image fits a description's general-to-specific region | Directional and compositional matching | Higher is better for all three. Their numerical scales differ. The SugarCrepe result above uses cone scoring.