Zero-Shot Image Classification
OpenCLIP
ONNX
Safetensors
Transformers
Transformers.js
English
siglip
clip
e-commerce
fashion
multimodal retrieval
custom_code
Instructions to use Marqo/marqo-fashionSigLIP with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- OpenCLIP
How to use Marqo/marqo-fashionSigLIP with OpenCLIP:
import open_clip model, preprocess_train, preprocess_val = open_clip.create_model_and_transforms('hf-hub:Marqo/marqo-fashionSigLIP') tokenizer = open_clip.get_tokenizer('hf-hub:Marqo/marqo-fashionSigLIP') - Transformers
How to use Marqo/marqo-fashionSigLIP with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("zero-shot-image-classification", model="Marqo/marqo-fashionSigLIP", trust_remote_code=True) pipe( "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/hub/parrots.png", candidate_labels=["animals", "humans", "landscape"], )# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Marqo/marqo-fashionSigLIP", trust_remote_code=True, device_map="auto") - Transformers.js
How to use Marqo/marqo-fashionSigLIP with Transformers.js:
// npm i @huggingface/transformers import { pipeline } from '@huggingface/transformers'; // Allocate pipeline const pipe = await pipeline('zero-shot-image-classification', 'Marqo/marqo-fashionSigLIP'); - Notebooks
- Google Colab
- Kaggle
Sentence Transformers Support
#6
by mdaya - opened
Hello, any plans to make this model compatible with sentence-transformers? https://github.com/UKPLab/sentence-transformers/
Thanks
Ok. Looks like all we need to do is to update config.json:
(config values were extracted from loaded model)
{
"architectures": ["SigLIP"],
"auto_map": {
"AutoConfig": "marqo_fashionSigLIP.MarqoFashionSigLIPConfig",
"AutoModel": "marqo_fashionSigLIP.MarqoFashionSigLIP",
"AutoProcessor": "marqo_fashionSigLIP.MarqoFashionSigLIPProcessor"
},
"open_clip_model_name": "hf-hub:Marqo/marqo-fashionSigLIP",
"model_type": "siglip",
"hidden_size": 768,
"projection_dim": 768,
"text_config": {
"attention_dropout": 0.0,
"bos_token_id": 49406,
"eos_token_id": 49407,
"hidden_act": "gelu_pytorch_tanh",
"hidden_size": 768,
"intermediate_size": 3072,
"layer_norm_eps": 1e-6,
"max_position_embeddings": 64,
"model_type": "siglip_text_model",
"num_attention_heads": 12,
"num_hidden_layers": 12,
"pad_token_id": 1,
"transformers_version": "4.47.1",
"vocab_size": 32000
},
"vision_config": {
"attention_dropout": 0.0,
"hidden_act": "gelu_pytorch_tanh",
"hidden_size": 768,
"image_size": 224,
"intermediate_size": 3072,
"layer_norm_eps": 1e-6,
"model_type": "siglip_vision_model",
"num_attention_heads": 12,
"num_channels": 3,
"num_hidden_layers": 12,
"patch_size": 16,
"transformers_version": "4.47.1"
},
"initializer_factor": 1.0,
"return_dict": true,
"output_hidden_states": false,
"output_attentions": false,
"torchscript": false,
"use_bfloat16": false,
"tie_word_embeddings": true
}