Instructions to use agentsea/paligemma-3b-ft-widgetcap-waveui-448 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use agentsea/paligemma-3b-ft-widgetcap-waveui-448 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="agentsea/paligemma-3b-ft-widgetcap-waveui-448")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("agentsea/paligemma-3b-ft-widgetcap-waveui-448") model = AutoModelForMultimodalLM.from_pretrained("agentsea/paligemma-3b-ft-widgetcap-waveui-448", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use agentsea/paligemma-3b-ft-widgetcap-waveui-448 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "agentsea/paligemma-3b-ft-widgetcap-waveui-448" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "agentsea/paligemma-3b-ft-widgetcap-waveui-448", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/agentsea/paligemma-3b-ft-widgetcap-waveui-448
- SGLang
How to use agentsea/paligemma-3b-ft-widgetcap-waveui-448 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "agentsea/paligemma-3b-ft-widgetcap-waveui-448" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "agentsea/paligemma-3b-ft-widgetcap-waveui-448", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "agentsea/paligemma-3b-ft-widgetcap-waveui-448" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "agentsea/paligemma-3b-ft-widgetcap-waveui-448", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use agentsea/paligemma-3b-ft-widgetcap-waveui-448 with Docker Model Runner:
docker model run hf.co/agentsea/paligemma-3b-ft-widgetcap-waveui-448
|
Download README.md from agentsea/paligemma-3b-ft-widgetcap-waveui-448: direct link, hf CLI and curl.
- Browser
- Download file 2.22 kB
-
https://huggingface.co/agentsea/paligemma-3b-ft-widgetcap-waveui-448/resolve/main/README.md
- Command line
-
hf download hf://agentsea/paligemma-3b-ft-widgetcap-waveui-448/README.md
-
curl -L -o README.md https://huggingface.co/agentsea/paligemma-3b-ft-widgetcap-waveui-448/resolve/main/README.md
2.22 kB
| datasets: | |
| - agentsea/wave-ui | |
| language: | |
| - en | |
| library_name: transformers | |
| # Paligemma WaveUI | |
| Transformers [PaliGemma 3B 448-res weights](https://huggingface.co/google/paligemma-3b-pt-448), fine-tuned on the [WaveUI](https://huggingface.co/datasets/agentsea/wave-ui) dataset for object-detection. | |
| ## Model Details | |
| ### Model Description | |
| This fine-tune was done atop of the [Paligemma 448 Widgetcap](https://huggingface.co/google/paligemma-3b-ft-widgetcap-448) model, using the [WaveUI](https://huggingface.co/datasets/agentsea/wave-ui) dataset, which contains ~80k examples of labeled UI elements. | |
| The fine-tune was done for the object detection task. Specifically, this model aims to perform well at UI element detection, as part of a wider effort to enable our open-source toolkit for building agents at [AgentSea](https://www.agentsea.ai/). | |
| - **Developed by:** https://agentsea.ai/ | |
| - **Language(s) (NLP):** en | |
| - **Finetuned from model:** https://huggingface.co/google/paligemma-3b-ft-widgetcap-448 | |
| ### Demo | |
| You can find a **demo** for this model [here](https://huggingface.co/spaces/agentsea/paligemma-waveui). | |
| ## Notes | |
| - The only task used in the fine-tune was the object detection task, so it might not perform well in other types of tasks. | |
| ## Usage | |
| To start using this model, run the following: | |
| ```python | |
| from transformers import AutoProcessor, PaliGemmaForConditionalGeneration | |
| model = PaliGemmaForConditionalGeneration.from_pretrained("agentsea/paligemma-3b-ft-widgetcap-waveui-448").eval() | |
| processor = AutoProcessor.from_pretrained("agentsea/paligemma-3b-ft-widgetcap-waveui-448") | |
| ``` | |
| ## Data | |
| We used the [WaveUI](https://huggingface.co/datasets/agentsea/wave-ui) dataset for this fine-tune. Before using it, we preprocessed the data to use the Paligemma bounding-box format. | |
| ## Evaluation | |
| We calculated the mean IoU over 1024 examples of the test set using 3 different closed-source models: Gemini 1.5 Pro, Claude 3.5 Sonnet and GPT 4o. We also ran this same calculation using the PaliGemma WaveUI fine-tunes. We obtained the following values: | |
| - Gemini 1.5 Pro: 0.12 | |
| - Claude 3.5 Sonnet: 0.05 | |
| - GPT 4o: 0.05 | |
| - **PaliGemma Widgetcap+WaveUI 448: 0.40** | |
| - PaliGemma WaveUI 896: 0.49 |