How to use from
vLLM
Install from pip and serve model
# Install vLLM from pip:
pip install vllm
# Start the vLLM server:
vllm serve "mixer3d/translategemma-4b-it-gguf"
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:8000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "mixer3d/translategemma-4b-it-gguf",
		"messages": [
			{
				"role": "user",
				"content": [
					{
						"type": "text",
						"text": "Describe this image in one sentence."
					},
					{
						"type": "image_url",
						"image_url": {
							"url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg"
						}
					}
				]
			}
		]
	}'
Use Docker
docker model run hf.co/mixer3d/translategemma-4b-it-gguf:Q5_K_M
Quick Links

translategemma-4b-it-q5_k_m.gguf

This repo contains GGUF weights for google/translategemma-4b-it.

Quantization Details

  • Method: llama-quantize
  • Llama.cpp Version: 7770 (fe44d3557)
  • Original Model Precision: BF16

Files Provided

File Quant Method Size Description
translategemma-4b-it-q5_k_m.gguf Q5_K_M 3.2 GB High quality, recommended for most uses.

Usage

You can use these models with llama.cpp

./llama-server -m translategemma-4b-it-q5_k_m.gguf -no-mmap -ngl 99 --port 8080 -c 8192 -fa 1 --jinja
Downloads last month
7
GGUF
Model size
4B params
Architecture
gemma3
Hardware compatibility
Log In to add your hardware

5-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mixer3d/translategemma-4b-it-gguf

Quantized
(44)
this model