Instructions to use unsloth/gemma-3n-E2B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use unsloth/gemma-3n-E2B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="unsloth/gemma-3n-E2B")# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("unsloth/gemma-3n-E2B") model = AutoModelForMultimodalLM.from_pretrained("unsloth/gemma-3n-E2B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use unsloth/gemma-3n-E2B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "unsloth/gemma-3n-E2B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "unsloth/gemma-3n-E2B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/unsloth/gemma-3n-E2B
- SGLang
How to use unsloth/gemma-3n-E2B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "unsloth/gemma-3n-E2B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "unsloth/gemma-3n-E2B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "unsloth/gemma-3n-E2B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "unsloth/gemma-3n-E2B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Unsloth Desktop
- Docker Model Runner
How to use unsloth/gemma-3n-E2B with Docker Model Runner:
docker model run hf.co/unsloth/gemma-3n-E2B
|
Download README.md from unsloth/gemma-3n-E2B: direct link, hf CLI and curl.
- Browser
- Download file 3.99 kB
-
https://huggingface.co/unsloth/gemma-3n-E2B/resolve/main/README.md
- Command line
-
hf download hf://unsloth/gemma-3n-E2B/README.md
-
curl -L -o README.md https://huggingface.co/unsloth/gemma-3n-E2B/resolve/main/README.md
3.99 kB
metadata
base_model: google/gemma-3n-E2B
language:
- en
pipeline_tag: image-text-to-text
library_name: transformers
license: gemma
tags:
- gemma3
- unsloth
- transformers
- gemma
- google
Learn how to run & fine-tune Gemma 3n correctly - Read our Guide.
See our collection for all versions of Gemma 3n including GGUF, 4-bit & 16-bit formats.
Unsloth Dynamic 2.0 achieves SOTA accuracy & performance versus other quants.
✨ Gemma 3n Usage Guidelines
- Currently only text is supported.
- Ollama:
ollama run hf.co/unsloth/gemma-3n-E4B-it:Q4_K_XL- auto-sets correct chat template and settings - Set temperature = 1.0, top_k = 64, top_p = 0.95, min_p = 0.0
- Gemma 3n max tokens (context length): 32K. Gemma 3n chat template:
<bos><start_of_turn>user\nHello!<end_of_turn>\n<start_of_turn>model\nHey there!<end_of_turn>\n<start_of_turn>user\nWhat is 1+1?<end_of_turn>\n<start_of_turn>model\n
- For complete detailed instructions, see our step-by-step guide.
🦥 Fine-tune Gemma 3n with Unsloth
- Fine-tune Gemma 3n (4B) for free using our Google Colab notebook here!
- Read our Blog about Gemma 3n support: unsloth.ai/blog/gemma-3n
- View the rest of our notebooks in our docs here.
| Unsloth supports | Free Notebooks | Performance | Memory use |
|---|---|---|---|
| Gemma-3n-E4B | ▶️ Start on Colab | 2x faster | 80% less |
| GRPO with Gemma 3 (1B) | ▶️ Start on Colab | 2x faster | 80% less |
| Gemma 3 (4B) | ▶️ Start on Colab | 2x faster | 60% less |
| Qwen3 (14B) | ▶️ Start on Colab | 2x faster | 60% less |
| DeepSeek-R1-0528-Qwen3-8B (14B) | ▶️ Start on Colab | 2x faster | 80% less |
| Llama-3.2 (3B) | ▶️ Start on Colab | 2.4x faster | 58% less |