Instructions to use Radamanthys11/Gemma-4-E2B-it-assistant-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Radamanthys11/Gemma-4-E2B-it-assistant-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Radamanthys11/Gemma-4-E2B-it-assistant-GGUF:F16 # Run inference directly in the terminal: llama cli -hf Radamanthys11/Gemma-4-E2B-it-assistant-GGUF:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Radamanthys11/Gemma-4-E2B-it-assistant-GGUF:F16 # Run inference directly in the terminal: llama cli -hf Radamanthys11/Gemma-4-E2B-it-assistant-GGUF:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Radamanthys11/Gemma-4-E2B-it-assistant-GGUF:F16 # Run inference directly in the terminal: ./llama-cli -hf Radamanthys11/Gemma-4-E2B-it-assistant-GGUF:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Radamanthys11/Gemma-4-E2B-it-assistant-GGUF:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf Radamanthys11/Gemma-4-E2B-it-assistant-GGUF:F16
Use Docker
docker model run hf.co/Radamanthys11/Gemma-4-E2B-it-assistant-GGUF:F16
- LM Studio
- Jan
- Ollama
How to use Radamanthys11/Gemma-4-E2B-it-assistant-GGUF with Ollama:
ollama run hf.co/Radamanthys11/Gemma-4-E2B-it-assistant-GGUF:F16
- Unsloth Desktop
- Docker Model Runner
How to use Radamanthys11/Gemma-4-E2B-it-assistant-GGUF with Docker Model Runner:
docker model run hf.co/Radamanthys11/Gemma-4-E2B-it-assistant-GGUF:F16
- Lemonade
How to use Radamanthys11/Gemma-4-E2B-it-assistant-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Radamanthys11/Gemma-4-E2B-it-assistant-GGUF:F16
Run and chat with the model
lemonade run user.Gemma-4-E2B-it-assistant-GGUF-F16
List all available models
lemonade list
- Atomic Chat
unknown speculative decoding type without draft model
I'm getting this ik_llama.cpp error and cannot run the model with mtp spec dec using your drafter:
unknown speculative decoding type without draft model
ik_llama.cpp built today (2025-05-07), do I have to use any specific branch? Could you add instructions about how to successfully run your drafter? Thx.
Yes, I forgot to mention in the description that these versions currently only work in open PR 1744, so you should build on top of them first.
I've try to make E4B assistant with your instructions but when I use it hangs on 271 output token , i get Nice 2x speed increase until 200 token output, please can you help me to figure this stuck on 271st token. Ik_llama of course the best for my A16
@agatazit , thank you for pointing that out. There is a specific bug in E4B that causes the model to freeze during runtime. I am analyzing the issue to find the best way to fix it. I will let you know when I submit a PR for you to test.
great and thanks for your response. I'll follow you on github ik_llama.cpp .