Instructions to use roelfrenkema/gemma-4-E4B-it.Q8_0-med.gguf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use roelfrenkema/gemma-4-E4B-it.Q8_0-med.gguf with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf roelfrenkema/gemma-4-E4B-it.Q8_0-med.gguf:Q8_0 # Run inference directly in the terminal: llama cli -hf roelfrenkema/gemma-4-E4B-it.Q8_0-med.gguf:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf roelfrenkema/gemma-4-E4B-it.Q8_0-med.gguf:Q8_0 # Run inference directly in the terminal: llama cli -hf roelfrenkema/gemma-4-E4B-it.Q8_0-med.gguf:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf roelfrenkema/gemma-4-E4B-it.Q8_0-med.gguf:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf roelfrenkema/gemma-4-E4B-it.Q8_0-med.gguf:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf roelfrenkema/gemma-4-E4B-it.Q8_0-med.gguf:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf roelfrenkema/gemma-4-E4B-it.Q8_0-med.gguf:Q8_0
Use Docker
docker model run hf.co/roelfrenkema/gemma-4-E4B-it.Q8_0-med.gguf:Q8_0
- LM Studio
- Jan
- Ollama
How to use roelfrenkema/gemma-4-E4B-it.Q8_0-med.gguf with Ollama:
ollama run hf.co/roelfrenkema/gemma-4-E4B-it.Q8_0-med.gguf:Q8_0
- Unsloth Desktop
- Pi
How to use roelfrenkema/gemma-4-E4B-it.Q8_0-med.gguf with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf roelfrenkema/gemma-4-E4B-it.Q8_0-med.gguf:Q8_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "roelfrenkema/gemma-4-E4B-it.Q8_0-med.gguf:Q8_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use roelfrenkema/gemma-4-E4B-it.Q8_0-med.gguf with Docker Model Runner:
docker model run hf.co/roelfrenkema/gemma-4-E4B-it.Q8_0-med.gguf:Q8_0
- Lemonade
How to use roelfrenkema/gemma-4-E4B-it.Q8_0-med.gguf with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull roelfrenkema/gemma-4-E4B-it.Q8_0-med.gguf:Q8_0
Run and chat with the model
lemonade run user.gemma-4-E4B-it.Q8_0-med.gguf-Q8_0
List all available models
lemonade list
- Hermes Agent
How to use roelfrenkema/gemma-4-E4B-it.Q8_0-med.gguf with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf roelfrenkema/gemma-4-E4B-it.Q8_0-med.gguf:Q8_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default roelfrenkema/gemma-4-E4B-it.Q8_0-med.gguf:Q8_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use roelfrenkema/gemma-4-E4B-it.Q8_0-med.gguf with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf roelfrenkema/gemma-4-E4B-it.Q8_0-med.gguf:Q8_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "roelfrenkema/gemma-4-E4B-it.Q8_0-med.gguf:Q8_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Model Card for Model ID gemma-4-E4B-it.Q8_0-med.gguf
Gemma-4-E4B with a medical layer.
Model Details
Model Description
Gemma 4 E4B is getraind met een in het Nederlands vertaalde MedAlpaca
Het is verbazingwekkend acuraat en bij uitstek geschikt voor gebruik met RAG.
Ik heb ook een E2B versie beschikbaar te vinden onder
Dit is werk dat ik als vrijwilliger doe voor Impuls en Woortblind
- Developed by: @roelfrenkema
- Funded by [optional]: Ik zoek funding.
- Model type: Gemma 4 E4B
- Language(s) (NLP): NL
- License: Apache
- Finetuned from model [optional]: Gemma 4 E4B
Uses
Gebruik de Modelfile voor Ollama (nieuwste versie) of llama.cpp etc.
Bias, Risks, and Limitations
Training Details
Training Data
training:
max_seq_length: 2048
num_epochs: 1
learning_rate: 0.0001
batch_size: 1
gradient_accumulation_steps: 8
warmup_steps: 100
max_steps: -1
save_steps: 500
eval_steps: 0
weight_decay: 0.001
random_seed: 3407
packing: false
train_on_completions: true
gradient_checkpointing: unsloth
optim: adamw_8bit
lr_scheduler_type: cosine
lora:
lora_r: 16
lora_alpha: 32
lora_dropout: 0
target_modules:
- q_proj
- k_proj
- v_proj
- o_proj
- gate_proj
- up_proj
- down_proj
use_rslora: false
use_loftq: false
finetune_vision_layers: true
finetune_language_layers: true
finetune_attention_modules: true
finetune_mlp_modules: true
Training Procedure
UNSLOTH
- Downloads last month
- 42
8-bit
