Instructions to use mertkayacs/Wahler-4B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use mertkayacs/Wahler-4B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf mertkayacs/Wahler-4B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf mertkayacs/Wahler-4B-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf mertkayacs/Wahler-4B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf mertkayacs/Wahler-4B-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf mertkayacs/Wahler-4B-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf mertkayacs/Wahler-4B-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf mertkayacs/Wahler-4B-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf mertkayacs/Wahler-4B-GGUF:Q4_K_M
Use Docker
docker model run hf.co/mertkayacs/Wahler-4B-GGUF:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use mertkayacs/Wahler-4B-GGUF with Ollama:
ollama run hf.co/mertkayacs/Wahler-4B-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use mertkayacs/Wahler-4B-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf mertkayacs/Wahler-4B-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "mertkayacs/Wahler-4B-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use mertkayacs/Wahler-4B-GGUF with Docker Model Runner:
docker model run hf.co/mertkayacs/Wahler-4B-GGUF:Q4_K_M
- Lemonade
How to use mertkayacs/Wahler-4B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull mertkayacs/Wahler-4B-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Wahler-4B-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use mertkayacs/Wahler-4B-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf mertkayacs/Wahler-4B-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default mertkayacs/Wahler-4B-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use mertkayacs/Wahler-4B-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf mertkayacs/Wahler-4B-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "mertkayacs/Wahler-4B-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Download README.md from mertkayacs/Wahler-4B-GGUF: direct link, hf CLI and curl.
- Browser
- Download file 4.25 kB
-
https://huggingface.co/mertkayacs/Wahler-4B-GGUF/resolve/main/README.md
- Command line
-
hf download hf://mertkayacs/Wahler-4B-GGUF/README.md
-
curl -L -o README.md https://huggingface.co/mertkayacs/Wahler-4B-GGUF/resolve/main/README.md
license: apache-2.0
language:
- de
- en
base_model: mertkayacs/Wahler-4B
base_model_relation: quantized
quantized_by: mertkayacs
pipeline_tag: text-classification
tags:
- gguf
- llama.cpp
- decision-model
- calibration
- local-ai
- german
- small-language-model
- local-llm
- on-device
- text-classification
- german-llm
widget:
- text: >-
{"state":"Hallo, mir wurde das März-Abo doppelt berechnet: zwei Zahlungen
über 29 € am 3. März. Bitte erstatten Sie den doppelten Betrag noch heute,
sonst kündige ich.\nViele Grüße,
Daniel","questions":{"entscheidung":{"type":"choice","instructions":"Welches
Team soll dieses Ticket bearbeiten?","criteria":{"Abrechnung":"Zahlungen,
Rechnungen, Erstattungen","Technischer Support":"Fehler, Störungen,
Ausfälle","Vertrieb":"Preise, Upgrades, neue Verträge","Konto":"Anmeldung,
Passwort, Profiländerungen"}}},"reasoning":"off","abstain":false}
example_title: 'Recorded full-precision Wähler-4B: support ticket, 1 October 2026'
output:
- label: Abrechnung
score: 0.917556
- label: Technischer Support
score: 0.049459
- label: Vertrieb
score: 0.020136
- label: Konto
score: 0.012848
library_name: gguf
datasets:
- mertkayacs/jevalt-data
Wähler-4B GGUF: small German language model for local decisions
GGUF builds of Wähler-4B, a small German language model for text classification and decisions on a CPU. Full-precision Wähler-4B: 92.0% accuracy; Kev-4B: 81.1% on held-out German decisions from the training data pipeline. This is a full-precision comparison; GGUF agreement is measured separately below.
Run with JevAlt and llama.cpp
Q4_K_M is the default, with 3.03 GB peak RAM at a 4k context.
pip install "jevalt[serve,gguf] @ git+https://github.com/mertkayacs/jevalt" && jevalt serve --model mertkayacs/Wahler-4B-GGUF --file Wahler-4B-Q4_K_M.gguf
Send requests to http://127.0.0.1:8000/v1/systemone. The main card has a complete request. The JevAlt server reads next-token decision probabilities through llama-cpp-python and applies calibration.json.
In Ollama: ollama run mertkayacs/wahler-4b (library page) answers with the option letter.
Plain llama.cpp, Ollama and LM Studio chat endpoints return generated text; use the JevAlt server for the Jev API and calibrated decisions.
Files and quantization checks
Same top choice means agreement with full-precision Wähler-4B on 40 validation decisions per quantization, not benchmark accuracy. The Q4_K_M build agrees on 97.5% of those decisions. More precision does not guarantee higher agreement on this small sample.
export-report.json records probability gaps, speed and RAM. RAM was measured in a fresh process on an 8-vCPU HF machine, 4 threads, 4k context, with no mmap and repacking enabled.
Use and limits
Use for routing, tagging and triage in German. Refit calibration on your own data before setting thresholds. Planted instructions, long irrelevant text and date arithmetic remain failure cases; keep authorization outside the model. Leave reasoning off for German: it did not improve accuracy on the reported date, number and policy test.
The widget records a full-precision answer. Full results and training | Try the Space | Code | Project page.
Apache-2.0. Quantized from mertkayacs/Wahler-4B.
An Eschatia Labs project. Mert Kaya.