Wähler-4B GGUF

Quantized GGUF files for Wähler-4B, the German decision model, for CPUs and small machines. The Q4_K_M file is the default; Q5_K_M and Q8_0 are higher fidelity at the cost of speed and memory.

Files · Use with jevalt · Code and links · Examples and results: Wähler-4B card · Try it: Space

Emberwick: every villager asks Deem-4B what to do next.

The 53-second film, sound on: two mistakes small decision models make and how JevAlt fixes each one.

Deutsche Version

Der 53-Sekunden-Film, mit Ton. Emberwick auf Deutsch: Jeder Dorfbewohner fragt Wähler-4B, was als Nächstes zu tun ist. GIF (4K) · leichtes GIF

Files

File Size Argmax agreement Max prob gap Sec/request Peak RAM Load setting
Q4_K_M 2.71 GB 97.5% 0.3419 10.03 s 3.03 GB no mmap, repack on (jevalt default)
Q5_K_M 3.08 GB 95.0% 0.1494 17.33 s 3.39 GB no mmap, repack on (jevalt default)
Q8_0 4.48 GB 97.5% 0.0315 11.58 s 4.76 GB no mmap, repack on (jevalt default)

Measured by the export job on an HF cpu-upgrade machine (8 vCPU): argmax agreement and probability gap against the full-precision model on held-out validation requests, seconds per request with 8 threads, peak RAM of a fresh process serving the file at 4k context with 4 threads. The load setting column shows the llama.cpp configuration behind the RAM figure.

Use with jevalt

pip install "jevalt[serve,gguf] @ git+https://github.com/mertkayacs/jevalt"
jevalt serve --model mertkayacs/Wahler-4B-GGUF --file Wahler-4B-Q4_K_M.gguf

Then send the Jev request body to http://127.0.0.1:8000/v1/systemone. The answers come from the model's next-token probabilities at each decision marker, which the jevalt server reads through llama-cpp-python. llama.cpp's own llama-server and LM Studio can load the file, but their chat endpoints return generated text, so they give you neither the Jev API nor calibrated probabilities.

Code and links

If this is useful to you, a star on GitHub helps other people find it.

Downloads last month
288
GGUF
Model size
4B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mertkayacs/Wahler-4B-GGUF

Finetuned
Qwen/Qwen3.5-4B
Quantized
(2)
this model

Collection including mertkayacs/Wahler-4B-GGUF