Wahler-4B-GGUF / README.md
mertkayacs's picture
docs: add the Ollama run command
cfdc2b6 verified
|
Raw History Blame Contribute Delete
4.25 kB
metadata
license: apache-2.0
language:
  - de
  - en
base_model: mertkayacs/Wahler-4B
base_model_relation: quantized
quantized_by: mertkayacs
pipeline_tag: text-classification
tags:
  - gguf
  - llama.cpp
  - decision-model
  - calibration
  - local-ai
  - german
  - small-language-model
  - local-llm
  - on-device
  - text-classification
  - german-llm
widget:
  - text: >-
      {"state":"Hallo, mir wurde das März-Abo doppelt berechnet: zwei Zahlungen
      über 29 € am 3. März. Bitte erstatten Sie den doppelten Betrag noch heute,
      sonst kündige ich.\nViele Grüße,
      Daniel","questions":{"entscheidung":{"type":"choice","instructions":"Welches
      Team soll dieses Ticket bearbeiten?","criteria":{"Abrechnung":"Zahlungen,
      Rechnungen, Erstattungen","Technischer Support":"Fehler, Störungen,
      Ausfälle","Vertrieb":"Preise, Upgrades, neue Verträge","Konto":"Anmeldung,
      Passwort, Profiländerungen"}}},"reasoning":"off","abstain":false}
    example_title: 'Recorded full-precision Wähler-4B: support ticket, 1 October 2026'
    output:
      - label: Abrechnung
        score: 0.917556
      - label: Technischer Support
        score: 0.049459
      - label: Vertrieb
        score: 0.020136
      - label: Konto
        score: 0.012848
library_name: gguf
datasets:
  - mertkayacs/jevalt-data

Wähler-4B GGUF: small German language model for local decisions

GGUF builds of Wähler-4B, a small German language model for text classification and decisions on a CPU. Full-precision Wähler-4B: 92.0% accuracy; Kev-4B: 81.1% on held-out German decisions from the training data pipeline. This is a full-precision comparison; GGUF agreement is measured separately below.

Run with JevAlt and llama.cpp

Q4_K_M is the default, with 3.03 GB peak RAM at a 4k context.

pip install "jevalt[serve,gguf] @ git+https://github.com/mertkayacs/jevalt" && jevalt serve --model mertkayacs/Wahler-4B-GGUF --file Wahler-4B-Q4_K_M.gguf

Send requests to http://127.0.0.1:8000/v1/systemone. The main card has a complete request. The JevAlt server reads next-token decision probabilities through llama-cpp-python and applies calibration.json.

In Ollama: ollama run mertkayacs/wahler-4b (library page) answers with the option letter.

Plain llama.cpp, Ollama and LM Studio chat endpoints return generated text; use the JevAlt server for the Jev API and calibrated decisions.

Files and quantization checks

File Size Same top choice
Q4_K_M 2.71 GB 97.5%
Q5_K_M 3.08 GB 95.0%
Q8_0 4.48 GB 97.5%

Same top choice means agreement with full-precision Wähler-4B on 40 validation decisions per quantization, not benchmark accuracy. The Q4_K_M build agrees on 97.5% of those decisions. More precision does not guarantee higher agreement on this small sample.

export-report.json records probability gaps, speed and RAM. RAM was measured in a fresh process on an 8-vCPU HF machine, 4 threads, 4k context, with no mmap and repacking enabled.

Use and limits

Use for routing, tagging and triage in German. Refit calibration on your own data before setting thresholds. Planted instructions, long irrelevant text and date arithmetic remain failure cases; keep authorization outside the model. Leave reasoning off for German: it did not improve accuracy on the reported date, number and policy test.

The widget records a full-precision answer. Full results and training | Try the Space | Code | Project page.

Apache-2.0. Quantized from mertkayacs/Wahler-4B.

An Eschatia Labs project. Mert Kaya.