plumb-4b-GGUF / README.md
crh225's picture
Plumb-4B v5 GGUF (Q8_0, Q4_K_M)
71c1c57
|
Raw History Blame Contribute Delete
1.74 kB
---
license: apache-2.0
base_model: crh225/plumb-4b
language: [en]
tags: [gguf, llama.cpp, ollama, decision-model, calibration]
---
# Plumb-4B GGUF
[Plumb-4B](https://huggingface.co/crh225/plumb-4b) for llama.cpp and Ollama, on almost any GPU or a CPU.
| File | Quantisation | Size |
|---|---|---:|
| `plumb-4b-v5-Q8_0.gguf` | 8-bit, closest to the full model | ~4.5 GB |
| `plumb-4b-v5-Q4_K_M.gguf` | 4-bit | 2.7 GB |
Plumb-4B is a decision model, not a chat model: it answers with one option letter, and its value is in
the **probabilities of those letters**, softened by its calibration temperature **T = 2.07**.
## With Ollama
`Modelfile` in this repo holds the decision prompt as the template, so an ordinary chat request carrying
the question as JSON (evidence, criterion, lettered options) renders the right prompt:
```bash
ollama create plumb-4b -f Modelfile # next to the downloaded .gguf
```
Ask for one token with `logprobs` and `top_logprobs`, send `reasoning_effort: "none"`, read the option
letters' log-probabilities, divide by 2.07 and renormalise.
## With llama.cpp
```bash
llama-server -m plumb-4b-v5-Q8_0.gguf -c 8192 -ngl 99
pip install --no-deps "jevk5 @ git+https://github.com/allebee/jevk5@v0.2.1" # standard library only
```
```python
from jevk5 import JevK5GGUF
model = JevK5GGUF(temperature=2.07) # llama-server on :8080
model.decide("Order #7120 shows delivered to No. 17; the customer lives at No. 71.",
{"type": "choice", "instructions": "What happened to the parcel?",
"criteria": ["delivered", "misdelivered", "unknown"]})
```
Converted with llama.cpp (`convert_hf_to_gguf.py --no-mtp`, then `llama-quantize`). Apache-2.0; see the
main model's card for credits.