jbrashear commited on
Commit
ffe1d99
·
verified ·
1 Parent(s): 0fda8e6

Model card

Browse files
Files changed (1) hide show
  1. README.md +80 -0
README.md ADDED
@@ -0,0 +1,80 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: frontier-infra/jebadiah-4b-v2
4
+ base_model_relation: quantized
5
+ library_name: gguf
6
+ language:
7
+ - en
8
+ tags:
9
+ - gguf
10
+ - llama.cpp
11
+ - decision-model
12
+ - system-one
13
+ - typed-decisions
14
+ - calibrated-probabilities
15
+ - ainode
16
+ ---
17
+
18
+ # Jebadiah 4B v2 GGUF
19
+
20
+ GGUF builds of [Jebadiah 4B v2](https://huggingface.co/frontier-infra/jebadiah-4b-v2) for [llama.cpp](https://github.com/ggml-org/llama.cpp), which runs on NVIDIA, AMD and
21
+ Apple GPUs and on plain CPUs. Jebadiah answers a typed question (choice, noul or score) with a probability for
22
+ every option, read from one forward pass. Nothing is generated. Code, trainer and evals:
23
+ [getainode/jebadiah](https://github.com/getainode/jebadiah).
24
+
25
+ ## Files
26
+
27
+ Every file was checked on the 260 held-out questions the merged weights were checked on, and compared with
28
+ the merged bf16 weights and with the training run's own eval records.
29
+
30
+ | File | Size | Same answer as bf16 | Same as the run | choice + noul | score | Prob. diff median / max |
31
+ |---|---:|---:|---:|---:|---:|---:|
32
+ | `jebadiah-4b-v2-Q8_0.gguf` | 4.6 GB | **256 / 260** | 257 / 260 | 171 / 173 | 86 / 87 | 0.004 / 0.053 |
33
+ | `jebadiah-4b-v2-Q5_K_M.gguf` | 3.2 GB | **239 / 260** | 240 / 260 | 166 / 173 | 74 / 87 | 0.016 / 0.119 |
34
+ | `jebadiah-4b-v2-Q4_K_M.gguf` | 2.8 GB | **236 / 260** | 237 / 260 | 164 / 173 | 73 / 87 | 0.022 / 0.355 |
35
+ | *bf16 weights* | | | 259 / 260 | 172 / 173 | 87 / 87 | 0.002 / 0.019 |
36
+
37
+ Which one: `Q8_0` if it fits (it changed 4 answers here); `Q4_K_M` when memory is short. A file
38
+ needs about its own size in GPU or unified memory, plus about 1 GB for a 4k context.
39
+
40
+ **`Q5_K_M` changes 21 of 260 answers** against bf16 (13 on score questions) and moves probabilities more (median 0.016, max 0.12). Use it only when a larger build does not fit.
41
+ **`Q4_K_M` changes 24 of 260 answers** against bf16 (14 on score questions) and moves probabilities more (median 0.022, max 0.36). Use it only when a larger build does not fit.
42
+
43
+ ## Run it
44
+
45
+ The answer is the log probability of each option label ("A", "B", ...) at the answer position, which
46
+ llama-server's `/completion` returns. The script renders the prompt exactly as AINode does, sends the raw
47
+ text (so the server's own chat template is never used), renormalises over the labels and applies
48
+ `temperatures.json` (choice 1.1167, noul 1.3319, score 1.1974). You need a llama.cpp that knows the `qwen35` architecture: we checked
49
+ v0.5.0 (older builds refuse the file).
50
+
51
+ ```bash
52
+ hf download frontier-infra/jebadiah-4b-v2-GGUF --include "*Q8_0.gguf" "scripts/*" "*.json" "*.jinja" "*.txt" --local-dir jebadiah-4b-v2-GGUF
53
+ cd jebadiah-4b-v2-GGUF
54
+ llama-server -m jebadiah-4b-v2-Q8_0.gguf -c 4096 -np 1 --port 8080
55
+ pip install transformers # the tokenizer only, no torch
56
+ python scripts/decide_gguf.py --server http://127.0.0.1:8080 --request scripts/example-request.json
57
+ ```
58
+
59
+ `--no-temperatures` returns the raw probabilities. We checked llama-server only. LM Studio or Ollama will
60
+ load the file, but a decision needs the log probability of every option label at one position; if your
61
+ runtime cannot return those, use llama-server.
62
+
63
+ On `example-request.json` (jebadiah-4b-v2-Q8_0.gguf):
64
+
65
+ ```json
66
+ {
67
+ "route": {"type": "choice", "choice": "billing", "confidence": 0.397366, "probabilities": {"billing": 0.598244, "support": 0.082792, "sales": 0.318963}},
68
+ "urgent": {"type": "noul", "noul": 0.153519}
69
+ }
70
+ ```
71
+
72
+ ## How it was measured
73
+
74
+ Jevals PubMedQA, Banking77 (77 options) and HelpSteer2, plus Nimble: the merge check's fixed sample (seed
75
+ 20260925), the run's option order and temperatures. "Same answer" is the top option; "prob. diff" is the
76
+ largest change on any option against the run's CUDA record. llama-server ran on Metal (M3 Ultra) with the same tokens as the Python renderer on every prompt. Records: `eval/agreement-*.json`.
77
+
78
+ ## License
79
+
80
+ Apache-2.0, as the base model. Made in Texas.