jbrashear commited on
Commit
4aa880b
·
verified ·
1 Parent(s): 331bb66

Add scripts/decide_ollama.py and a Use it in Ollama section; fix the hf download commands

Browse files
Files changed (2) hide show
  1. README.md +27 -5
  2. scripts/decide_ollama.py +84 -0
README.md CHANGED
@@ -55,16 +55,14 @@ text (so the server's own chat template is never used), renormalises over the la
55
  v0.5.0 (older builds refuse the file).
56
 
57
  ```bash
58
- hf download frontier-infra/jebadiah-27b-GGUF --include "*Q8_0.gguf" "scripts/*" "*.json" "*.jinja" "*.txt" --local-dir jebadiah-27b-GGUF
59
  cd jebadiah-27b-GGUF
60
  llama-server -m jebadiah-27b-Q8_0.gguf -c 4096 -np 1 --port 8080
61
  pip install transformers # the tokenizer only, no torch
62
  python scripts/decide_gguf.py --server http://127.0.0.1:8080 --request scripts/example-request.json
63
  ```
64
 
65
- `--no-temperatures` returns the raw probabilities. For LM Studio, see the next section. Ollama was not checked: a
66
- decision needs the log probability of every option label at one position; if your
67
- runtime cannot return those, use llama-server.
68
 
69
  On `example-request.json` (jebadiah-27b-Q8_0.gguf):
70
 
@@ -75,6 +73,30 @@ On `example-request.json` (jebadiah-27b-Q8_0.gguf):
75
  }
76
  ```
77
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
78
  ## Use it in LM Studio
79
 
80
  Jeb works in LM Studio through its local server, not the chat window: chat runs with thinking on and shows
@@ -90,7 +112,7 @@ stops with an error if LM Studio's prompt length differs from the local tokenize
90
  3. In a terminal:
91
 
92
  ```bash
93
- hf download frontier-infra/jebadiah-27b-GGUF --include "scripts/*" "*.json" "*.jinja" "*.txt" --local-dir jebadiah-27b-GGUF
94
  pip install transformers # the tokenizer only, no torch
95
  python jebadiah-27b-GGUF/scripts/decide_lmstudio.py --model jebadiah-27b --request jebadiah-27b-GGUF/scripts/example-request.json
96
  ```
 
55
  v0.5.0 (older builds refuse the file).
56
 
57
  ```bash
58
+ hf download frontier-infra/jebadiah-27b-GGUF --include "*Q8_0.gguf" --include "scripts/*" --include "*.json" --include "*.jinja" --include "*.txt" --local-dir jebadiah-27b-GGUF
59
  cd jebadiah-27b-GGUF
60
  llama-server -m jebadiah-27b-Q8_0.gguf -c 4096 -np 1 --port 8080
61
  pip install transformers # the tokenizer only, no torch
62
  python scripts/decide_gguf.py --server http://127.0.0.1:8080 --request scripts/example-request.json
63
  ```
64
 
65
+ `--no-temperatures` returns the raw probabilities. For Ollama and LM Studio, see the next sections.
 
 
66
 
67
  On `example-request.json` (jebadiah-27b-Q8_0.gguf):
68
 
 
73
  }
74
  ```
75
 
76
+ ## Use it in Ollama
77
+
78
+ Jeb works in Ollama through its API, not `ollama run`: the chat window shows text, while a decision needs the
79
+ probability of every option label with thinking off. `scripts/decide_ollama.py` takes the same arguments and prints
80
+ the same output as `decide_gguf.py`. It sends the rendered prompt to `/api/generate` with `"raw": true` (so Ollama's
81
+ own template is never used) and `"think": false`, asks for one token with the top 20 log probabilities, and stops
82
+ with an error if Ollama's prompt token count differs from the local tokenizer's.
83
+
84
+ ```bash
85
+ ollama pull hf.co/frontier-infra/jebadiah-27b-GGUF:Q8_0
86
+ hf download frontier-infra/jebadiah-27b-GGUF --include "scripts/*" --include "*.json" --include "*.jinja" --include "*.txt" --local-dir jebadiah-27b-GGUF
87
+ pip install transformers # the tokenizer only, no torch
88
+ python jebadiah-27b-GGUF/scripts/decide_ollama.py --model hf.co/frontier-infra/jebadiah-27b-GGUF:Q8_0 --request jebadiah-27b-GGUF/scripts/example-request.json
89
+ ```
90
+
91
+ Tested with Ollama 0.34.4 and `jebadiah-27b-Q8_0` pulled from this repository: 260 of 260 answers the same as bf16,
92
+ and the same answer as llama-server on the same file on 260 of 260. On the questions with 20 options or
93
+ fewer the probabilities match llama-server to 0.0005 at most, so `example-request.json` prints the numbers above.
94
+
95
+ **At most 20 options per question.** Ollama returns at most the top 20 log probabilities, the same cap as LM Studio
96
+ and AINode's own route. On the 70 Banking77 questions (77 options) the pick matched llama-server on
97
+ 70 of 70, but the probabilities moved, so do not rely on them past 20
98
+ options.
99
+
100
  ## Use it in LM Studio
101
 
102
  Jeb works in LM Studio through its local server, not the chat window: chat runs with thinking on and shows
 
112
  3. In a terminal:
113
 
114
  ```bash
115
+ hf download frontier-infra/jebadiah-27b-GGUF --include "scripts/*" --include "*.json" --include "*.jinja" --include "*.txt" --local-dir jebadiah-27b-GGUF
116
  pip install transformers # the tokenizer only, no torch
117
  python jebadiah-27b-GGUF/scripts/decide_lmstudio.py --model jebadiah-27b --request jebadiah-27b-GGUF/scripts/example-request.json
118
  ```
scripts/decide_ollama.py ADDED
@@ -0,0 +1,84 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Run a Jebadiah GGUF inside Ollama and print typed answers.
2
+
3
+ Same interface and output as decide_gguf.py, but it talks to Ollama's /api/generate.
4
+
5
+ The prompt is rendered in Python exactly as AINode's /v1/systemone does (jebadiah_prompt.py, the chat
6
+ template with thinking off) and sent with "raw": true, so Ollama's own template never touches it. One
7
+ token is requested with logprobs on and top_logprobs 20 (Ollama's cap). The answer is read off the log
8
+ probabilities of the option labels at that position by exact token text, renormalised over the labels,
9
+ with the per-type temperature from temperatures.json applied. Ollama reports full-vocabulary log
10
+ probabilities from the raw logits (before temperature), which is what the temperatures were fitted on.
11
+
12
+ Every call checks that Ollama counted the same number of prompt tokens as the local tokenizer.
13
+
14
+ ollama pull hf.co/frontier-infra/jebadiah-9b-v2-GGUF:Q8_0
15
+ python scripts/decide_ollama.py --model hf.co/frontier-infra/jebadiah-9b-v2-GGUF:Q8_0 --request scripts/example-request.json
16
+
17
+ At most 20 options per question for exact probabilities: a label outside Ollama's top 20 gets the
18
+ smallest returned value (an upper bound), and the script reports how many labels that hit. Needs
19
+ `transformers` (the tokenizer only, no torch) and nothing else outside the standard library.
20
+ """
21
+ from __future__ import annotations
22
+
23
+ import argparse
24
+ import json
25
+ import os
26
+ import sys
27
+ import urllib.request
28
+
29
+ HERE = os.path.dirname(os.path.abspath(__file__))
30
+ sys.path.insert(0, HERE)
31
+ from decide_gguf import load_renderer, option_probs, read_temperatures # noqa: E402
32
+ from jebadiah_prompt import answer_from_probs # noqa: E402
33
+
34
+ TOP_MAX = 20 # Ollama rejects top_logprobs above 20 (server/routes.go)
35
+
36
+
37
+ def post(server: str, path: str, body: dict, timeout: float = 900) -> dict:
38
+ req = urllib.request.Request(server.rstrip("/") + path, data=json.dumps(body).encode(),
39
+ headers={"Content-Type": "application/json"})
40
+ with urllib.request.urlopen(req, timeout=timeout) as r:
41
+ return json.loads(r.read())
42
+
43
+
44
+ def label_logprobs(server: str, model: str, rd) -> tuple[list[float], int, int]:
45
+ r = post(server, "/api/generate", {
46
+ "model": model, "prompt": rd.prompt, "raw": True, "stream": False, "think": False,
47
+ "logprobs": True, "top_logprobs": TOP_MAX,
48
+ "options": {"num_predict": 1, "temperature": 0}})
49
+ top = r["logprobs"][0]["top_logprobs"]
50
+ lp = {}
51
+ for t in top:
52
+ lp.setdefault(t["token"], t["logprob"])
53
+ floor = min(lp.values())
54
+ return [lp.get(L, floor) for L in rd.letters], sum(1 for L in rd.letters if L not in lp), r["prompt_eval_count"]
55
+
56
+
57
+ def main():
58
+ ap = argparse.ArgumentParser()
59
+ ap.add_argument("--server", default="http://127.0.0.1:11434", help="Ollama's server")
60
+ ap.add_argument("--model", required=True, help="the Ollama model name for the Jebadiah GGUF")
61
+ ap.add_argument("--request", required=True)
62
+ ap.add_argument("--tokenizer", default=os.path.dirname(HERE))
63
+ ap.add_argument("--temperatures", default=os.path.join(os.path.dirname(HERE), "temperatures.json"))
64
+ ap.add_argument("--no-temperatures", action="store_true")
65
+ a = ap.parse_args()
66
+ req = json.load(open(a.request))
67
+ renderer = load_renderer(a.tokenizer)
68
+ temps = {} if a.no_temperatures else read_temperatures(a.temperatures)
69
+ out = {"temperatures_applied": temps, "answers": {}}
70
+ for qid, q in req["questions"].items():
71
+ rd = renderer.render(req["state"], q)
72
+ n_local = len(renderer.tok.encode(rd.prompt, add_special_tokens=False))
73
+ lps, missing, n_prompt = label_logprobs(a.server, a.model, rd)
74
+ if n_prompt != n_local:
75
+ sys.exit(f"{qid}: Ollama counted {n_prompt} prompt tokens, the local tokenizer {n_local}.")
76
+ if missing:
77
+ print(f"{qid}: {missing} of {len(rd.letters)} labels were outside Ollama's top {TOP_MAX}", file=sys.stderr)
78
+ probs = option_probs(lps, float(temps.get(q["type"], 1.0)))
79
+ out["answers"][qid] = answer_from_probs(q, rd.keys, probs)
80
+ print(json.dumps(out, indent=1))
81
+
82
+
83
+ if __name__ == "__main__":
84
+ main()