jbrashear commited on
Commit
7cd0c19
·
verified ·
1 Parent(s): ffe1d99

Add decide_lmstudio.py and a "Use it in LM Studio" section

Browse files

LM Studio 0.4.21 returns the top 20 label log probabilities on /v1/chat/completions with reasoning_effort "none". Checked on the 9B Q8_0: 257/260 the same as bf16, 260/260 the same as llama-server.

Files changed (2) hide show
  1. README.md +34 -2
  2. scripts/decide_lmstudio.py +143 -0
README.md CHANGED
@@ -56,8 +56,8 @@ pip install transformers # the tokenizer only, no torch
56
  python scripts/decide_gguf.py --server http://127.0.0.1:8080 --request scripts/example-request.json
57
  ```
58
 
59
- `--no-temperatures` returns the raw probabilities. We checked llama-server only. LM Studio or Ollama will
60
- load the file, but a decision needs the log probability of every option label at one position; if your
61
  runtime cannot return those, use llama-server.
62
 
63
  On `example-request.json` (jebadiah-4b-v2-Q8_0.gguf):
@@ -69,6 +69,38 @@ On `example-request.json` (jebadiah-4b-v2-Q8_0.gguf):
69
  }
70
  ```
71
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
72
  ## How it was measured
73
 
74
  Jevals PubMedQA, Banking77 (77 options) and HelpSteer2, plus Nimble: the merge check's fixed sample (seed
 
56
  python scripts/decide_gguf.py --server http://127.0.0.1:8080 --request scripts/example-request.json
57
  ```
58
 
59
+ `--no-temperatures` returns the raw probabilities. For LM Studio, see the next section. Ollama was not checked: a
60
+ decision needs the log probability of every option label at one position; if your
61
  runtime cannot return those, use llama-server.
62
 
63
  On `example-request.json` (jebadiah-4b-v2-Q8_0.gguf):
 
69
  }
70
  ```
71
 
72
+ ## Use it in LM Studio
73
+
74
+ Jeb works in LM Studio through its local server, not the chat window: chat runs with thinking on and shows
75
+ text, while a decision needs the probability of every option label. `scripts/decide_lmstudio.py` takes the same
76
+ arguments and prints the same output as `decide_gguf.py`. It sends AINode's messages to LM Studio's
77
+ `/v1/chat/completions` with thinking off (`"reasoning_effort": "none"`), where LM Studio renders the same prompt
78
+ text the llama-server path sends, and reads the option labels from the top log probabilities that come back. It
79
+ stops with an error if LM Studio's prompt length differs from the local tokenizer's.
80
+
81
+ 1. In LM Studio, search for `jebadiah-4b-v2` and download `jebadiah-4b-v2-Q8_0.gguf` from this repository.
82
+ 2. Open the **Developer** tab, start the server and load the model. Note the identifier LM Studio shows for it
83
+ (for example `jebadiah-4b-v2`).
84
+ 3. In a terminal:
85
+
86
+ ```bash
87
+ hf download frontier-infra/jebadiah-4b-v2-GGUF --include "scripts/*" "*.json" "*.jinja" "*.txt" --local-dir jebadiah-4b-v2-GGUF
88
+ pip install transformers # the tokenizer only, no torch
89
+ python jebadiah-4b-v2-GGUF/scripts/decide_lmstudio.py --model jebadiah-4b-v2 --request jebadiah-4b-v2-GGUF/scripts/example-request.json
90
+ ```
91
+
92
+ If **Require Authentication** is on in LM Studio's server settings, create a token there and
93
+ `export LM_API_TOKEN=...` first.
94
+
95
+ Tested on the 9B only: LM Studio 0.4.21 with `jebadiah-9b-v2-Q8_0`, 257 of 260 answers the same as bf16, and the
96
+ same answer as llama-server on the same file on 260 of 260. This 4B build uses the same script and the same prompt, but it has not been run in LM Studio.
97
+
98
+ **At most 20 options per question.** LM Studio returns only the top 20 log probabilities, the same cap AINode's own
99
+ route has. On the 77-option Banking77 questions the pick was still right, but the probabilities moved by up to
100
+ 0.16, so do not rely on them past 20 options.
101
+
102
+ The vision file (mmproj) that some third-party GGUF repositories ship is not needed; Jeb was not trained on images.
103
+
104
  ## How it was measured
105
 
106
  Jevals PubMedQA, Banking77 (77 options) and HelpSteer2, plus Nimble: the merge check's fixed sample (seed
scripts/decide_lmstudio.py ADDED
@@ -0,0 +1,143 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Run a Jebadiah GGUF inside LM Studio and print typed answers.
2
+
3
+ Same interface and output as decide_gguf.py, but it talks to LM Studio's local server
4
+ (Developer tab, or `lms server start`) instead of llama-server.
5
+
6
+ LM Studio applies the chat template itself, so this script sends the chat MESSAGES that AINode's
7
+ /v1/systemone builds (jebadiah_prompt.py) with reasoning off ("reasoning_effort": "none"). With
8
+ thinking off, LM Studio renders the same bytes that decide_gguf.py sends raw: the Qwen template
9
+ with an empty think block and the generation prompt. The script checks this on every call, by
10
+ comparing the prompt token count LM Studio reports with the count from the local tokenizer. The
11
+ answer is read off the log probabilities of the single-token option labels ("A", "B", ...) at the
12
+ first generated position, renormalised over those labels, with the per-type temperature from
13
+ temperatures.json applied.
14
+
15
+ LM Studio returns at most the top 20 log probabilities. That is enough for any question AINode's
16
+ own route accepts (20 options at most). On a wider question, a label outside the top 20 gets the
17
+ smallest returned value, which is an upper bound; the script reports how many labels that hit.
18
+
19
+ lms load jebadiah-9b-v2 --identifier jebadiah
20
+ python scripts/decide_lmstudio.py --model jebadiah --request scripts/example-request.json
21
+
22
+ If "Require Authentication" is on in LM Studio's server settings, pass a token with --api-key or
23
+ set LM_API_TOKEN. Needs `transformers` (the tokenizer only, no torch) and nothing else outside
24
+ the standard library.
25
+ """
26
+ from __future__ import annotations
27
+
28
+ import argparse
29
+ import json
30
+ import math
31
+ import os
32
+ import sys
33
+ import urllib.request
34
+
35
+ HERE = os.path.dirname(os.path.abspath(__file__))
36
+ sys.path.insert(0, HERE)
37
+ from ainode_prompt_verbatim import build_messages, option_label # noqa: E402
38
+ from jebadiah_prompt import CHAT_TEMPLATE_KWARGS, Renderer, answer_from_probs # noqa: E402
39
+
40
+ TOP_MAX = 20 # LM Studio's cap on top_logprobs
41
+
42
+
43
+ class MessageRenderer(Renderer):
44
+ """jebadiah_prompt.Renderer that also keeps the chat messages it rendered, so they can be
45
+ sent to a server that applies the template itself."""
46
+
47
+ def render_messages(self, state_text, question, options, letters=None):
48
+ messages = build_messages(state_text, None, question, options)
49
+ if letters is not None:
50
+ user = messages[1]["content"]
51
+ head, _, rest = user.partition("\nOPTIONS:\n")
52
+ lines = rest.split("\n")
53
+ for i in range(len(options)):
54
+ assert lines[i].startswith(f"{option_label(i)}. ")
55
+ lines[i] = f"{letters[i]}. " + lines[i][len(option_label(i)) + 2:]
56
+ messages[1]["content"] = head + "\nOPTIONS:\n" + "\n".join(lines)
57
+ self.last_messages = messages
58
+ return self.tok.apply_chat_template(messages, tokenize=False, **CHAT_TEMPLATE_KWARGS)
59
+
60
+ def render(self, state, q, order=None):
61
+ rd = super().render(state, q, order)
62
+ rd.messages = self.last_messages
63
+ rd.n_tokens = len(self.tok.encode(rd.prompt, add_special_tokens=False))
64
+ return rd
65
+
66
+
67
+ def load_renderer(tokenizer: str, max_tokens: int = 2048) -> MessageRenderer:
68
+ from transformers import AutoTokenizer
69
+ tok = AutoTokenizer.from_pretrained(tokenizer)
70
+ if tok.pad_token_id is None:
71
+ tok.pad_token = tok.eos_token
72
+ return MessageRenderer(tok, max_tokens)
73
+
74
+
75
+ def read_temperatures(path: str | None) -> dict:
76
+ if not path or not os.path.exists(path):
77
+ return {}
78
+ return {k: float(v) for k, v in json.load(open(path))["temperatures"].items()}
79
+
80
+
81
+ def post(server: str, path: str, body: dict, api_key: str | None = None, timeout: float = 900) -> dict:
82
+ headers = {"Content-Type": "application/json"}
83
+ if api_key:
84
+ headers["Authorization"] = "Bearer " + api_key
85
+ req = urllib.request.Request(server.rstrip("/") + path, data=json.dumps(body).encode(), headers=headers)
86
+ with urllib.request.urlopen(req, timeout=timeout) as r:
87
+ return json.loads(r.read())
88
+
89
+
90
+ def label_logprobs(server: str, model: str, rd, api_key: str | None = None) -> tuple[list[float], int, int]:
91
+ """Log probabilities of each option label at the first answer position, from LM Studio's
92
+ OpenAI-compatible chat endpoint with thinking off. A label outside the returned top 20 gets the
93
+ smallest returned value (an upper bound). Returns (logprobs, labels not returned, prompt tokens
94
+ LM Studio counted)."""
95
+ r = post(server, "/v1/chat/completions", {
96
+ "model": model, "messages": rd.messages, "max_tokens": 1, "temperature": 0,
97
+ "logprobs": True, "top_logprobs": TOP_MAX, "reasoning_effort": "none", "stream": False}, api_key)
98
+ top = r["choices"][0]["logprobs"]["content"][0]["top_logprobs"]
99
+ lp = {}
100
+ for t in top:
101
+ lp.setdefault(t["token"], t["logprob"]) # exact token text: "A", not " A"
102
+ floor = min(lp.values())
103
+ return [lp.get(L, floor) for L in rd.letters], sum(1 for L in rd.letters if L not in lp), r["usage"]["prompt_tokens"]
104
+
105
+
106
+ def option_probs(logprobs: list[float], temperature: float = 1.0) -> list[float]:
107
+ """Softmax over the labels only, after dividing by the temperature (see decide_gguf.py)."""
108
+ z = [x / temperature for x in logprobs]
109
+ m = max(z)
110
+ e = [math.exp(x - m) for x in z]
111
+ s = sum(e)
112
+ return [x / s for x in e]
113
+
114
+
115
+ def main():
116
+ ap = argparse.ArgumentParser()
117
+ ap.add_argument("--server", default="http://127.0.0.1:1234", help="LM Studio's local server")
118
+ ap.add_argument("--model", required=True, help="the model identifier LM Studio shows for the loaded Jebadiah")
119
+ ap.add_argument("--api-key", default=os.environ.get("LM_API_TOKEN"), help="LM Studio API token, if auth is on")
120
+ ap.add_argument("--request", required=True, help="JSON file: {state, questions: {id: {type, instructions, criteria}}}")
121
+ ap.add_argument("--tokenizer", default=os.path.dirname(HERE), help="folder or Hub repo id with the tokenizer and chat template")
122
+ ap.add_argument("--temperatures", default=os.path.join(os.path.dirname(HERE), "temperatures.json"))
123
+ ap.add_argument("--no-temperatures", action="store_true", help="raw probabilities, as the served route returns today")
124
+ a = ap.parse_args()
125
+ req = json.load(open(a.request))
126
+ renderer = load_renderer(a.tokenizer)
127
+ temps = {} if a.no_temperatures else read_temperatures(a.temperatures)
128
+ out = {"temperatures_applied": temps, "answers": {}}
129
+ for qid, q in req["questions"].items():
130
+ rd = renderer.render(req["state"], q)
131
+ lps, missing, n_prompt = label_logprobs(a.server, a.model, rd, a.api_key)
132
+ if n_prompt != rd.n_tokens:
133
+ sys.exit(f"{qid}: LM Studio rendered {n_prompt} prompt tokens, the local template {rd.n_tokens}. "
134
+ "Is thinking off, and is the loaded model's prompt template the GGUF's own?")
135
+ if missing:
136
+ print(f"{qid}: {missing} of {len(rd.letters)} labels were outside LM Studio's top {TOP_MAX}", file=sys.stderr)
137
+ probs = option_probs(lps, float(temps.get(q["type"], 1.0)))
138
+ out["answers"][qid] = answer_from_probs(q, rd.keys, probs)
139
+ print(json.dumps(out, indent=1))
140
+
141
+
142
+ if __name__ == "__main__":
143
+ main()