GGUF
English
llama.cpp
decision-model
system-one
typed-decisions
calibrated-probabilities
ainode
conversational
Instructions to use frontier-infra/jebadiah-4b-v2-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use frontier-infra/jebadiah-4b-v2-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf frontier-infra/jebadiah-4b-v2-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf frontier-infra/jebadiah-4b-v2-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf frontier-infra/jebadiah-4b-v2-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf frontier-infra/jebadiah-4b-v2-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf frontier-infra/jebadiah-4b-v2-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf frontier-infra/jebadiah-4b-v2-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf frontier-infra/jebadiah-4b-v2-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf frontier-infra/jebadiah-4b-v2-GGUF:Q4_K_M
Use Docker
docker model run hf.co/frontier-infra/jebadiah-4b-v2-GGUF:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use frontier-infra/jebadiah-4b-v2-GGUF with Ollama:
ollama run hf.co/frontier-infra/jebadiah-4b-v2-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use frontier-infra/jebadiah-4b-v2-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf frontier-infra/jebadiah-4b-v2-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "frontier-infra/jebadiah-4b-v2-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use frontier-infra/jebadiah-4b-v2-GGUF with Docker Model Runner:
docker model run hf.co/frontier-infra/jebadiah-4b-v2-GGUF:Q4_K_M
- Lemonade
How to use frontier-infra/jebadiah-4b-v2-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull frontier-infra/jebadiah-4b-v2-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.jebadiah-4b-v2-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use frontier-infra/jebadiah-4b-v2-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf frontier-infra/jebadiah-4b-v2-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default frontier-infra/jebadiah-4b-v2-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use frontier-infra/jebadiah-4b-v2-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf frontier-infra/jebadiah-4b-v2-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "frontier-infra/jebadiah-4b-v2-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Add decide_lmstudio.py and a "Use it in LM Studio" section
Browse filesLM Studio 0.4.21 returns the top 20 label log probabilities on /v1/chat/completions with reasoning_effort "none". Checked on the 9B Q8_0: 257/260 the same as bf16, 260/260 the same as llama-server.
- README.md +34 -2
- scripts/decide_lmstudio.py +143 -0
README.md
CHANGED
|
@@ -56,8 +56,8 @@ pip install transformers # the tokenizer only, no torch
|
|
| 56 |
python scripts/decide_gguf.py --server http://127.0.0.1:8080 --request scripts/example-request.json
|
| 57 |
```
|
| 58 |
|
| 59 |
-
`--no-temperatures` returns the raw probabilities.
|
| 60 |
-
|
| 61 |
runtime cannot return those, use llama-server.
|
| 62 |
|
| 63 |
On `example-request.json` (jebadiah-4b-v2-Q8_0.gguf):
|
|
@@ -69,6 +69,38 @@ On `example-request.json` (jebadiah-4b-v2-Q8_0.gguf):
|
|
| 69 |
}
|
| 70 |
```
|
| 71 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 72 |
## How it was measured
|
| 73 |
|
| 74 |
Jevals PubMedQA, Banking77 (77 options) and HelpSteer2, plus Nimble: the merge check's fixed sample (seed
|
|
|
|
| 56 |
python scripts/decide_gguf.py --server http://127.0.0.1:8080 --request scripts/example-request.json
|
| 57 |
```
|
| 58 |
|
| 59 |
+
`--no-temperatures` returns the raw probabilities. For LM Studio, see the next section. Ollama was not checked: a
|
| 60 |
+
decision needs the log probability of every option label at one position; if your
|
| 61 |
runtime cannot return those, use llama-server.
|
| 62 |
|
| 63 |
On `example-request.json` (jebadiah-4b-v2-Q8_0.gguf):
|
|
|
|
| 69 |
}
|
| 70 |
```
|
| 71 |
|
| 72 |
+
## Use it in LM Studio
|
| 73 |
+
|
| 74 |
+
Jeb works in LM Studio through its local server, not the chat window: chat runs with thinking on and shows
|
| 75 |
+
text, while a decision needs the probability of every option label. `scripts/decide_lmstudio.py` takes the same
|
| 76 |
+
arguments and prints the same output as `decide_gguf.py`. It sends AINode's messages to LM Studio's
|
| 77 |
+
`/v1/chat/completions` with thinking off (`"reasoning_effort": "none"`), where LM Studio renders the same prompt
|
| 78 |
+
text the llama-server path sends, and reads the option labels from the top log probabilities that come back. It
|
| 79 |
+
stops with an error if LM Studio's prompt length differs from the local tokenizer's.
|
| 80 |
+
|
| 81 |
+
1. In LM Studio, search for `jebadiah-4b-v2` and download `jebadiah-4b-v2-Q8_0.gguf` from this repository.
|
| 82 |
+
2. Open the **Developer** tab, start the server and load the model. Note the identifier LM Studio shows for it
|
| 83 |
+
(for example `jebadiah-4b-v2`).
|
| 84 |
+
3. In a terminal:
|
| 85 |
+
|
| 86 |
+
```bash
|
| 87 |
+
hf download frontier-infra/jebadiah-4b-v2-GGUF --include "scripts/*" "*.json" "*.jinja" "*.txt" --local-dir jebadiah-4b-v2-GGUF
|
| 88 |
+
pip install transformers # the tokenizer only, no torch
|
| 89 |
+
python jebadiah-4b-v2-GGUF/scripts/decide_lmstudio.py --model jebadiah-4b-v2 --request jebadiah-4b-v2-GGUF/scripts/example-request.json
|
| 90 |
+
```
|
| 91 |
+
|
| 92 |
+
If **Require Authentication** is on in LM Studio's server settings, create a token there and
|
| 93 |
+
`export LM_API_TOKEN=...` first.
|
| 94 |
+
|
| 95 |
+
Tested on the 9B only: LM Studio 0.4.21 with `jebadiah-9b-v2-Q8_0`, 257 of 260 answers the same as bf16, and the
|
| 96 |
+
same answer as llama-server on the same file on 260 of 260. This 4B build uses the same script and the same prompt, but it has not been run in LM Studio.
|
| 97 |
+
|
| 98 |
+
**At most 20 options per question.** LM Studio returns only the top 20 log probabilities, the same cap AINode's own
|
| 99 |
+
route has. On the 77-option Banking77 questions the pick was still right, but the probabilities moved by up to
|
| 100 |
+
0.16, so do not rely on them past 20 options.
|
| 101 |
+
|
| 102 |
+
The vision file (mmproj) that some third-party GGUF repositories ship is not needed; Jeb was not trained on images.
|
| 103 |
+
|
| 104 |
## How it was measured
|
| 105 |
|
| 106 |
Jevals PubMedQA, Banking77 (77 options) and HelpSteer2, plus Nimble: the merge check's fixed sample (seed
|
scripts/decide_lmstudio.py
ADDED
|
@@ -0,0 +1,143 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Run a Jebadiah GGUF inside LM Studio and print typed answers.
|
| 2 |
+
|
| 3 |
+
Same interface and output as decide_gguf.py, but it talks to LM Studio's local server
|
| 4 |
+
(Developer tab, or `lms server start`) instead of llama-server.
|
| 5 |
+
|
| 6 |
+
LM Studio applies the chat template itself, so this script sends the chat MESSAGES that AINode's
|
| 7 |
+
/v1/systemone builds (jebadiah_prompt.py) with reasoning off ("reasoning_effort": "none"). With
|
| 8 |
+
thinking off, LM Studio renders the same bytes that decide_gguf.py sends raw: the Qwen template
|
| 9 |
+
with an empty think block and the generation prompt. The script checks this on every call, by
|
| 10 |
+
comparing the prompt token count LM Studio reports with the count from the local tokenizer. The
|
| 11 |
+
answer is read off the log probabilities of the single-token option labels ("A", "B", ...) at the
|
| 12 |
+
first generated position, renormalised over those labels, with the per-type temperature from
|
| 13 |
+
temperatures.json applied.
|
| 14 |
+
|
| 15 |
+
LM Studio returns at most the top 20 log probabilities. That is enough for any question AINode's
|
| 16 |
+
own route accepts (20 options at most). On a wider question, a label outside the top 20 gets the
|
| 17 |
+
smallest returned value, which is an upper bound; the script reports how many labels that hit.
|
| 18 |
+
|
| 19 |
+
lms load jebadiah-9b-v2 --identifier jebadiah
|
| 20 |
+
python scripts/decide_lmstudio.py --model jebadiah --request scripts/example-request.json
|
| 21 |
+
|
| 22 |
+
If "Require Authentication" is on in LM Studio's server settings, pass a token with --api-key or
|
| 23 |
+
set LM_API_TOKEN. Needs `transformers` (the tokenizer only, no torch) and nothing else outside
|
| 24 |
+
the standard library.
|
| 25 |
+
"""
|
| 26 |
+
from __future__ import annotations
|
| 27 |
+
|
| 28 |
+
import argparse
|
| 29 |
+
import json
|
| 30 |
+
import math
|
| 31 |
+
import os
|
| 32 |
+
import sys
|
| 33 |
+
import urllib.request
|
| 34 |
+
|
| 35 |
+
HERE = os.path.dirname(os.path.abspath(__file__))
|
| 36 |
+
sys.path.insert(0, HERE)
|
| 37 |
+
from ainode_prompt_verbatim import build_messages, option_label # noqa: E402
|
| 38 |
+
from jebadiah_prompt import CHAT_TEMPLATE_KWARGS, Renderer, answer_from_probs # noqa: E402
|
| 39 |
+
|
| 40 |
+
TOP_MAX = 20 # LM Studio's cap on top_logprobs
|
| 41 |
+
|
| 42 |
+
|
| 43 |
+
class MessageRenderer(Renderer):
|
| 44 |
+
"""jebadiah_prompt.Renderer that also keeps the chat messages it rendered, so they can be
|
| 45 |
+
sent to a server that applies the template itself."""
|
| 46 |
+
|
| 47 |
+
def render_messages(self, state_text, question, options, letters=None):
|
| 48 |
+
messages = build_messages(state_text, None, question, options)
|
| 49 |
+
if letters is not None:
|
| 50 |
+
user = messages[1]["content"]
|
| 51 |
+
head, _, rest = user.partition("\nOPTIONS:\n")
|
| 52 |
+
lines = rest.split("\n")
|
| 53 |
+
for i in range(len(options)):
|
| 54 |
+
assert lines[i].startswith(f"{option_label(i)}. ")
|
| 55 |
+
lines[i] = f"{letters[i]}. " + lines[i][len(option_label(i)) + 2:]
|
| 56 |
+
messages[1]["content"] = head + "\nOPTIONS:\n" + "\n".join(lines)
|
| 57 |
+
self.last_messages = messages
|
| 58 |
+
return self.tok.apply_chat_template(messages, tokenize=False, **CHAT_TEMPLATE_KWARGS)
|
| 59 |
+
|
| 60 |
+
def render(self, state, q, order=None):
|
| 61 |
+
rd = super().render(state, q, order)
|
| 62 |
+
rd.messages = self.last_messages
|
| 63 |
+
rd.n_tokens = len(self.tok.encode(rd.prompt, add_special_tokens=False))
|
| 64 |
+
return rd
|
| 65 |
+
|
| 66 |
+
|
| 67 |
+
def load_renderer(tokenizer: str, max_tokens: int = 2048) -> MessageRenderer:
|
| 68 |
+
from transformers import AutoTokenizer
|
| 69 |
+
tok = AutoTokenizer.from_pretrained(tokenizer)
|
| 70 |
+
if tok.pad_token_id is None:
|
| 71 |
+
tok.pad_token = tok.eos_token
|
| 72 |
+
return MessageRenderer(tok, max_tokens)
|
| 73 |
+
|
| 74 |
+
|
| 75 |
+
def read_temperatures(path: str | None) -> dict:
|
| 76 |
+
if not path or not os.path.exists(path):
|
| 77 |
+
return {}
|
| 78 |
+
return {k: float(v) for k, v in json.load(open(path))["temperatures"].items()}
|
| 79 |
+
|
| 80 |
+
|
| 81 |
+
def post(server: str, path: str, body: dict, api_key: str | None = None, timeout: float = 900) -> dict:
|
| 82 |
+
headers = {"Content-Type": "application/json"}
|
| 83 |
+
if api_key:
|
| 84 |
+
headers["Authorization"] = "Bearer " + api_key
|
| 85 |
+
req = urllib.request.Request(server.rstrip("/") + path, data=json.dumps(body).encode(), headers=headers)
|
| 86 |
+
with urllib.request.urlopen(req, timeout=timeout) as r:
|
| 87 |
+
return json.loads(r.read())
|
| 88 |
+
|
| 89 |
+
|
| 90 |
+
def label_logprobs(server: str, model: str, rd, api_key: str | None = None) -> tuple[list[float], int, int]:
|
| 91 |
+
"""Log probabilities of each option label at the first answer position, from LM Studio's
|
| 92 |
+
OpenAI-compatible chat endpoint with thinking off. A label outside the returned top 20 gets the
|
| 93 |
+
smallest returned value (an upper bound). Returns (logprobs, labels not returned, prompt tokens
|
| 94 |
+
LM Studio counted)."""
|
| 95 |
+
r = post(server, "/v1/chat/completions", {
|
| 96 |
+
"model": model, "messages": rd.messages, "max_tokens": 1, "temperature": 0,
|
| 97 |
+
"logprobs": True, "top_logprobs": TOP_MAX, "reasoning_effort": "none", "stream": False}, api_key)
|
| 98 |
+
top = r["choices"][0]["logprobs"]["content"][0]["top_logprobs"]
|
| 99 |
+
lp = {}
|
| 100 |
+
for t in top:
|
| 101 |
+
lp.setdefault(t["token"], t["logprob"]) # exact token text: "A", not " A"
|
| 102 |
+
floor = min(lp.values())
|
| 103 |
+
return [lp.get(L, floor) for L in rd.letters], sum(1 for L in rd.letters if L not in lp), r["usage"]["prompt_tokens"]
|
| 104 |
+
|
| 105 |
+
|
| 106 |
+
def option_probs(logprobs: list[float], temperature: float = 1.0) -> list[float]:
|
| 107 |
+
"""Softmax over the labels only, after dividing by the temperature (see decide_gguf.py)."""
|
| 108 |
+
z = [x / temperature for x in logprobs]
|
| 109 |
+
m = max(z)
|
| 110 |
+
e = [math.exp(x - m) for x in z]
|
| 111 |
+
s = sum(e)
|
| 112 |
+
return [x / s for x in e]
|
| 113 |
+
|
| 114 |
+
|
| 115 |
+
def main():
|
| 116 |
+
ap = argparse.ArgumentParser()
|
| 117 |
+
ap.add_argument("--server", default="http://127.0.0.1:1234", help="LM Studio's local server")
|
| 118 |
+
ap.add_argument("--model", required=True, help="the model identifier LM Studio shows for the loaded Jebadiah")
|
| 119 |
+
ap.add_argument("--api-key", default=os.environ.get("LM_API_TOKEN"), help="LM Studio API token, if auth is on")
|
| 120 |
+
ap.add_argument("--request", required=True, help="JSON file: {state, questions: {id: {type, instructions, criteria}}}")
|
| 121 |
+
ap.add_argument("--tokenizer", default=os.path.dirname(HERE), help="folder or Hub repo id with the tokenizer and chat template")
|
| 122 |
+
ap.add_argument("--temperatures", default=os.path.join(os.path.dirname(HERE), "temperatures.json"))
|
| 123 |
+
ap.add_argument("--no-temperatures", action="store_true", help="raw probabilities, as the served route returns today")
|
| 124 |
+
a = ap.parse_args()
|
| 125 |
+
req = json.load(open(a.request))
|
| 126 |
+
renderer = load_renderer(a.tokenizer)
|
| 127 |
+
temps = {} if a.no_temperatures else read_temperatures(a.temperatures)
|
| 128 |
+
out = {"temperatures_applied": temps, "answers": {}}
|
| 129 |
+
for qid, q in req["questions"].items():
|
| 130 |
+
rd = renderer.render(req["state"], q)
|
| 131 |
+
lps, missing, n_prompt = label_logprobs(a.server, a.model, rd, a.api_key)
|
| 132 |
+
if n_prompt != rd.n_tokens:
|
| 133 |
+
sys.exit(f"{qid}: LM Studio rendered {n_prompt} prompt tokens, the local template {rd.n_tokens}. "
|
| 134 |
+
"Is thinking off, and is the loaded model's prompt template the GGUF's own?")
|
| 135 |
+
if missing:
|
| 136 |
+
print(f"{qid}: {missing} of {len(rd.letters)} labels were outside LM Studio's top {TOP_MAX}", file=sys.stderr)
|
| 137 |
+
probs = option_probs(lps, float(temps.get(q["type"], 1.0)))
|
| 138 |
+
out["answers"][qid] = answer_from_probs(q, rd.keys, probs)
|
| 139 |
+
print(json.dumps(out, indent=1))
|
| 140 |
+
|
| 141 |
+
|
| 142 |
+
if __name__ == "__main__":
|
| 143 |
+
main()
|