jbrashear commited on
Commit
071adba
·
verified ·
1 Parent(s): 5f73fe5

Model card: results links and community quants

Browse files
Files changed (1) hide show
  1. README.md +19 -0
README.md CHANGED
@@ -73,6 +73,16 @@ This repository holds full bf16 weights: the v2 LoRA merged into `Qwen/Qwen3.5-9
73
 
74
  Made in Texas.
75
 
 
 
 
 
 
 
 
 
 
 
76
  ## Results
77
 
78
  Measured by us with AINode's bench, one logit read per question, the same rendered prompt for every model. Accuracy is the share of questions whose top label is the human label. Headline is the macro over the zero-shot public sets. These are our numbers on the public suites, not rows on the Jevals board.
@@ -148,6 +158,15 @@ Then send the same body to `https://<your-ainode>/v1/systemone` with `"model": "
148
  - **GGUF** for llama.cpp: [jebadiah-9b-v2-GGUF](https://huggingface.co/frontier-infra/jebadiah-9b-v2-GGUF). Same answer as these weights on 257 of 260 held-out questions (Q8_0) and 240 of 260 (Q4_K_M).
149
  - **MLX** for Apple silicon: [jebadiah-9b-v2-MLX](https://huggingface.co/frontier-infra/jebadiah-9b-v2-MLX). Same answer on 256 of 260 (8-bit; the 4-bit build fell under 90% and was not published).
150
 
 
 
 
 
 
 
 
 
 
151
  ## How it decides
152
 
153
  The prompt is AINode's own decide rendering (source commit `e5c08938`, hash in `prompt_contract.json`), through the chat template with thinking off. The option labels are single tokens; the answer is the distribution over those label tokens at the last prompt position, read in fp32 and temperature scaled per type (`temperatures.json`: choice 1.19, noul 1.09, score 1.22). Nothing is generated. Those are the `train` fit, which minimises NLL against the soft or ordinal target the model was taught; it softens, and it costs some ECE against hard labels on the calibration split. A sharper `hard` fit (0.79 / 0.67 / 0.83) is in the same file for a consumer who gates on the argmax. Thresholds belong to the caller: act on a high probability, confirm or escalate on a middle one, hand a low one to a person or a bigger model. The model never refuses.
 
73
 
74
  Made in Texas.
75
 
76
+ **Results and docs**
77
+
78
+ - Project site, with every result and how to run the models: [jebadiah.ai](https://jebadiah.ai).
79
+ - JevBench v1.4.2: on its 231 public items, run through its own harness, Jebadiah 9B v2 scores 0.818 (Jebadiah 27B: 0.866, the same as Jev 1.13.0). This is my own run on the public items, not the official board, which also uses sealed items. [Details and caveats](https://github.com/getainode/jebadiah#where-we-stand-on-jevbench).
80
+ - [JDE](https://github.com/Titanium-Devops/jde) blind test (as of 2026-09-26): on 290 real decisions from Titanium Computing's production decision engine, Jebadiah 9B v2 got 277 against Jev's 282, with 0 flips across 5,800 repeat calls. Jev still leads on coverage checks.
81
+ - For the family: [Decision Index 0.2.1](https://huggingface.co/spaces/multimodalart/jev-decision-index): the 27B scores 54.67, #5 of 67 open models (as of 2026-09-26). This is the board's own number; the maintainer validated my run and put it on the leaderboard. [Run record](https://huggingface.co/datasets/frontier-infra/jebadiah-decision-index-results/tree/main/runs/jebadiah-27b-1c0d794f).
82
+ - All sizes: the [Hugging Face collection](https://huggingface.co/collections/frontier-infra/jebadiah-open-system-one-decision-models-6ab80765ddd3fa0b3eba5213), mirrored on [ModelScope](https://www.modelscope.ai/profile/JasonBrashear).
83
+
84
+ **Which one should I use?** For local use, start with [Jebadiah 9B v2 GGUF](https://huggingface.co/frontier-infra/jebadiah-9b-v2-GGUF). On Apple silicon, use an MLX build: [27B](https://huggingface.co/frontier-infra/jebadiah-27b-MLX), [9B v2](https://huggingface.co/frontier-infra/jebadiah-9b-v2-MLX) or [4B v2](https://huggingface.co/frontier-infra/jebadiah-4b-v2-MLX). For vLLM or fine-tuning, use the full weights: [27B](https://huggingface.co/frontier-infra/jebadiah-27b), [9B v2](https://huggingface.co/frontier-infra/jebadiah-9b-v2) or [4B v2](https://huggingface.co/frontier-infra/jebadiah-4b-v2).
85
+
86
  ## Results
87
 
88
  Measured by us with AINode's bench, one logit read per question, the same rendered prompt for every model. Accuracy is the share of questions whose top label is the human label. Headline is the macro over the zero-shot public sets. These are our numbers on the public suites, not rows on the Jevals board.
 
158
  - **GGUF** for llama.cpp: [jebadiah-9b-v2-GGUF](https://huggingface.co/frontier-infra/jebadiah-9b-v2-GGUF). Same answer as these weights on 257 of 260 held-out questions (Q8_0) and 240 of 260 (Q4_K_M).
159
  - **MLX** for Apple silicon: [jebadiah-9b-v2-MLX](https://huggingface.co/frontier-infra/jebadiah-9b-v2-MLX). Same answer on 256 of 260 (8-bit; the 4-bit build fell under 90% and was not published).
160
 
161
+ ## Community quantizations
162
+
163
+ Thanks to bartowski and mradermacher for building GGUF quantizations of this model:
164
+
165
+ - [bartowski/frontier-infra_jebadiah-9b-v2-GGUF](https://huggingface.co/bartowski/frontier-infra_jebadiah-9b-v2-GGUF)
166
+ - [mradermacher/jebadiah-9b-v2-GGUF](https://huggingface.co/mradermacher/jebadiah-9b-v2-GGUF)
167
+
168
+ These are independent builds. The agreement numbers under Other formats are for my builds in [jebadiah-9b-v2-GGUF](https://huggingface.co/frontier-infra/jebadiah-9b-v2-GGUF), not for these.
169
+
170
  ## How it decides
171
 
172
  The prompt is AINode's own decide rendering (source commit `e5c08938`, hash in `prompt_contract.json`), through the chat template with thinking off. The option labels are single tokens; the answer is the distribution over those label tokens at the last prompt position, read in fp32 and temperature scaled per type (`temperatures.json`: choice 1.19, noul 1.09, score 1.22). Nothing is generated. Those are the `train` fit, which minimises NLL against the soft or ordinal target the model was taught; it softens, and it costs some ECE against hard labels on the calibration split. A sharper `hard` fit (0.79 / 0.67 / 0.83) is in the same file for a consumer who gates on the argmax. Thresholds belong to the caller: act on a high probability, confirm or escalate on a middle one, hand a low one to a person or a bigger model. The model never refuses.