SeaWolf-AI commited on
Commit
91fad24
·
verified ·
1 Parent(s): 746fa6a

card: ladder position, per-domain vs baseline, both probes

Browse files
Files changed (1) hide show
  1. README.md +174 -0
README.md ADDED
@@ -0,0 +1,174 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ language:
4
+ - en
5
+ - ko
6
+ tags:
7
+ - ztc
8
+ - answer-verification
9
+ - hallucination-detection
10
+ - zero-token
11
+ - confidence-estimation
12
+ library_name: transformers
13
+ pipeline_tag: text-classification
14
+ ---
15
+
16
+ # ZTC-Judge-9B
17
+
18
+ **Answer verification from a single forward pass, with zero generated tokens — at 9B.**
19
+
20
+ ZTC-Judge-9B takes a question and an answer written by *any* model and scores whether that
21
+ answer can be trusted. It is the third rung of a four-point size ladder measured under one
22
+ identical protocol, and it is published so the shape of that ladder can be checked rather than
23
+ asserted.
24
+
25
+ > **ZTC** — Zero-Token Confidence · **Judge** — it evaluates *someone else's* answer, not its own
26
+
27
+ ---
28
+
29
+ ## Read this before deploying
30
+
31
+ This model is **not** the strongest member of the family, and the card says so with numbers.
32
+
33
+ | Model | Leaderboard AUC |
34
+ |---|---|
35
+ | Darwin-397B-ZTC | 0.7364 |
36
+ | ZTC-Judge-27B | 0.7282 |
37
+ | **ZTC-Judge-9B** | **0.6506** |
38
+ | **ZTC-Judge-4B** | **0.6360** |
39
+ | *Answer length and formatting only* | *0.6223* |
40
+
41
+ **The ladder does not decline smoothly — it steps.** Between 9B and 27B the score moves 0.078,
42
+ while between 4B and 9B it moves 0.015. Whatever carries verification quality is largely absent
43
+ below 27B on this axis.
44
+
45
+ **Where this model is worth deploying is one specific place**, and it is a real one:
46
+
47
+ | Domain | Surface baseline | **9B** | Margin |
48
+ |---|---|---|---|
49
+ | **Professional exams (law · math · biology)** | 0.7138 | **0.7623** | ****+0.0485**** |
50
+ | Scientific reasoning | 0.7272 | 0.6764 | 🔴 -0.0508 |
51
+ | Biology & medicine | 0.5908 | 0.6310 | **+0.0402** |
52
+ | Disaster & safety procedures | 0.5949 | 0.5862 | 🔴 -0.0087 |
53
+ | General multi-step reasoning | 0.5420 | 0.5881 | **+0.0461** |
54
+ | **Size-weighted mean** | **0.6223** | **0.6506** | ****+0.0283**** |
55
+
56
+ 🔴 **Do not use this model for disaster and safety content.** In that domain it does not clear the
57
+ surface baseline, which means it is reading answer shape rather than correctness there.
58
+
59
+ ✅ **Professional-exam style content is where it earns its size.** It runs on a laptop, on CPU, and
60
+ inside networks that never reach the internet — places a hosted API cannot go.
61
+
62
+ ## How it works
63
+
64
+ ```
65
+ [question + answer] → one forward pass
66
+ → final-layer hidden state at the last position (4096-d)
67
+ → probe
68
+ → score
69
+ ```
70
+
71
+ **Generated tokens: 0.** No access to the answering model's weights or logits is required; the text
72
+ of the answer is the only input. Latency is one forward pass, and batching converts directly into
73
+ throughput.
74
+
75
+ ## Usage
76
+
77
+ ```python
78
+ import json
79
+ import numpy as np, torch
80
+ from huggingface_hub import hf_hub_download, snapshot_download
81
+ from transformers import AutoModel, AutoTokenizer
82
+
83
+ REPO = "FINAL-Bench/ZTC-Judge-9B"
84
+ cfg = json.load(open(hf_hub_download(REPO, "ztc_config.json"), encoding="utf-8"))
85
+ path = snapshot_download(REPO)
86
+
87
+ tok = AutoTokenizer.from_pretrained(path)
88
+ model = AutoModel.from_pretrained(path, dtype=torch.bfloat16, low_cpu_mem_usage=True).eval()
89
+
90
+ def hidden(question, answer):
91
+ b = tok([cfg["template"] % (question, answer)], return_tensors="pt",
92
+ truncation=True, max_length=cfg["max_length"])
93
+ dev = next(model.parameters()).device
94
+ with torch.no_grad():
95
+ h = model(input_ids=b["input_ids"].to(dev),
96
+ attention_mask=b["attention_mask"].to(dev)).last_hidden_state
97
+ return h[0, int(b["attention_mask"].sum()) - 1].float().numpy().astype(np.float64)
98
+
99
+ # linear probe — one dot product
100
+ p = np.load(hf_hub_download(REPO, "ztc_probe.npz"))
101
+ v = hidden("Which defensive chemical does an insect release?", "C. Allomone")
102
+ print(float(((v - p["mu"]) / p["sd"]) @ p["w"]))
103
+ ```
104
+
105
+ The score is an unbounded real number; higher means more likely correct. It is a **ranking signal**,
106
+ not a calibrated probability — choose a threshold from your own review budget.
107
+
108
+ ## Two probes ship with this model
109
+
110
+ | File | Produces |
111
+ |---|---|
112
+ | `ztc_probe.npz` | linear readout — one dot product |
113
+ | `ztc_curve_probe.npz` | **the reported figure 0.6506** — 256 anchors, RBF kernel |
114
+
115
+ Both read the same input. The curved probe is the one to use when the number matters.
116
+
117
+ ## Protocol
118
+
119
+ | | |
120
+ |---|---|
121
+ | Items | 2,018 · 508 incorrect · 5 domains · answers written by 4 different models |
122
+ | Metric | AUC — how well wrong answers sort to the bottom. Threshold-free. 0.5 = coin flip |
123
+ | Selection | **Leave-one-domain-out.** Every figure comes from a domain the probe never saw; hyper-parameters are chosen inside the training domains only |
124
+ | Aggregation | Per domain, then size-weighted. Pooling all items into one AUC inflates the result |
125
+
126
+ The same protocol, item set and grading code are applied to every rung of the ladder and to the
127
+ other systems on the independent leaderboard:
128
+ <https://huggingface.co/spaces/mayafree/typed-decision-leaderboard>
129
+
130
+ ## Out of scope
131
+
132
+ - **Not a grounding checker.** It does not take a source document and decide whether the answer
133
+ follows from it.
134
+ - **Not a safety, toxicity or policy classifier.**
135
+ - **Not a calibrated probability.** Use it to rank and threshold.
136
+ - **Not a general-purpose verifier at this size.** See the domain table above.
137
+
138
+ ## Limitations
139
+
140
+ - **Domain coverage.** Scores are meaningful only for the five domains listed. Outside them nothing
141
+ has been measured and no guarantee is published.
142
+ - **Below the surface baseline on disaster and safety.** Stated in the table rather than omitted.
143
+ - **Revision lock.** The probe is fitted to one specific revision of the base model. Running it on a
144
+ different revision produces **no error and silently wrong scores**; this repository ships the
145
+ matching weights so that failure mode cannot occur.
146
+ - **It reports the verifier's judgement**, which is not the answering model's own confidence — that
147
+ quantity measures 0.5000 on this set.
148
+
149
+ ## Lineage
150
+
151
+ | | |
152
+ |---|---|
153
+ | Base model | `Qwen/Qwen3.5-9B`, Apache-2.0 |
154
+ | Modification to base weights | none — the probes are separate files |
155
+ | Added by FINAL-Bench | probes, inference code, evaluation protocol and tables |
156
+
157
+ ## What this repository contains
158
+
159
+ | | |
160
+ |---|---|
161
+ | Included | Base weights · tokenizer · linear probe · curved probe · configuration |
162
+ | Not included | Training corpus · hidden-state matrices · fitting pipeline |
163
+
164
+ ## The rest of the ladder
165
+
166
+ [ZTC-Judge-27B](https://huggingface.co/FINAL-Bench/ZTC-Judge-27B) ·
167
+ [Darwin-397B-ZTC](https://huggingface.co/FINAL-Bench/Darwin-397B-ZTC) ·
168
+ [ZTC-Judge-9B](https://huggingface.co/FINAL-Bench/ZTC-Judge-9B) ·
169
+ [ZTC-Judge-4B](https://huggingface.co/FINAL-Bench/ZTC-Judge-4B)
170
+
171
+ ## License
172
+
173
+ The base model is Apache-2.0 and redistributable. The probes, the inference code and the evaluation
174
+ tables are assets of FINAL-Bench / VIDRAFT.