Text Classification
Transformers
Safetensors
English
Korean
qwen3_5
image-text-to-text
ztc
answer-verification
hallucination-detection
zero-token
confidence-estimation
Instructions to use FINAL-Bench/ZTC-Judge-9B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use FINAL-Bench/ZTC-Judge-9B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="FINAL-Bench/ZTC-Judge-9B")# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("FINAL-Bench/ZTC-Judge-9B") model = AutoModelForMultimodalLM.from_pretrained("FINAL-Bench/ZTC-Judge-9B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
card: ladder position, per-domain vs baseline, both probes
Browse files
README.md
ADDED
|
@@ -0,0 +1,174 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
language:
|
| 4 |
+
- en
|
| 5 |
+
- ko
|
| 6 |
+
tags:
|
| 7 |
+
- ztc
|
| 8 |
+
- answer-verification
|
| 9 |
+
- hallucination-detection
|
| 10 |
+
- zero-token
|
| 11 |
+
- confidence-estimation
|
| 12 |
+
library_name: transformers
|
| 13 |
+
pipeline_tag: text-classification
|
| 14 |
+
---
|
| 15 |
+
|
| 16 |
+
# ZTC-Judge-9B
|
| 17 |
+
|
| 18 |
+
**Answer verification from a single forward pass, with zero generated tokens — at 9B.**
|
| 19 |
+
|
| 20 |
+
ZTC-Judge-9B takes a question and an answer written by *any* model and scores whether that
|
| 21 |
+
answer can be trusted. It is the third rung of a four-point size ladder measured under one
|
| 22 |
+
identical protocol, and it is published so the shape of that ladder can be checked rather than
|
| 23 |
+
asserted.
|
| 24 |
+
|
| 25 |
+
> **ZTC** — Zero-Token Confidence · **Judge** — it evaluates *someone else's* answer, not its own
|
| 26 |
+
|
| 27 |
+
---
|
| 28 |
+
|
| 29 |
+
## Read this before deploying
|
| 30 |
+
|
| 31 |
+
This model is **not** the strongest member of the family, and the card says so with numbers.
|
| 32 |
+
|
| 33 |
+
| Model | Leaderboard AUC |
|
| 34 |
+
|---|---|
|
| 35 |
+
| Darwin-397B-ZTC | 0.7364 |
|
| 36 |
+
| ZTC-Judge-27B | 0.7282 |
|
| 37 |
+
| **ZTC-Judge-9B** | **0.6506** |
|
| 38 |
+
| **ZTC-Judge-4B** | **0.6360** |
|
| 39 |
+
| *Answer length and formatting only* | *0.6223* |
|
| 40 |
+
|
| 41 |
+
**The ladder does not decline smoothly — it steps.** Between 9B and 27B the score moves 0.078,
|
| 42 |
+
while between 4B and 9B it moves 0.015. Whatever carries verification quality is largely absent
|
| 43 |
+
below 27B on this axis.
|
| 44 |
+
|
| 45 |
+
**Where this model is worth deploying is one specific place**, and it is a real one:
|
| 46 |
+
|
| 47 |
+
| Domain | Surface baseline | **9B** | Margin |
|
| 48 |
+
|---|---|---|---|
|
| 49 |
+
| **Professional exams (law · math · biology)** | 0.7138 | **0.7623** | ****+0.0485**** |
|
| 50 |
+
| Scientific reasoning | 0.7272 | 0.6764 | 🔴 -0.0508 |
|
| 51 |
+
| Biology & medicine | 0.5908 | 0.6310 | **+0.0402** |
|
| 52 |
+
| Disaster & safety procedures | 0.5949 | 0.5862 | 🔴 -0.0087 |
|
| 53 |
+
| General multi-step reasoning | 0.5420 | 0.5881 | **+0.0461** |
|
| 54 |
+
| **Size-weighted mean** | **0.6223** | **0.6506** | ****+0.0283**** |
|
| 55 |
+
|
| 56 |
+
🔴 **Do not use this model for disaster and safety content.** In that domain it does not clear the
|
| 57 |
+
surface baseline, which means it is reading answer shape rather than correctness there.
|
| 58 |
+
|
| 59 |
+
✅ **Professional-exam style content is where it earns its size.** It runs on a laptop, on CPU, and
|
| 60 |
+
inside networks that never reach the internet — places a hosted API cannot go.
|
| 61 |
+
|
| 62 |
+
## How it works
|
| 63 |
+
|
| 64 |
+
```
|
| 65 |
+
[question + answer] → one forward pass
|
| 66 |
+
→ final-layer hidden state at the last position (4096-d)
|
| 67 |
+
→ probe
|
| 68 |
+
→ score
|
| 69 |
+
```
|
| 70 |
+
|
| 71 |
+
**Generated tokens: 0.** No access to the answering model's weights or logits is required; the text
|
| 72 |
+
of the answer is the only input. Latency is one forward pass, and batching converts directly into
|
| 73 |
+
throughput.
|
| 74 |
+
|
| 75 |
+
## Usage
|
| 76 |
+
|
| 77 |
+
```python
|
| 78 |
+
import json
|
| 79 |
+
import numpy as np, torch
|
| 80 |
+
from huggingface_hub import hf_hub_download, snapshot_download
|
| 81 |
+
from transformers import AutoModel, AutoTokenizer
|
| 82 |
+
|
| 83 |
+
REPO = "FINAL-Bench/ZTC-Judge-9B"
|
| 84 |
+
cfg = json.load(open(hf_hub_download(REPO, "ztc_config.json"), encoding="utf-8"))
|
| 85 |
+
path = snapshot_download(REPO)
|
| 86 |
+
|
| 87 |
+
tok = AutoTokenizer.from_pretrained(path)
|
| 88 |
+
model = AutoModel.from_pretrained(path, dtype=torch.bfloat16, low_cpu_mem_usage=True).eval()
|
| 89 |
+
|
| 90 |
+
def hidden(question, answer):
|
| 91 |
+
b = tok([cfg["template"] % (question, answer)], return_tensors="pt",
|
| 92 |
+
truncation=True, max_length=cfg["max_length"])
|
| 93 |
+
dev = next(model.parameters()).device
|
| 94 |
+
with torch.no_grad():
|
| 95 |
+
h = model(input_ids=b["input_ids"].to(dev),
|
| 96 |
+
attention_mask=b["attention_mask"].to(dev)).last_hidden_state
|
| 97 |
+
return h[0, int(b["attention_mask"].sum()) - 1].float().numpy().astype(np.float64)
|
| 98 |
+
|
| 99 |
+
# linear probe — one dot product
|
| 100 |
+
p = np.load(hf_hub_download(REPO, "ztc_probe.npz"))
|
| 101 |
+
v = hidden("Which defensive chemical does an insect release?", "C. Allomone")
|
| 102 |
+
print(float(((v - p["mu"]) / p["sd"]) @ p["w"]))
|
| 103 |
+
```
|
| 104 |
+
|
| 105 |
+
The score is an unbounded real number; higher means more likely correct. It is a **ranking signal**,
|
| 106 |
+
not a calibrated probability — choose a threshold from your own review budget.
|
| 107 |
+
|
| 108 |
+
## Two probes ship with this model
|
| 109 |
+
|
| 110 |
+
| File | Produces |
|
| 111 |
+
|---|---|
|
| 112 |
+
| `ztc_probe.npz` | linear readout — one dot product |
|
| 113 |
+
| `ztc_curve_probe.npz` | **the reported figure 0.6506** — 256 anchors, RBF kernel |
|
| 114 |
+
|
| 115 |
+
Both read the same input. The curved probe is the one to use when the number matters.
|
| 116 |
+
|
| 117 |
+
## Protocol
|
| 118 |
+
|
| 119 |
+
| | |
|
| 120 |
+
|---|---|
|
| 121 |
+
| Items | 2,018 · 508 incorrect · 5 domains · answers written by 4 different models |
|
| 122 |
+
| Metric | AUC — how well wrong answers sort to the bottom. Threshold-free. 0.5 = coin flip |
|
| 123 |
+
| Selection | **Leave-one-domain-out.** Every figure comes from a domain the probe never saw; hyper-parameters are chosen inside the training domains only |
|
| 124 |
+
| Aggregation | Per domain, then size-weighted. Pooling all items into one AUC inflates the result |
|
| 125 |
+
|
| 126 |
+
The same protocol, item set and grading code are applied to every rung of the ladder and to the
|
| 127 |
+
other systems on the independent leaderboard:
|
| 128 |
+
<https://huggingface.co/spaces/mayafree/typed-decision-leaderboard>
|
| 129 |
+
|
| 130 |
+
## Out of scope
|
| 131 |
+
|
| 132 |
+
- **Not a grounding checker.** It does not take a source document and decide whether the answer
|
| 133 |
+
follows from it.
|
| 134 |
+
- **Not a safety, toxicity or policy classifier.**
|
| 135 |
+
- **Not a calibrated probability.** Use it to rank and threshold.
|
| 136 |
+
- **Not a general-purpose verifier at this size.** See the domain table above.
|
| 137 |
+
|
| 138 |
+
## Limitations
|
| 139 |
+
|
| 140 |
+
- **Domain coverage.** Scores are meaningful only for the five domains listed. Outside them nothing
|
| 141 |
+
has been measured and no guarantee is published.
|
| 142 |
+
- **Below the surface baseline on disaster and safety.** Stated in the table rather than omitted.
|
| 143 |
+
- **Revision lock.** The probe is fitted to one specific revision of the base model. Running it on a
|
| 144 |
+
different revision produces **no error and silently wrong scores**; this repository ships the
|
| 145 |
+
matching weights so that failure mode cannot occur.
|
| 146 |
+
- **It reports the verifier's judgement**, which is not the answering model's own confidence — that
|
| 147 |
+
quantity measures 0.5000 on this set.
|
| 148 |
+
|
| 149 |
+
## Lineage
|
| 150 |
+
|
| 151 |
+
| | |
|
| 152 |
+
|---|---|
|
| 153 |
+
| Base model | `Qwen/Qwen3.5-9B`, Apache-2.0 |
|
| 154 |
+
| Modification to base weights | none — the probes are separate files |
|
| 155 |
+
| Added by FINAL-Bench | probes, inference code, evaluation protocol and tables |
|
| 156 |
+
|
| 157 |
+
## What this repository contains
|
| 158 |
+
|
| 159 |
+
| | |
|
| 160 |
+
|---|---|
|
| 161 |
+
| Included | Base weights · tokenizer · linear probe · curved probe · configuration |
|
| 162 |
+
| Not included | Training corpus · hidden-state matrices · fitting pipeline |
|
| 163 |
+
|
| 164 |
+
## The rest of the ladder
|
| 165 |
+
|
| 166 |
+
[ZTC-Judge-27B](https://huggingface.co/FINAL-Bench/ZTC-Judge-27B) ·
|
| 167 |
+
[Darwin-397B-ZTC](https://huggingface.co/FINAL-Bench/Darwin-397B-ZTC) ·
|
| 168 |
+
[ZTC-Judge-9B](https://huggingface.co/FINAL-Bench/ZTC-Judge-9B) ·
|
| 169 |
+
[ZTC-Judge-4B](https://huggingface.co/FINAL-Bench/ZTC-Judge-4B)
|
| 170 |
+
|
| 171 |
+
## License
|
| 172 |
+
|
| 173 |
+
The base model is Apache-2.0 and redistributable. The probes, the inference code and the evaluation
|
| 174 |
+
tables are assets of FINAL-Bench / VIDRAFT.
|