Darwin-180B-RSI / ztc /README.md
SeaWolf-AI's picture
ztc: ship early ZTC probe (AUROC 0.64) + handler
7b875ae verified
|
Raw History Blame Contribute Delete
1.61 kB

ZTC probe for Darwin-180B-RSI (early release)

Zero-Token Confidence (ZTC) estimates, before any answer token is generated, how likely the model is to answer a question correctly. It reads the model's hidden state at the last prompt token and maps it to a probability with a small linear probe.

Files

File What it is
ztc_probe_darwin180rsi.npz Probe weights: ridge regression on the final-layer hidden state of the last prompt token, followed by Platt scaling
handler.py Loads the probe and returns a confidence score for a prompt
usage.py Minimal usage example

How it works

  1. Run the prompt through Darwin-180B-RSI once, with no generation.
  2. Take the final-layer hidden state at the last prompt token.
  3. Apply the ridge weights and bias, then the Platt calibration, to get a probability in [0, 1].

Cost is one forward pass over the prompt. No answer tokens are produced.

Status and measured quality

This is an early probe. On our held-out validation split it reaches an AUROC of 0.64. That is a weak signal, useful for coarse routing (for example, flagging questions for a second pass or for review), not for deciding on its own whether an answer is right.

A retrained version with more training data is planned and will replace this file. The numbers above will be updated when it ships.

Intended use

  • Ranking or filtering questions by expected difficulty
  • Deciding where to spend extra samples or a longer thinking budget
  • A gate that sends low-confidence cases to a human or a stronger check

Not intended as a correctness guarantee.