File size: 3,287 Bytes
154a764
c22a16c
 
 
 
154a764
 
 
c22a16c
09b08d4
b0ccac6
 
 
 
 
09b08d4
 
 
154a764
 
09b08d4
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
---
title: Jev-style typed decisions
emoji: 🎚️
colorFrom: gray
colorTo: gray
sdk: gradio
sdk_version: 6.28.0
app_file: app.py
license: apache-2.0
short_description: Cached causal typed scorer, probabilities per typed question
models:
  - abidlabs/jev-typed-decisions-causal-0.6b
  - Qwen/Qwen3-0.6B-Base
datasets:
  - pngwn/typed-decisions-v2
tags:
  - decision-model
  - calibration
---

# Jev-style typed decisions (arm B)

A demo of the **cached causal typed scorer** trained in
[abidlabs/jev-typed-decisions-causal-0.6b](https://huggingface.co/abidlabs/jev-typed-decisions-causal-0.6b) — the arm-B
model from [pngwn/typed-decisions-causal-experiment](https://huggingface.co/datasets/pngwn/typed-decisions-causal-experiment/blob/main/REPORT.md),
trained with that report's exact recipe (Qwen3-0.6B-Base + LoRA r16, frozen LM head, restricted candidate-letter
cross-entropy, branch encoding).

You give it one document (the **state**) and a set of typed questions; it returns a probability for every option of
every question. No text is generated. The layout follows [jaredpalmer/kev](https://huggingface.co/spaces/jaredpalmer/kev)
(the Kev decision-model Space); the architecture is the report's arm-B "cached causal typed scorer".

## How it runs

- Questions are rendered into the corpus' `fmt=multi` serialization (`### State: … ### Questions: 1) q [A) o1, …]`),
  exactly the format the model was trained on.
- The shared prompt is **prefilled once** (KV cache), then every question is scored in **one batched branch forward**
  against that cache, reading the next-token distribution restricted to the candidate option letters from the frozen
  LM head's alphabet rows — the full-vocabulary logits are never computed.
- The header under the answers reports the timing/token accounting: 1 prefill + 1 branch vs. the naive
  re-encode (P·k tokens).

## Differences from Kev

- One model (the 0.6B arm-B scorer), not a family; no pointer head — options are the letter tokens of the frozen LM head.
- Questions are **not isolated** from each other: the corpus' multi-question format puts all questions in the shared
  prompt, so every answer conditions on the full question list (Kev's Isolation probe property does not hold here).
- Calibration: the toggle applies T = 1.59, the temperature fitted on the corpus' cal split for this arm
  (REPORT.md §2). Argmax unchanged.
- Corpus options never exceeded 10; up to 26 are accepted (letters A–Z) but expect degraded quality beyond 10 and on
  out-of-distribution states — measure on your own inputs.

## API

```python
from gradio_client import Client
c = Client("abidlabs/jev-typed-decisions-demo")
rendered, response, report = c.predict(
    "Ticket TD-25184\nProduct: auth-service\nCustomer tier: silver\nLoad reading: 40\nRegion: us-east-1",
    '{"severity": {"type": "score", "instructions": "What severity does the true load fall in?", "criteria": ["1", "2", "3", "4", "5"]}, "escalate": {"type": "noul", "instructions": "Should the ticket be escalated now?"}}',
    False, False, 4,
    api_name="/decide",
)
print(response["answers"])
```

Credits: layout follows jaredpalmer/kev (Apache-2.0); corpus, recipe and architecture from
pngwn/typed-decisions-causal-experiment. Base model by Qwen (Apache-2.0).