Text Classification
PEFT
Safetensors
English
lora
qlora
calibration
decision-model
system-one
typesafe
Instructions to use vagmi/jev-lite with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use vagmi/jev-lite with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Update model card to improve readability . move usage up front.
Browse files
README.md
CHANGED
|
@@ -18,9 +18,9 @@ language:
|
|
| 18 |
|
| 19 |
A QLoRA adapter that turns Gemma 4 E4B into a **System One** decision model: it reads a
|
| 20 |
state, reads a typed question about it, and returns a calibrated probability distribution
|
| 21 |
-
over the allowed answers
|
| 22 |
|
| 23 |
-
|
| 24 |
outside the options it was given. There is no parsing, no retry loop, and no
|
| 25 |
"as an AI language model".
|
| 26 |
|
|
@@ -32,50 +32,6 @@ It implements the three [TypeSafe](https://docs.typesafe.ai) primitives:
|
|
| 32 |
| `score` | rate against ordered levels | expected level, legend, probabilities, confidence |
|
| 33 |
| `noul` | is this true? | a single probability |
|
| 34 |
|
| 35 |
-
## Results
|
| 36 |
-
|
| 37 |
-
Evaluated on 1,898 held-out rows — 31 Super-NaturalInstructions **tasks** the model never
|
| 38 |
-
saw, 78 synthetic **states** it never saw, and SST-2 held out in its entirety. The split is
|
| 39 |
-
by task and by state, never by row, so these are questions of kinds it was not trained on.
|
| 40 |
-
|
| 41 |
-
| | accuracy | ECE | NLL | Brier |
|
| 42 |
-
|---|---|---|---|---|
|
| 43 |
-
| **overall** | **0.816** | **0.019** | 0.430 | 0.222 |
|
| 44 |
-
| choice (n=1196) | 0.809 | — | 0.431 | 0.238 |
|
| 45 |
-
| noul (n=547) | 0.819 | — | — | — |
|
| 46 |
-
| score (n=155) | 0.852 | — | — | — |
|
| 47 |
-
|
| 48 |
-
Score answers are off by **0.216 levels** on average (mean absolute error of the expected
|
| 49 |
-
level).
|
| 50 |
-
|
| 51 |
-
**Calibration is the point.** ECE of 0.019 means that when it reports 80% confidence it is
|
| 52 |
-
right about 80% of the time. Accuracy was flat from step 250 to the end of training while
|
| 53 |
-
ECE fell 0.086 → 0.019: the model did not learn to be right more often, it learned to be
|
| 54 |
-
honest about when it isn't.
|
| 55 |
-
|
| 56 |
-
Operationally, gating on confidence:
|
| 57 |
-
|
| 58 |
-
| band | share of traffic | accuracy |
|
| 59 |
-
|---|---|---|
|
| 60 |
-
| ≥ 0.80 — act automatically | 62% | 93% |
|
| 61 |
-
| 0.50–0.80 — confirm or review | 36% | 60% |
|
| 62 |
-
| < 0.50 — route to a human | 2% | 32% |
|
| 63 |
-
|
| 64 |
-
## Confidence
|
| 65 |
-
|
| 66 |
-
`confidence` is the model's probability that the answer it returned is the correct one:
|
| 67 |
-
|
| 68 |
-
- **choice** — the probability of the selected option (`p_max`)
|
| 69 |
-
- **score** — the probability mass that rounds to the reported expected level
|
| 70 |
-
- **noul** — no confidence field; the probability *is* the answer
|
| 71 |
-
|
| 72 |
-
This was chosen by measurement, not taste. Over 1,430 gold-labelled held-out rows,
|
| 73 |
-
`p_max` beat normalized entropy, margin, chance-corrected top and collision entropy on
|
| 74 |
-
**both** AUROC (0.816, ranking right answers above wrong ones) and ECE (0.036). Normalized
|
| 75 |
-
entropy — the obvious first guess — was the worst of the lot: it is dominated by small
|
| 76 |
-
probabilities, so it reads a decisive 0.85/0.08/0.07 as low confidence and routes 45% of
|
| 77 |
-
traffic to human review at 64% accuracy.
|
| 78 |
-
|
| 79 |
## Usage
|
| 80 |
|
| 81 |
The adapter is only meaningful with the exact prompt format it was trained on, so
|
|
@@ -155,6 +111,51 @@ a thinking budget the teacher commits to its answer with probability 1.0 on esse
|
|
| 155 |
every row, which is useless for distillation — the measured mean entropy was 0.000 with
|
| 156 |
reasoning versus 0.211 without.
|
| 157 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 158 |
## Limitations
|
| 159 |
|
| 160 |
- **English only.**
|
|
|
|
| 18 |
|
| 19 |
A QLoRA adapter that turns Gemma 4 E4B into a **System One** decision model: it reads a
|
| 20 |
state, reads a typed question about it, and returns a calibrated probability distribution
|
| 21 |
+
over the allowed answers from a single **forward pass**, with no generation.
|
| 22 |
|
| 23 |
+
Since the answer is read from the logits at the option letters, it **cannot** answer
|
| 24 |
outside the options it was given. There is no parsing, no retry loop, and no
|
| 25 |
"as an AI language model".
|
| 26 |
|
|
|
|
| 32 |
| `score` | rate against ordered levels | expected level, legend, probabilities, confidence |
|
| 33 |
| `noul` | is this true? | a single probability |
|
| 34 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 35 |
## Usage
|
| 36 |
|
| 37 |
The adapter is only meaningful with the exact prompt format it was trained on, so
|
|
|
|
| 111 |
every row, which is useless for distillation — the measured mean entropy was 0.000 with
|
| 112 |
reasoning versus 0.211 without.
|
| 113 |
|
| 114 |
+
|
| 115 |
+
## Results
|
| 116 |
+
|
| 117 |
+
Evaluated on 1,898 held-out rows — 31 Super-NaturalInstructions **tasks** the model never
|
| 118 |
+
saw, 78 synthetic **states** it never saw, and SST-2 held out in its entirety.
|
| 119 |
+
|
| 120 |
+
| | accuracy | ECE | NLL | Brier |
|
| 121 |
+
|---|---|---|---|---|
|
| 122 |
+
| **overall** | **0.816** | **0.019** | 0.430 | 0.222 |
|
| 123 |
+
| choice (n=1196) | 0.809 | — | 0.431 | 0.238 |
|
| 124 |
+
| noul (n=547) | 0.819 | — | — | — |
|
| 125 |
+
| score (n=155) | 0.852 | — | — | — |
|
| 126 |
+
|
| 127 |
+
Score answers are off by **0.216 levels** on average (mean absolute error of the expected
|
| 128 |
+
level).
|
| 129 |
+
|
| 130 |
+
**Calibration is the point.** ECE of 0.019 means that when it reports 80% confidence it is
|
| 131 |
+
right about 80% of the time. Accuracy was flat from step 250 to the end of training while
|
| 132 |
+
ECE fell 0.086 → 0.019: the model did not learn to be right more often, it learned to be
|
| 133 |
+
honest about when it isn't.
|
| 134 |
+
|
| 135 |
+
Operationally, gating on confidence:
|
| 136 |
+
|
| 137 |
+
| band | share of traffic | accuracy |
|
| 138 |
+
|---|---|---|
|
| 139 |
+
| ≥ 0.80 — act automatically | 62% | 93% |
|
| 140 |
+
| 0.50–0.80 — confirm or review | 36% | 60% |
|
| 141 |
+
| < 0.50 — route to a human | 2% | 32% |
|
| 142 |
+
|
| 143 |
+
## Confidence
|
| 144 |
+
|
| 145 |
+
`confidence` is the model's probability that the answer it returned is the correct one:
|
| 146 |
+
|
| 147 |
+
- **choice** — the probability of the selected option (`p_max`)
|
| 148 |
+
- **score** — the probability mass that rounds to the reported expected level
|
| 149 |
+
- **noul** — no confidence field; the probability *is* the answer
|
| 150 |
+
|
| 151 |
+
This was chosen by measurement, not taste. Over 1,430 gold-labelled held-out rows,
|
| 152 |
+
`p_max` beat normalized entropy, margin, chance-corrected top and collision entropy on
|
| 153 |
+
**both** AUROC (0.816, ranking right answers above wrong ones) and ECE (0.036). Normalized
|
| 154 |
+
entropy — the obvious first guess — was the worst of the lot: it is dominated by small
|
| 155 |
+
probabilities, so it reads a decisive 0.85/0.08/0.07 as low confidence and routes 45% of
|
| 156 |
+
traffic to human review at 64% accuracy.
|
| 157 |
+
|
| 158 |
+
|
| 159 |
## Limitations
|
| 160 |
|
| 161 |
- **English only.**
|