Text Classification
PEFT
Safetensors
English
lora
qlora
calibration
decision-model
system-one
typesafe
Instructions to use vagmi/jev-lite with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use vagmi/jev-lite with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
add README.md
Browse files
README.md
ADDED
|
@@ -0,0 +1,200 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: gemma
|
| 3 |
+
base_model: google/gemma-4-E4B-it
|
| 4 |
+
library_name: peft
|
| 5 |
+
pipeline_tag: text-classification
|
| 6 |
+
tags:
|
| 7 |
+
- lora
|
| 8 |
+
- qlora
|
| 9 |
+
- calibration
|
| 10 |
+
- decision-model
|
| 11 |
+
- system-one
|
| 12 |
+
- typesafe
|
| 13 |
+
language:
|
| 14 |
+
- en
|
| 15 |
+
---
|
| 16 |
+
|
| 17 |
+
# jev-lite
|
| 18 |
+
|
| 19 |
+
A QLoRA adapter that turns Gemma 4 E4B into a **System One** decision model: it reads a
|
| 20 |
+
state, reads a typed question about it, and returns a calibrated probability distribution
|
| 21 |
+
over the allowed answers — from **one forward pass**, with no generation.
|
| 22 |
+
|
| 23 |
+
Because the answer is read from the logits at the option letters, it **cannot** answer
|
| 24 |
+
outside the options it was given. There is no parsing, no retry loop, and no
|
| 25 |
+
"as an AI language model".
|
| 26 |
+
|
| 27 |
+
It implements the three [TypeSafe](https://docs.typesafe.ai) primitives:
|
| 28 |
+
|
| 29 |
+
| type | question | returns |
|
| 30 |
+
|---|---|---|
|
| 31 |
+
| `choice` | pick one of several named options | the option, probabilities, confidence |
|
| 32 |
+
| `score` | rate against ordered levels | expected level, legend, probabilities, confidence |
|
| 33 |
+
| `noul` | is this true? | a single probability |
|
| 34 |
+
|
| 35 |
+
## Results
|
| 36 |
+
|
| 37 |
+
Evaluated on 1,898 held-out rows — 31 Super-NaturalInstructions **tasks** the model never
|
| 38 |
+
saw, 78 synthetic **states** it never saw, and SST-2 held out in its entirety. The split is
|
| 39 |
+
by task and by state, never by row, so these are questions of kinds it was not trained on.
|
| 40 |
+
|
| 41 |
+
| | accuracy | ECE | NLL | Brier |
|
| 42 |
+
|---|---|---|---|---|
|
| 43 |
+
| **overall** | **0.816** | **0.019** | 0.430 | 0.222 |
|
| 44 |
+
| choice (n=1196) | 0.809 | — | 0.431 | 0.238 |
|
| 45 |
+
| noul (n=547) | 0.819 | — | — | — |
|
| 46 |
+
| score (n=155) | 0.852 | — | — | — |
|
| 47 |
+
|
| 48 |
+
Score answers are off by **0.216 levels** on average (mean absolute error of the expected
|
| 49 |
+
level).
|
| 50 |
+
|
| 51 |
+
**Calibration is the point.** ECE of 0.019 means that when it reports 80% confidence it is
|
| 52 |
+
right about 80% of the time. Accuracy was flat from step 250 to the end of training while
|
| 53 |
+
ECE fell 0.086 → 0.019: the model did not learn to be right more often, it learned to be
|
| 54 |
+
honest about when it isn't.
|
| 55 |
+
|
| 56 |
+
Operationally, gating on confidence:
|
| 57 |
+
|
| 58 |
+
| band | share of traffic | accuracy |
|
| 59 |
+
|---|---|---|
|
| 60 |
+
| ≥ 0.80 — act automatically | 62% | 93% |
|
| 61 |
+
| 0.50–0.80 — confirm or review | 36% | 60% |
|
| 62 |
+
| < 0.50 — route to a human | 2% | 32% |
|
| 63 |
+
|
| 64 |
+
## Confidence
|
| 65 |
+
|
| 66 |
+
`confidence` is the model's probability that the answer it returned is the correct one:
|
| 67 |
+
|
| 68 |
+
- **choice** — the probability of the selected option (`p_max`)
|
| 69 |
+
- **score** — the probability mass that rounds to the reported expected level
|
| 70 |
+
- **noul** — no confidence field; the probability *is* the answer
|
| 71 |
+
|
| 72 |
+
This was chosen by measurement, not taste. Over 1,430 gold-labelled held-out rows,
|
| 73 |
+
`p_max` beat normalized entropy, margin, chance-corrected top and collision entropy on
|
| 74 |
+
**both** AUROC (0.816, ranking right answers above wrong ones) and ECE (0.036). Normalized
|
| 75 |
+
entropy — the obvious first guess — was the worst of the lot: it is dominated by small
|
| 76 |
+
probabilities, so it reads a decisive 0.85/0.08/0.07 as low confidence and routes 45% of
|
| 77 |
+
traffic to human review at 64% accuracy.
|
| 78 |
+
|
| 79 |
+
## Usage
|
| 80 |
+
|
| 81 |
+
The adapter is only meaningful with the exact prompt format it was trained on, so
|
| 82 |
+
[`primitives.py`](./primitives.py) is included in this repo and **must** be used to render
|
| 83 |
+
questions. Criteria descriptions are part of that format.
|
| 84 |
+
|
| 85 |
+
```python
|
| 86 |
+
import torch, primitives
|
| 87 |
+
from peft import PeftModel
|
| 88 |
+
from transformers import AutoModelForImageTextToText, AutoTokenizer, BitsAndBytesConfig
|
| 89 |
+
|
| 90 |
+
tok = AutoTokenizer.from_pretrained("vagmi/jev-lite")
|
| 91 |
+
base = AutoModelForImageTextToText.from_pretrained(
|
| 92 |
+
"google/gemma-4-E4B-it", device_map={"": 0}, dtype=torch.bfloat16,
|
| 93 |
+
quantization_config=BitsAndBytesConfig(
|
| 94 |
+
load_in_4bit=True, bnb_4bit_quant_type="nf4",
|
| 95 |
+
bnb_4bit_compute_dtype=torch.bfloat16, bnb_4bit_use_double_quant=True))
|
| 96 |
+
model = PeftModel.from_pretrained(base, "vagmi/jev-lite").eval()
|
| 97 |
+
|
| 98 |
+
row = primitives.normalize({
|
| 99 |
+
"type": "choice",
|
| 100 |
+
"state": "Help! My payouts have been failing for 3 days.",
|
| 101 |
+
"question": "Which team should handle this?",
|
| 102 |
+
"options": ["billing", "technical", "sales"],
|
| 103 |
+
"criteria": {"billing": "Payments, invoicing, refunds",
|
| 104 |
+
"technical": "Bugs, outages, integrations",
|
| 105 |
+
"sales": "Pricing, upgrades, new accounts"},
|
| 106 |
+
})
|
| 107 |
+
|
| 108 |
+
prompt = primitives.PREFIX + row["state"] + "\n</state>\n\n" + \
|
| 109 |
+
primitives.question_block(row) + "\nAnswer:"
|
| 110 |
+
ids = torch.tensor([[tok.bos_token_id] + tok.encode(prompt, add_special_tokens=False)])
|
| 111 |
+
letters = [tok.encode(" " + c, add_special_tokens=False)[0] for c in primitives.LETTERS]
|
| 112 |
+
|
| 113 |
+
with torch.no_grad():
|
| 114 |
+
logits = model(input_ids=ids.to(model.device)).logits[0, -1].float()
|
| 115 |
+
probs = torch.softmax(logits[letters[:len(row["options"])]], -1).tolist()
|
| 116 |
+
print(primitives.answer(row, probs))
|
| 117 |
+
# {'type': 'choice', 'choice': 'billing', 'probabilities': {...}, 'confidence': 0.58}
|
| 118 |
+
```
|
| 119 |
+
|
| 120 |
+
A server speaking the TypeSafe System One wire API (`POST /v1/systemone`), for both
|
| 121 |
+
transformers and vLLM backends, is at
|
| 122 |
+
[github.com/vagmi/jevlite](https://github.com/vagmi/jevlite).
|
| 123 |
+
|
| 124 |
+
## Training
|
| 125 |
+
|
| 126 |
+
QLoRA on `google/gemma-4-E4B-it`, 4-bit NF4 base with bf16 compute.
|
| 127 |
+
|
| 128 |
+
| | |
|
| 129 |
+
|---|---|
|
| 130 |
+
| LoRA rank / alpha | 16 / 32, dropout 0.05 |
|
| 131 |
+
| target modules | q, k, v, o, gate, up, down |
|
| 132 |
+
| trainable params | 34.9M of 7.98B (0.44%) |
|
| 133 |
+
| objective | soft-label cross-entropy over option letters |
|
| 134 |
+
| epochs / LR / schedule | 1 / 2e-4 / cosine, 3% warmup |
|
| 135 |
+
| effective batch | 16 (grad accum) |
|
| 136 |
+
| hardware | one RTX 4090 (24 GB), ~1h40m |
|
| 137 |
+
|
| 138 |
+
Options are shuffled every epoch so the model cannot learn "A is usually right" — except
|
| 139 |
+
for `score`, where the order of levels carries meaning and is preserved.
|
| 140 |
+
|
| 141 |
+
## Training data
|
| 142 |
+
|
| 143 |
+
23,632 rows across 301 tasks: 13,492 choice, 6,368 noul, 3,772 score. 92% carry criteria
|
| 144 |
+
descriptions; 35.5% carry soft labels.
|
| 145 |
+
|
| 146 |
+
| source | rows | labels |
|
| 147 |
+
|---|---|---|
|
| 148 |
+
| Super-NaturalInstructions (288 tasks) | 8,629 | gold, with criteria written by the teacher |
|
| 149 |
+
| MNLI / ANLI / BoolQ / RACE / Yelp | 9,605 | gold, hand-written criteria |
|
| 150 |
+
| synthetic states (978, 8 domains) | 5,398 | soft labels from the teacher |
|
| 151 |
+
|
| 152 |
+
Soft labels came from **Qwen3.6-35B-A3B** read at the answer token with option order
|
| 153 |
+
reversed and averaged to cancel position bias. Reasoning was deliberately left **off**: with
|
| 154 |
+
a thinking budget the teacher commits to its answer with probability 1.0 on essentially
|
| 155 |
+
every row, which is useless for distillation — the measured mean entropy was 0.000 with
|
| 156 |
+
reasoning versus 0.211 without.
|
| 157 |
+
|
| 158 |
+
## Limitations
|
| 159 |
+
|
| 160 |
+
- **English only.**
|
| 161 |
+
- **Score evaluation rests on teacher labels, not gold.** All 155 held-out score rows are
|
| 162 |
+
synthetic; SNI has no rating scales and the HF eval split is SST-2. Choice and noul
|
| 163 |
+
numbers (n=1,430) are against real gold labels.
|
| 164 |
+
- **Serving precision matters.** The adapter was trained against a 4-bit NF4 base and
|
| 165 |
+
learned to correct that quantization. Served in bf16 (e.g. vLLM), accuracy holds (0.807
|
| 166 |
+
vs 0.799) but ECE roughly doubles, 0.036 → 0.069. A single temperature of 1.4, fitted on
|
| 167 |
+
half the eval set and measured on the other half, restores it to 0.033. Serve 4-bit, or
|
| 168 |
+
apply the correction.
|
| 169 |
+
- **Tokenization is part of the contract.** Feeding a differently-tokenized prompt (a
|
| 170 |
+
missing BOS, say) moved 16% of argmaxes in testing. Build ids the way `primitives.py`
|
| 171 |
+
does.
|
| 172 |
+
- **The teacher disagreed with gold on 24.9%** of a 3,000-row sample — mostly noisy
|
| 173 |
+
annotation in SBIC/civil-comments-style tasks plus adversarial ANLI. Those rows were
|
| 174 |
+
blended toward gold at 0.5 rather than dropped.
|
| 175 |
+
- Not evaluated for fairness, toxicity, or adversarial robustness. Several source tasks
|
| 176 |
+
(stereotype and offensiveness classification) carry the biases of their annotators.
|
| 177 |
+
|
| 178 |
+
## Licensing and provenance
|
| 179 |
+
|
| 180 |
+
The adapter is a derivative of Gemma and is governed by the
|
| 181 |
+
[Gemma Terms of Use](https://ai.google.dev/gemma/terms), including its use restrictions.
|
| 182 |
+
|
| 183 |
+
Training data licenses are **mixed, and some are non-commercial**:
|
| 184 |
+
|
| 185 |
+
| dataset | license |
|
| 186 |
+
|---|---|
|
| 187 |
+
| Super-NaturalInstructions | Apache-2.0 |
|
| 188 |
+
| ANLI | CC BY-NC 4.0 — **non-commercial** |
|
| 189 |
+
| RACE | research use only — **non-commercial** |
|
| 190 |
+
| BoolQ | CC BY-SA 3.0 |
|
| 191 |
+
| MNLI (GLUE) | mixed, per-genre |
|
| 192 |
+
| Yelp Review Full | Yelp Dataset Terms of Use |
|
| 193 |
+
|
| 194 |
+
Because ANLI and RACE rows are in the training mix, **this adapter should be treated as
|
| 195 |
+
non-commercial / research use** unless retrained without them. `build_data.py` in the
|
| 196 |
+
source repo takes `--sets` to select sources, so a commercially-clean rebuild is a flag
|
| 197 |
+
change plus a retrain.
|
| 198 |
+
|
| 199 |
+
Synthetic rows and soft labels were generated with Qwen3.6-35B-A3B, an open-weight model
|
| 200 |
+
whose license permits using outputs to train other models.
|