jev-lite / README.md
vagmi's picture
add README.md
d54f613 verified
|
Raw History Blame
8.61 kB
---
license: gemma
base_model: google/gemma-4-E4B-it
library_name: peft
pipeline_tag: text-classification
tags:
- lora
- qlora
- calibration
- decision-model
- system-one
- typesafe
language:
- en
---
# jev-lite
A QLoRA adapter that turns Gemma 4 E4B into a **System One** decision model: it reads a
state, reads a typed question about it, and returns a calibrated probability distribution
over the allowed answers — from **one forward pass**, with no generation.
Because the answer is read from the logits at the option letters, it **cannot** answer
outside the options it was given. There is no parsing, no retry loop, and no
"as an AI language model".
It implements the three [TypeSafe](https://docs.typesafe.ai) primitives:
| type | question | returns |
|---|---|---|
| `choice` | pick one of several named options | the option, probabilities, confidence |
| `score` | rate against ordered levels | expected level, legend, probabilities, confidence |
| `noul` | is this true? | a single probability |
## Results
Evaluated on 1,898 held-out rows — 31 Super-NaturalInstructions **tasks** the model never
saw, 78 synthetic **states** it never saw, and SST-2 held out in its entirety. The split is
by task and by state, never by row, so these are questions of kinds it was not trained on.
| | accuracy | ECE | NLL | Brier |
|---|---|---|---|---|
| **overall** | **0.816** | **0.019** | 0.430 | 0.222 |
| choice (n=1196) | 0.809 | — | 0.431 | 0.238 |
| noul (n=547) | 0.819 | — | — | — |
| score (n=155) | 0.852 | — | — | — |
Score answers are off by **0.216 levels** on average (mean absolute error of the expected
level).
**Calibration is the point.** ECE of 0.019 means that when it reports 80% confidence it is
right about 80% of the time. Accuracy was flat from step 250 to the end of training while
ECE fell 0.086 → 0.019: the model did not learn to be right more often, it learned to be
honest about when it isn't.
Operationally, gating on confidence:
| band | share of traffic | accuracy |
|---|---|---|
| ≥ 0.80 — act automatically | 62% | 93% |
| 0.50–0.80 — confirm or review | 36% | 60% |
| < 0.50 — route to a human | 2% | 32% |
## Confidence
`confidence` is the model's probability that the answer it returned is the correct one:
- **choice** — the probability of the selected option (`p_max`)
- **score** — the probability mass that rounds to the reported expected level
- **noul** — no confidence field; the probability *is* the answer
This was chosen by measurement, not taste. Over 1,430 gold-labelled held-out rows,
`p_max` beat normalized entropy, margin, chance-corrected top and collision entropy on
**both** AUROC (0.816, ranking right answers above wrong ones) and ECE (0.036). Normalized
entropy — the obvious first guess — was the worst of the lot: it is dominated by small
probabilities, so it reads a decisive 0.85/0.08/0.07 as low confidence and routes 45% of
traffic to human review at 64% accuracy.
## Usage
The adapter is only meaningful with the exact prompt format it was trained on, so
[`primitives.py`](./primitives.py) is included in this repo and **must** be used to render
questions. Criteria descriptions are part of that format.
```python
import torch, primitives
from peft import PeftModel
from transformers import AutoModelForImageTextToText, AutoTokenizer, BitsAndBytesConfig
tok = AutoTokenizer.from_pretrained("vagmi/jev-lite")
base = AutoModelForImageTextToText.from_pretrained(
"google/gemma-4-E4B-it", device_map={"": 0}, dtype=torch.bfloat16,
quantization_config=BitsAndBytesConfig(
load_in_4bit=True, bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.bfloat16, bnb_4bit_use_double_quant=True))
model = PeftModel.from_pretrained(base, "vagmi/jev-lite").eval()
row = primitives.normalize({
"type": "choice",
"state": "Help! My payouts have been failing for 3 days.",
"question": "Which team should handle this?",
"options": ["billing", "technical", "sales"],
"criteria": {"billing": "Payments, invoicing, refunds",
"technical": "Bugs, outages, integrations",
"sales": "Pricing, upgrades, new accounts"},
})
prompt = primitives.PREFIX + row["state"] + "\n</state>\n\n" + \
primitives.question_block(row) + "\nAnswer:"
ids = torch.tensor([[tok.bos_token_id] + tok.encode(prompt, add_special_tokens=False)])
letters = [tok.encode(" " + c, add_special_tokens=False)[0] for c in primitives.LETTERS]
with torch.no_grad():
logits = model(input_ids=ids.to(model.device)).logits[0, -1].float()
probs = torch.softmax(logits[letters[:len(row["options"])]], -1).tolist()
print(primitives.answer(row, probs))
# {'type': 'choice', 'choice': 'billing', 'probabilities': {...}, 'confidence': 0.58}
```
A server speaking the TypeSafe System One wire API (`POST /v1/systemone`), for both
transformers and vLLM backends, is at
[github.com/vagmi/jevlite](https://github.com/vagmi/jevlite).
## Training
QLoRA on `google/gemma-4-E4B-it`, 4-bit NF4 base with bf16 compute.
| | |
|---|---|
| LoRA rank / alpha | 16 / 32, dropout 0.05 |
| target modules | q, k, v, o, gate, up, down |
| trainable params | 34.9M of 7.98B (0.44%) |
| objective | soft-label cross-entropy over option letters |
| epochs / LR / schedule | 1 / 2e-4 / cosine, 3% warmup |
| effective batch | 16 (grad accum) |
| hardware | one RTX 4090 (24 GB), ~1h40m |
Options are shuffled every epoch so the model cannot learn "A is usually right" — except
for `score`, where the order of levels carries meaning and is preserved.
## Training data
23,632 rows across 301 tasks: 13,492 choice, 6,368 noul, 3,772 score. 92% carry criteria
descriptions; 35.5% carry soft labels.
| source | rows | labels |
|---|---|---|
| Super-NaturalInstructions (288 tasks) | 8,629 | gold, with criteria written by the teacher |
| MNLI / ANLI / BoolQ / RACE / Yelp | 9,605 | gold, hand-written criteria |
| synthetic states (978, 8 domains) | 5,398 | soft labels from the teacher |
Soft labels came from **Qwen3.6-35B-A3B** read at the answer token with option order
reversed and averaged to cancel position bias. Reasoning was deliberately left **off**: with
a thinking budget the teacher commits to its answer with probability 1.0 on essentially
every row, which is useless for distillation — the measured mean entropy was 0.000 with
reasoning versus 0.211 without.
## Limitations
- **English only.**
- **Score evaluation rests on teacher labels, not gold.** All 155 held-out score rows are
synthetic; SNI has no rating scales and the HF eval split is SST-2. Choice and noul
numbers (n=1,430) are against real gold labels.
- **Serving precision matters.** The adapter was trained against a 4-bit NF4 base and
learned to correct that quantization. Served in bf16 (e.g. vLLM), accuracy holds (0.807
vs 0.799) but ECE roughly doubles, 0.036 → 0.069. A single temperature of 1.4, fitted on
half the eval set and measured on the other half, restores it to 0.033. Serve 4-bit, or
apply the correction.
- **Tokenization is part of the contract.** Feeding a differently-tokenized prompt (a
missing BOS, say) moved 16% of argmaxes in testing. Build ids the way `primitives.py`
does.
- **The teacher disagreed with gold on 24.9%** of a 3,000-row sample — mostly noisy
annotation in SBIC/civil-comments-style tasks plus adversarial ANLI. Those rows were
blended toward gold at 0.5 rather than dropped.
- Not evaluated for fairness, toxicity, or adversarial robustness. Several source tasks
(stereotype and offensiveness classification) carry the biases of their annotators.
## Licensing and provenance
The adapter is a derivative of Gemma and is governed by the
[Gemma Terms of Use](https://ai.google.dev/gemma/terms), including its use restrictions.
Training data licenses are **mixed, and some are non-commercial**:
| dataset | license |
|---|---|
| Super-NaturalInstructions | Apache-2.0 |
| ANLI | CC BY-NC 4.0 — **non-commercial** |
| RACE | research use only — **non-commercial** |
| BoolQ | CC BY-SA 3.0 |
| MNLI (GLUE) | mixed, per-genre |
| Yelp Review Full | Yelp Dataset Terms of Use |
Because ANLI and RACE rows are in the training mix, **this adapter should be treated as
non-commercial / research use** unless retrained without them. `build_data.py` in the
source repo takes `--sets` to select sources, so a commercially-clean rebuild is a flag
change plus a retrain.
Synthetic rows and soft labels were generated with Qwen3.6-35B-A3B, an open-weight model
whose license permits using outputs to train other models.