Text Classification
PEFT
Safetensors
English
lora
qlora
calibration
decision-model
system-one
typesafe
Instructions to use vagmi/jev-lite with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use vagmi/jev-lite with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
|
Download README.md from vagmi/jev-lite: direct link, hf CLI and curl.
- Browser
- Download file 8.61 kB
-
https://huggingface.co/vagmi/jev-lite/resolve/d54f613cf2fefe0ad4f29c833b682a8a42094327/README.md
- Command line
-
hf download hf://vagmi/jev-lite@d54f613cf2fefe0ad4f29c833b682a8a42094327/README.md
-
curl -L -o README.md https://huggingface.co/vagmi/jev-lite/resolve/d54f613cf2fefe0ad4f29c833b682a8a42094327/README.md
8.61 kB
| license: gemma | |
| base_model: google/gemma-4-E4B-it | |
| library_name: peft | |
| pipeline_tag: text-classification | |
| tags: | |
| - lora | |
| - qlora | |
| - calibration | |
| - decision-model | |
| - system-one | |
| - typesafe | |
| language: | |
| - en | |
| # jev-lite | |
| A QLoRA adapter that turns Gemma 4 E4B into a **System One** decision model: it reads a | |
| state, reads a typed question about it, and returns a calibrated probability distribution | |
| over the allowed answers — from **one forward pass**, with no generation. | |
| Because the answer is read from the logits at the option letters, it **cannot** answer | |
| outside the options it was given. There is no parsing, no retry loop, and no | |
| "as an AI language model". | |
| It implements the three [TypeSafe](https://docs.typesafe.ai) primitives: | |
| | type | question | returns | | |
| |---|---|---| | |
| | `choice` | pick one of several named options | the option, probabilities, confidence | | |
| | `score` | rate against ordered levels | expected level, legend, probabilities, confidence | | |
| | `noul` | is this true? | a single probability | | |
| ## Results | |
| Evaluated on 1,898 held-out rows — 31 Super-NaturalInstructions **tasks** the model never | |
| saw, 78 synthetic **states** it never saw, and SST-2 held out in its entirety. The split is | |
| by task and by state, never by row, so these are questions of kinds it was not trained on. | |
| | | accuracy | ECE | NLL | Brier | | |
| |---|---|---|---|---| | |
| | **overall** | **0.816** | **0.019** | 0.430 | 0.222 | | |
| | choice (n=1196) | 0.809 | — | 0.431 | 0.238 | | |
| | noul (n=547) | 0.819 | — | — | — | | |
| | score (n=155) | 0.852 | — | — | — | | |
| Score answers are off by **0.216 levels** on average (mean absolute error of the expected | |
| level). | |
| **Calibration is the point.** ECE of 0.019 means that when it reports 80% confidence it is | |
| right about 80% of the time. Accuracy was flat from step 250 to the end of training while | |
| ECE fell 0.086 → 0.019: the model did not learn to be right more often, it learned to be | |
| honest about when it isn't. | |
| Operationally, gating on confidence: | |
| | band | share of traffic | accuracy | | |
| |---|---|---| | |
| | ≥ 0.80 — act automatically | 62% | 93% | | |
| | 0.50–0.80 — confirm or review | 36% | 60% | | |
| | < 0.50 — route to a human | 2% | 32% | | |
| ## Confidence | |
| `confidence` is the model's probability that the answer it returned is the correct one: | |
| - **choice** — the probability of the selected option (`p_max`) | |
| - **score** — the probability mass that rounds to the reported expected level | |
| - **noul** — no confidence field; the probability *is* the answer | |
| This was chosen by measurement, not taste. Over 1,430 gold-labelled held-out rows, | |
| `p_max` beat normalized entropy, margin, chance-corrected top and collision entropy on | |
| **both** AUROC (0.816, ranking right answers above wrong ones) and ECE (0.036). Normalized | |
| entropy — the obvious first guess — was the worst of the lot: it is dominated by small | |
| probabilities, so it reads a decisive 0.85/0.08/0.07 as low confidence and routes 45% of | |
| traffic to human review at 64% accuracy. | |
| ## Usage | |
| The adapter is only meaningful with the exact prompt format it was trained on, so | |
| [`primitives.py`](./primitives.py) is included in this repo and **must** be used to render | |
| questions. Criteria descriptions are part of that format. | |
| ```python | |
| import torch, primitives | |
| from peft import PeftModel | |
| from transformers import AutoModelForImageTextToText, AutoTokenizer, BitsAndBytesConfig | |
| tok = AutoTokenizer.from_pretrained("vagmi/jev-lite") | |
| base = AutoModelForImageTextToText.from_pretrained( | |
| "google/gemma-4-E4B-it", device_map={"": 0}, dtype=torch.bfloat16, | |
| quantization_config=BitsAndBytesConfig( | |
| load_in_4bit=True, bnb_4bit_quant_type="nf4", | |
| bnb_4bit_compute_dtype=torch.bfloat16, bnb_4bit_use_double_quant=True)) | |
| model = PeftModel.from_pretrained(base, "vagmi/jev-lite").eval() | |
| row = primitives.normalize({ | |
| "type": "choice", | |
| "state": "Help! My payouts have been failing for 3 days.", | |
| "question": "Which team should handle this?", | |
| "options": ["billing", "technical", "sales"], | |
| "criteria": {"billing": "Payments, invoicing, refunds", | |
| "technical": "Bugs, outages, integrations", | |
| "sales": "Pricing, upgrades, new accounts"}, | |
| }) | |
| prompt = primitives.PREFIX + row["state"] + "\n</state>\n\n" + \ | |
| primitives.question_block(row) + "\nAnswer:" | |
| ids = torch.tensor([[tok.bos_token_id] + tok.encode(prompt, add_special_tokens=False)]) | |
| letters = [tok.encode(" " + c, add_special_tokens=False)[0] for c in primitives.LETTERS] | |
| with torch.no_grad(): | |
| logits = model(input_ids=ids.to(model.device)).logits[0, -1].float() | |
| probs = torch.softmax(logits[letters[:len(row["options"])]], -1).tolist() | |
| print(primitives.answer(row, probs)) | |
| # {'type': 'choice', 'choice': 'billing', 'probabilities': {...}, 'confidence': 0.58} | |
| ``` | |
| A server speaking the TypeSafe System One wire API (`POST /v1/systemone`), for both | |
| transformers and vLLM backends, is at | |
| [github.com/vagmi/jevlite](https://github.com/vagmi/jevlite). | |
| ## Training | |
| QLoRA on `google/gemma-4-E4B-it`, 4-bit NF4 base with bf16 compute. | |
| | | | | |
| |---|---| | |
| | LoRA rank / alpha | 16 / 32, dropout 0.05 | | |
| | target modules | q, k, v, o, gate, up, down | | |
| | trainable params | 34.9M of 7.98B (0.44%) | | |
| | objective | soft-label cross-entropy over option letters | | |
| | epochs / LR / schedule | 1 / 2e-4 / cosine, 3% warmup | | |
| | effective batch | 16 (grad accum) | | |
| | hardware | one RTX 4090 (24 GB), ~1h40m | | |
| Options are shuffled every epoch so the model cannot learn "A is usually right" — except | |
| for `score`, where the order of levels carries meaning and is preserved. | |
| ## Training data | |
| 23,632 rows across 301 tasks: 13,492 choice, 6,368 noul, 3,772 score. 92% carry criteria | |
| descriptions; 35.5% carry soft labels. | |
| | source | rows | labels | | |
| |---|---|---| | |
| | Super-NaturalInstructions (288 tasks) | 8,629 | gold, with criteria written by the teacher | | |
| | MNLI / ANLI / BoolQ / RACE / Yelp | 9,605 | gold, hand-written criteria | | |
| | synthetic states (978, 8 domains) | 5,398 | soft labels from the teacher | | |
| Soft labels came from **Qwen3.6-35B-A3B** read at the answer token with option order | |
| reversed and averaged to cancel position bias. Reasoning was deliberately left **off**: with | |
| a thinking budget the teacher commits to its answer with probability 1.0 on essentially | |
| every row, which is useless for distillation — the measured mean entropy was 0.000 with | |
| reasoning versus 0.211 without. | |
| ## Limitations | |
| - **English only.** | |
| - **Score evaluation rests on teacher labels, not gold.** All 155 held-out score rows are | |
| synthetic; SNI has no rating scales and the HF eval split is SST-2. Choice and noul | |
| numbers (n=1,430) are against real gold labels. | |
| - **Serving precision matters.** The adapter was trained against a 4-bit NF4 base and | |
| learned to correct that quantization. Served in bf16 (e.g. vLLM), accuracy holds (0.807 | |
| vs 0.799) but ECE roughly doubles, 0.036 → 0.069. A single temperature of 1.4, fitted on | |
| half the eval set and measured on the other half, restores it to 0.033. Serve 4-bit, or | |
| apply the correction. | |
| - **Tokenization is part of the contract.** Feeding a differently-tokenized prompt (a | |
| missing BOS, say) moved 16% of argmaxes in testing. Build ids the way `primitives.py` | |
| does. | |
| - **The teacher disagreed with gold on 24.9%** of a 3,000-row sample — mostly noisy | |
| annotation in SBIC/civil-comments-style tasks plus adversarial ANLI. Those rows were | |
| blended toward gold at 0.5 rather than dropped. | |
| - Not evaluated for fairness, toxicity, or adversarial robustness. Several source tasks | |
| (stereotype and offensiveness classification) carry the biases of their annotators. | |
| ## Licensing and provenance | |
| The adapter is a derivative of Gemma and is governed by the | |
| [Gemma Terms of Use](https://ai.google.dev/gemma/terms), including its use restrictions. | |
| Training data licenses are **mixed, and some are non-commercial**: | |
| | dataset | license | | |
| |---|---| | |
| | Super-NaturalInstructions | Apache-2.0 | | |
| | ANLI | CC BY-NC 4.0 — **non-commercial** | | |
| | RACE | research use only — **non-commercial** | | |
| | BoolQ | CC BY-SA 3.0 | | |
| | MNLI (GLUE) | mixed, per-genre | | |
| | Yelp Review Full | Yelp Dataset Terms of Use | | |
| Because ANLI and RACE rows are in the training mix, **this adapter should be treated as | |
| non-commercial / research use** unless retrained without them. `build_data.py` in the | |
| source repo takes `--sets` to select sources, so a commercially-clean rebuild is a flag | |
| change plus a retrain. | |
| Synthetic rows and soft labels were generated with Qwen3.6-35B-A3B, an open-weight model | |
| whose license permits using outputs to train other models. | |