Text Classification
PEFT
Safetensors
English
decision-model
calibration
lora
multiple-choice
typesafe
qwen3.5
Eval Results (legacy)
Instructions to use jaredpalmer/kev-4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use jaredpalmer/kev-4b with PEFT:
from peft import PeftModel from transformers import AutoModel base_model = AutoModel.from_pretrained("Qwen/Qwen3.5-4B-Base") model = PeftModel.from_pretrained(base_model, "jaredpalmer/kev-4b") - Notebooks
- Google Colab
- Kaggle
Card: product-example caveat
Browse files
README.md
CHANGED
|
@@ -78,6 +78,7 @@ Seeds: the recipe was run at two seeds on this suite (transfer 0.759 / 0.758) an
|
|
| 78 |
## Known limits
|
| 79 |
|
| 80 |
- Held-out policy reasoning (unseen rule compositions, date arithmetic with grace periods) is far from Jev.
|
|
|
|
| 81 |
- Out-of-domain probabilities are usable but not calibrated (raw ECE 0.096); temperature fitted in-domain does not transfer.
|
| 82 |
- 4B fp32 needs ~16 GB; on a 32 GB Mac use `KEV_DTYPE=bf16`. Latency on an H100 is ~45 ms per packed request; on an M5 several hundred ms.
|
| 83 |
|
|
|
|
| 78 |
## Known limits
|
| 79 |
|
| 80 |
- Held-out policy reasoning (unseen rule compositions, date arithmetic with grace periods) is far from Jev.
|
| 81 |
+
- Product-shaped questions with no training analogue are not guaranteed: on the TypeSafe docs example ("two charges on my card" → *Is there a billing problem?*) this checkpoint answers 0.22 while kev-0.6b answers 0.97 and picks the return reason (wrong size, 0.54) correctly. Lower drift from the base means fewer task-specific priors; measure on your own inputs.
|
| 82 |
- Out-of-domain probabilities are usable but not calibrated (raw ECE 0.096); temperature fitted in-domain does not transfer.
|
| 83 |
- 4B fp32 needs ~16 GB; on a 32 GB Mac use `KEV_DTYPE=bf16`. Latency on an H100 is ~45 ms per packed request; on an M5 several hundred ms.
|
| 84 |
|