Text Classification
Transformers
Safetensors
English
Turkish
German
qwen3_5
image-text-to-text
decision-model
calibration
conformal-prediction
uncertainty
reasoning
routing
triage
jev
typesafe
qwen3.5
english
small-language-model
local-llm
on-device
english-llm
Instructions to use mertkayacs/Deem-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use mertkayacs/Deem-4B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="mertkayacs/Deem-4B")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("mertkayacs/Deem-4B") model = AutoModelForMultimodalLM.from_pretrained("mertkayacs/Deem-4B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
card: results as charts (Jev 1.13, Kev-4B, Laya), 4K share card
Browse files- README.md +3 -3
- assets/jev.png +2 -2
README.md
CHANGED
|
@@ -20,7 +20,7 @@ An English decision model with the Jev API. You send a state and typed questions
|
|
| 20 |
|
| 21 |
<video controls playsinline preload="none" poster="https://huggingface.co/datasets/mertkayacs/emberwick-videos/resolve/main/film/jevalt-film-en.jpg" src="https://huggingface.co/datasets/mertkayacs/emberwick-videos/resolve/main/film/jevalt-film-en-1080p.mp4"></video>
|
| 22 |
|
| 23 |
-
*The
|
| 24 |
|
| 25 |
## Try it
|
| 26 |
|
|
@@ -68,12 +68,12 @@ client = TypeSafeClient(api_key="local", base_url="http://127.0.0.1:8000")
|
|
| 68 |
|
| 69 |

|
| 70 |
|
| 71 |
-
Same items and client for every model, each as shipped: [Kev-4B](https://huggingface.co/jaredpalmer/kev-4b) r10 and [Laya](https://huggingface.co/convaiinnovations/laya) 0.3.22 on their own servers with their own calibration. Jev 1.13 rows come from
|
| 72 |
|
| 73 |
<details>
|
| 74 |
<summary><b>Significance and caveats</b></summary>
|
| 75 |
|
| 76 |
-

|
| 70 |
|
| 71 |
+
Same items and client for every model, each as shipped: [Kev-4B](https://huggingface.co/jaredpalmer/kev-4b) r10 and [Laya](https://huggingface.co/convaiinnovations/laya) 0.3.22 on their own servers with their own calibration. Jev 1.13 rows come from TypeSafe's [API reference](https://docs.typesafe.ai/api) and [Models page](https://docs.typesafe.ai/models). The held-out tests come from JevAlt's own data pipeline, so they favour JevAlt. Kev-4B and Laya both do better on long padding; Laya is far smaller and faster. Every number and every decision: [results/comparison](https://huggingface.co/datasets/mertkayacs/jevalt-bench/tree/main/results/comparison).
|
| 72 |
|
| 73 |
<details>
|
| 74 |
<summary><b>Significance and caveats</b></summary>
|
| 75 |
|
| 76 |
+

|
| 77 |
|
| 78 |
The three test splits went through the same pipeline as the training rows, so they measure what the training aimed at. On the English split Deem-4B gains 4.4 accuracy points (paired bootstrap, 95% interval +3.4 to +5.5) and lowers Brier by 0.075; the Turkish and German splits move by +4.8 and +12.1 points. On JevBench-hard, TurkishMMLU and GermEval, which the training never saw, and on the typed-decisions test split (its train split was in the mix), accuracy does not change significantly, and Brier gets slightly worse on typed-decisions (+0.013) and 10kGNAD (+0.047). With `reasoning: "auto"` the English date, number and policy test rows go from 0.761 to 0.769 accuracy, an interval that touches zero.
|
| 79 |
|
assets/jev.png
CHANGED
|
Git LFS Details
|
|
Git LFS Details
|