Text Classification
Transformers
Safetensors
English
qwen3_5
image-text-to-text
typed-decisions
calibrated-classification
system-one
classification
structured-prediction
candidate-logit
jev
single-forward-pass
commercial-use
Instructions to use Raymond1122/metask-jev-4b-policy-mix with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Raymond1122/metask-jev-4b-policy-mix with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="Raymond1122/metask-jev-4b-policy-mix")# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Raymond1122/metask-jev-4b-policy-mix") model = AutoModelForMultimodalLM.from_pretrained("Raymond1122/metask-jev-4b-policy-mix", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Laya-style model card: benchmark figures, 13-subset table w/ CI, honest limits
Browse files- .gitattributes +2 -0
- README.md +10 -0
- eval/figs/fig6_board_style.png +3 -0
- eval/figs/fig7_scatter.png +3 -0
.gitattributes
CHANGED
|
@@ -34,3 +34,5 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
| 36 |
tokenizer.json filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
|
|
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
| 36 |
tokenizer.json filter=lfs diff=lfs merge=lfs -text
|
| 37 |
+
eval/figs/fig6_board_style.png filter=lfs diff=lfs merge=lfs -text
|
| 38 |
+
eval/figs/fig7_scatter.png filter=lfs diff=lfs merge=lfs -text
|
README.md
CHANGED
|
@@ -99,6 +99,16 @@ Single forward pass over the prompt, one softmax over ≤26 candidate logits.
|
|
| 99 |
|
| 100 |
Objective and prompt format are unchanged from the official Nimble protocol; the recipe card with reproduction commands lives in the [GitHub repo](https://github.com/metask-ai/metask-jev).
|
| 101 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 102 |
## Honest limits
|
| 103 |
|
| 104 |
- **summeval-relevance (26.7%)** is the one clear regression vs 9B (49.2%): a 5-level rubric with a systematic 3↔4 boundary shift. NLL and expected-score error are actually *better* than 9B — the argmax metric amplifies the boundary shift. If your use case is fine-grained relevance scoring, evaluate this subset yourself first.
|
|
|
|
| 99 |
|
| 100 |
Objective and prompt format are unchanged from the official Nimble protocol; the recipe card with reproduction commands lives in the [GitHub repo](https://github.com/metask-ai/metask-jev).
|
| 101 |
|
| 102 |
+
## On the JevBench board
|
| 103 |
+
|
| 104 |
+
Self-measured axes inserted into the published v1.2.7 ranking (16 official entrants + this model). Official run pending — axes here use our 231-decision protocol for Intelligence, val-fit temperature for Calibration, self-hosted 4090 for Speed/Cost.
|
| 105 |
+
|
| 106 |
+
<img src="eval/figs/fig6_board_style.png" width="660" alt="JevBench board with metask-jev-4b">
|
| 107 |
+
|
| 108 |
+
Would rank **#5** — ahead of GPT-5.6 Luna and DeepSeek V4.1 Flash, behind djev — with the top-right quadrant of the Intelligence×Speed plane to itself among open weights:
|
| 109 |
+
|
| 110 |
+
<img src="eval/figs/fig7_scatter.png" width="660" alt="Intelligence vs Speed scatter">
|
| 111 |
+
|
| 112 |
## Honest limits
|
| 113 |
|
| 114 |
- **summeval-relevance (26.7%)** is the one clear regression vs 9B (49.2%): a 5-level rubric with a systematic 3↔4 boundary shift. NLL and expected-score error are actually *better* than 9B — the argmax metric amplifies the boundary shift. If your use case is fine-grained relevance scoring, evaluate this subset yourself first.
|
eval/figs/fig6_board_style.png
ADDED
|
Git LFS Details
|
eval/figs/fig7_scatter.png
ADDED
|
Git LFS Details
|