Text Classification
Transformers
Safetensors
qwen3_5
image-text-to-text
typed-decisions
calibrated-classification
system-one
classification
structured-prediction
candidate-logit
jev
single-forward-pass
multilingual
commercial-use
Instructions to use wayfind/metask-jev-4b-policy-mix with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use wayfind/metask-jev-4b-policy-mix with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="wayfind/metask-jev-4b-policy-mix")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("wayfind/metask-jev-4b-policy-mix") model = AutoModelForMultimodalLM.from_pretrained("wayfind/metask-jev-4b-policy-mix", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Upload README.md with huggingface_hub
Browse files
README.md
CHANGED
|
@@ -35,15 +35,15 @@ pipeline_tag: text-classification
|
|
| 35 |
|
| 36 |
# Metask-Jev-4B
|
| 37 |
|
| 38 |
-
A calibrated **typed-decision model** in **16 languages**: give it a state (text, ticket, policy, JSON) and a typed question — `choice`, `boolean`, or rubric `score` — and it returns a probability for every option in a **single forward pass (~
|
| 39 |
|
| 40 |
## On the JevBench board
|
| 41 |
|
| 42 |
-
Self-measured axes inserted into the published v1.2.7 ranking (26 official entrants + this model). Official run pending — axes use our 231-decision protocol for Intelligence, val-fit temperature for Calibration, and production
|
| 43 |
|
| 44 |
<img src="eval/figs/fig6_board_style.png" width="660" alt="JevBench board with metask-jev-4b">
|
| 45 |
|
| 46 |
-
Would rank **#1** — ahead of Jev 1.13.0 itself — under this estimate, and occupies the top-right corner of the Intelligence×Speed plane outright (no ranked system, open or closed, beats I 88.5 / S
|
| 47 |
|
| 48 |
<img src="eval/figs/fig7_scatter.png" width="660" alt="Intelligence vs Speed scatter">
|
| 49 |
|
|
@@ -147,11 +147,11 @@ Per-kind temperatures: **choice 1.9 / noul 2.375 / score 2.3**. Answer tokens A
|
|
| 147 |
|---|---:|---:|---:|
|
| 148 |
| 13 human-labeled subsets (3,880 items), macro | **78.9%** | 74.8% | 76.0% |
|
| 149 |
| JevBench v1.2 public 231 @4096 ctx | **80.1%** | 63.5% | 75.3 |
|
| 150 |
-
| JevBench Score (official-methodology estimate) | **
|
| 151 |
| MASSIVE 14-locale dev held-out macro | **85.9%** | — | — |
|
| 152 |
| ECE after per-kind temperature | **0.040** | — | — |
|
| 153 |
-
| p50 latency (
|
| 154 |
-
| serving cost (owned
|
| 155 |
|
| 156 |
**12 of 13 subsets exceed Bespoke Nimble-9B** — a model 2.2× its size — same prompt format, same scoring protocol.
|
| 157 |
|
|
|
|
| 35 |
|
| 36 |
# Metask-Jev-4B
|
| 37 |
|
| 38 |
+
A calibrated **typed-decision model** in **16 languages**: give it a state (text, ticket, policy, JSON) and a typed question — `choice`, `boolean`, or rubric `score` — and it returns a probability for every option in a **single forward pass (~63 ms measured p50 on a 4090)**. No generation, no parsing, nothing to hallucinate.
|
| 39 |
|
| 40 |
## On the JevBench board
|
| 41 |
|
| 42 |
+
Self-measured axes inserted into the published v1.2.7 ranking (26 official entrants + this model). Official run pending — axes use our 231-decision protocol for Intelligence, val-fit temperature for Calibration, and measured production numbers for Speed/Cost: JevBench-231 p50 **62.8 ms** on a 4090 → adjusted 0.276 s (official ×2 + 0.15 s self-hosted formula) → **S 91.2**; owned-hardware cost ¥6,000/month for an 8×4090 server (this model fits twice on one card — 2×9.1 GB weights — and sustains **~20 QPS per card with dual replicas**) → $105/card/month ÷ (20 QPS × 70% utilization) ≈ **$0.0029 per 1,000 decisions → K 86.2** — an order of magnitude below every ranked system.
|
| 43 |
|
| 44 |
<img src="eval/figs/fig6_board_style.png" width="660" alt="JevBench board with metask-jev-4b">
|
| 45 |
|
| 46 |
+
Would rank **#1** — ahead of Jev 1.13.0 itself — under this estimate, and occupies the top-right corner of the Intelligence×Speed plane outright (no ranked system, open or closed, beats I 88.5 / S 91.2 on both axes):
|
| 47 |
|
| 48 |
<img src="eval/figs/fig7_scatter.png" width="660" alt="Intelligence vs Speed scatter">
|
| 49 |
|
|
|
|
| 147 |
|---|---:|---:|---:|
|
| 148 |
| 13 human-labeled subsets (3,880 items), macro | **78.9%** | 74.8% | 76.0% |
|
| 149 |
| JevBench v1.2 public 231 @4096 ctx | **80.1%** | 63.5% | 75.3 |
|
| 150 |
+
| JevBench Score (official-methodology estimate) | **86.4** (would rank #1) | 61.8 | 75.4 |
|
| 151 |
| MASSIVE 14-locale dev held-out macro | **85.9%** | — | — |
|
| 152 |
| ECE after per-kind temperature | **0.040** | — | — |
|
| 153 |
+
| p50 latency (JevBench 231, 4090) | **62.8 ms** | ~190 ms | 236–276 ms |
|
| 154 |
+
| serving cost (owned 8×4090, dual-replica 20 QPS) | **$0.0029/1k** | $0.166 | $0.040 |
|
| 155 |
|
| 156 |
**12 of 13 subsets exceed Bespoke Nimble-9B** — a model 2.2× its size — same prompt format, same scoring protocol.
|
| 157 |
|