Jev-Omni general decision head v3
This repository contains a 256-slot decision head for akhilaaa3/Jev-Omni, not the 12B backbone. The backbone stayed frozen. The v2 general head was retrained with full replay of 47,680 mixed text, image, and video decisions plus 7,962 labeled-control computer-use decisions, totaling 55,642 training examples. The new data came from Multimodal Mind2Web and BrowserGym MiniWoB action-only trajectories. Source media and datasets are not redistributed.
Held-out evaluation
| Derived task / split | v2 head | v3 head |
|---|---|---|
| Mind2Web official test-task source split, four named candidates | 851/1,086 (78.36%) | 853/1,086 (78.55%) |
| Existing general test, all tasks | 7,983/9,284 (85.99%) | 8,051/9,284 (86.72%) |
| Existing general test, vision | 5,677/6,754 (84.05%) | 5,749/6,754 (85.12%) |
| Existing general test, text | 2,306/2,530 (91.15%) | 2,302/2,530 (90.99%) |
The Mind2Web set uses the official test-task source split, but the four-candidate marked-screenshot task is derived and is not the published Mind2Web leaderboard metric. Its target is guaranteed among four preselected named controls. It does not measure full-page retrieval or completed browser tasks. The MiniWoB supervision is teacher generated and only retained click rows. See metrics.json for per-operation, per-family, log-loss, and checkpoint details. Raw probabilities and confidence are not demonstrated to be calibrated for deployment.
Load
Run inference on a remote CUDA GPU such as Modal:
from pathlib import Path
import sys
import torch
from huggingface_hub import hf_hub_download, snapshot_download
base = snapshot_download("akhilaaa3/Jev-Omni", revision="5addda86ddee081a68fb067477ea100c221b8917")
source = hf_hub_download("akhilaaa3/Jev-Omni", "jev_omni.py", revision="5addda86ddee081a68fb067477ea100c221b8917")
sys.path.insert(0, str(Path(source).parent))
import jev_omni
jev_omni.snapshot_download = lambda *_args, **_kwargs: base
classifier = jev_omni.load_jev_omni()
head = hf_hub_download("ferdinandl007/jev-omni-general-head-v3", "decision_head.pt")
classifier.head.load_state_dict(torch.load(head, map_location=classifier.device, weights_only=True))
classifier.head.eval()
print(classifier.predict(state="A meeting starts at 10:00; it is 09:00.",
question="Has it started?", options=["No", "Yes"]))
The training and evaluation scripts at GitHub commit 565e72dfc471 include a Python adapter for Choice, Noul, and Score response types. This model is independent of TypeSafe AI's Jev service. Code is MIT licensed; upstream model and dataset rights are separate. The Mind2Web source is research-only under OpenRAIL; inspect each upstream source before other uses.
- Downloads last month
- 14