Jev-Omni general decision head v3

This repository contains a 256-slot decision head for akhilaaa3/Jev-Omni, not the 12B backbone. The backbone stayed frozen. The v2 general head was retrained with full replay of 47,680 mixed text, image, and video decisions plus 7,962 labeled-control computer-use decisions, totaling 55,642 training examples. The new data came from Multimodal Mind2Web and BrowserGym MiniWoB action-only trajectories. Source media and datasets are not redistributed.

Held-out evaluation

Derived task / split v2 head v3 head
Mind2Web official test-task source split, four named candidates 851/1,086 (78.36%) 853/1,086 (78.55%)
Existing general test, all tasks 7,983/9,284 (85.99%) 8,051/9,284 (86.72%)
Existing general test, vision 5,677/6,754 (84.05%) 5,749/6,754 (85.12%)
Existing general test, text 2,306/2,530 (91.15%) 2,302/2,530 (90.99%)

The Mind2Web set uses the official test-task source split, but the four-candidate marked-screenshot task is derived and is not the published Mind2Web leaderboard metric. Its target is guaranteed among four preselected named controls. It does not measure full-page retrieval or completed browser tasks. The MiniWoB supervision is teacher generated and only retained click rows. See metrics.json for per-operation, per-family, log-loss, and checkpoint details. Raw probabilities and confidence are not demonstrated to be calibrated for deployment.

Load

Run inference on a remote CUDA GPU such as Modal:

from pathlib import Path
import sys
import torch
from huggingface_hub import hf_hub_download, snapshot_download

base = snapshot_download("akhilaaa3/Jev-Omni", revision="5addda86ddee081a68fb067477ea100c221b8917")
source = hf_hub_download("akhilaaa3/Jev-Omni", "jev_omni.py", revision="5addda86ddee081a68fb067477ea100c221b8917")
sys.path.insert(0, str(Path(source).parent))
import jev_omni
jev_omni.snapshot_download = lambda *_args, **_kwargs: base
classifier = jev_omni.load_jev_omni()
head = hf_hub_download("ferdinandl007/jev-omni-general-head-v3", "decision_head.pt")
classifier.head.load_state_dict(torch.load(head, map_location=classifier.device, weights_only=True))
classifier.head.eval()
print(classifier.predict(state="A meeting starts at 10:00; it is 09:00.",
                         question="Has it started?", options=["No", "Yes"]))

The training and evaluation scripts at GitHub commit 565e72dfc471 include a Python adapter for Choice, Noul, and Score response types. This model is independent of TypeSafe AI's Jev service. Code is MIT licensed; upstream model and dataset rights are separate. The Mind2Web source is research-only under OpenRAIL; inspect each upstream source before other uses.

Downloads last month
14
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ferdinandl007/jev-omni-general-head-v3

Finetuned
(3)
this model