File size: 6,876 Bytes
eb84ce2
 
 
 
 
 
8504f62
 
 
eb84ce2
 
 
 
 
 
 
 
 
 
 
 
 
8504f62
eb84ce2
39a7e39
 
7418ae8
8504f62
7418ae8
eb84ce2
 
 
 
c8f9e9c
3ef38c6
eb84ce2
 
 
 
 
c8f9e9c
eb84ce2
 
 
 
8504f62
eb84ce2
 
 
 
 
 
 
 
 
 
 
 
 
 
8504f62
eb84ce2
 
 
 
d0d93a2
7418ae8
d0d93a2
eb84ce2
 
 
 
 
7418ae8
eb84ce2
 
 
 
 
7418ae8
 
eb84ce2
 
 
 
 
 
 
7418ae8
eb84ce2
98ac595
eb84ce2
 
 
7418ae8
eb84ce2
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
c8f9e9c
eb84ce2
 
 
 
 
 
 
 
 
 
 
 
d0d93a2
 
 
eb84ce2
 
 
 
 
 
 
 
a1bd162
 
7418ae8
 
a1bd162
 
 
3ef38c6
 
 
 
 
 
a1bd162
 
 
eb84ce2
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
---
license: apache-2.0
base_model: Qwen/Qwen3-1.7B-Base
language:
- en
tags:
- jev
- jev-alternative
- open-source-jev
- decision-model
- browser-agent
- system-one
- lora
datasets:
- osunlp/Mind2Web
- stanfordnlp/nnetnav-live
- LocalLLaMA/typed-decisions
- tasksource/tasksource-jev
- SargeDev/jev-distill-corpus-v3
- n4ze3m/typed-decisions-synth
---

# OpenJev-1.7B

**Jun Huang\*, Xin Ren\*** · University of Electronic Science and Technology of China · \*Equal contribution

**OpenJev** is an open-source, local alternative to Jev's System One decision API: same request shape, your own GPU, no API key.

**A local decision model: typed questions in, calibrated probabilities out, in one forward pass.** OpenJev-1.7B answers
the `POST /v1/systemone` request shape (choice, yes/no and score questions over a free-form state), for general
decisions and for browser-agent steps (*which operation? which element?*). It runs on your own GPU: no API key,
no per-call cost, nothing generated.

Code, training and evaluation: [alanhuangyoo/OpenJev](https://github.com/alanhuangyoo/OpenJev). Paper:
[doi:10.5281/zenodo.22941164](https://zenodo.org/records/22941164). Independent project; not affiliated
with TypeSafe AI.

## Quickstart

```bash
pip install "wev-ai[serve]"   # or: pip install "openjev-ai[serve]"
```

```python
import wev
m = wev.load("alanhuangya/OpenJev-1.7B")
out = m.predict(
    state="Refund request: order #4411 arrived damaged, customer attached photos, first refund this year.",
    questions={
        "action": {"type": "choice", "instructions": "What should support do?",
                   "criteria": {"refund": "Refund the order.", "replace": "Ship a replacement.",
                                "escalate": "Send to a human agent."}},
        "fraud_risk": {"type": "noul", "instructions": "This request looks fraudulent.",
                       "criteria": {"true": "Likely fraud.", "false": "No sign of fraud."}},
    },
)
print(out["answers"])
```

```bash
wev serve --model alanhuangya/OpenJev-1.7B --port 8009   # drop-in POST /v1/systemone, e.g. for jev-ultrafast
```

## Results

Test splits, held out from training; every other model was run on the same requests and scored the same way
(per-question accuracy; `scripts/compare.py`). For OpenJev-4B and OpenJev-8B, two candidates each were read on test (see the
repository README).

**General typed decisions**

| model | kev decision-v7 | kev transfer-v4 | typed-decisions |
|---|---|---|---|
| **OpenJev-1.7B** | 81.1 | 65.5 | **79.5** |
| Kev-4B | **88.2** | **82.1** | 65.1 |
| Kev-8B | 88.1 | 76.8 | 62.7 |
| Laya (typed-decisions) | 65.7 | 62.8 | 76.8 |
| Laya | 64.3 | 63.7 | 36.2 |

OpenJev-1.7B trains on 80% of the typed-decisions train split, like the Laya (typed-decisions) specialist; Kev and Laya
do not, so on that column they are generalists. kev decision-v7 is Kev's own training suite (OpenJev-1.7B also trains on
its train split); transfer-v4 is out-of-domain for every model here.

**Browser steps** (Mind2Web test split: websites unseen in training, jev-ultrafast request format; step success =
operation and target element both right)

| model | step success | operation |
|---|---|---|
| **OpenJev-1.7B** | **68.2** | **88.1** |
| Kev-4B | 21.2 | 35.7 |
| Kev-8B | 19.0 | 73.3 |
| Laya (typed-decisions) | 0.7 | 13.1 |
| Laya | 0.0 | 2.5 |

873 requests; 11 exceed the context OpenJev-1.7B is evaluated with and count as wrong for it.

NNetNav test split (live-web steps, DONE judged by an LLM): step success 61.0, DONE recall
80.4, premature DONE 9.4.

## Model

- Backbone: `Qwen/Qwen3-1.7B-Base` without its vocabulary head, LoRA r=16 on every attention and MLP projection, merged
  into the weights of this export; 28 layers, bf16.
- Readout: a pointer head scores each option's `</opt>` state against the question's `<decide>` state.
- Each question sees the state and itself only (block-causal branches, positions restart after the state), so a
  request with many questions costs one pass and answers never depend on question order.
- Context: state up to 4096 tokens, each question up to 8192 tokens (trained with
  2048); longer page states are shrunk before encoding.

## Training

1 epoch, lr 0.0001, one-cycle schedule, soft-label cross-entropy where the source has soft labels.
Recipe and data builders: [alanhuangyoo/OpenJev](https://github.com/alanhuangyoo/OpenJev).

| source | license | what it adds |
|---|---|---|
| [Mind2Web](https://huggingface.co/datasets/osunlp/Mind2Web) | CC BY 4.0 | human browser steps: click, type, select |
| [NNetNav-live](https://huggingface.co/datasets/stanfordnlp/nnetnav-live) | Apache-2.0 | live-web steps; DONE relabelled by an LLM judge |
| teacher episodes | outputs of qwen3-max | jev-ultrafast on live sites with qwen3-max as System One, success judge-verified |
| [kev decision-v7](https://github.com/jaredpalmer/kev) | per source | ten public classification / QA sources plus rule records |
| [typed-decisions](https://huggingface.co/datasets/LocalLLaMA/typed-decisions) | Apache-2.0 | agent / ops workflows, 5 questions per case (80% of train) |
| [tasksource-jev](https://huggingface.co/datasets/tasksource/tasksource-jev) | mixed (per source task; some research-only) | hundreds of classification tasks as decisions |
| [jev-distill-corpus-v3](https://huggingface.co/datasets/SargeDev/jev-distill-corpus-v3) | Apache-2.0 | synthetic operational scenarios, soft labels |
| [typed-decisions-synth](https://huggingface.co/datasets/n4ze3m/typed-decisions-synth) | MIT | multi-question cases over 149 domains |

**Use terms.** Some training data carries its own terms: several tasksource-jev source tasks are research-only, and the
teacher episodes are qwen3-max outputs subject to its provider's terms. Treat this model as a research artifact and
check those terms before any commercial use.

## Limitations

- English only. Decisions, not text: TYPE values come from a separate text model, as in jev-ultrafast.
- Browser targets are scored among the candidates the agent lists (8–40 per step), not every element on the page.
- DONE and BLOCKED are the hardest operations; gate DONE on its probability when early stops are costly.
- Not compared with Jev itself (no API access).

## Citation

Jun Huang and Xin Ren contributed equally (University of Electronic Science and Technology of China). The paper
describes OpenJev under its earlier name, wev.

```bibtex
@misc{huang2026wev,
  title     = {wev: Distilling LLM Browser Agents into Open, Local System-One Decision Models},
  author    = {Huang, Jun and Ren, Xin},
  year      = {2026},
  publisher = {Zenodo},
  doi       = {10.5281/zenodo.22941164},
  url       = {https://doi.org/10.5281/zenodo.22941164}
}
```

## License

Apache-2.0, like the base model. Architecture code adapted from [kev](https://github.com/jaredpalmer/kev) (Apache-2.0).