Model card: end-to-end results
Browse files
README.md
ADDED
|
@@ -0,0 +1,129 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
base_model: Qwen/Qwen3-1.7B-Base
|
| 4 |
+
language:
|
| 5 |
+
- en
|
| 6 |
+
tags:
|
| 7 |
+
- decision-model
|
| 8 |
+
- browser-agent
|
| 9 |
+
- system-one
|
| 10 |
+
- lora
|
| 11 |
+
datasets:
|
| 12 |
+
- osunlp/Mind2Web
|
| 13 |
+
- stanfordnlp/nnetnav-live
|
| 14 |
+
- LocalLLaMA/typed-decisions
|
| 15 |
+
- tasksource/tasksource-jev
|
| 16 |
+
- SargeDev/jev-distill-corpus-v3
|
| 17 |
+
- n4ze3m/typed-decisions-synth
|
| 18 |
+
---
|
| 19 |
+
|
| 20 |
+
# wev-1.7b
|
| 21 |
+
|
| 22 |
+
**A local decision model: typed questions in, calibrated probabilities out, in one forward pass.** `wev-1.7b` answers
|
| 23 |
+
the `POST /v1/systemone` request shape (choice, yes/no and score questions over a free-form state), for general
|
| 24 |
+
decisions and for browser-agent steps (*which operation? which element?*). It runs on your own GPU: no API key,
|
| 25 |
+
no per-call cost, nothing generated.
|
| 26 |
+
|
| 27 |
+
Code, training and evaluation: [alanhuangyoo/wev](https://github.com/alanhuangyoo/wev). Independent project; not affiliated
|
| 28 |
+
with TypeSafe AI.
|
| 29 |
+
|
| 30 |
+
## Quickstart
|
| 31 |
+
|
| 32 |
+
```bash
|
| 33 |
+
pip install "wev-ai[serve]"
|
| 34 |
+
```
|
| 35 |
+
|
| 36 |
+
```python
|
| 37 |
+
import wev
|
| 38 |
+
m = wev.load("alanhuangya/wev-1.7b")
|
| 39 |
+
out = m.predict(
|
| 40 |
+
state="Refund request: order #4411 arrived damaged, customer attached photos, first refund this year.",
|
| 41 |
+
questions={
|
| 42 |
+
"action": {"type": "choice", "instructions": "What should support do?",
|
| 43 |
+
"criteria": {"refund": "Refund the order.", "replace": "Ship a replacement.",
|
| 44 |
+
"escalate": "Send to a human agent."}},
|
| 45 |
+
"fraud_risk": {"type": "noul", "instructions": "This request looks fraudulent.",
|
| 46 |
+
"criteria": {"true": "Likely fraud.", "false": "No sign of fraud."}},
|
| 47 |
+
},
|
| 48 |
+
)
|
| 49 |
+
print(out["answers"])
|
| 50 |
+
```
|
| 51 |
+
|
| 52 |
+
```bash
|
| 53 |
+
wev serve --model alanhuangya/wev-1.7b --port 8009 # drop-in POST /v1/systemone, e.g. for jev-ultrafast
|
| 54 |
+
```
|
| 55 |
+
|
| 56 |
+
## Results
|
| 57 |
+
|
| 58 |
+
One read of the locked test splits; every other model was run on the same requests and scored the same way
|
| 59 |
+
(per-question accuracy; `scripts/compare.py`).
|
| 60 |
+
|
| 61 |
+
**General typed decisions**
|
| 62 |
+
|
| 63 |
+
| model | kev decision-v7 | kev transfer-v4 | typed-decisions |
|
| 64 |
+
|---|---|---|---|
|
| 65 |
+
| **wev-1.7b** | 81.1 | 65.5 | **79.5** |
|
| 66 |
+
| Kev-4B | **88.2** | **82.1** | 65.1 |
|
| 67 |
+
| Kev-8B | 88.1 | 76.8 | 62.7 |
|
| 68 |
+
| Laya (typed-decisions) | 65.7 | 62.8 | 76.8 |
|
| 69 |
+
| Laya | 64.3 | 63.7 | 36.2 |
|
| 70 |
+
|
| 71 |
+
`wev-1.7b` trains on 80% of the typed-decisions train split, like the Laya (typed-decisions) specialist; Kev and Laya
|
| 72 |
+
do not, so on that column they are generalists. kev decision-v7 is Kev's own training suite (`wev-1.7b` also trains on
|
| 73 |
+
its train split); transfer-v4 is out-of-domain for every model here.
|
| 74 |
+
|
| 75 |
+
**Browser steps** (Mind2Web test split: websites unseen in training, jev-ultrafast request format; step success =
|
| 76 |
+
operation and target element both right)
|
| 77 |
+
|
| 78 |
+
| model | step success | operation |
|
| 79 |
+
|---|---|---|
|
| 80 |
+
| **wev-1.7b** | **68.2** | **88.1** |
|
| 81 |
+
| Kev-4B | 21.2 | 35.7 |
|
| 82 |
+
| Kev-8B | — | — |
|
| 83 |
+
| Laya (typed-decisions) | 0.7 | 13.1 |
|
| 84 |
+
| Laya | 0.0 | 2.5 |
|
| 85 |
+
|
| 86 |
+
873 requests; 11 exceed the context `wev-1.7b` is evaluated with and count as wrong for it.
|
| 87 |
+
|
| 88 |
+
NNetNav test split (live-web steps, DONE judged by an LLM): step success 61.0, DONE recall
|
| 89 |
+
80.4, premature DONE 9.4.
|
| 90 |
+
|
| 91 |
+
## Model
|
| 92 |
+
|
| 93 |
+
- Backbone: `Qwen/Qwen3-1.7B-Base` without its vocabulary head, LoRA r=16 on every attention and MLP projection, merged
|
| 94 |
+
into the weights of this export; 28 layers, bf16.
|
| 95 |
+
- Readout: a pointer head scores each option's `</opt>` state against the question's `<decide>` state.
|
| 96 |
+
- Each question sees the state and itself only (block-causal branches, positions restart after the state), so a
|
| 97 |
+
request with many questions costs one pass and answers never depend on question order.
|
| 98 |
+
- Context: state up to 4096 tokens, each question up to 8192 tokens (trained with
|
| 99 |
+
2048); longer page states are shrunk before encoding.
|
| 100 |
+
|
| 101 |
+
## Training
|
| 102 |
+
|
| 103 |
+
1 epoch, lr 0.0001, one-cycle schedule, soft-label cross-entropy where the source has soft labels.
|
| 104 |
+
Recipe and data builders: [alanhuangyoo/wev](https://github.com/alanhuangyoo/wev).
|
| 105 |
+
|
| 106 |
+
| source | license | what it adds |
|
| 107 |
+
|---|---|---|
|
| 108 |
+
| [Mind2Web](https://huggingface.co/datasets/osunlp/Mind2Web) | CC BY 4.0 | human browser steps: click, type, select |
|
| 109 |
+
| [NNetNav-live](https://huggingface.co/datasets/stanfordnlp/nnetnav-live) | Apache-2.0 | live-web steps; DONE relabelled by an LLM judge |
|
| 110 |
+
| teacher episodes | outputs of qwen3-max | jev-ultrafast on live sites with qwen3-max as System One, success judge-verified |
|
| 111 |
+
| [kev decision-v7](https://github.com/jaredpalmer/kev) | per source | ten public classification / QA sources plus rule records |
|
| 112 |
+
| [typed-decisions](https://huggingface.co/datasets/LocalLLaMA/typed-decisions) | Apache-2.0 | agent / ops workflows, 5 questions per case (80% of train) |
|
| 113 |
+
| [tasksource-jev](https://huggingface.co/datasets/tasksource/tasksource-jev) | mixed (per source task; some research-only) | hundreds of classification tasks as decisions |
|
| 114 |
+
| [jev-distill-corpus-v3](https://huggingface.co/datasets/SargeDev/jev-distill-corpus-v3) | Apache-2.0 | synthetic operational scenarios, soft labels |
|
| 115 |
+
| [typed-decisions-synth](https://huggingface.co/datasets/n4ze3m/typed-decisions-synth) | MIT | multi-question cases over 149 domains |
|
| 116 |
+
|
| 117 |
+
Some sources carry their own terms (research-only tasks in tasksource-jev, model-output terms of the teacher); check
|
| 118 |
+
them for your use.
|
| 119 |
+
|
| 120 |
+
## Limitations
|
| 121 |
+
|
| 122 |
+
- English only. Decisions, not text: TYPE values come from a separate text model, as in jev-ultrafast.
|
| 123 |
+
- Browser targets are scored among the candidates the agent lists (8–40 per step), not every element on the page.
|
| 124 |
+
- DONE and BLOCKED are the hardest operations; gate DONE on its probability when early stops are costly.
|
| 125 |
+
- Not compared with Jev itself (no API access).
|
| 126 |
+
|
| 127 |
+
## License
|
| 128 |
+
|
| 129 |
+
Apache-2.0, like the base model. Architecture code adapted from [kev](https://github.com/jaredpalmer/kev) (Apache-2.0).
|