alanhuangya commited on
Commit
eb84ce2
·
verified ·
1 Parent(s): fcffd28

Model card: end-to-end results

Browse files
Files changed (1) hide show
  1. README.md +129 -0
README.md ADDED
@@ -0,0 +1,129 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: Qwen/Qwen3-1.7B-Base
4
+ language:
5
+ - en
6
+ tags:
7
+ - decision-model
8
+ - browser-agent
9
+ - system-one
10
+ - lora
11
+ datasets:
12
+ - osunlp/Mind2Web
13
+ - stanfordnlp/nnetnav-live
14
+ - LocalLLaMA/typed-decisions
15
+ - tasksource/tasksource-jev
16
+ - SargeDev/jev-distill-corpus-v3
17
+ - n4ze3m/typed-decisions-synth
18
+ ---
19
+
20
+ # wev-1.7b
21
+
22
+ **A local decision model: typed questions in, calibrated probabilities out, in one forward pass.** `wev-1.7b` answers
23
+ the `POST /v1/systemone` request shape (choice, yes/no and score questions over a free-form state), for general
24
+ decisions and for browser-agent steps (*which operation? which element?*). It runs on your own GPU: no API key,
25
+ no per-call cost, nothing generated.
26
+
27
+ Code, training and evaluation: [alanhuangyoo/wev](https://github.com/alanhuangyoo/wev). Independent project; not affiliated
28
+ with TypeSafe AI.
29
+
30
+ ## Quickstart
31
+
32
+ ```bash
33
+ pip install "wev-ai[serve]"
34
+ ```
35
+
36
+ ```python
37
+ import wev
38
+ m = wev.load("alanhuangya/wev-1.7b")
39
+ out = m.predict(
40
+ state="Refund request: order #4411 arrived damaged, customer attached photos, first refund this year.",
41
+ questions={
42
+ "action": {"type": "choice", "instructions": "What should support do?",
43
+ "criteria": {"refund": "Refund the order.", "replace": "Ship a replacement.",
44
+ "escalate": "Send to a human agent."}},
45
+ "fraud_risk": {"type": "noul", "instructions": "This request looks fraudulent.",
46
+ "criteria": {"true": "Likely fraud.", "false": "No sign of fraud."}},
47
+ },
48
+ )
49
+ print(out["answers"])
50
+ ```
51
+
52
+ ```bash
53
+ wev serve --model alanhuangya/wev-1.7b --port 8009 # drop-in POST /v1/systemone, e.g. for jev-ultrafast
54
+ ```
55
+
56
+ ## Results
57
+
58
+ One read of the locked test splits; every other model was run on the same requests and scored the same way
59
+ (per-question accuracy; `scripts/compare.py`).
60
+
61
+ **General typed decisions**
62
+
63
+ | model | kev decision-v7 | kev transfer-v4 | typed-decisions |
64
+ |---|---|---|---|
65
+ | **wev-1.7b** | 81.1 | 65.5 | **79.5** |
66
+ | Kev-4B | **88.2** | **82.1** | 65.1 |
67
+ | Kev-8B | 88.1 | 76.8 | 62.7 |
68
+ | Laya (typed-decisions) | 65.7 | 62.8 | 76.8 |
69
+ | Laya | 64.3 | 63.7 | 36.2 |
70
+
71
+ `wev-1.7b` trains on 80% of the typed-decisions train split, like the Laya (typed-decisions) specialist; Kev and Laya
72
+ do not, so on that column they are generalists. kev decision-v7 is Kev's own training suite (`wev-1.7b` also trains on
73
+ its train split); transfer-v4 is out-of-domain for every model here.
74
+
75
+ **Browser steps** (Mind2Web test split: websites unseen in training, jev-ultrafast request format; step success =
76
+ operation and target element both right)
77
+
78
+ | model | step success | operation |
79
+ |---|---|---|
80
+ | **wev-1.7b** | **68.2** | **88.1** |
81
+ | Kev-4B | 21.2 | 35.7 |
82
+ | Kev-8B | — | — |
83
+ | Laya (typed-decisions) | 0.7 | 13.1 |
84
+ | Laya | 0.0 | 2.5 |
85
+
86
+ 873 requests; 11 exceed the context `wev-1.7b` is evaluated with and count as wrong for it.
87
+
88
+ NNetNav test split (live-web steps, DONE judged by an LLM): step success 61.0, DONE recall
89
+ 80.4, premature DONE 9.4.
90
+
91
+ ## Model
92
+
93
+ - Backbone: `Qwen/Qwen3-1.7B-Base` without its vocabulary head, LoRA r=16 on every attention and MLP projection, merged
94
+ into the weights of this export; 28 layers, bf16.
95
+ - Readout: a pointer head scores each option's `</opt>` state against the question's `<decide>` state.
96
+ - Each question sees the state and itself only (block-causal branches, positions restart after the state), so a
97
+ request with many questions costs one pass and answers never depend on question order.
98
+ - Context: state up to 4096 tokens, each question up to 8192 tokens (trained with
99
+ 2048); longer page states are shrunk before encoding.
100
+
101
+ ## Training
102
+
103
+ 1 epoch, lr 0.0001, one-cycle schedule, soft-label cross-entropy where the source has soft labels.
104
+ Recipe and data builders: [alanhuangyoo/wev](https://github.com/alanhuangyoo/wev).
105
+
106
+ | source | license | what it adds |
107
+ |---|---|---|
108
+ | [Mind2Web](https://huggingface.co/datasets/osunlp/Mind2Web) | CC BY 4.0 | human browser steps: click, type, select |
109
+ | [NNetNav-live](https://huggingface.co/datasets/stanfordnlp/nnetnav-live) | Apache-2.0 | live-web steps; DONE relabelled by an LLM judge |
110
+ | teacher episodes | outputs of qwen3-max | jev-ultrafast on live sites with qwen3-max as System One, success judge-verified |
111
+ | [kev decision-v7](https://github.com/jaredpalmer/kev) | per source | ten public classification / QA sources plus rule records |
112
+ | [typed-decisions](https://huggingface.co/datasets/LocalLLaMA/typed-decisions) | Apache-2.0 | agent / ops workflows, 5 questions per case (80% of train) |
113
+ | [tasksource-jev](https://huggingface.co/datasets/tasksource/tasksource-jev) | mixed (per source task; some research-only) | hundreds of classification tasks as decisions |
114
+ | [jev-distill-corpus-v3](https://huggingface.co/datasets/SargeDev/jev-distill-corpus-v3) | Apache-2.0 | synthetic operational scenarios, soft labels |
115
+ | [typed-decisions-synth](https://huggingface.co/datasets/n4ze3m/typed-decisions-synth) | MIT | multi-question cases over 149 domains |
116
+
117
+ Some sources carry their own terms (research-only tasks in tasksource-jev, model-output terms of the teacher); check
118
+ them for your use.
119
+
120
+ ## Limitations
121
+
122
+ - English only. Decisions, not text: TYPE values come from a separate text model, as in jev-ultrafast.
123
+ - Browser targets are scored among the candidates the agent lists (8–40 per step), not every element on the page.
124
+ - DONE and BLOCKED are the hardest operations; gate DONE on its probability when early stops are costly.
125
+ - Not compared with Jev itself (no API access).
126
+
127
+ ## License
128
+
129
+ Apache-2.0, like the base model. Architecture code adapted from [kev](https://github.com/jaredpalmer/kev) (Apache-2.0).