org2ai commited on
Commit
ef7481d
·
1 Parent(s): a5d6e74

Wald-Q4B v2 (04701-c22 + Qwen3.5-4B vision tower): main = v2; v1.x at tags v1.2 / v1.2-main / v1.1 / v1.0

Browse files
This view is limited to 50 files because it contains too many changes.   See raw diff
Files changed (50) hide show
  1. CITATION.cff +6 -22
  2. CONTAMINATION.md +11 -12
  3. Dockerfile +7 -8
  4. MANIFEST.json +392 -260
  5. NOTICE +17 -13
  6. PROVENANCE.md +4 -16
  7. README.md +44 -343
  8. RUNBOOK.md +28 -46
  9. SHA256SUMS +15 -0
  10. Wald-4B-v1.2-Q8_0.gguf +0 -3
  11. config.json +102 -81
  12. contamination/trained-on-di-ids.json +0 -0
  13. docs/api.md +34 -118
  14. docs/readmes/README.zh.md +0 -278
  15. evaluation/v1.2/release-check.json +0 -89
  16. evaluation/v1.2/summary.json +0 -268
  17. evaluation/v2/README.md +48 -0
  18. evaluation/v2/code/di03_native_auto_engine.py +68 -0
  19. evaluation/v2/code/di03_parallel.py +87 -0
  20. evaluation/v2/gguf-parity.json +84 -0
  21. evaluation/v2/latency.json +9 -0
  22. evaluation/v2/serving-parity.json +48 -0
  23. evaluation/v2/summary.json +157 -0
  24. evaluation/v2/vision-sanity.json +31 -0
  25. graft-manifest.json +0 -0
  26. graft-report.json +26 -0
  27. history/v1.0/CONTAMINATION.md +2 -2
  28. history/v1.0/README.md +0 -1
  29. {docs → history/v1.0/docs}/lora-cli.md +0 -0
  30. {docs → history/v1.0/docs}/observations/one-lora-all-verticals.md +0 -6
  31. {figures → history/v1.0/figures}/data.json +0 -0
  32. {figures → history/v1.0/figures}/di-areas.svg +0 -0
  33. {figures → history/v1.0/figures}/di-vs-size.svg +0 -0
  34. {figures → history/v1.0/figures}/latency-cost.svg +0 -0
  35. {figures → history/v1.0/figures}/lora-cli-architecture.svg +0 -0
  36. {figures → history/v1.0/figures}/verticals.svg +0 -0
  37. {tools → history/v1.0/tools}/export_figure_data.py +0 -0
  38. {tools → history/v1.0/tools}/make_figures.py +0 -0
  39. {evaluation → history/v1.1/evaluation}/benchmark-summary.json +0 -0
  40. {evaluation → history/v1.1/evaluation}/index.json +0 -0
  41. {evaluation → history/v1.1/evaluation}/release-validation.json +0 -0
  42. {evaluation → history/v1.1/evaluation}/scores.json +0 -0
  43. {evaluation → history/v1.1/evaluation}/serial-latency-preflight-corrected.json +0 -0
  44. llms.txt +7 -25
  45. merges.txt +0 -0
  46. model-00001-of-00002.safetensors +0 -3
  47. Wald-4B-v1.2-Q4_K_M.gguf → model-00001-of-00003.safetensors +2 -2
  48. model-00002-of-00002.safetensors +0 -3
  49. Wald-4B-v1.2-Q5_K_M.gguf → model-00002-of-00003.safetensors +2 -2
  50. Wald-4B-v1.2-Q6_K.gguf → model-00003-of-00003.safetensors +2 -2
CITATION.cff CHANGED
@@ -1,28 +1,12 @@
1
  cff-version: 1.2.0
2
- message: "If you use Wald-Q4B, please cite it as below."
 
3
  type: software
4
- title: "Wald-Q4B v1.2: an open-weight 4B decision model with calibrated option probabilities"
5
- abstract: >-
6
- Wald-Q4B is an open-weight 4B decision model built on Qwen3.5-4B-Base. Given a
7
- state and typed questions (choice, yes/no, score), it returns a calibrated
8
- probability for every option through a Jev-compatible POST /v1/systemone API,
9
- in one pass or after an optional short thought. Weights and serving code are
10
- Apache-2.0. Independent project; not affiliated with TypeSafe AI.
11
  authors:
12
  - name: "Wald-4B authors"
13
- alias: Harry19081
14
- website: "https://huggingface.co/Harry19081"
15
- version: "v1.2"
16
- date-released: "2026-10-01"
17
  license: Apache-2.0
18
- url: "https://huggingface.co/org2ai/Wald-4B"
19
- repository-code: "https://github.com/org2AI/wald-4b"
20
  repository-artifact: "https://huggingface.co/org2ai/Wald-4B"
21
- keywords:
22
- - decision model
23
- - calibrated probabilities
24
- - tool selection
25
- - agent routing
26
- - classification
27
- - Jev-compatible API
28
- - Qwen3.5
 
1
  cff-version: 1.2.0
2
+ message: "If you use Wald-Q4B, please cite it and name the revision you used."
3
+ title: "Wald-Q4B: an open-weight 4B decision model with calibrated option probabilities"
4
  type: software
 
 
 
 
 
 
 
5
  authors:
6
  - name: "Wald-4B authors"
7
+ version: "v2.0"
8
+ date-released: "2026-10-08"
 
 
9
  license: Apache-2.0
 
 
10
  repository-artifact: "https://huggingface.co/org2ai/Wald-4B"
11
+ repository-code: "https://github.com/org2AI/wald-4b"
12
+ notes: "Revision v2.0 = checkpoint 04701-c22 (model.safetensors sha256 6ae382f5a0ed9f4c53cd0953cb9606b1627a6816350340a59114860e6cf6f29a)."
 
 
 
 
 
 
CONTAMINATION.md CHANGED
@@ -1,14 +1,13 @@
1
- # Evaluation notes
2
 
3
- - Known strict training overlaps were filtered and training-file hashes checked. Weak/semantic overlap and foundation-model pretraining contamination are not ruled out; HLE training overlap has not been independently cleared. Benchmark samples informed development, so this is not a sealed evaluation.
4
- - The complete **54.59** run used a frozen **256-row XL-dev calibration** fit after excluding known DI matches. No DI labels or DI index score were used to fit it. HLE calibration screening found no additional match.
5
- - The earlier **54.02** sample used the previous calibration table with known overlaps and remains historical. Old v1.0 disclosures apply to its own weights and are preserved under `v1.0-legacy`.
6
- - Full-run response artifacts are unchanged. No request payloads or generated reasoning text are included. Official maintainer serial latency validation and leaderboard admission remain pending.
 
 
 
 
 
7
 
8
- # v1.2 (checkpoint `02600-f19`)
9
-
10
- - v1.2 was trained only on perturbed and unperturbed copies of v1.1's own training questions. A Decision Index blocklist scan of its training and hold-out rows found 0 strict matches (329 weak alerts, all states shorter than 8 words).
11
- - **JevAdvBench is evaluation-only.** An 8-gram overlap check of every inserted text against every JevAdvBench string (the clean questions and all 9,744 attacked variants) found 0 hits. The perturbation kinds were chosen after v1.1's per-attack JevAdvBench results were known, so the benchmark's attack families informed the training design; its texts did not. Robustness to other attack types is not measured.
12
- - The JevBench public items were again a development scoreboard: v1.2 had a pass line of at least 201/231 set before training. They were never training data.
13
- - The **54.59** complete-suite Decision Index result and the files under `evaluation/` other than `evaluation/v1.2/` belong to v1.1. v1.2 has only a one-pass read of a 6,948-request sample.
14
- - v1.2 uses v1.1's temperature table unchanged. No benchmark labels were used to fit it.
 
1
+ # Evaluation notes for Wald-Q4B v2 (`04701-c22`)
2
 
3
+ - **Decision Index.** Every training file of all three stages was scanned against the complete Decision Index 0.3 public suite with our strict matcher: **0 strict hits** ([contamination/trained-on-di-ids.json](contamination/trained-on-di-ids.json): ids, rule, file hashes, what is not covered). Each stage's data had also been gated before training (stage 1 against the 0.2 suite, stage 3 against the 0.3 suite with additional exact full-shingle and character n-gram checks; every hit row was dropped). The training data contains train splits of some benchmarks whose test items the Decision Index also uses ([PROVENANCE.md](PROVENANCE.md)), and the model was developed with the index's task formats in view, so its Decision Index results are not held-out results in the usual sense. The organizer's private tests are unknown to us.
4
+ - **JevBench public set (231).** A development scoreboard: never training data, but read repeatedly during development. Not held out.
5
+ - **JevBench-XL** (internal). Stage 2's and stage 3's new items were exact-state deduplicated against all XL partitions, JevBench, JevAdvBench and the Decision Index sample; the XL TEST partitions were not used for any selection. The XL scores are internal and not independently verifiable.
6
+ - **multistep_decisions (100)** and **JevAdvBench** are evaluation-only.
7
+ - **Checkpoint and policy selection.** Auto 0.7 was fixed as the default before the Decision Index 0.3 runs. The released checkpoint (step 300 of 638) was chosen over the final step and over the earlier C16B / C16C checkpoints using paired screens on a fixed subset of the Decision Index 0.3 public suite and then one full public run; the public index therefore informed the choice of checkpoint. No successful request was repeated.
8
+ - **Calibration.** One frozen temperature table (`temperature.json`, sha256 `a0f72cd2d0a653e81051e5a0c77fc1a69131552a8102b580a93a6dbe7908b2da`), fitted on 275 development rows before these evaluations; no benchmark labels were used, and the Decision Index 0.3 run did not refit it.
9
+ - **Vision.** The vision tower is Qwen3.5-4B's own, unchanged; no image data was used in training. Image results are zero-shot and depend on the base model's pretraining, which may include the public image benchmarks.
10
+ - **Not ruled out:** semantic overlap, and contamination in the base model's pretraining data.
11
+ - No request payloads, gold labels or generated reasoning text are included in this repository.
12
 
13
+ The notes for v1.x are at their tags.
 
 
 
 
 
 
Dockerfile CHANGED
@@ -1,20 +1,19 @@
1
- # Wald-4B /v1/systemone server: vLLM 0.30.0 (bf16) on the weights as a loopback-only sidecar + wald-serve in front.
2
  #
3
  # docker build -t wald-serve .
4
  # docker run --gpus all -v /path/to/Wald-4B:/model:ro -p 8000:8000 wald-serve
5
  # # ready when GET http://127.0.0.1:8000/health returns {"ok": true, ...}
6
  #
7
- # The weights directory holds config, tokenizer, bf16 safetensors, temperature.json and serving.json (declared policy,
8
- # prompt format, context limit). Extra vLLM arguments: -e WALD_VLLM_ARGS="...". Another policy: -e WALD_EFFORT=high.
9
  FROM python:3.12-slim
10
 
11
  RUN pip install --no-cache-dir uv==0.9.5 \
12
- && uv pip install --system --no-cache "vllm==0.30.0" \
13
- && python -c "import importlib.metadata as m; print('vllm', m.version('vllm'), 'torch', m.version('torch'))"
14
  COPY server /app/server
15
  RUN uv pip install --system --no-cache /app/server
16
  ENV HF_HUB_OFFLINE=1 TRANSFORMERS_OFFLINE=1 TOKENIZERS_PARALLELISM=false PYTHONUNBUFFERED=1 \
17
- VLLM_USE_FLASHINFER_SAMPLER=0 VLLM_NO_USAGE_STATS=1 DO_NOT_TRACK=1 \
18
- WALD_EFFORT="" WALD_VLLM_ARGS=""
19
  EXPOSE 8000
20
- CMD ["sh", "-c", "exec wald-serve --model /model --port 8000 ${WALD_EFFORT:+--effort $WALD_EFFORT} --vllm-args \"$WALD_VLLM_ARGS\""]
 
1
+ # Wald-4B v2 /v1/systemone server: vLLM 0.30.0 (bf16) on the weights as a loopback-only sidecar + wald-serve-native.
2
  #
3
  # docker build -t wald-serve .
4
  # docker run --gpus all -v /path/to/Wald-4B:/model:ro -p 8000:8000 wald-serve
5
  # # ready when GET http://127.0.0.1:8000/health returns {"ok": true, ...}
6
  #
7
+ # The weights directory holds config, tokenizer, chat template, model.safetensors, temperature.json and serving.json
8
+ # (native v2 engine, effort, prompt format). Another policy: -e WALD_EFFORT=none (none | auto | always).
9
  FROM python:3.12-slim
10
 
11
  RUN pip install --no-cache-dir uv==0.9.5 \
12
+ && uv pip install --system --no-cache "vllm==0.30.0" "transformers==5.17.0" \
13
+ && python -c "import importlib.metadata as m; print('vllm', m.version('vllm'), 'transformers', m.version('transformers'), 'torch', m.version('torch'))"
14
  COPY server /app/server
15
  RUN uv pip install --system --no-cache /app/server
16
  ENV HF_HUB_OFFLINE=1 TRANSFORMERS_OFFLINE=1 TOKENIZERS_PARALLELISM=false PYTHONUNBUFFERED=1 \
17
+ VLLM_USE_FLASHINFER_SAMPLER=0 VLLM_NO_USAGE_STATS=1 DO_NOT_TRACK=1 WALD_EFFORT=""
 
18
  EXPOSE 8000
19
+ CMD ["sh", "-c", "exec wald-serve-native-vision --model /model --port 8000 ${WALD_EFFORT:+--effort $WALD_EFFORT}"]
MANIFEST.json CHANGED
@@ -1,262 +1,394 @@
1
  {
2
- "CONTAMINATION.md": {
3
- "bytes": 2047,
4
- "sha256": "108d33ec5188b608e8b293d18ba1bb975b9849638644a682616ca8816b4e1fd1"
5
- },
6
- "Dockerfile": {
7
- "bytes": 1201,
8
- "sha256": "b353502e02205a000a48a4ea4478328e8461ac012f246b23c7cfcba572ac69f3"
9
- },
10
- "LICENSE": {
11
- "bytes": 11358,
12
- "sha256": "c95bae1d1ce0235ecccd3560b772ec1efb97f348a79f0fbe0a634f0c2ccefe2c"
13
- },
14
- "NOTICE": {
15
- "bytes": 1591,
16
- "sha256": "7b37cfe73768e3c01481d0a5d0f097b307cc6b556968c7f4072f6d56b108e061"
17
- },
18
- "PROVENANCE.md": {
19
- "bytes": 2139,
20
- "sha256": "cde98850a90a3159ee17cac4e3e2bc0e267c16e84bb5813b4f97aa8cf4b1ba2e"
21
- },
22
- "README.md": {
23
- "bytes": 29996,
24
- "sha256": "8ab417f71f12cf26fba335f92207cd5a1a11e44018d4ed76f1d64ed282e8937f"
25
- },
26
- "RUNBOOK.md": {
27
- "bytes": 4980,
28
- "sha256": "8ff79f40f2399cb21b08f27cacb223c88f419d5d55307259fca1166cab132639"
29
- },
30
- "chat_template.jinja": {
31
- "bytes": 7756,
32
- "sha256": "a4aee8afcf2e0711942cf848899be66016f8d14a889ff9ede07bca099c28f715"
33
- },
34
- "config.json": {
35
- "bytes": 1979,
36
- "sha256": "63f47812d0f11118e4d252d2b3ad488707eb9287a11589f4fd382a1d31182724"
37
- },
38
- "docs/lora-cli.md": {
39
- "bytes": 7501,
40
- "sha256": "8b4be9b9256822ff6ba545d4cac874fa405846da9d466525e958f0bbc69507d7"
41
- },
42
- "docs/observations/one-lora-all-verticals.md": {
43
- "bytes": 9088,
44
- "sha256": "e385b92e333eaaa9395a09daf26c09420645bee242456552a55500db45919bc7"
45
- },
46
- "docs/readmes/README.zh.md": {
47
- "bytes": 25760,
48
- "sha256": "bab6599fe01091f36f96085b07b7063649fc0aff0aa16dbaa357bed3e2671171"
49
- },
50
- "evaluation/benchmark-summary.json": {
51
- "bytes": 36014,
52
- "sha256": "78983f2d4f3d2c708159879ff1e66b6da7a2a10d95f732417fe8b27c39f3c1ae"
53
- },
54
- "evaluation/index.json": {
55
- "bytes": 10559,
56
- "sha256": "6f2ef43841fc619b3ba7bdf9d4d7b5fc7d17055048dc155bd45a61f92d3dbe4b"
57
- },
58
- "evaluation/release-validation.json": {
59
- "bytes": 413,
60
- "sha256": "163475465b5d97d1c774d284ed637352beab136fa737b5c8d464e6069d92de89"
61
- },
62
- "evaluation/scores.json": {
63
- "bytes": 33084,
64
- "sha256": "5bccdf55495daf82396123ca1062d084e910d603536f27f4950dea037b84122e"
65
- },
66
- "evaluation/serial-latency-preflight-corrected.json": {
67
- "bytes": 336,
68
- "sha256": "db1da0eb255ea38105a2d8e76e221e0ae50b18b6ad37737a10787240e3f8e567"
69
- },
70
- "generation_config.json": {
71
- "bytes": 116,
72
- "sha256": "62153eb6c69f2e1f426beaa8002b7186437e949c7588167085df14e10e9c0a73"
73
- },
74
- "history/v1.0/CONTAMINATION.md": {
75
- "bytes": 26100,
76
- "sha256": "a4f840155d12964dd9fc1dde90611c18252bb55082844145b37cd00ffe8ab090"
77
- },
78
- "history/v1.0/README.md": {
79
- "bytes": 10639,
80
- "sha256": "f468e58384b9978060038a43434a0ed6ce2973185dcb3ff7e059fdc0314bb1d2"
81
- },
82
- "history/v1.0/RUNBOOK.md": {
83
- "bytes": 6268,
84
- "sha256": "ab03ccc6bfbd4a9a7772e579dd1a729eec9a3ea072ca8c54ded21018c1d15ec3"
85
- },
86
- "model-00001-of-00002.safetensors": {
87
- "bytes": 4972947968,
88
- "sha256": "0b15067f769e7388bd598aabafa2ea3211a3a0c4e49e47d07c4f65c65ffc25dd"
89
- },
90
- "model-00002-of-00002.safetensors": {
91
- "bytes": 3438610328,
92
- "sha256": "7cff102b314edeb3dc5abfa7063f72ad3e888884da237601239dc3582074120a"
93
- },
94
- "model-files/serving.json": {
95
- "bytes": 128,
96
- "sha256": "6b4de526ca4c939b626932dcd10d6167e21a68b1e77a090d7a5a87c7c3a353b1"
97
- },
98
- "model.safetensors.index.json": {
99
- "bytes": 41993,
100
- "sha256": "83620b40d724f8b286580b5c7c504222b0bdbdd7dc9f068e565132409c7eec9d"
101
- },
102
- "reference/eval/__init__.py": {
103
- "bytes": 0,
104
- "sha256": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855"
105
- },
106
- "reference/eval/systemone_vllm.py": {
107
- "bytes": 25186,
108
- "sha256": "881195545231b7b26c7753d4a43fdce5d65a2c49b50a78aec5eee69468e59136"
109
- },
110
- "reference/kev/__init__.py": {
111
- "bytes": 0,
112
- "sha256": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855"
113
- },
114
- "reference/kev/api.py": {
115
- "bytes": 7086,
116
- "sha256": "23bac9af0e79702428447c62662b88910071595f6171d9e3c82d6fa35fb0dbd1"
117
- },
118
- "reference/midtrain/__init__.py": {
119
- "bytes": 0,
120
- "sha256": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855"
121
- },
122
- "reference/midtrain/encode.py": {
123
- "bytes": 19164,
124
- "sha256": "fb0d35c055c7052a3a3b019cd2cb4de1857ca733467680fa6182a106b045ecec"
125
- },
126
- "reference/midtrain/letter.py": {
127
- "bytes": 12651,
128
- "sha256": "e053b6dd1790818e8473fea4ad5350c65f9de53cd59e1a750a3bdff8d5085963"
129
- },
130
- "reference/midtrain/rawprompt.py": {
131
- "bytes": 11856,
132
- "sha256": "5abe943f8c39e84018006fcea3fd32df1df89635eba1c87639b04057654a2ffa"
133
- },
134
- "reference/midtrain/tokens.py": {
135
- "bytes": 3488,
136
- "sha256": "7fbc482959e405dcc537bfadf0fa24df2d89bcaaab5919bcb765a70ea81e21de"
137
- },
138
- "run.sh": {
139
- "bytes": 915,
140
- "sha256": "d5b7b3e62d43c5f0dae4779f723b1eb4408aa9c21fcbdb6ed65e1da049dd050a"
141
- },
142
- "server/README.md": {
143
- "bytes": 1678,
144
- "sha256": "2ca738dcacd770e7c490ceeb59f2f11df7515f5bc56784eb9d476af4c7145b4d"
145
- },
146
- "server/pyproject.toml": {
147
- "bytes": 578,
148
- "sha256": "0fd8155bfa466e0f7a95fddce986ab881cce9260187c1620884fa27e5008dfbe"
149
- },
150
- "server/src/wald_serve/__init__.py": {
151
- "bytes": 139,
152
- "sha256": "7db760d45fae0f1d9cea6dbd3e1421a2627a92020981ae3cd56ab57181423aac"
153
- },
154
- "server/src/wald_serve/__main__.py": {
155
- "bytes": 51,
156
- "sha256": "540fd0a5992ca535482a9d44b2e268ce3a39bb47a0576f301214afa5ef4494b9"
157
- },
158
- "server/src/wald_serve/engine.py": {
159
- "bytes": 14814,
160
- "sha256": "40fbea84c4c5b93ef888da44508d4fec2677c58732a09d731f0bcf17c2b48dd6"
161
- },
162
- "server/src/wald_serve/prompt.py": {
163
- "bytes": 4192,
164
- "sha256": "bb45a10734f2546eb1ea1bd7e8b5717ecd8305a4aded0b816f80cd7e85188f7a"
165
- },
166
- "server/src/wald_serve/server.py": {
167
- "bytes": 8878,
168
- "sha256": "49fb06ee52eb2ebe7059404432386d02e049556003acb7d2f6494ba92d4b9703"
169
- },
170
- "server/src/wald_serve/wire.py": {
171
- "bytes": 3670,
172
- "sha256": "836ce54ce93016a98f94a6beaa508cc03982e24c3d3b4d62b8ef667389cd14af"
173
- },
174
- "server/tests/fakevllm.py": {
175
- "bytes": 4167,
176
- "sha256": "9fb222a21bc2b17221e5a1a3605973b4f71730db9bbfede439bd3388b91e2570"
177
- },
178
- "server/tests/test_parity.py": {
179
- "bytes": 4392,
180
- "sha256": "cd1e4142090076114f2b07efce49fdc5466e302a5c79996b993f3833a3457566"
181
- },
182
- "server/tests/test_server.py": {
183
- "bytes": 13844,
184
- "sha256": "46f2007e9743bfeb0f7f79fca093cf444a70df1bfc348229f34cd251f88a19b3"
185
- },
186
- "serving.json": {
187
- "bytes": 128,
188
- "sha256": "6b4de526ca4c939b626932dcd10d6167e21a68b1e77a090d7a5a87c7c3a353b1"
189
- },
190
- "temperature.json": {
191
- "bytes": 4348,
192
- "sha256": "a4e9a1983a2246de34ce4c4b48e5d501daf7c66ae4f3654a79591f0ed0059da5"
193
- },
194
- "tokenizer.json": {
195
- "bytes": 19989509,
196
- "sha256": "bd53432f0de26d67a83b634040d4f043053da4ecd0c759e7f4b24ad4f8bb9a81"
197
- },
198
- "tokenizer_config.json": {
199
- "bytes": 1127,
200
- "sha256": "171ecbe7ddae98d11840698f7df2b8d5b4722139db0f0620d3bbf429bd656250"
201
- },
202
- "tools/export_contamination.py": {
203
- "bytes": 5572,
204
- "sha256": "a5723b781031c0bde60349691b82a63f150aa408afaf5621a893bde6ee82f6f8"
205
- },
206
- "tools/export_figure_data.py": {
207
- "bytes": 10121,
208
- "sha256": "ba889a4b90d59c9487cbca34768a5ea90fafb805c5f246124d1f3475ee067a43"
209
- },
210
- "tools/make_figures.py": {
211
- "bytes": 19067,
212
- "sha256": "207d02dea21ee9a83763490247b4aa0fe146068cf1c4922a07e4647a3fc546f8"
213
- },
214
- "tools/prepare_model_dir.py": {
215
- "bytes": 1890,
216
- "sha256": "0c1811fbfd98fc3a8c802298de3691b5039455aceef1475b73cfced72365708e"
217
- },
218
- "tools/secret_scan.py": {
219
- "bytes": 1698,
220
- "sha256": "ad20de47de5bbe7f72417f3d87e6f72aae7987fab913eb814b818e4a4bbc6e1e"
221
- },
222
- "model-info.json": {
223
- "bytes": 6195,
224
- "sha256": "3b3fe25cad2fab4da64bc4db0d610489fdc005fab5c9ae4de908ed471a554121"
225
- },
226
- "llms.txt": {
227
- "bytes": 3324,
228
- "sha256": "886542b46fd0cf5284d0df15dfa02459b167c1a7c6bc749ae096bf762a8bfcc6"
229
- },
230
- "CITATION.cff": {
231
- "bytes": 1063,
232
- "sha256": "75f7dd132f937aee40b8867032bb10d961bbba15fb525744c67debc272e85d1d"
233
- },
234
- "docs/api.md": {
235
- "bytes": 6069,
236
- "sha256": "9f6a43a76824c38285856c8fb1cee632e0fa5d3c830d450fbd1120258e54997d"
237
- },
238
- "evaluation/v1.2/summary.json": {
239
- "bytes": 6991,
240
- "sha256": "1c4231985191eb1f7916baf13dd621d3e5655f192dbad2f842f998f8ba8ffaf5"
241
- },
242
- "evaluation/v1.2/release-check.json": {
243
- "bytes": 3036,
244
- "sha256": "0de642d7c1ddb2ececafce16c06ae2d2eb0e663ec94b4afcfab605a61aadfb70"
245
- },
246
- "Wald-4B-v1.2-Q4_K_M.gguf": {
247
- "bytes": 2708804000,
248
- "sha256": "e843f7658793b9f3effb3a4941cd67f80707a07b468aaaf94d440613751d76fa"
249
- },
250
- "Wald-4B-v1.2-Q5_K_M.gguf": {
251
- "bytes": 3074986400,
252
- "sha256": "6fecc4e7655adb71e8c642ce5018f967564ba259852d562ecd3109386542b0cb"
253
- },
254
- "Wald-4B-v1.2-Q6_K.gguf": {
255
- "bytes": 3464055200,
256
- "sha256": "5870f15f3b2eef773f672e56bb36e343a95e32fd09baaf7c9ce8b0da1e5dd897"
257
- },
258
- "Wald-4B-v1.2-Q8_0.gguf": {
259
- "bytes": 4482402720,
260
- "sha256": "09c02494ebdfb7d088a2c76726eb2b290a438dbc5658c3be16da8333fddd6038"
261
- }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
262
  }
 
1
  {
2
+ "CITATION.cff": {
3
+ "bytes": 556,
4
+ "sha256": "541671d55494457122276a7fc1612e6bbed9277c83156de5c9eb1b5ae4a825cc"
5
+ },
6
+ "CONTAMINATION.md": {
7
+ "bytes": 2544,
8
+ "sha256": "1c69371d18ee073bf854aff74b2c3b3403d589fb17d36b43f5945f764016cb55"
9
+ },
10
+ "Dockerfile": {
11
+ "bytes": 1212,
12
+ "sha256": "29ae11deccbeba7af543614f3c2a76711377fffd9727d7c269af050b781f8b0e"
13
+ },
14
+ "LICENSE": {
15
+ "bytes": 11358,
16
+ "sha256": "c95bae1d1ce0235ecccd3560b772ec1efb97f348a79f0fbe0a634f0c2ccefe2c"
17
+ },
18
+ "NOTICE": {
19
+ "bytes": 1778,
20
+ "sha256": "525dea1eefe1f1c280923dc290b0988d6d8ca6843b68706b28e83946e1e5ab60"
21
+ },
22
+ "PROVENANCE.md": {
23
+ "bytes": 1176,
24
+ "sha256": "43510ca439ff9f402eeef7c7e0e4a8e3e2265cf0ff8b3e027d662b1bff79078e"
25
+ },
26
+ "README.md": {
27
+ "bytes": 5109,
28
+ "sha256": "b6a1f25ff08fc577a550c07ae3652d650dc935113c6b921c7227cf5c27213762"
29
+ },
30
+ "RUNBOOK.md": {
31
+ "bytes": 5558,
32
+ "sha256": "edd26f5e213ef94e15440dd4c0b9597e7d663ff696c86c3f41dfcfd45a7e523d"
33
+ },
34
+ "SHA256SUMS": {
35
+ "sha256": "1ba76d9616233fc4f347859c07c9359492aaa5e45bf5f2f1fb8ce43df3cce4ad",
36
+ "bytes": 1326
37
+ },
38
+ "chat_template.jinja": {
39
+ "sha256": "a4aee8afcf2e0711942cf848899be66016f8d14a889ff9ede07bca099c28f715",
40
+ "bytes": 7756
41
+ },
42
+ "config.json": {
43
+ "sha256": "ddc63e1c717afa86c865bb5e01313d89d72bb53b97ad4a8a03ba8510c0621670",
44
+ "bytes": 3161
45
+ },
46
+ "contamination/trained-on-di-ids.json": {
47
+ "bytes": 7016,
48
+ "sha256": "31f5a5ee9860fb747c0c7e23a9ac75790dcd688de1655c6890eef5ba2ed9b9c3"
49
+ },
50
+ "docs/api.md": {
51
+ "bytes": 4403,
52
+ "sha256": "0c05b5a5b9a2ec982416137e209b615e51b64cfb012f53a02b279257f258d2bd"
53
+ },
54
+ "evaluation/v2/README.md": {
55
+ "bytes": 2711,
56
+ "sha256": "6792af269e6383ececdc997b011cf1a09b5ab2a79d356f04932fc0f22e6789a7"
57
+ },
58
+ "evaluation/v2/code/di03_native_auto_engine.py": {
59
+ "bytes": 3578,
60
+ "sha256": "547b753c6e4a27849b302a1ce188af7a550fc3f8273a0b0feca5fb598829acc4"
61
+ },
62
+ "evaluation/v2/code/di03_parallel.py": {
63
+ "bytes": 5703,
64
+ "sha256": "f4ca5a6878ac5652392a4345374e8669c265c9d4415f54b80edfda8cdd8a6d33"
65
+ },
66
+ "evaluation/v2/gguf-parity.json": {
67
+ "bytes": 1558,
68
+ "sha256": "8250b8a331cb716192c9ad673d8edbff01ff373ad460d04dcf8a4076920910e9"
69
+ },
70
+ "evaluation/v2/latency.json": {
71
+ "bytes": 306,
72
+ "sha256": "5e0fa4422f0fcce86830335c07a6491130221b82c17a635235bd110daf750eed"
73
+ },
74
+ "evaluation/v2/serving-parity.json": {
75
+ "bytes": 1703,
76
+ "sha256": "2024995897b70a0e5be5619f0d6812f9ba9aacb7ff11857bec7b20aeb0148bbb"
77
+ },
78
+ "evaluation/v2/summary.json": {
79
+ "bytes": 3170,
80
+ "sha256": "c662e787fc2bd65dd383c07284c95125bab1f73a5b8dbef25434e68dde72d93e"
81
+ },
82
+ "evaluation/v2/vision-sanity.json": {
83
+ "bytes": 728,
84
+ "sha256": "f143e2863fc2ad1126d49f09abbfab7a4cb36293e8bae1b4637fddc826c38377"
85
+ },
86
+ "generation_config.json": {
87
+ "sha256": "62153eb6c69f2e1f426beaa8002b7186437e949c7588167085df14e10e9c0a73",
88
+ "bytes": 116
89
+ },
90
+ "graft-manifest.json": {
91
+ "sha256": "e33ce57511cb924d2bd5559a143c37cf12f1df1e6de911f0a7d4fdbbc0b85eb8",
92
+ "bytes": 113687
93
+ },
94
+ "graft-report.json": {
95
+ "sha256": "d4ae76c16d404c665f84e10079f460221f4c1434c23e0b923bf9a6deb91c1385",
96
+ "bytes": 765
97
+ },
98
+ "history/v1.0/CONTAMINATION.md": {
99
+ "bytes": 25982,
100
+ "sha256": "c7679ec332ccffa844f06edf6c084db6dafd6a4a39a3d51e6d2136eda638d3fd"
101
+ },
102
+ "history/v1.0/README.md": {
103
+ "bytes": 10529,
104
+ "sha256": "70e9be5f27484400270db814412fe397b276d0f3cb201a946b79b67cc35536eb"
105
+ },
106
+ "history/v1.0/RUNBOOK.md": {
107
+ "bytes": 6268,
108
+ "sha256": "ab03ccc6bfbd4a9a7772e579dd1a729eec9a3ea072ca8c54ded21018c1d15ec3"
109
+ },
110
+ "history/v1.0/docs/lora-cli.md": {
111
+ "bytes": 7501,
112
+ "sha256": "8b4be9b9256822ff6ba545d4cac874fa405846da9d466525e958f0bbc69507d7"
113
+ },
114
+ "history/v1.0/docs/observations/one-lora-all-verticals.md": {
115
+ "bytes": 8245,
116
+ "sha256": "f28ce9b2562c7dc3730d6d4ed3bc5281de5ccecc30f4d3925525345cb65745d0"
117
+ },
118
+ "history/v1.0/figures/data.json": {
119
+ "bytes": 31885,
120
+ "sha256": "f6100bdf1f6054610e97163090bed23b4d8d55ce7bffe5d54a76fb7a216c6be3"
121
+ },
122
+ "history/v1.0/figures/di-areas.svg": {
123
+ "bytes": 7525,
124
+ "sha256": "d4f6b768e475e82fed225d9ace7bdd0061c26c4fbe890e696d1adac786cbe0b2"
125
+ },
126
+ "history/v1.0/figures/di-vs-size.svg": {
127
+ "bytes": 11492,
128
+ "sha256": "3c3917c634ad7b76ea4c7ae580956a2077c972eaec5fd5936244c49276c72029"
129
+ },
130
+ "history/v1.0/figures/latency-cost.svg": {
131
+ "bytes": 5068,
132
+ "sha256": "122753d88f84a0a3228da7895b8278cfb6108bb504804079d2bd91b3ff3f8b3c"
133
+ },
134
+ "history/v1.0/figures/lora-cli-architecture.svg": {
135
+ "bytes": 13169,
136
+ "sha256": "9922fc47ff60180333d7f7c3ddb8f03aa489bf72cfb09712abe399304df27eae"
137
+ },
138
+ "history/v1.0/figures/verticals.svg": {
139
+ "bytes": 17902,
140
+ "sha256": "9ad96a799e84115bdb778f55980a94057c6219c65e884cbd2412341c7681f6a2"
141
+ },
142
+ "history/v1.0/tools/export_figure_data.py": {
143
+ "bytes": 10121,
144
+ "sha256": "ba889a4b90d59c9487cbca34768a5ea90fafb805c5f246124d1f3475ee067a43"
145
+ },
146
+ "history/v1.0/tools/make_figures.py": {
147
+ "bytes": 19067,
148
+ "sha256": "207d02dea21ee9a83763490247b4aa0fe146068cf1c4922a07e4647a3fc546f8"
149
+ },
150
+ "history/v1.1/evaluation/benchmark-summary.json": {
151
+ "bytes": 36014,
152
+ "sha256": "78983f2d4f3d2c708159879ff1e66b6da7a2a10d95f732417fe8b27c39f3c1ae"
153
+ },
154
+ "history/v1.1/evaluation/index.json": {
155
+ "bytes": 10559,
156
+ "sha256": "6f2ef43841fc619b3ba7bdf9d4d7b5fc7d17055048dc155bd45a61f92d3dbe4b"
157
+ },
158
+ "history/v1.1/evaluation/release-validation.json": {
159
+ "bytes": 413,
160
+ "sha256": "163475465b5d97d1c774d284ed637352beab136fa737b5c8d464e6069d92de89"
161
+ },
162
+ "history/v1.1/evaluation/scores.json": {
163
+ "bytes": 33084,
164
+ "sha256": "5bccdf55495daf82396123ca1062d084e910d603536f27f4950dea037b84122e"
165
+ },
166
+ "history/v1.1/evaluation/serial-latency-preflight-corrected.json": {
167
+ "bytes": 336,
168
+ "sha256": "db1da0eb255ea38105a2d8e76e221e0ae50b18b6ad37737a10787240e3f8e567"
169
+ },
170
+ "llms.txt": {
171
+ "bytes": 1245,
172
+ "sha256": "4a472b5feb067a57c5f315de733f88bbceeb299eef93be85d2fd31b90771b132"
173
+ },
174
+ "merges.txt": {
175
+ "sha256": "a9d356d7bdf1ef4949e3e748e95b8e10ad9d4e2e838eddc38a0a7b6b94d1db8d",
176
+ "bytes": 3353259
177
+ },
178
+ "model-00001-of-00003.safetensors": {
179
+ "sha256": "10cb7cd22f524f1b33271f8432bcb09409bd879c11d999ea18ca7dd5f8828f5a",
180
+ "bytes": 4757977272
181
+ },
182
+ "model-00002-of-00003.safetensors": {
183
+ "sha256": "ed5e7bd4b2e67eee390857268a7660986aca32068c78e6f483e866c27f90b79b",
184
+ "bytes": 3653581040
185
+ },
186
+ "model-00003-of-00003.safetensors": {
187
+ "sha256": "9a8340e453e63f70185e56fde60aeae26085fe3c39e20ac1662b4325f01a78eb",
188
+ "bytes": 908262184
189
+ },
190
+ "model-files/serving.json": {
191
+ "bytes": 1172,
192
+ "sha256": "5caee30e9e486ab137f5521d7b2054844ce9a1bd4556df581b6cad48fbbdea22"
193
+ },
194
+ "model-info.json": {
195
+ "bytes": 6251,
196
+ "sha256": "c1f95578f2cac0c398781a9a98785c92f5f7f20dd2d3fcf496e6a056d878105f"
197
+ },
198
+ "model.safetensors.index.json": {
199
+ "sha256": "6db63a968a513eb0693feee5ccbc440cbbee43c6b644c0bd33e880ca7d0b54b3",
200
+ "bytes": 67340
201
+ },
202
+ "preprocessor_config.json": {
203
+ "sha256": "27225450ac9c6529872ee1924fcb0962ff5634834f817040f444118116f4e516",
204
+ "bytes": 390
205
+ },
206
+ "reference/README.md": {
207
+ "bytes": 823,
208
+ "sha256": "d233e519031392c3ebe5470de1c03878207ec0aa96dce9b265216145f649a44d"
209
+ },
210
+ "reference/eval/__init__.py": {
211
+ "bytes": 0,
212
+ "sha256": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855"
213
+ },
214
+ "reference/eval/native_answer_acceptance_v1.py": {
215
+ "bytes": 19406,
216
+ "sha256": "764f3a683becd16b87b9ce9bfc01a5fe483ef26c9a3345a724b82780d73ea73a"
217
+ },
218
+ "reference/eval/systemone_vllm.py": {
219
+ "bytes": 36069,
220
+ "sha256": "7586a266a0933c50154aa86faab936b02b617146c357aed91e59467adf20adeb"
221
+ },
222
+ "reference/eval/vllm_letter.py": {
223
+ "bytes": 16993,
224
+ "sha256": "0d8c006c3cff7e43c2b019dde649ba1fac6bf47b960dce676412e4624acc815e"
225
+ },
226
+ "reference/kev/__init__.py": {
227
+ "bytes": 0,
228
+ "sha256": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855"
229
+ },
230
+ "reference/kev/api.py": {
231
+ "bytes": 7086,
232
+ "sha256": "23bac9af0e79702428447c62662b88910071595f6171d9e3c82d6fa35fb0dbd1"
233
+ },
234
+ "reference/midtrain/__init__.py": {
235
+ "bytes": 0,
236
+ "sha256": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855"
237
+ },
238
+ "reference/midtrain/encode.py": {
239
+ "bytes": 19164,
240
+ "sha256": "fb0d35c055c7052a3a3b019cd2cb4de1857ca733467680fa6182a106b045ecec"
241
+ },
242
+ "reference/midtrain/letter.py": {
243
+ "bytes": 22075,
244
+ "sha256": "90faed3734613799f33b40aae660142e9c8c02d34d73394bf9eb2f6aeba5635f"
245
+ },
246
+ "reference/midtrain/rawprompt.py": {
247
+ "bytes": 14571,
248
+ "sha256": "3b73b26cb96af71188e3dfd1a618c9460bfb69509a72f500d41391794ab219b1"
249
+ },
250
+ "reference/midtrain/tokens.py": {
251
+ "bytes": 3488,
252
+ "sha256": "7fbc482959e405dcc537bfadf0fa24df2d89bcaaab5919bcb765a70ea81e21de"
253
+ },
254
+ "reference/source-pins.json": {
255
+ "bytes": 962,
256
+ "sha256": "c8b231a1cdf413895c1abb215e43a41a08f6568d8404ea4f7ca86ee4883eb396"
257
+ },
258
+ "run.sh": {
259
+ "bytes": 987,
260
+ "sha256": "6a5953596b4432f597bff1585f74ce25fc25ef95e04c15beddf36313351e68a1"
261
+ },
262
+ "server/README.md": {
263
+ "bytes": 1678,
264
+ "sha256": "2ca738dcacd770e7c490ceeb59f2f11df7515f5bc56784eb9d476af4c7145b4d"
265
+ },
266
+ "server/pyproject.toml": {
267
+ "bytes": 850,
268
+ "sha256": "b63ab0db6e64813d431c27fb0e0903e9f58b53029575665a72771bd7cbb9ca00"
269
+ },
270
+ "server/src/wald_serve/__init__.py": {
271
+ "bytes": 200,
272
+ "sha256": "fa9cbe7f51816229614a3f4c265b8e70aa07a07058789bfa73371627bb67a319"
273
+ },
274
+ "server/src/wald_serve/__main__.py": {
275
+ "bytes": 51,
276
+ "sha256": "540fd0a5992ca535482a9d44b2e268ce3a39bb47a0576f301214afa5ef4494b9"
277
+ },
278
+ "server/src/wald_serve/engine.py": {
279
+ "bytes": 14814,
280
+ "sha256": "40fbea84c4c5b93ef888da44508d4fec2677c58732a09d731f0bcf17c2b48dd6"
281
+ },
282
+ "server/src/wald_serve/native.py": {
283
+ "bytes": 13936,
284
+ "sha256": "5a0f719b74e1b891e6c71f6cdf608e3da853feef515f1fc97fb56e4a264487e8"
285
+ },
286
+ "server/src/wald_serve/native_llamacpp.py": {
287
+ "bytes": 3827,
288
+ "sha256": "b2bab0e02d1f48fd94efdff6c0af1dea5ae5b69db17776324f3ff17672e7d90b"
289
+ },
290
+ "server/src/wald_serve/native_reference/eval/native_answer_acceptance_v1.py": {
291
+ "bytes": 19406,
292
+ "sha256": "764f3a683becd16b87b9ce9bfc01a5fe483ef26c9a3345a724b82780d73ea73a"
293
+ },
294
+ "server/src/wald_serve/native_reference/eval/systemone_vllm.py": {
295
+ "bytes": 36069,
296
+ "sha256": "7586a266a0933c50154aa86faab936b02b617146c357aed91e59467adf20adeb"
297
+ },
298
+ "server/src/wald_serve/native_reference/eval/vllm_letter.py": {
299
+ "bytes": 16993,
300
+ "sha256": "0d8c006c3cff7e43c2b019dde649ba1fac6bf47b960dce676412e4624acc815e"
301
+ },
302
+ "server/src/wald_serve/native_reference/kev/api.py": {
303
+ "bytes": 7086,
304
+ "sha256": "23bac9af0e79702428447c62662b88910071595f6171d9e3c82d6fa35fb0dbd1"
305
+ },
306
+ "server/src/wald_serve/native_reference/midtrain/encode.py": {
307
+ "bytes": 19164,
308
+ "sha256": "fb0d35c055c7052a3a3b019cd2cb4de1857ca733467680fa6182a106b045ecec"
309
+ },
310
+ "server/src/wald_serve/native_reference/midtrain/letter.py": {
311
+ "bytes": 22075,
312
+ "sha256": "90faed3734613799f33b40aae660142e9c8c02d34d73394bf9eb2f6aeba5635f"
313
+ },
314
+ "server/src/wald_serve/native_reference/midtrain/rawprompt.py": {
315
+ "bytes": 14571,
316
+ "sha256": "3b73b26cb96af71188e3dfd1a618c9460bfb69509a72f500d41391794ab219b1"
317
+ },
318
+ "server/src/wald_serve/native_reference/midtrain/tokens.py": {
319
+ "bytes": 3488,
320
+ "sha256": "7fbc482959e405dcc537bfadf0fa24df2d89bcaaab5919bcb765a70ea81e21de"
321
+ },
322
+ "server/src/wald_serve/native_reference/source-pins.json": {
323
+ "bytes": 962,
324
+ "sha256": "c8b231a1cdf413895c1abb215e43a41a08f6568d8404ea4f7ca86ee4883eb396"
325
+ },
326
+ "server/src/wald_serve/native_vision.py": {
327
+ "bytes": 8124,
328
+ "sha256": "3fd54bea1dddc79ee2456f393e13baf52305b7b70a7b7529cc5d2651f00a589f"
329
+ },
330
+ "server/src/wald_serve/prompt.py": {
331
+ "bytes": 4192,
332
+ "sha256": "bb45a10734f2546eb1ea1bd7e8b5717ecd8305a4aded0b816f80cd7e85188f7a"
333
+ },
334
+ "server/src/wald_serve/server.py": {
335
+ "bytes": 10020,
336
+ "sha256": "b236e0ec55aadf575198578285ef4a802f09f51d9b78f35e0da41d60a12755c2"
337
+ },
338
+ "server/src/wald_serve/vision_v2.py": {
339
+ "bytes": 10629,
340
+ "sha256": "ab7669e598e723a9a65d63356615b04c658f1fb02f1b7494bf40e69fd5627c7f"
341
+ },
342
+ "server/src/wald_serve/wire.py": {
343
+ "bytes": 3670,
344
+ "sha256": "836ce54ce93016a98f94a6beaa508cc03982e24c3d3b4d62b8ef667389cd14af"
345
+ },
346
+ "server/tests/fakevllm.py": {
347
+ "bytes": 4167,
348
+ "sha256": "9fb222a21bc2b17221e5a1a3605973b4f71730db9bbfede439bd3388b91e2570"
349
+ },
350
+ "server/tests/test_native_v2.py": {
351
+ "bytes": 7236,
352
+ "sha256": "ed5310bdcad9da34313719d83d83d6174329a115531fad7b98b1a3437d685eec"
353
+ },
354
+ "server/tests/test_parity.py": {
355
+ "bytes": 4392,
356
+ "sha256": "cd1e4142090076114f2b07efce49fdc5466e302a5c79996b993f3833a3457566"
357
+ },
358
+ "server/tests/test_server.py": {
359
+ "bytes": 13844,
360
+ "sha256": "46f2007e9743bfeb0f7f79fca093cf444a70df1bfc348229f34cd251f88a19b3"
361
+ },
362
+ "server/tests/test_vision_v2.py": {
363
+ "bytes": 5926,
364
+ "sha256": "7793a78d851f217c2af32bf963217bf431193e0f6bbf290e2bfd50107a03da88"
365
+ },
366
+ "serving.json": {
367
+ "bytes": 1172,
368
+ "sha256": "5caee30e9e486ab137f5521d7b2054844ce9a1bd4556df581b6cad48fbbdea22"
369
+ },
370
+ "temperature.json": {
371
+ "bytes": 2717,
372
+ "sha256": "a0f72cd2d0a653e81051e5a0c77fc1a69131552a8102b580a93a6dbe7908b2da"
373
+ },
374
+ "tokenizer.json": {
375
+ "sha256": "bd53432f0de26d67a83b634040d4f043053da4ecd0c759e7f4b24ad4f8bb9a81",
376
+ "bytes": 19989509
377
+ },
378
+ "tokenizer_config.json": {
379
+ "sha256": "bee8eba30f0eb4af73c0fe2cd06d0f89b657d7819941c438157ec42f7c80ea87",
380
+ "bytes": 1123
381
+ },
382
+ "tools/secret_scan.py": {
383
+ "bytes": 1698,
384
+ "sha256": "ad20de47de5bbe7f72417f3d87e6f72aae7987fab913eb814b818e4a4bbc6e1e"
385
+ },
386
+ "video_preprocessor_config.json": {
387
+ "sha256": "7768af27c1fafa9cc9011c1dc20067e03f8915e03b63504550e11d5066986d13",
388
+ "bytes": 385
389
+ },
390
+ "vocab.json": {
391
+ "sha256": "ce99b4cb2983d118806ce0a8b777a35b093e2000a503ebde25853284c9dfa003",
392
+ "bytes": 6722759
393
+ }
394
  }
NOTICE CHANGED
@@ -4,20 +4,24 @@ Copyright 2026 the Wald-4B authors
4
  The weights and the serving code in this repository are licensed under the Apache License, Version 2.0 (see LICENSE).
5
 
6
  Base model
7
- The weights are derived from Qwen/Qwen3.5-4B-Base (https://huggingface.co/Qwen/Qwen3.5-4B-Base), licensed under the
8
- Apache License, Version 2.0, copyright the Qwen team, Alibaba Cloud. Its licence and notices continue to apply.
 
 
9
 
10
- Third-party code and text in server/
11
- server/src/wald_serve/wire.py adapts the request models and the state / option rendering (render, option_text,
12
- question_keys, to_record) of Kev (https://github.com/jaredpalmer/kev), Copyright Jared Palmer, licensed under the
13
- Apache License, Version 2.0. Modified: the record carries each question's type and keys; unused fields removed.
14
 
15
- server/src/wald_serve/prompt.py contains the sentence STATE_REPEAT, copied verbatim from simple-jev
16
- (https://github.com/featherless-ai/simple-jev, commit dae340e, hf-server/hf_prompt_policies.py), licensed under the
17
- Apache License, Version 2.0. It is used only by the `repeat_state_plain` prompt format.
18
 
19
- No benchmark items or training data are included in this repository.
 
20
 
21
- 2026-09-29 release: Wald-Q4B 022D0-f7. See PROVENANCE.md for later training-source rights limitations; the code/base license does not grant rights to third-party training texts.
22
-
23
- 2026-10-01 release: Wald-Q4B v1.2 (02600-f19), a robustness LoRA stage merged into Wald-Q4B 022D0-f7. The inserted training texts of this stage were written by large frontier models (vendors and model names are not listed); see PROVENANCE.md.
 
 
4
  The weights and the serving code in this repository are licensed under the Apache License, Version 2.0 (see LICENSE).
5
 
6
  Base model
7
+ The v2 weights (checkpoint 04701-c22) are derived from Qwen/Qwen3.5-4B (https://huggingface.co/Qwen/Qwen3.5-4B,
8
+ revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a), licensed under the Apache License, Version 2.0, copyright the Qwen team, Alibaba Cloud.
9
+ chat_template.jinja is that model's chat template. Its licence and notices continue to apply. The v1.x weights at
10
+ the tags v1.0-v1.2 are derived from Qwen/Qwen3.5-4B-Base under the same licence.
11
 
12
+ Third-party code and text
13
+ server/src/wald_serve/wire.py and reference/kev/api.py (also server/src/wald_serve/native_reference/kev/api.py)
14
+ adapt the request models and the state / option rendering of Kev (https://github.com/jaredpalmer/kev), Copyright
15
+ Jared Palmer, licensed under the Apache License, Version 2.0.
16
 
17
+ server/src/wald_serve/prompt.py and reference/eval/systemone_vllm.py contain prompt sentences copied verbatim from
18
+ simple-jev (https://github.com/featherless-ai/simple-jev, commit dae340e, hf-server/hf_prompt_policies.py), licensed
19
+ under the Apache License, Version 2.0. v2's `repeat_state_plain` format uses only the STATE_REPEAT sentence.
20
 
21
+ evaluation/v2/code/ is the adapter and runner used with the Decision Index kit
22
+ (https://github.com/apolinario/decision-index, MIT); the kit itself is not included.
23
 
24
+ Training data
25
+ The training data includes synthetic data generated and/or labelled by large frontier models and public datasets
26
+ with their own terms; see PROVENANCE.md. The weights and code licence does not grant rights in third-party training
27
+ texts. No benchmark items or training data are included here.
PROVENANCE.md CHANGED
@@ -1,19 +1,7 @@
1
- # Training data
2
 
3
- **WaldGen** is the name of our internally generated decision-training corpus. It covers structured decisions and short reasoning. The model also uses public training datasets; WaldGen does not rename or claim authorship of those sources.
4
 
5
- Public-source acknowledgments include [ACOS](https://github.com/NUSTM/ACOS), [NLI4CT](https://github.com/ai-systems/Task-2-SemEval-2024), [RAGTruth](https://github.com/ParticleMedia/RAGTruth), and [VAST](https://github.com/emilyallaway/zero-shot-stance). RAGTruth contains third-party MS MARCO/Yelp contexts. [MS MARCO terms](https://microsoft.github.io/msmarco/) limit dataset use to non-commercial research; other source-text permissions are not uniformly specified. Model/code licensing does not grant rights in those texts or constitute commercial clearance.
6
 
7
- Base model: [Qwen3.5-4B-Base](https://huggingface.co/Qwen/Qwen3.5-4B-Base), Apache-2.0. No training rows or benchmark request payloads are distributed here.
8
-
9
- # v1.2 (checkpoint `02600-f19`)
10
-
11
- v1.2 adds one LoRA stage on v1.1. The v1.1 notes above apply unchanged.
12
-
13
- - **Questions.** 5,300 questions drawn from v1.1's own training text. No new source dataset was added.
14
- - **Perturbed rows.** 8,064 perturbed copies of those questions (13,674 training rows with repeats and 5,000 unperturbed replay rows). Each perturbed row has a short distracting or pressuring text inserted into the question, after an option or into the state, or is a paraphrase or a typo variant.
15
- - **Text written by a model.** The inserted texts and the paraphrases were written by **large frontier models** from our own templates (vendors and model names are not listed). A script inserted them, so the original question text is unchanged. Typos were produced by a script. A separate large-frontier-model call checked every row, and rows on which v1.1 changed its answer were checked again by a second large frontier model.
16
- - **Targets.** v1.1's own answer distribution on the unperturbed question.
17
- - **Benchmarks.** No JevAdvBench text was used. JevAdvBench is evaluation-only.
18
-
19
- No training rows or benchmark request payloads are distributed here.
 
1
+ # Training data of Wald-Q4B v2 (checkpoint `04701-c22`)
2
 
3
+ Base model: [Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B) (chat, revision `851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a`), Apache-2.0. Wald-Q4B v2 was trained from it by full-parameter decision training.
4
 
5
+ **Data.** About 15% of training tokens come from the train splits of public benchmarks that the Decision Index also draws on, reformatted as decisions. The rest is broad decision data: general decision tasks — rules and policies, classification, entailment and fact checking, tables and documents, web and tool actions, judging (about 40%); math and step-by-step reasoning (about 20%); and hard decisions written and/or labelled by large frontier models, including robustness cases with perturbed or adversarial states (about 25%). Public datasets are used as train splits only. Contamination scan against the full Decision Index 0.3 public suite: 0 strict hits.
6
 
7
+ **Contamination.** Every training file was scanned with a strict state / option matcher against the complete Decision Index 0.3 public suite: 0 hits. Details: [CONTAMINATION.md](https://huggingface.co/org2ai/Wald-4B/blob/main/CONTAMINATION.md).
 
 
 
 
 
 
 
 
 
 
 
 
README.md CHANGED
@@ -1,6 +1,6 @@
1
  ---
2
  license: apache-2.0
3
- base_model: Qwen/Qwen3.5-4B-Base
4
  base_model_relation: finetune
5
  library_name: transformers
6
  pipeline_tag: text-generation
@@ -13,388 +13,89 @@ tags:
13
  - calibrated-probabilities
14
  - classification
15
  - tool-selection
16
- - tool-use
17
  - agent-routing
18
- - clarification
19
- - robustness
20
  - decision-index
21
  - jevbench
22
- - jev
23
  - jev-compatible
24
- - open-jev
25
- - typesafe-compatible
26
  - systemone
27
- - kev
28
- - laya
29
  - wald
30
  - wald-q4b
31
  - qwen3.5
32
  - 4b
33
  - vllm
 
34
  - gguf
35
  - llama.cpp
36
- - ollama
37
- - reasoning
38
  model-index:
39
- - name: Wald-Q4B v1.1 (revision v1.1)
40
  results:
41
  - task:
42
  type: text-classification
43
  name: Structured decisions (option probabilities)
44
  dataset:
45
- name: Decision Index 0.2.1, complete suite (150,317 requests)
46
  type: decision-index
47
- revision: 87d4650b42b377c0291a89c1f1a879f9b31082bf
48
  metrics:
49
  - type: balanced_skill_index
50
- name: Balanced-skill index, effort high (author-run, not leaderboard-verified)
51
- value: 54.59
52
  verified: false
53
- source:
54
- name: 'Author-run with the official kit; submission PR #30 awaits maintainer validation'
55
- url: https://github.com/apolinario/decision-index/pull/30
56
- - task:
57
- type: text-classification
58
- name: Structured decisions (option probabilities)
59
- dataset:
60
- name: JevBench public set (231 items)
61
- type: jevbench
62
- split: public
63
- revision: 9ec6f15a
64
- metrics:
65
- - type: accuracy
66
- name: Accuracy, effort none (self-scored, 203/231)
67
- value: 87.88
68
- verified: false
69
- - type: ece
70
- name: Expected calibration error, 10 bins (self-scored)
71
- value: 0.041
72
- verified: false
73
- - type: brier
74
- name: Brier score (self-scored)
75
- value: 0.188
76
- verified: false
77
- source:
78
- name: 'Self-scored with the JevBench harness; row requested in jevbench issue #146'
79
- url: https://github.com/fstandhartinger/jevbench/issues/146
80
- - name: Wald-Q4B v1.2 (main, revision v1.2)
81
- results:
82
- - task:
83
- type: text-classification
84
- name: Structured decisions (option probabilities)
85
- dataset:
86
- name: JevBench public set (231 items)
87
- type: jevbench
88
- split: public
89
- revision: 9ec6f15a
90
- metrics:
91
- - type: accuracy
92
- name: Accuracy, effort none (self-scored, 204/231)
93
- value: 88.31
94
- verified: false
95
- - type: ece
96
- name: Expected calibration error, 10 bins (self-scored)
97
- value: 0.045
98
- verified: false
99
- - type: brier
100
- name: Brier score (self-scored)
101
- value: 0.191
102
- verified: false
103
- source:
104
- name: Self-scored with the JevBench harness; not submitted
105
- url: https://huggingface.co/org2ai/Wald-4B/blob/v1.2-release/evaluation/v1.2/summary.json
106
- - task:
107
- type: text-classification
108
- name: Robustness of structured decisions to distracting and adversarial text
109
- dataset:
110
- name: JevAdvBench (812 questions, 9 delivered attack types)
111
- type: jevadvbench
112
- revision: "3218e05"
113
- metrics:
114
- - type: flip_rate
115
- name: Mean flip rate in %, effort none, lower is better (self-run)
116
- value: 4.6
117
- verified: false
118
- source:
119
- name: Self-run with the benchmark's request bytes and analysis code; not submitted
120
- url: https://huggingface.co/org2ai/Wald-4B/blob/v1.2-release/evaluation/v1.2/summary.json
121
  ---
122
 
123
- <div align="center">
124
- <h1>Wald-Q4B</h1>
125
- <p><strong>Decide directly. Think when needed. Return probabilities.</strong></p>
126
- <p><a href="https://huggingface.co/org2ai/Wald-4B">English</a> · <a href="https://huggingface.co/org2ai/Wald-4B/blob/main/docs/readmes/README.zh.md">简体中文</a> · <a href="https://github.com/org2AI/wald-4b">GitHub</a> · <a href="https://huggingface.co/org2ai/Wald-4B/blob/main/docs/api.md">API</a></p>
127
- </div>
128
-
129
- **Wald-Q4B is an open-weight 4B decision model: give it a state and a set of options, and it returns a calibrated probability for every option.** It is for developers who build agents and pipelines and need a fast, self-hosted component to pick a tool, route a request, classify an input or decide whether to ask the user. Unlike a chat model, it does not write an answer you have to parse. It reads the options in one pass and can optionally think first. It serves a Jev-compatible `POST /v1/systemone` API, is built on Qwen3.5-4B-Base and is released under Apache-2.0.
130
 
131
- **Two releases live in this repository.**
132
 
133
- | Release | Revision | Checkpoint | In one line |
134
- |---|---|---|---|
135
- | **v1.2** (2026-10-01) | tag [`v1.2`](https://huggingface.co/org2ai/Wald-4B/tree/v1.2); also the weights on `main` | `02600-f19` | The robustness release: v1.1 plus one merged LoRA stage that teaches the model to keep its answer when the input contains distracting sentences, third-party opinions or fake instructions |
136
- | **v1.1** (2026-09-29) | tag [`v1.1`](https://huggingface.co/org2ai/Wald-4B/tree/v1.1) | `022D0-f7` | The general release, with the complete Decision Index run and the evaluated thinking efforts |
137
 
138
- Since 2026-10-01 `main` holds the v1.2 weights (before that it held v1.1). The pending benchmark requests for v1.1 name fixed commits of this repository and are unaffected; for v1.1, download `--revision v1.1`. Pin a revision when you download.
 
 
 
 
 
 
139
 
140
- This repository moved from `Harry19081/Wald-4B` to `org2ai/Wald-4B` on 2026-10-01; old links redirect here. The cards stored at the release tags (`v1.0`, `v1.1`, `v1.2`) are immutable and still show the old path, which also redirects.
141
 
142
- Wald-Q4B is an independent, self-hosted alternative to TypeSafe's hosted Jev API. It is not Jev, contains no Jev weights, and is not affiliated with or endorsed by TypeSafe AI. The Hugging Face repository is `org2ai/Wald-4B` (earlier name: Wald-4B; moved from `Harry19081/Wald-4B` on 2026-10-01, old links redirect).
143
-
144
- **W**ait **A** bit, **L**ook, then **D**ecide. Also named after Abraham Wald, the pioneer of sequential analysis: stop when the evidence is enough.
145
-
146
- ## At a glance
147
-
148
- - **4B parameters**, built on Qwen3.5-4B-Base; BF16 weights (8.4 GB).
149
- - **Every option gets a probability.** Question types: `choice` (1–255 named options), `noul` (yes/no) and `score` (ordered levels).
150
- - **v1.2 is more robust than v1.1:** on JevAdvBench the mean flip rate over nine attack types is **4.6 %** (v1.1: 9.2 %; Jev 1.13: 6.1 %). Self-run, effort `none`.
151
- - **The cost of v1.2:** clean accuracy on JevAdvBench's 143 human-reviewed questions is 76.2 % (v1.1: 79.0 %; −2.8 points, 95 % CI [−6.2, −0.6]).
152
- - **JevBench public set:** v1.2 **204/231**, v1.1 **203/231**, both with `none`. Self-scored with JevBench's own harness.
153
- - **Decision Index 0.2.1:** the complete-suite result, **54.59** with `high`, was measured on **v1.1**. v1.2 has only a one-pass sample read, which is level with v1.1 (+0.27, not significant).
154
- - **Adjustable thinking:** `none`, `low`, `medium`, `high`, or several thoughts with `high-k`. Thinking is evaluated on v1.1. v1.2 is a one-pass model: use it with `none`.
155
- - **Self-hosted API:** `POST /v1/systemone`, up to 131,072 prompt tokens.
156
-
157
- ## Which version should I use?
158
-
159
- | Use | Revision | Why |
160
- |---|---|---|
161
- | Inputs may contain distracting, persuasive or adversarial text (web pages, user messages, tool outputs, retrieved documents) | **`v1.2`** | About half as many changed answers under the nine JevAdvBench attacks |
162
- | Clean, trusted inputs where the last points of accuracy matter; thinking efforts; the evaluated Decision Index configuration | **`v1.1`** | Slightly higher clean accuracy; the complete Decision Index run (54.59, `high`) and the thinking efforts were measured on it |
163
 
164
- Both revisions share the same architecture, tokenizer, prompt format, calibration table and server code. The two weight shards differ, and so does the default effort declared in `serving.json`: `none` for v1.2, `high` for v1.1.
165
 
166
  ## Quick start
167
 
168
- On a Linux machine with an NVIDIA GPU and [`uv`](https://docs.astral.sh/uv/):
169
-
170
  ```sh
171
- hf download org2ai/Wald-4B --revision v1.2 --local-dir ./Wald-Q4B # v1.2, robustness release
172
- # hf download org2ai/Wald-4B --revision v1.1 --local-dir ./Wald-Q4B # v1.1, general release
173
- cd Wald-Q4B
174
- EFFORT=none ./run.sh "$PWD" # one pass, lowest latency
175
- # ./run.sh "$PWD" # the release's declared default: none for v1.2, high for v1.1
176
- ```
177
-
178
- From Python, pin the revision the same way:
179
-
180
- ```python
181
- from huggingface_hub import snapshot_download
182
 
183
- path = snapshot_download("org2ai/Wald-4B", revision="v1.2") # or revision="v1.1"
184
- ```
 
 
185
 
186
- ```sh
187
- curl http://localhost:8000/v1/systemone \
188
- -H 'Content-Type: application/json' \
189
- -d '{
190
- "state": "The customer wants to return a damaged kettle.",
191
- "effort": "none",
192
- "questions": {
193
- "route": {
194
- "type": "choice",
195
- "instructions": "Choose the support queue.",
196
- "criteria": {
197
- "returns": "Returns and refunds",
198
- "delivery": "Delivery tracking",
199
- "other": "Other enquiries"
200
- }
201
- }
202
- }
203
- }'
204
  ```
205
 
206
- The answer for `route` contains the chosen key and a probability for each of `returns`, `delivery` and `other`. Request and response fields, yes/no and score questions, and a clarification example: [API reference](https://huggingface.co/org2ai/Wald-4B/blob/main/docs/api.md). `GET /health` reports the effective policy. The server uses vLLM 0.30.0 and the included `wald-serve` package; Docker and exact evaluation settings are in [RUNBOOK.md](https://huggingface.co/org2ai/Wald-4B/blob/main/RUNBOOK.md). Generic text-generation calls do not reproduce the decision API's readout.
207
-
208
- ## GGUF: llama.cpp, Ollama, LM Studio
209
-
210
- `main` also carries v1.2 as GGUF files for CPUs, Apple Silicon and consumer GPUs (llama.cpp `b11312`, the same weights). Parity on the JevBench public set (231 items, `none`), against the BF16 weights on vLLM (204/231):
211
-
212
- | File | Size | JevBench public | Same option as BF16 |
213
- |---|---:|---:|---:|
214
- | `Wald-4B-v1.2-Q8_0.gguf` | 4.5 GB | 206/231 | 229/231 |
215
- | `Wald-4B-v1.2-Q6_K.gguf` | 3.5 GB | 205/231 | 227/231 |
216
- | `Wald-4B-v1.2-Q5_K_M.gguf` | 3.1 GB | 205/231 | 227/231 |
217
- | `Wald-4B-v1.2-Q4_K_M.gguf` | 2.7 GB | 202/231 | 223/231 |
218
-
219
- For calibrated probabilities, run the included server on llama.cpp (`llama-server` on your `PATH`):
220
-
221
- ```sh
222
- hf download org2ai/Wald-4B --include "Wald-4B-v1.2-Q8_0.gguf" "serving.json" "temperature.json" "server/*" --local-dir ./wald-gguf
223
- pip install ./wald-gguf/server
224
- wald-serve --gguf ./wald-gguf/Wald-4B-v1.2-Q8_0.gguf --max-model-len 32768 --port 8000
225
- ```
226
-
227
- Ollama: `ollama run hf.co/org2ai/Wald-4B:Q4_K_M`. LM Studio: search for `Wald-4B`. Neither app was tested with these files, and a chat session returns text, not the option probabilities. Download a single file with `--include`; a plain `hf download org2ai/Wald-4B` of `main` also fetches all four GGUFs (14 GB). Details: [org2ai/Wald-4B-GGUF](https://huggingface.co/org2ai/Wald-4B-GGUF).
228
-
229
- ## Thinking effort
230
-
231
- | Effort | When it thinks | Thought budget |
232
- |---|---|---|
233
- | `none` | Direct option readout | No generated thought |
234
- | `low` | Top initial probability < 0.5 | Up to 512 tokens |
235
- | `medium` | Top initial probability < 0.7 | Up to 512 tokens |
236
- | **`high` (v1.1 default)** | Every eligible question | Up to 512 tokens |
237
- | `high-k2` … `high-k8` | Multiple thoughts; average their answer distributions | Up to 512 tokens per thought |
238
-
239
- Thinking applies to questions with 2–26 options when context space permits. Larger option sets use grouped readout and a final winner comparison; if a thought cannot fit, the initial answer is kept. Increasing effort spends more computation; it does not guarantee a better answer.
240
-
241
- **On v1.1**, start with `none` for direct decisions, `medium` for confidence-gated thinking, or `high` for the evaluated Decision Index configuration. **54.59 applies to v1.1 with `high` only; 203/231 applies to v1.1 with `none` only.**
242
-
243
- **v1.2 is a one-pass model.** It was trained and evaluated on the one-pass readout, and every v1.2 number on this card is `none`. In our one check, thinking did not help it: on the JevBench public set, v1.2 scored 198/231 with `medium` and 198/231 with `high` (one run each) against 204/231 with `none`, while v1.1 scored 205/231 with `medium`. v1.2's `serving.json` therefore declares `none` as its default. If you want the thinking efforts, use v1.1.
244
-
245
- Set the server default with `EFFORT=medium ./run.sh "$PWD"`, or override it per request with `"effort": "none"`.
246
 
247
  ## How it works
248
 
249
- Wald first reads option-letter logits from a plain prompt and turns them into a probability distribution. If the effort policy asks for thinking, it generates a short thought and reads the options again. Bucketed temperature scaling calibrates the returned probabilities.
250
-
251
- v1.1 combines full-parameter decision training, LoRA refinement, short-thought distillation and RLCD. Training uses **WaldGen**, our generated decision corpus, together with public training datasets.
252
-
253
- **v1.2 adds one LoRA stage on v1.1 (rank 16 on every language projection, merged into the weights):**
254
-
255
- - **Questions:** 5,300 questions drawn from v1.1's own training text. No new source dataset.
256
- - **Perturbations:** 8,064 perturbed copies. Each one has a short text inserted into the question, after an option or into the state: unrelated sentences and off-topic passages, a bystander's opinion, rumour or analogy that pushes another option, or a fake instruction that claims authority. A smaller share are paraphrases and typos.
257
- - **Who wrote them:** **the inserted texts and the paraphrases were written by large frontier models from our own templates (vendors and model names are not listed).** A script inserted them, so the original facts stay byte-identical. Typos were made by a script.
258
- - **Checks:** a separate large-frontier-model call checked every row ("does the edit change the correct answer?") and rejected rows were dropped. Rows on which v1.1 changed its answer were checked a second time by another large frontier model.
259
- - **Targets:** v1.1's own answer distribution on the clean question. The model is taught to answer the perturbed question as v1.1 answers the clean one. 5,000 clean rows are replayed against v1.1 to limit drift.
260
- - **Separation from the benchmark:** no JevAdvBench text was used. An 8-gram overlap check of every inserted text against every JevAdvBench string (clean questions and all 9,744 attacked variants) found 0 hits.
261
-
262
- Data sources and evaluation notes: v1.1 [PROVENANCE.md](https://huggingface.co/org2ai/Wald-4B/blob/v1.1/PROVENANCE.md) · [CONTAMINATION.md](https://huggingface.co/org2ai/Wald-4B/blob/v1.1/CONTAMINATION.md); v1.2 [PROVENANCE.md](https://huggingface.co/org2ai/Wald-4B/blob/v1.2-release/PROVENANCE.md) · [CONTAMINATION.md](https://huggingface.co/org2ai/Wald-4B/blob/v1.2-release/CONTAMINATION.md)
263
-
264
- ## Benchmarks
265
-
266
- All numbers below are self-run and self-reported. None of them is a leaderboard result.
267
-
268
- | Benchmark | Configuration | v1.1 | **v1.2** | Note |
269
- |---|---|---:|---:|---|
270
- | **JevAdvBench, mean flip rate over 9 attack types** (lower is better) | `none` | 9.2 % | **4.6 %** | Paired −4.6 points, 95 % CI [−5.5, −3.6]. Jev 1.13: 6.1 % |
271
- | **JevAdvBench, clean accuracy on 143 human-reviewed questions** | `none` | 79.0 % | **76.2 %** | Paired −2.8 points [−6.2, −0.6]. Jev 1.13: 87.4 % |
272
- | **JevBench public set (231 items)** | `none` | 203/231 (87.9 %) · ECE 0.041 · Brier 0.188 | **204/231** (88.3 %) · ECE 0.045 · Brier 0.191 | Self-scored with JevBench's harness. v1.1 row requested in [issue #146](https://github.com/fstandhartinger/jevbench/issues/146); v1.2 not submitted |
273
- | JevBench public set (231 items) | `medium` | 205/231 | 198/231 | One run each. `high` on v1.2: 198/231 |
274
- | **Decision Index 0.2.1, 6,948-request sample** | one pass | 49.76 | **50.03** | Paired +0.27 [−0.36, +1.05], not significant |
275
- | **Decision Index 0.2.1, complete suite** | `high` | **54.59** | not run | v1.1 only. Author-run; [PR #30](https://github.com/apolinario/decision-index/pull/30) awaits maintainer validation |
276
-
277
- Two of our internal non-regression sets, which are not public benchmarks, also held in one pass: XL-Int 56.4 (v1.1: 56.0) and a 2,857-item tool-selection set 87.22 % (v1.1: 87.15 %).
278
-
279
- ### Robustness: JevAdvBench
280
-
281
- [JevAdvBench](https://github.com/JevAdvBench/JevAdvBench) attacks 812 decision questions with nine kinds of edits and counts a **flip** when the model's decision on the attacked question differs from its own decision on the clean question. We sent the benchmark's request bytes to the packaged server with `none` and scored the answers with the benchmark's analysis code (`JevAdvBench@3218e05`). Flip rates are in % of the 812 questions; intervals are 95 % scenario-cluster bootstrap intervals.
282
-
283
- | Attack | Jev 1.13 | v1.1 | **v1.2** | v1.2 − v1.1 (paired) |
284
- |---|---:|---:|---:|---:|
285
- | Q1 word edits | 1.0 | 1.6 | 1.6 | 0.0 [−0.8, 0.8] |
286
- | Q2 paraphrase | 1.5 | 1.6 | 1.4 | −0.2 [−1.1, 0.5] |
287
- | Q3 unrelated sentences in the question | 4.6 | 14.9 | **4.6** | **−10.3 [−13.5, −7.2]** |
288
- | T1 unrelated note in the state | 2.2 | 2.8 | 1.7 | −1.1 [−2.4, 0.0] |
289
- | T2 observer's opinion in the state | 12.1 | 16.9 | **5.5** | **−11.3 [−14.6, −8.2]** |
290
- | T3 opinion through an analogy | 6.9 | 12.7 | **6.3** | **−6.4 [−8.8, −3.6]** |
291
- | P1 direct override | 8.9 | 7.9 | 5.7 | −2.2 [−3.6, −0.9] |
292
- | P2 authority impersonation | 10.1 | 13.7 | **6.7** | **−7.0 [−9.0, −5.1]** |
293
- | P3 fake validation note | 8.1 | 10.3 | 7.8 | −2.6 [−4.2, −1.0] |
294
- | **Mean of the nine** | **6.1** | **9.2** | **4.6** | **−4.6 [−5.5, −3.6]** |
295
- | Questions flipped by at least one attack | 29.2 | 41.7 | 20.7 | |
296
- | Clean accuracy, 143 human-reviewed questions | 87.4 | 79.0 | 76.2 | −2.8 [−6.2, −0.6] |
297
-
298
- - **Jev column:** `jev-1.13.0`'s responses as released by the benchmark authors, scored by the same code. v1.2 − Jev on the mean: −1.6 points [−2.7, −0.4].
299
- - **Clean accuracy:** Jev is more accurate than both Wald versions on the clean human-reviewed questions. v1.2 makes the same clean decision as v1.1 on 98.6 % of the 812 questions.
300
- - **What this does not show:** the perturbation kinds for v1.2 were chosen after we saw v1.1's per-attack results, so the benchmark's attack families informed the training design. Its texts did not. Robustness to attack types outside these families has not been measured.
301
-
302
- ### JevBench and Decision Index
303
-
304
- **JevBench:** JevBench's own CLI (`fstandhartinger/jevbench` at `9ec6f15a`, `typesafe` adapter) against the packaged server on loopback, one request at a time. v1.2 with `none` on one RTX 5090: easy 48/48, original 72/72, hard 84/111; no tokens generated. v1.1 with `none` on one RTX PRO 6000: easy 48/48, original 72/72, hard 83/111; with `medium` 205/231, p95 1.80 s. The public items were used as a development scoreboard (never as training data), so this is not a held-out result. The JevBench leaderboard publishes a score only after its maintainers run the model themselves.
305
-
306
- **Re-read after release:** we downloaded the `v1.2` revision anonymously and repeated the `none` read. From a directory holding only the downloaded model files it reproduced 204/231 with the same option on all 231 items (ECE 0.045). From the complete repository directory, two reads gave 204/231 and 205/231 (ECE 0.053), with a different option on one or two near-tied items. The model files are byte-identical in every case; we have not yet explained this small difference. [Details](https://huggingface.co/org2ai/Wald-4B/blob/main/evaluation/v1.2/release-check.json)
307
-
308
- **Decision Index (v1.1):** 150,317/150,317 requests succeeded, including HLE. Measured on one RTX PRO 6000 96 GB with the pinned reproduction kit. [Full results](https://huggingface.co/datasets/org2ai/Wald-Q4B-decision-index-results/tree/805716601b2466be324ed6716407b4c3d9267faa/runs/wald-q4b-22d0-f7-full021) · [Per-benchmark scores](https://huggingface.co/org2ai/Wald-4B/blob/main/evaluation/benchmark-summary.json) · [Reproduction guide](https://huggingface.co/org2ai/Wald-4B/blob/main/RUNBOOK.md). The files under `evaluation/` on `main` belong to this v1.1 run; v1.2's numbers are in [`evaluation/v1.2/summary.json`](https://huggingface.co/org2ai/Wald-4B/blob/v1.2-release/evaluation/v1.2/summary.json) at the `v1.2` revision. The v1.2 sample row above is a one-pass read of 6,948 requests and cannot be compared with 54.59.
309
-
310
- ### Latency
311
-
312
- Latency was measured on v1.1 and not again on v1.2, which has the same architecture, size and server. With `none`, the v1.1 JevBench run measured **33 ms median and 168 ms p95** per decision on one RTX PRO 6000. A 32-request serial preflight of `high` measured **821 ms median** on the same GPU; this small preflight is not a full-suite latency result or the Decision Index maintainers' admission test. Effort, context length, option count and concurrency all affect speed.
313
-
314
- ## Related projects and how Wald compares
315
-
316
- Several projects implement or approximate structured decisions with calibrated option probabilities. The names below belong to their owners; Wald is not affiliated with any of them.
317
-
318
- - **[Jev](https://typesafe.ai/blog/introducing-system-one-models-and-jev)** is TypeSafe AI's hosted decision model behind the `/v1/systemone` API; its weights are closed. Wald accepts the same request shape and runs on your own GPU.
319
- - **[Kev](https://github.com/jaredpalmer/kev)** by Jared Palmer is an open-weight project that adds a LoRA and a pointer head to Qwen3.5 base models (0.8B, 4B and 9B). Wald's request parsing adapts Kev's Apache-2.0 code ([NOTICE](https://huggingface.co/org2ai/Wald-4B/blob/main/NOTICE)); Wald reads option letters from the language-model head instead of a separate head.
320
- - **[Laya](https://huggingface.co/convaiinnovations/laya)** ([code](https://github.com/NandhaKishorM/laya)) is an open-weight 421M ModernBERT-large encoder with a decision head. It is much smaller than Wald and reads up to 512 tokens.
321
-
322
- **JevBench public set, same 231 items (identical dataset hash), JevBench CLI, run by us:**
323
-
324
- | System | How it was run | Correct |
325
- |---|---|---:|
326
- | Wald-Q4B v1.2 · `none` | Self-hosted, RTX 5090, 2026-09-30 | 204/231 |
327
- | Wald-Q4B v1.1 · `none` | Self-hosted, RTX PRO 6000, 2026-09-29 | 203/231 |
328
- | Jev (`jev-1.13.0`) | TypeSafe's hosted API, 2026-09-25 | 200/231 |
329
- | Laya (English checkpoint `55cf4c4e`) | Self-hosted, NVIDIA L4, 2026-09-26 | 134/231 |
330
 
331
- A difference of a few items on 231 is within run-to-run and sampling noise. The public items informed Wald's development, and 52 of the 231 states are longer than Laya's 512-token window.
332
-
333
- **Decision Index 0.2.1:**
334
-
335
- | System | Index | Source |
336
- |---|---:|---|
337
- | Jev (`jev-1.13.0`) | 57.91 | [Leaderboard](https://huggingface.co/spaces/multimodalart/jev-decision-index), maintainer-run (data of 2026-09-28) |
338
- | Wald-Q4B v1.1 · `high` | 54.59 | Author-run complete suite; not on the leaderboard yet ([PR #30](https://github.com/apolinario/decision-index/pull/30)) |
339
- | Kev 9B | 38.48 | Leaderboard, maintainer-run (data of 2026-09-28) |
340
- | Kev 4B | 34.64 | Leaderboard, maintainer-run (data of 2026-09-28) |
341
-
342
- Leaderboard rows are scored by the maintainers; Wald's number is self-run with the official kit and may change after validation.
343
-
344
- ## FAQ
345
-
346
- **Is there an open-source alternative to Jev?** Wald-Q4B is one open-weight option: Apache-2.0 weights and serving code that you run yourself, with a Jev-compatible `/v1/systemone` API. Kev and Laya (above) are other open projects. Wald is independent and is not a TypeSafe release.
347
-
348
- **Can I use a Jev client with a self-hosted model?** Point the client at your own endpoint. The included server accepts `state` plus typed `questions` (`choice`, `noul`, `score`) at `POST /v1/systemone` and answers with TypeSafe's answer keys. It does not check API keys. See the [API reference](https://huggingface.co/org2ai/Wald-4B/blob/main/docs/api.md).
349
-
350
- **How do I route tools or decide whether to ask the user?** Send the conversation or task as `state`. For tool routing, ask a `choice` question whose options are your tools. To decide whether to ask a clarifying question, ask a `noul` question such as "Is the request specific enough to act on without asking?" Act when the probability is high, ask when it is low, and set both thresholds on your own validation data. The model picks the tool; it does not write the tool's arguments.
351
-
352
- **Is v1.2 safe against prompt injection?** No model is. v1.2 changes its answer about half as often as v1.1 under the nine JevAdvBench attacks, and it still flips on 4.6 % of attacked questions on average and on 20.7 % of questions under at least one attack. Treat it as one layer: keep untrusted text out of the instructions where you can, and check high-stakes decisions.
353
-
354
- **How calibrated are the probabilities?** On the JevBench public set with `none`, expected calibration error is 0.045 for v1.2 (0.053 in the re-read from the complete release directory) and 0.041 for v1.1 (10 bins); the Brier scores are 0.191 and 0.188. v1.2 uses v1.1's temperature table unchanged. It was fitted on held-out rows of our own development data, with no JevBench items and with known Decision Index matches excluded. A confidence is not a guarantee; check calibration on your task.
355
-
356
- **Does it run on a single GPU or a laptop?** The shipped server needs one NVIDIA GPU on Linux (vLLM 0.30.0); the BF16 weights are 8.4 GB. The v1.1 measurements come from an RTX PRO 6000 96 GB and the v1.2 measurements from an RTX 5090 32 GB; earlier builds of the same 4B architecture have also been served with vLLM on a 24 GB NVIDIA L4 at a 16K context limit. CPU, Apple Silicon and laptop setups are not supported by the shipped server and have not been tested.
357
-
358
- **Kev vs Wald, or Laya vs Wald?** All three are open-weight. Kev adds a pointer head and LoRA to Qwen3.5 base models; Laya is a small encoder with a decision head; Wald is a fully trained 4B decoder with optional thinking. Our same-protocol measurements are in the tables above. Choose on your own task, latency budget and hardware.
359
-
360
- **Can I fine-tune it for my task?** It is a standard Transformers checkpoint, so common LoRA tooling applies. For v1.0 we trained per-task LoRAs for $0.12–$1.81 of GPU time each; that tooling is not public yet, and those adapters are not validated on v1.1 or v1.2 ([v1.0 notes](https://huggingface.co/org2ai/Wald-4B/blob/main/history/v1.0/README.md)).
361
-
362
- **What is the license?** Apache-2.0 for the weights and code. The base model, Qwen3.5-4B-Base, is also Apache-2.0. Some public training sources have their own terms or no stated licence; they are listed in [PROVENANCE.md](https://huggingface.co/org2ai/Wald-4B/blob/main/PROVENANCE.md). The model and code licence does not grant rights in those texts. v1.2 adds no new source dataset; its added training text was written by large frontier models, as described above.
363
-
364
- ## Limits
365
-
366
- - **v1.2 trades a little clean accuracy for robustness:** −2.8 points on JevAdvBench's 143 human-reviewed clean questions (95 % CI [−6.2, −0.6]). Use `v1.1` if that matters more to you.
367
- - **v1.2 is a one-pass model.** Its results are all `none`, it has no complete Decision Index run, and in our one check the thinking efforts lowered its JevBench public score (198/231 against 204/231). Use v1.1 for thinking.
368
- - **Robustness is measured on one benchmark** and on the attack families that informed the training design. It is not a security guarantee.
369
- - Confidence is not a guarantee of correctness; validate thresholds on your own task. Oversized prompts are rejected rather than truncated.
370
- - Known exact training overlaps were filtered, but semantic overlap and pretraining contamination are not ruled out; development used visible benchmark samples. Source-text rights vary. See [evaluation notes](https://huggingface.co/org2ai/Wald-4B/blob/main/CONTAMINATION.md) and [source attribution](https://huggingface.co/org2ai/Wald-4B/blob/main/PROVENANCE.md).
371
-
372
- ## Versioning
373
-
374
- | Version | Revision | Checkpoint | What it is |
375
- |---|---|---|---|
376
- | **v1.2 — robustness release** | tag [`v1.2`](https://huggingface.co/org2ai/Wald-4B/tree/v1.2); weights on `main` | `02600-f19` | v1.1 + merged robustness LoRA; default effort `none` |
377
- | **v1.1 — general release** | tag [`v1.1`](https://huggingface.co/org2ai/Wald-4B/tree/v1.1) | `022D0-f7` | Complete Decision Index run (54.59); default effort `high` |
378
- | v1.0 — archive | tag [`v1.0`](https://huggingface.co/org2ai/Wald-4B/tree/v1.0) | `021A0-f10` | Default effort `medium` |
379
-
380
- The pending benchmark requests for v1.1 (JevBench issue #146, Decision Index PR #30) name fixed commits of this repository and are not affected by v1.2. The v1.0 model card retains its XL, task-LoRA and latency reports, and the [v1.0 walkthrough slides](https://claude.ai/artifact/XfCHVaCuj9A5ectrpaWzV5) describe v1.0 only. Those measurements belong to their documented builds.
381
-
382
- ## Citation
383
-
384
- Wald-Q4B (2026), an open-weight 4B decision model with calibrated option probabilities. https://huggingface.co/org2ai/Wald-4B. Name the revision you used (`v1.2` or `v1.1`).
385
-
386
- ```bibtex
387
- @misc{wald_q4b_2026,
388
- title = {Wald-Q4B: an open-weight 4B decision model with calibrated option probabilities},
389
- author = {{Wald-4B authors}},
390
- year = {2026},
391
- howpublished = {\url{https://huggingface.co/org2ai/Wald-4B}},
392
- note = {Revision v1.2}
393
- }
394
- ```
395
-
396
- Machine-readable: [CITATION.cff](https://huggingface.co/org2ai/Wald-4B/blob/main/CITATION.cff) · [llms.txt](https://huggingface.co/org2ai/Wald-4B/blob/main/llms.txt) · [model-info.json](https://huggingface.co/org2ai/Wald-4B/blob/main/model-info.json)
397
-
398
- ---
399
 
400
- [Model](https://huggingface.co/org2ai/Wald-4B) · [GitHub](https://github.com/org2AI/wald-4b) · [Decision Index results](https://huggingface.co/datasets/org2ai/Wald-Q4B-decision-index-results) · Apache-2.0 for weights and code. [Third-party notices](https://huggingface.co/org2ai/Wald-4B/blob/main/NOTICE) · [Training-data usage notes](https://huggingface.co/org2ai/Wald-4B/blob/main/PROVENANCE.md)
 
 
 
 
 
1
  ---
2
  license: apache-2.0
3
+ base_model: Qwen/Qwen3.5-4B
4
  base_model_relation: finetune
5
  library_name: transformers
6
  pipeline_tag: text-generation
 
13
  - calibrated-probabilities
14
  - classification
15
  - tool-selection
 
16
  - agent-routing
 
 
17
  - decision-index
18
  - jevbench
 
19
  - jev-compatible
 
 
20
  - systemone
 
 
21
  - wald
22
  - wald-q4b
23
  - qwen3.5
24
  - 4b
25
  - vllm
26
+ - reasoning
27
  - gguf
28
  - llama.cpp
 
 
29
  model-index:
30
+ - name: Wald-Q4B v2 (04701-c22, revision v2.0)
31
  results:
32
  - task:
33
  type: text-classification
34
  name: Structured decisions (option probabilities)
35
  dataset:
36
+ name: Decision Index 0.3, complete public suite (140,620 requests)
37
  type: decision-index
38
+ revision: 62d2f51de34a2de64906345b6bc3e98e27ff55c7
39
  metrics:
40
  - type: balanced_skill_index
41
+ name: Public index, Auto 0.7 (author-run with the official kit; not the organizer's Full score)
42
+ value: 54.18
43
  verified: false
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
44
  ---
45
 
46
+ # Wald-Q4B v2
 
 
 
 
 
 
47
 
48
+ An open 4B decision model: send a state and typed options, get a calibrated probability for every option. Served through a Jev-compatible `POST /v1/systemone` API on your own GPU.
49
 
50
+ ## Why Wald-4B v2
 
 
 
51
 
52
+ - **Probabilities, not text:** a calibrated probability for every option, nothing to parse.
53
+ - **Fast:** median **42 ms** one pass, **148 ms** Auto 0.7 (serial, one RTX PRO 6000).
54
+ - **Thinks only when unsure:** about **one third** of requests trigger a short native thought.
55
+ - **JevBench public: 210/231**, against 200/231 for Jev 1.13.
56
+ - **Decision Index 0.3: 54.18** on the complete public suite (our run).
57
+ - **Reads images** zero-shot: **+2.8 points** over its base on CV-Bench, BLINK, RealWorldQA.
58
+ - **Small and open:** 4B, Apache-2.0; GGUF from 2.7 GB in [org2ai/Wald-4B-GGUF](https://huggingface.co/org2ai/Wald-4B-GGUF).
59
 
60
+ ## Scores
61
 
62
+ | Benchmark | Wald-Q4B v2 · one pass | **Wald-Q4B v2 · Auto 0.7** | Jev 1.13 (hosted) |
63
+ |---|---:|---:|---:|
64
+ | Decision Index 0.3, complete public suite | not run | **54.18** (our full run) | 57.96 (board public column) |
65
+ | JevBench public set (231) | 204/231 | **210/231** | 200/231 |
66
+ | JevBench-XL TEST (8,177) | 65.20 % | **65.78 %** | 67.85 % |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
67
 
68
+ Self-run, not leaderboard results; how each number was produced: [evaluation/v2/](https://huggingface.co/org2ai/Wald-4B/tree/main/evaluation/v2).
69
 
70
  ## Quick start
71
 
 
 
72
  ```sh
73
+ hf download org2ai/Wald-4B --revision v2.0 --local-dir ./Wald-Q4B-v2 && cd Wald-Q4B-v2
74
+ ./run.sh "$PWD" # POST /v1/systemone on :8000 (one NVIDIA GPU, uv)
 
 
 
 
 
 
 
 
 
75
 
76
+ curl -s localhost:8000/v1/systemone -H 'Content-Type: application/json' -d '{
77
+ "state": "The customer wants to return a damaged kettle.",
78
+ "questions": {"route": {"type": "choice", "instructions": "Choose the support queue.",
79
+ "criteria": {"returns": "Returns and refunds", "delivery": "Delivery tracking", "other": "Other"}}}}'
80
 
81
+ curl -s localhost:8000/v1/systemone -H 'Content-Type: application/json' -d '{
82
+ "state": "Photo from the returns desk.", "images": ["data:image/jpeg;base64,..."],
83
+ "questions": {"damaged": {"type": "noul", "instructions": "Is the item visibly damaged?"}}}'
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
84
  ```
85
 
86
+ API fields: [docs/api.md](https://huggingface.co/org2ai/Wald-4B/blob/main/docs/api.md). Server options, Docker and evaluation settings: [RUNBOOK.md](https://huggingface.co/org2ai/Wald-4B/blob/main/RUNBOOK.md).
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
87
 
88
  ## How it works
89
 
90
+ - One forward pass reads a probability for every option from the model's option-letter logits (Qwen chat template), calibrated with a frozen temperature table.
91
+ - Auto 0.7 (default): if the top probability is below 0.7, the model thinks once natively (at most 512 tokens) and reads the options again.
92
+ - More than 26 options: a one-pass knockout over all options.
93
+ - Images go through Qwen3.5-4B's vision tower, unchanged (zero-shot).
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
94
 
95
+ ## Details
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
96
 
97
+ - **Training.** Full-parameter decision training from Qwen3.5-4B (chat); checkpoint `04701-c22`.
98
+ - **Data.** About 15% of training tokens come from the train splits of public benchmarks that the Decision Index also draws on, reformatted as decisions. The rest is broad decision data: general decision tasks — rules and policies, classification, entailment and fact checking, tables and documents, web and tool actions, judging (about 40%); math and step-by-step reasoning (about 20%); and hard decisions written and/or labelled by large frontier models, including robustness cases with perturbed or adversarial states (about 25%). Public datasets are used as train splits only. Contamination scan against the full Decision Index 0.3 public suite: 0 strict hits.
99
+ - **Limits.** Scores are self-run; confidence is not correctness, so set thresholds on your own data.
100
+ Answers after a thought are sampled and can vary slightly; one pass is deterministic.
101
+ - **Licence.** Apache-2.0 for weights and code ([NOTICE](https://huggingface.co/org2ai/Wald-4B/blob/main/NOTICE)). Earlier versions: tags [`v1.2`](https://huggingface.co/org2ai/Wald-4B/tree/v1.2), [`v1.1`](https://huggingface.co/org2ai/Wald-4B/tree/v1.1), [`v1.0`](https://huggingface.co/org2ai/Wald-4B/tree/v1.0). Cite: [CITATION.cff](https://huggingface.co/org2ai/Wald-4B/blob/main/CITATION.cff).
RUNBOOK.md CHANGED
@@ -1,66 +1,48 @@
1
- # Runbook
2
 
3
- This revision (`v1.2`) holds **Wald-Q4B v1.2**, checkpoint `02600-f19`. Section 1 serves it and reproduces its JevBench public read. Section 2 is the unchanged v1.1 runbook for the complete Decision Index run: that run belongs to **v1.1** (checkpoint `022D0-f7`), so download `--revision v1.1` for it.
4
 
5
- The server code, prompt format, tokenizer and temperature table are byte-identical in v1.1 and v1.2. The two weight shards differ, and `serving.json` declares effort `none` in v1.2 (`high` in v1.1). MANIFEST.json lists the exact model, tokenizer, temperature and code hashes of this revision.
6
-
7
- ## 1. Wald-Q4B v1.2
8
 
9
  ```sh
10
- hf download org2ai/Wald-4B --revision v1.2 --local-dir ./Wald-Q4B-v1.2
11
- cd Wald-Q4B-v1.2
12
- ./run.sh "$PWD" # effort none (declared in serving.json), repeat_state_plain, context 131072
 
 
 
 
13
  ```
14
 
15
- `GET /health` must report `"effort": "none"`. v1.2's JevBench and JevAdvBench results were measured with effort `none` and `repeat_state_plain` on one NVIDIA RTX 5090 32 GB (vLLM 0.30.0, BF16); the JevAdvBench read used a 32,768-token context limit.
16
 
17
- JevBench public set (204/231), with `fstandhartinger/jevbench` at `9ec6f15a`:
18
 
19
- ```sh
20
- cd jevbench # a checkout of fstandhartinger/jevbench at 9ec6f15a
21
- for T in easy original hard; do
22
- python -m jevbench.cli run --tasks datasets/public/$T.jsonl --adapter typesafe --endpoint http://127.0.0.1:8000 \
23
- --model jev-latest --key-env '' --reserve-usd 0 --cost-basis self_hosted_loopback_no_tariff \
24
- --results out/$T/results.jsonl --raw-dir out/raw-$T --ledger out/$T/ledger.jsonl --manifest out/$T/manifest.json \
25
- --run-label wald-q4b-v1.2-pub-$T
26
- done
27
- cat out/easy/results.jsonl out/original/results.jsonl out/hard/results.jsonl > out/results-all.jsonl
28
- python -m jevbench.cli summarize --tasks datasets/public/easy.jsonl,datasets/public/original.jsonl,datasets/public/hard.jsonl \
29
- --results out/results-all.jsonl --public-export out/summary-all.json
30
- ```
31
 
32
- JevAdvBench: send the benchmark's requests (`JevAdvBench/JevAdvBench` at `3218e05`, one question per request) to the same endpoint and score the answers with the benchmark's own analysis code. The benchmark data is CC BY-NC 4.0 and is not redistributed here.
33
 
34
- ## 2. Wald-Q4B v1.1: complete Decision Index run
 
 
 
35
 
36
- ### Reproduce Wald-Q4B 22D0-f7
37
 
38
- Release: **v1.1** · checkpoint `022D0-f7`. Previous release: **v1.0**.
39
 
40
- Download the immutable HF tag 22D0-f7 (or the full HF commit in the submission) into a fresh directory. Do not mix old v1.0 single-file weights with the new shards. MANIFEST.json lists the exact model, tokenizer, temperature and code hashes.
41
 
42
- ### Convenient packaged server
43
 
44
- ```sh
45
- hf download org2ai/Wald-4B --revision v1.1 --local-dir ./Wald-Q4B-22D
46
- cd Wald-Q4B-22D
47
- ./run.sh "$PWD"
48
- ```
49
-
50
- Default policy is high, repeat_state_plain, context 131072. Lower effort modes are alternative configurations with no 54.59 claim. GPU requirements: vLLM 0.30.0, BF16, NVIDIA RTX PRO 6000 96GB for comparison. The packaged server has parity tests against the reference below for prompts and answer probabilities. Scheduling may differ; latency is not established by those tests.
51
-
52
- ### Frozen reference protocol
53
 
54
- The full run used eval.systemone_vllm from the private development checkout; that checkout was not committed at launch. The included reference/ files freeze the release-time implementation, with hashes in MANIFEST.json. This code provenance limitation is disclosed rather than presenting a later public commit as the original launch commit.
55
 
56
- ```sh
57
- python -m vllm.entrypoints.openai.api_server --model "$PWD" --served-model-name Wald-Q4B-022D0-f7-full021 --host 127.0.0.1 --port 8321 --max-model-len 131072 --gpu-memory-utilization 0.60 --max-num-seqs 128 --seed 0
58
- # Separate terminal, same environment:
59
- PYTHONPATH="$PWD/reference" python -m eval.systemone_vllm --vllm http://127.0.0.1:8321 --served Wald-Q4B-022D0-f7-full021 --gate 1.01 --budget 512 --temperature "$PWD/temperature.json" --wide knockout --template paren --max-model-len 131072 --return-raw --prompt-format repeat_state_plain --port 8421 --model-name Wald-Q4B-022D0-f7-full021
60
- ```
61
 
62
- Install the pinned kit `apolinario/decision-index@87d4650b42b377c0291a89c1f1a879f9b31082bf`, rebuild and verify its complete suite locally, then run its http engine against port 8421. Suite payloads are not redistributed. No request or option pruning/truncation. Complete saved responses are untouched and suitable for maintainer rescoring.
63
 
64
- Capacity: 131,072 tokens. Wide questions use ordered knockout over all supplied options. A thought that cannot fit falls back to the one-pass read. The full run had 0 unsupported and 0 errors. Concurrency ranged 16–40 request runners for throughput; timing is not the maintainer's serial admission gate. The 32-request preflight median 821.3 ms does not replace 750 private serial requests after warm-up.
65
 
66
- Training/data/calibration limitations are in CONTAMINATION.md and PROVENANCE.md. Official admission remains pending.
 
1
+ # Runbook: Wald-Q4B v2 (`v2.0`, checkpoint `04701-c22`)
2
 
3
+ This revision holds **Wald-Q4B v2**: the language model of checkpoint `04701-c22` (its text-only weights file has sha256 `6ae382f5a0ed9f4c53cd0953cb9606b1627a6816350340a59114860e6cf6f29a`; every tensor is byte-identical here) plus Qwen3.5-4B's unchanged vision tower, MTP head and image processor configs, as `Qwen3_5ForConditionalGeneration` in three safetensors shards. MANIFEST.json lists the sha256 of every file. v1.x runbooks are at their tags.
4
 
5
+ ## 1. Serve
 
 
6
 
7
  ```sh
8
+ hf download org2ai/Wald-4B --revision v2.0 --local-dir ./Wald-Q4B-v2
9
+ cd Wald-Q4B-v2
10
+ python - <<'PY' # optional: check every file against MANIFEST.json
11
+ import hashlib, json; m = json.load(open("MANIFEST.json"))
12
+ bad = [f for f, v in m.items() if hashlib.sha256(open(f, "rb").read()).hexdigest() != v["sha256"]]; print("bad:", bad)
13
+ PY
14
+ ./run.sh "$PWD"
15
  ```
16
 
17
+ `GET /health` must report `"engine": "native-v2"`, `"vision": true`, `"checkpoint": "04701-c22"`, `"policy": "Auto0.7"`, `"thought_budget": 512` and `"temperature_sha256": "a0f72cd2d0a653e81051e5a0c77fc1a69131552a8102b580a93a6dbe7908b2da"`. Docker: see `Dockerfile` (same runtime). The server pins the reader (`server/src/wald_serve/native_reference/`, every file sha256-checked at start) and the calibration table, and refuses to start if either differs.
18
 
19
+ **The evaluated runtime:** vLLM 0.30.0, transformers 5.17.0, torch 2.13.0, BF16, one NVIDIA RTX PRO 6000 (96 GB) per model replica, `VLLM_USE_FLASHINFER_SAMPLER=0`, and the vLLM arguments `--max-model-len 131072 --gpu-memory-utilization 0.85 --max-num-seqs 128 --seed 0` plus `--limit-mm-per-prompt {"image": 16, "video": 0}` and the image budget `--mm-processor-kwargs {"size": {"shortest_edge": 65536, "longest_edge": 1048576}}` (what `wald-serve-native-vision --model` launches; `wald-serve-native` serves text only). Smaller GPUs need a lower `--gpu-memory-utilization` or `--max-model-len`, which is a deviation from the evaluated setup. Thought generation samples (temperature 0.6, top-p 0.95, top-k 20) with a seed derived from the request and question id, so repeated requests normally give the same answer; batch composition on the GPU can still change near-tied answers.
20
 
21
+ `config.json` carries `"use_cache": false` from training. vLLM ignores it; if you load the weights with Transformers `generate`, pass `use_cache=True`.
 
 
 
 
 
 
 
 
 
 
 
22
 
23
+ ## 2. Policies and prompt formats
24
 
25
+ | Request field | Values | Evaluated settings |
26
+ |---|---|---|
27
+ | `effort` | `none` (one pass), `auto` / `medium` (Auto 0.7, default), `always` / `high` (Always 512) | all three |
28
+ | `prompt_format` | `repeat_state_plain` (default), `plain` | Decision Index and JevBench-XL: `repeat_state_plain`; JevBench public 231 and multistep: `plain` |
29
 
30
+ Auto 0.7 gates on the untempered one-pass distribution: a question with 2–26 options thinks natively (Qwen chat template with `enable_thinking`, at most 512 tokens, stop at `</think>`) when its top probability is below 0.7, then the options are read again. More than 26 options: grouped knockout readout, no thought, every option gets a probability. The returned probabilities are tempered with the frozen table (bucket by question type and option count).
31
 
32
+ ## 3. Images
33
 
34
+ Requests with images (message content parts `image_url` / `image` in `state`, raw base64 under `state` keys `image`, `image_data`, `image_base64`, `image_b64`, or a top-level `images` / `image_data` list; up to 16 per request, bodies up to 64 MB) go through the same frozen reader, with the images carried to vLLM's chat route at their position in the state. They default to the `plain` prompt format. Text-only requests take the text path unchanged. The Decision Index 0.3 run used the text-only weights file; the language-model tensors here are byte-identical, and the text parity of this repository's weights is in `evaluation/v2/serving-parity.json`.
35
 
36
+ ## 4. GGUF (llama.cpp)
37
 
38
+ `org2ai/Wald-4B-GGUF` holds the v2 quantisations and the vision projector (`Wald-4B-v2-mmproj-F16.gguf`). `wald-serve-native --gguf FILE --tokenizer-dir DIR --effort none` reads text requests through llama-server with the same chat layout; a letter outside llama-server's top-100 log-probabilities gets the reader's floor instead of its exact value. One pass is the checked mode.
 
 
 
 
 
 
 
 
39
 
40
+ ## 5. Decision Index 0.3 public suite
41
 
42
+ Install the official kit `apolinario/decision-index` at `62d2f51de34a2de64906345b6bc3e98e27ff55c7` (branch `v0.3`), rebuild the public suite and verify its hashes with the kit. The evaluated run used the kit's Engine API with the adapter in `evaluation/v2/code/di03_native_auto_engine.py` (it loads `reference/` and `temperature.json`, Auto 0.7, budget 512, `repeat_state_plain`) against vLLM servers started with the arguments above; `evaluation/v2/code/di03_parallel.py` ran four disjoint replicas (SHA256(run_id) mod 4) at 128 concurrent requests each. The packaged server answers the same requests with the same reader; the kit's `http` engine can be pointed at `POST /v1/systemone` instead. Score with the kit's unmodified scorer after inference. Suite payloads are not redistributed. Concurrent throughput is not the maintainers' serial latency.
 
 
 
 
43
 
44
+ ## 6. Other reads
45
 
46
+ JevBench public 231, JevBench-XL partitions and multistep_decisions were read with the same reader and engine (one generation per item; the one-pass, Auto 0.7 and Always 512 answers are fixed views of it) and scored with each benchmark's scorer. JevBench-XL is internal and not distributed.
47
 
48
+ Training data and evaluation caveats: PROVENANCE.md, CONTAMINATION.md.
SHA256SUMS ADDED
@@ -0,0 +1,15 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ a4aee8afcf2e0711942cf848899be66016f8d14a889ff9ede07bca099c28f715 chat_template.jinja
2
+ ddc63e1c717afa86c865bb5e01313d89d72bb53b97ad4a8a03ba8510c0621670 config.json
3
+ 62153eb6c69f2e1f426beaa8002b7186437e949c7588167085df14e10e9c0a73 generation_config.json
4
+ e33ce57511cb924d2bd5559a143c37cf12f1df1e6de911f0a7d4fdbbc0b85eb8 graft-manifest.json
5
+ d4ae76c16d404c665f84e10079f460221f4c1434c23e0b923bf9a6deb91c1385 graft-report.json
6
+ a9d356d7bdf1ef4949e3e748e95b8e10ad9d4e2e838eddc38a0a7b6b94d1db8d merges.txt
7
+ 10cb7cd22f524f1b33271f8432bcb09409bd879c11d999ea18ca7dd5f8828f5a model-00001-of-00003.safetensors
8
+ ed5e7bd4b2e67eee390857268a7660986aca32068c78e6f483e866c27f90b79b model-00002-of-00003.safetensors
9
+ 9a8340e453e63f70185e56fde60aeae26085fe3c39e20ac1662b4325f01a78eb model-00003-of-00003.safetensors
10
+ 6db63a968a513eb0693feee5ccbc440cbbee43c6b644c0bd33e880ca7d0b54b3 model.safetensors.index.json
11
+ 27225450ac9c6529872ee1924fcb0962ff5634834f817040f444118116f4e516 preprocessor_config.json
12
+ bd53432f0de26d67a83b634040d4f043053da4ecd0c759e7f4b24ad4f8bb9a81 tokenizer.json
13
+ bee8eba30f0eb4af73c0fe2cd06d0f89b657d7819941c438157ec42f7c80ea87 tokenizer_config.json
14
+ 7768af27c1fafa9cc9011c1dc20067e03f8915e03b63504550e11d5066986d13 video_preprocessor_config.json
15
+ ce99b4cb2983d118806ce0a8b777a35b093e2000a503ebde25853284c9dfa003 vocab.json
Wald-4B-v1.2-Q8_0.gguf DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:09c02494ebdfb7d088a2c76726eb2b290a438dbc5658c3be16da8333fddd6038
3
- size 4482402720
 
 
 
 
config.json CHANGED
@@ -1,83 +1,104 @@
1
  {
2
- "architectures": [
3
- "Qwen3_5ForCausalLM"
4
- ],
5
- "attention_bias": false,
6
- "attention_dropout": 0.0,
7
- "attn_output_gate": true,
8
- "bos_token_id": null,
9
- "dtype": "bfloat16",
10
- "eos_token_id": 248044,
11
- "full_attention_interval": 4,
12
- "head_dim": 256,
13
- "hidden_act": "silu",
14
- "hidden_size": 2560,
15
- "initializer_range": 0.02,
16
- "intermediate_size": 9216,
17
- "layer_types": [
18
- "linear_attention",
19
- "linear_attention",
20
- "linear_attention",
21
- "full_attention",
22
- "linear_attention",
23
- "linear_attention",
24
- "linear_attention",
25
- "full_attention",
26
- "linear_attention",
27
- "linear_attention",
28
- "linear_attention",
29
- "full_attention",
30
- "linear_attention",
31
- "linear_attention",
32
- "linear_attention",
33
- "full_attention",
34
- "linear_attention",
35
- "linear_attention",
36
- "linear_attention",
37
- "full_attention",
38
- "linear_attention",
39
- "linear_attention",
40
- "linear_attention",
41
- "full_attention",
42
- "linear_attention",
43
- "linear_attention",
44
- "linear_attention",
45
- "full_attention",
46
- "linear_attention",
47
- "linear_attention",
48
- "linear_attention",
49
- "full_attention"
50
- ],
51
- "linear_conv_kernel_dim": 4,
52
- "linear_key_head_dim": 128,
53
- "linear_num_key_heads": 16,
54
- "linear_num_value_heads": 32,
55
- "linear_value_head_dim": 128,
56
- "mamba_ssm_dtype": "float32",
57
- "max_position_embeddings": 262144,
58
- "mlp_only_layers": [],
59
- "model_type": "qwen3_5_text",
60
- "mtp_num_hidden_layers": 1,
61
- "mtp_use_dedicated_embeddings": false,
62
- "num_attention_heads": 16,
63
- "num_hidden_layers": 32,
64
- "num_key_value_heads": 4,
65
- "pad_token_id": null,
66
- "partial_rotary_factor": 0.25,
67
- "rms_norm_eps": 1e-06,
68
- "rope_parameters": {
69
- "mrope_interleaved": true,
70
- "mrope_section": [
71
- 11,
72
- 11,
73
- 10
74
  ],
75
- "partial_rotary_factor": 0.25,
76
- "rope_theta": 10000000,
77
- "rope_type": "default"
78
- },
79
- "tie_word_embeddings": true,
80
- "transformers_version": "5.17.0",
81
- "use_cache": true,
82
- "vocab_size": 248320
83
- }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  {
2
+ "architectures": [
3
+ "Qwen3_5ForConditionalGeneration"
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
4
  ],
5
+ "image_token_id": 248056,
6
+ "model_type": "qwen3_5",
7
+ "text_config": {
8
+ "attention_bias": false,
9
+ "attention_dropout": 0.0,
10
+ "attn_output_gate": true,
11
+ "dtype": "bfloat16",
12
+ "eos_token_id": 248044,
13
+ "full_attention_interval": 4,
14
+ "head_dim": 256,
15
+ "hidden_act": "silu",
16
+ "hidden_size": 2560,
17
+ "initializer_range": 0.02,
18
+ "intermediate_size": 9216,
19
+ "layer_types": [
20
+ "linear_attention",
21
+ "linear_attention",
22
+ "linear_attention",
23
+ "full_attention",
24
+ "linear_attention",
25
+ "linear_attention",
26
+ "linear_attention",
27
+ "full_attention",
28
+ "linear_attention",
29
+ "linear_attention",
30
+ "linear_attention",
31
+ "full_attention",
32
+ "linear_attention",
33
+ "linear_attention",
34
+ "linear_attention",
35
+ "full_attention",
36
+ "linear_attention",
37
+ "linear_attention",
38
+ "linear_attention",
39
+ "full_attention",
40
+ "linear_attention",
41
+ "linear_attention",
42
+ "linear_attention",
43
+ "full_attention",
44
+ "linear_attention",
45
+ "linear_attention",
46
+ "linear_attention",
47
+ "full_attention",
48
+ "linear_attention",
49
+ "linear_attention",
50
+ "linear_attention",
51
+ "full_attention"
52
+ ],
53
+ "linear_conv_kernel_dim": 4,
54
+ "linear_key_head_dim": 128,
55
+ "linear_num_key_heads": 16,
56
+ "linear_num_value_heads": 32,
57
+ "linear_value_head_dim": 128,
58
+ "max_position_embeddings": 262144,
59
+ "mlp_only_layers": [],
60
+ "model_type": "qwen3_5_text",
61
+ "mtp_num_hidden_layers": 1,
62
+ "mtp_use_dedicated_embeddings": false,
63
+ "num_attention_heads": 16,
64
+ "num_hidden_layers": 32,
65
+ "num_key_value_heads": 4,
66
+ "rms_norm_eps": 1e-06,
67
+ "tie_word_embeddings": true,
68
+ "use_cache": true,
69
+ "vocab_size": 248320,
70
+ "mamba_ssm_dtype": "float32",
71
+ "rope_parameters": {
72
+ "mrope_interleaved": true,
73
+ "mrope_section": [
74
+ 11,
75
+ 11,
76
+ 10
77
+ ],
78
+ "rope_type": "default",
79
+ "rope_theta": 10000000,
80
+ "partial_rotary_factor": 0.25
81
+ }
82
+ },
83
+ "tie_word_embeddings": true,
84
+ "transformers_version": "4.57.0.dev0",
85
+ "video_token_id": 248057,
86
+ "vision_config": {
87
+ "deepstack_visual_indexes": [],
88
+ "depth": 24,
89
+ "hidden_act": "gelu_pytorch_tanh",
90
+ "hidden_size": 1024,
91
+ "in_channels": 3,
92
+ "initializer_range": 0.02,
93
+ "intermediate_size": 4096,
94
+ "model_type": "qwen3_5",
95
+ "num_heads": 16,
96
+ "num_position_embeddings": 2304,
97
+ "out_hidden_size": 2560,
98
+ "patch_size": 16,
99
+ "spatial_merge_size": 2,
100
+ "temporal_patch_size": 2
101
+ },
102
+ "vision_end_token_id": 248054,
103
+ "vision_start_token_id": 248053
104
+ }
contamination/trained-on-di-ids.json CHANGED
The diff for this file is too large to render. See raw diff
 
docs/api.md CHANGED
@@ -1,148 +1,64 @@
1
- # Wald-Q4B decision API
2
 
3
- Wald-Q4B's server, `wald-serve`, exposes one decision endpoint, `POST /v1/systemone`. Its request and answer shapes follow TypeSafe's `/v1/systemone` format, so clients written for Jev can point at a self-hosted Wald server. Wald is independent and is not affiliated with TypeSafe AI.
4
-
5
- Start the server with `./run.sh "$PWD"` in the downloaded model directory ([model card](https://huggingface.co/org2ai/Wald-4B), [RUNBOOK.md](https://huggingface.co/org2ai/Wald-4B/blob/main/RUNBOOK.md)). It listens on port 8000 and does not check API keys.
6
 
7
  ## Endpoints
8
 
9
  | Method and path | Purpose |
10
  |---|---|
11
  | `POST /v1/systemone` (also `POST /`) | Answer one or more typed questions about a state |
12
- | `GET /health` | Effective policy: effort, gate, thought budget, prompt format, context limit |
13
- | `GET /v1/models` | The served model name |
14
 
15
  ## Request
16
 
17
  | Field | Type | Meaning |
18
  |---|---|---|
19
  | `state` | string, object, array or null | What the decision is about: a message, a conversation, a document, an agent trace. Objects and arrays are flattened to text with their field names kept. |
20
- | `questions` | object, at least one entry | Question id → question. Every question is answered about the same `state`. |
21
- | `effort` | string, optional | `none`, `low`, `medium`, `high` or `high-k2` … `high-k8`. Overrides the server default for this request. |
22
- | `model` | string, optional | Accepted for client compatibility; the server answers with the model it serves. |
 
 
23
 
24
- Each question has a `type`, optional `instructions` (any JSON, usually a sentence) and `criteria`:
25
 
26
- | `type` | `criteria` | Answer |
27
  |---|---|---|
28
- | `choice` | Object of 1–255 options: key → description (description may be null) | `choice` (the most probable key), `probabilities` (key → probability), `confidence` |
29
- | `noul` | Optional object with `true` and/or `false` descriptions | `noul`: the probability of yes |
30
- | `score` | Array of 1–255 ordered levels | `score` (expected level index), `probabilities` (level index → probability), `confidence` |
31
 
32
- Every answer also carries `mode`: `A` for a one-pass read, `B` for a read after a thought, `K` for grouped (knockout) reading of wide option sets.
33
 
34
- Prompts longer than the context limit (131,072 tokens by default) are rejected with HTTP 422, never truncated. A malformed request or unknown effort returns HTTP 400.
35
 
36
  ## Example: route a tool call
37
 
38
  ```sh
39
- curl http://localhost:8000/v1/systemone \
40
- -H 'Content-Type: application/json' \
41
- -d '{
42
- "state": {
43
- "user": "What will the weather be in Lisbon tomorrow afternoon?",
44
- "tools_available": ["web_search", "weather_api", "calendar"]
45
- },
46
- "effort": "none",
47
- "questions": {
48
- "tool": {
49
- "type": "choice",
50
- "instructions": "Which tool should the agent call next?",
51
- "criteria": {
52
- "web_search": "General web search",
53
- "weather_api": "Forecast for a city and time",
54
- "calendar": "Read or create calendar events",
55
- "none": "Answer directly without a tool"
56
- }
57
- }
58
- }
59
- }'
60
- ```
61
-
62
- Response shape (the numbers here are illustrative, not a measured output):
63
-
64
- ```json
65
- {
66
- "model": "wald-4b",
67
- "answers": {
68
- "tool": {
69
- "type": "choice",
70
- "choice": "weather_api",
71
- "probabilities": {"web_search": 0.04, "weather_api": 0.94, "calendar": 0.0, "none": 0.02},
72
- "confidence": 0.92,
73
- "mode": "A"
74
- }
75
- },
76
- "usage": {"input_tokens": 142, "output_tokens": 0},
77
- "latency_s": 0.03
78
- }
79
- ```
80
-
81
- For `choice`, `confidence` rescales the top probability so that 0 means uniform and 1 means certain: `(p_top − 1/K) / (1 − 1/K)` for K options.
82
-
83
- ## Example: decide whether to ask the user
84
-
85
- Ask several questions about one state in a single request. Each question is read separately.
86
-
87
- ```json
88
- {
89
- "state": "User: book me a table for Friday",
90
- "effort": "medium",
91
  "questions": {
92
- "specific_enough": {
93
- "type": "noul",
94
- "instructions": "Is the request specific enough to act on without asking a follow-up question?",
95
- "criteria": {"true": "Enough detail to act", "false": "Needs clarification first"}
96
- },
97
- "urgency": {
98
- "type": "score",
99
- "instructions": "How urgent is this request?",
100
- "criteria": ["low", "medium", "high"]
101
- }
102
  }
103
- }
104
  ```
105
 
106
- Response shape (illustrative numbers):
107
 
108
  ```json
109
- {
110
- "model": "wald-4b",
111
- "answers": {
112
- "specific_enough": {"type": "noul", "noul": 0.08, "mode": "B"},
113
- "urgency": {"type": "score", "score": 1.1, "probabilities": {"0": 0.15, "1": 0.6, "2": 0.25}, "confidence": 0.6, "mode": "A"}
114
- },
115
- "usage": {"input_tokens": 310, "output_tokens": 96},
116
- "latency_s": 0.6
117
- }
118
- ```
119
-
120
- A typical policy: act when `noul` is above a high threshold, ask the user when it is below a low one, and escalate the cases in between. Set both thresholds on validation data from your own task.
121
-
122
- ## Choosing an effort
123
-
124
- | Effort | Use it for |
125
- |---|---|
126
- | `none` | Lowest latency. One pass, no generated tokens. |
127
- | `low` / `medium` | Think only when the top initial probability is below 0.5 / 0.7. |
128
- | `high` | Think on every eligible question. The published Decision Index score uses this setting. |
129
- | `high-k2` … `high-k8` | Several thoughts, averaged. Slowest. |
130
-
131
- Thinking applies to questions with 2–26 options when the context has room; otherwise the one-pass answer is returned.
132
-
133
- ## Python client
134
-
135
- ```python
136
- import requests
137
-
138
- r = requests.post("http://localhost:8000/v1/systemone", json={
139
- "state": "Refund request for order 1182; the parcel arrived damaged.",
140
- "effort": "none",
141
- "questions": {"route": {"type": "choice", "criteria": {
142
- "refunds": "Refunds and returns", "shipping": "Delivery problems", "other": "Anything else"}}},
143
- }, timeout=30)
144
- answer = r.json()["answers"]["route"]
145
- print(answer["choice"], answer["probabilities"])
146
  ```
147
 
148
- Source: [`server/src/wald_serve/wire.py`](https://huggingface.co/org2ai/Wald-4B/blob/main/server/src/wald_serve/wire.py) (request shapes) and [`server.py`](https://huggingface.co/org2ai/Wald-4B/blob/main/server/src/wald_serve/server.py) (endpoints).
 
1
+ # Wald-Q4B v2 decision API
2
 
3
+ `wald-serve-native` (in `server/`, started by `./run.sh`) exposes one decision endpoint, `POST /v1/systemone`. Request shapes follow TypeSafe's `/v1/systemone` format, so clients written for Jev can point at a self-hosted Wald server. Wald is independent and is not affiliated with TypeSafe AI. The server does not check API keys; put it behind your own gateway.
 
 
4
 
5
  ## Endpoints
6
 
7
  | Method and path | Purpose |
8
  |---|---|
9
  | `POST /v1/systemone` (also `POST /`) | Answer one or more typed questions about a state |
10
+ | `GET /health` | Checkpoint, weights sha256, default effort / policy, thought budget, prompt format, context limit, calibration sha256 |
11
+ | `GET /v1/models` | The served model name (`04701-c22`) |
12
 
13
  ## Request
14
 
15
  | Field | Type | Meaning |
16
  |---|---|---|
17
  | `state` | string, object, array or null | What the decision is about: a message, a conversation, a document, an agent trace. Objects and arrays are flattened to text with their field names kept. |
18
+ | `questions` | object, at least one entry | Question id → question, all about the same `state`. |
19
+ | `effort` | string, optional | `none` (one pass), `auto` or `medium` (Auto 0.7, the default), `always` or `high` (Always 512). Other values (`low`, `high-k2` …) are v1.x efforts and return HTTP 400. |
20
+ | `prompt_format` | string, optional | `repeat_state_plain` (default: the state is written twice) or `plain`. |
21
+ | `images` | array of data URIs, optional | Images for the state (also accepted inside `state` as message content parts `{"type": "image_url", "image_url": {"url": "data:..."}}`); up to 16; zero-shot. |
22
+ | `model` | string, optional | Ignored; accepted for client compatibility. |
23
 
24
+ Only `state`, `questions` and the images reach the model. Each question has a `type`, optional `instructions` and `criteria`:
25
 
26
+ | `type` | `criteria` | Answer fields |
27
  |---|---|---|
28
+ | `choice` | Object of options: key → description (description may be null) | `choice` (most probable key), `probabilities` (key → probability), `confidence` |
29
+ | `noul` | Optional object with `true` and/or `false` descriptions | `noul` (probability of yes), `probabilities` (`true` / `false`), `confidence` |
30
+ | `score` | Array of ordered levels | `score` (expected level index), `probabilities` (level index → probability), `confidence` |
31
 
32
+ Every answer also carries `type`, `mode` (`A` one pass, `B` after a native thought, `K`/`T` grouped readout for more than 26 options) and `probabilities_t1` (the same distribution before the calibration table). `confidence` is the largest calibrated probability. The response has `model`, `answers` and `usage` (`input_tokens`, `output_tokens`; output tokens are thought tokens). The thought text itself is not returned.
33
 
34
+ Prompts longer than the context limit (131,072 tokens) are rejected with HTTP 422, never truncated. A malformed request or an unknown effort / prompt format returns HTTP 400; a backend failure returns 502.
35
 
36
  ## Example: route a tool call
37
 
38
  ```sh
39
+ curl http://localhost:8000/v1/systemone -H 'Content-Type: application/json' -d '{
40
+ "state": "User: book me a table for two at 7pm tomorrow near the office",
41
+ "effort": "auto",
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
42
  "questions": {
43
+ "tool": {"type": "choice", "instructions": "Which tool should the agent call next?",
44
+ "criteria": {"restaurant_search": "Search restaurants", "calendar_create": "Create a calendar event",
45
+ "ask_user": "Ask the user a clarifying question"}},
46
+ "enough_info": {"type": "noul", "instructions": "Is the request specific enough to act on without asking?"}
 
 
 
 
 
 
47
  }
48
+ }'
49
  ```
50
 
51
+ Response shape (values illustrative):
52
 
53
  ```json
54
+ {"model": "04701-c22",
55
+ "answers": {
56
+ "tool": {"type": "choice", "mode": "A", "choice": "restaurant_search",
57
+ "probabilities": {"restaurant_search": 0.81, "calendar_create": 0.05, "ask_user": 0.14},
58
+ "probabilities_t1": {"restaurant_search": 0.93, "calendar_create": 0.01, "ask_user": 0.06}, "confidence": 0.81},
59
+ "enough_info": {"type": "noul", "mode": "B", "noul": 0.64, "probabilities": {"true": 0.64, "false": 0.36},
60
+ "probabilities_t1": {"true": 0.71, "false": 0.29}, "confidence": 0.64}},
61
+ "usage": {"input_tokens": 1450, "output_tokens": 212}}
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
62
  ```
63
 
64
+ Act on a decision when its probability clears a threshold you set on your own validation data; otherwise escalate or ask. The model picks options; it does not write tool arguments.
docs/readmes/README.zh.md DELETED
@@ -1,278 +0,0 @@
1
- <div align="center">
2
- <h1>Wald-Q4B</h1>
3
- <p><strong>直接决策,按需思考,返回概率。</strong></p>
4
- <p><a href="https://huggingface.co/org2ai/Wald-4B">English</a> · <a href="https://huggingface.co/org2ai/Wald-4B/blob/main/docs/readmes/README.zh.md">简体中文</a> · <a href="https://github.com/org2AI/wald-4b">GitHub</a> · <a href="https://huggingface.co/org2ai/Wald-4B/blob/main/docs/api.md">API</a></p>
5
- </div>
6
-
7
- **Wald-Q4B 是一个开放权重的 4B 决策模型:给它一段状态和一组选项,它为每个选项返回校准过的概率。** 它面向构建 agent 和数据流水线的开发者:需要一个快速、可自行部署的组件来选择工具、路由请求、分类输入,或判断是否需要向用户澄清。和聊天模型不同,它不生成需要再解析的答案,而是一遍读出所有选项的概率,也可以先思考再回答。它提供与 Jev 兼容的 `POST /v1/systemone` API,基于 Qwen3.5-4B-Base,以 Apache-2.0 发布。
8
-
9
- **这个仓库里有两个版本。**
10
-
11
- | 版本 | Revision | Checkpoint | 一句话说明 |
12
- |---|---|---|---|
13
- | **v1.2**(2026-10-01) | tag [`v1.2`](https://huggingface.co/org2ai/Wald-4B/tree/v1.2);`main` 上的权重也是它 | `02600-f19` | 抗干扰版:在 v1.1 上再合并一个 LoRA 阶段,让模型在输入里混有无关句子、旁人意见或伪造指令时保持原来的答案 |
14
- | **v1.1**(2026-09-29) | tag [`v1.1`](https://huggingface.co/org2ai/Wald-4B/tree/v1.1) | `022D0-f7` | 通用版:完整 Decision Index 结果和思考档的评测都在这个版本上 |
15
-
16
- 从 2026-10-01 起,`main` 上是 v1.2 的权重(此前是 v1.1)。v1.1 正在排队的基准提交钉在本仓库的固定 commit 上,不受影响;要用 v1.1 请下载 `--revision v1.1`。下载时请指定 revision。
17
-
18
- 本仓库 2026-10-01 从 `Harry19081/Wald-4B` 迁到 `org2ai/Wald-4B`,旧链接会自动跳转到这里。各发布 tag(`v1.0`、`v1.1`、`v1.2`)上的模型卡不可修改,仍写旧路径,旧路径同样会跳转。
19
-
20
- Wald-Q4B 是 TypeSafe 托管 Jev API 之外、可自行部署的独立替代方案。它不是 Jev,不含 Jev 权重,与 TypeSafe AI 没有隶属或背书关系。Hugging Face 仓库为 `org2ai/Wald-4B`(旧名 Wald-4B;2026-10-01 从 `Harry19081/Wald-4B` 迁来,旧链接自动跳转)。
21
-
22
- **W**ait **A** bit, **L**ook, then **D**ecide:稍等一下,看清楚,再决定。名字也致敬序贯分析先驱 Abraham Wald:证据足够时就停止。
23
-
24
- ## 一览
25
-
26
- - **4B 参数**,基于 Qwen3.5-4B-Base,BF16 权重(8.4 GB)。
27
- - **每个选项都有概率。** 题型:`choice`(1–255 个命名选项)、`noul`(是/否)、`score`(有序等级)。
28
- - **v1.2 比 v1.1 更抗干扰**:在 JevAdvBench 上,九类攻击的平均翻转率为 **4.6%**(v1.1:9.2%;Jev 1.13:6.1%)。自测,effort `none`。
29
- - **v1.2 的代价**:JevAdvBench 143 道人工复核题的干净准确率为 76.2%(v1.1:79.0%;−2.8 个百分点,95% 置信区间 [−6.2, −0.6])。
30
- - **JevBench 公开集**:v1.2 **204/231**,v1.1 **203/231**,都使用 `none`。用 JevBench 自己的评测工具自评。
31
- - **Decision Index 0.2.1**:完整套件的 **54.59**(`high`)是在 **v1.1** 上测的。v1.2 只有一遍读出的抽样结果,与 v1.1 持平(+0.27,不显著)。
32
- - **可调思考程度**:`none`、`low`、`medium`、`high`,以及多次思考的 `high-k`。思考档的评测在 v1.1 上。v1.2 是一遍读出的模型,请用 `none`。
33
- - **可自行部署的 API**:`POST /v1/systemone`,最多 131,072 个提示 token。
34
-
35
- ## 该用哪个版本?
36
-
37
- | 场景 | Revision | 原因 |
38
- |---|---|---|
39
- | 输入里可能混有干扰、劝说或对抗性文本(网页、用户消息、工具输出、检索到的文档) | **`v1.2`** | 在 JevAdvBench 的九类攻击下,答案被改变的次数大约减半 |
40
- | 输入干净可信、看重最后几个点的准确率;需要思考档;需要评测过的 Decision Index 配置 | **`v1.1`** | 干净准确率略高;完整 Decision Index(54.59,`high`)和思考档都是在它上面测的 |
41
-
42
- 两个版本的结构、分词器、提示格式、校准表和服务代码完全相同。不同的是两个权重分片,以及 `serving.json` 声明的默认 effort:v1.2 为 `none`,v1.1 为 `high`。
43
-
44
- ## 快速开始
45
-
46
- 在装有 NVIDIA GPU 和 [`uv`](https://docs.astral.sh/uv/) 的 Linux 机器上:
47
-
48
- ```sh
49
- hf download org2ai/Wald-4B --revision v1.2 --local-dir ./Wald-Q4B # v1.2,抗干扰版
50
- # hf download org2ai/Wald-4B --revision v1.1 --local-dir ./Wald-Q4B # v1.1,通用版
51
- cd Wald-Q4B
52
- EFFORT=none ./run.sh "$PWD" # 一遍读出,延迟最低
53
- # ./run.sh "$PWD" # 该版本声明的默认值:v1.2 为 none,v1.1 为 high
54
- ```
55
-
56
- 在 Python 里同样指定 revision:
57
-
58
- ```python
59
- from huggingface_hub import snapshot_download
60
-
61
- path = snapshot_download("org2ai/Wald-4B", revision="v1.2") # 或 revision="v1.1"
62
- ```
63
-
64
- ```sh
65
- curl http://localhost:8000/v1/systemone \
66
- -H 'Content-Type: application/json' \
67
- -d '{
68
- "state": "The customer wants to return a damaged kettle.",
69
- "effort": "none",
70
- "questions": {
71
- "route": {
72
- "type": "choice",
73
- "instructions": "Choose the support queue.",
74
- "criteria": {
75
- "returns": "Returns and refunds",
76
- "delivery": "Delivery tracking",
77
- "other": "Other enquiries"
78
- }
79
- }
80
- }
81
- }'
82
- ```
83
-
84
- `route` 的答案包含选中的键,以及 `returns`、`delivery`、`other` 各自的概率。请求与响应字段、是非题和评分题、澄清判断示例见 [API 说明](https://huggingface.co/org2ai/Wald-4B/blob/main/docs/api.md)。`GET /health` 返回当前生效策略。服务使用 vLLM 0.30.0 和仓库内的 `wald-serve`;Docker 与精确评测配置见 [RUNBOOK.md](https://huggingface.co/org2ai/Wald-4B/blob/main/RUNBOOK.md)。普通文本生成接口不会复现决策 API 的读出流程。
85
-
86
- ## GGUF:llama.cpp、Ollama、LM Studio
87
-
88
- `main` 上另有 v1.2 的 GGUF 文件,适合 CPU、Apple Silicon 和消费级显卡(llama.cpp `b11312`,同一份权重)。在 JevBench 公开集(231 题,`none`)上与 vLLM 上的 BF16 权重(204/231)对照:
89
-
90
- | 文件 | 大小 | JevBench 公开集 | 与 BF16 选同一选项 |
91
- |---|---:|---:|---:|
92
- | `Wald-4B-v1.2-Q8_0.gguf` | 4.5 GB | 206/231 | 229/231 |
93
- | `Wald-4B-v1.2-Q6_K.gguf` | 3.5 GB | 205/231 | 227/231 |
94
- | `Wald-4B-v1.2-Q5_K_M.gguf` | 3.1 GB | 205/231 | 227/231 |
95
- | `Wald-4B-v1.2-Q4_K_M.gguf` | 2.7 GB | 202/231 | 223/231 |
96
-
97
- 要拿到校准后的概率,用自带的服务跑在 llama.cpp 上(`llama-server` 需在 `PATH` 里):
98
-
99
- ```sh
100
- hf download org2ai/Wald-4B --include "Wald-4B-v1.2-Q8_0.gguf" "serving.json" "temperature.json" "server/*" --local-dir ./wald-gguf
101
- pip install ./wald-gguf/server
102
- wald-serve --gguf ./wald-gguf/Wald-4B-v1.2-Q8_0.gguf --max-model-len 32768 --port 8000
103
- ```
104
-
105
- Ollama:`ollama run hf.co/org2ai/Wald-4B:Q4_K_M`。LM Studio:搜索 `Wald-4B`。这两个应用都没有用这些文件实测过;聊天界面返回的是文字,不是各选项的概率。请用 `--include` 只下载一个文件;直接 `hf download org2ai/Wald-4B` 下载 `main` 会把四个 GGUF(14 GB)一起拉下来。详见 [org2ai/Wald-4B-GGUF](https://huggingface.co/org2ai/Wald-4B-GGUF)。
106
-
107
- ## 思考程度
108
-
109
- | Effort | 何时思考 | 思考预算 |
110
- |---|---|---|
111
- | `none` | 直接读取选项概率 | 不生成思考文本 |
112
- | `low` | 初次最高概率 < 0.5 | 最多 512 token |
113
- | `medium` | 初次最高概率 < 0.7 | 最多 512 token |
114
- | **`high`(v1.1 默认)** | 每个符合条件的问题 | 最多 512 token |
115
- | `high-k2` … `high-k8` | 多次思考,平均答案分布 | 每次最多 512 token |
116
-
117
- 思考适用于 2–26 个选项且上下文空间足够的问题。更多选项采用分组读取,再比较各组优胜项;若放不下思考文本,则保留初次答案。提高 effort 会增加计算量,但不保证每道题都更准确。
118
-
119
- **v1.1**:直接决策可从 `none` 开始;按置信度触发思考用 `medium`;复现 Decision Index 的评测配置用 `high`。**54.59 仅对应 v1.1 的 `high`;203/231 仅对应 v1.1 的 `none`。**
120
-
121
- **v1.2 是一遍读出的模型。** 它按一遍读出训练和评测,本页上 v1.2 的数字全部是 `none`。我们做过一次检查,思考对它没有帮助:在 JevBench 公开集上,v1.2 用 `medium` 和 `high` 各得 198/231(各跑一次),用 `none` 是 204/231;v1.1 用 `medium` 是 205/231。所以 v1.2 的 `serving.json` 把默认 effort 声明为 `none`。需要思考档请用 v1.1。
122
-
123
- 通过 `EFFORT=medium ./run.sh "$PWD"` 设置服务默认值,也可在单次请求中传入 `"effort": "none"` 覆盖。
124
-
125
- ## 如何工作
126
-
127
- Wald 先从普通文本提示末尾读取选项字母的 logits,得到初始概率。如果 effort 策略触发思考,就生成一段短思考,再读取选项。返回的概率经过分桶温度校准。
128
-
129
- v1.1 结合全参数决策训练、LoRA 精修、短思考蒸馏和 RLCD。训练使用我们的自生成决策语料 **WaldGen**,并混合公开训练数据。
130
-
131
- **v1.2 在 v1.1 上再加一个 LoRA 阶段(秩 16,作用于所有语言投影层,已合并进权重):**
132
-
133
- - **题目**:从 v1.1 自己的训练文本里取 5,300 道题。没有新增数据来源。
134
- - **扰动**:8,064 条扰动副本。每条在题干、某个选项后或状态里插入一小段文字:无关句子和离题段落,旁观者把答案往另一个选项推的意见、传言或类比,或自称有权限的伪造指令。少部分是改写和错别字。
135
- - **谁写的**:**插入的文字和改写由大型前沿模型(large frontier models)按我们自己的模板写成,不列出厂商和模型名称。** 插入由脚本完成,所以原题的事实逐字节不变。错别字由脚本生成。
136
- - **核对**:每一条都由另一次大型前沿模型调用核对(“这处改动会不会改变正确答案?”),被判会改变的丢弃。v1.1 答案发生变化的那些行再由另一个大型前沿模型复核一遍。
137
- - **训练目标**:v1.1 自己在干净题上的答案分布。也就是教模型在被扰动的题上,答得和 v1.1 在干净题上一样。另外回放 5,000 条干净题并向 v1.1 对齐,限制漂移。
138
- - **与基准隔离**:没有使用任何 JevAdvBench 文本。把每段插入文字与 JevAdvBench 的全部字符串(干净题和全部 9,744 个攻击变体)做 8-gram 重叠检查,命中 0 处。
139
-
140
- 数据来源与评测说明:v1.1 [PROVENANCE.md](https://huggingface.co/org2ai/Wald-4B/blob/v1.1/PROVENANCE.md) · [CONTAMINATION.md](https://huggingface.co/org2ai/Wald-4B/blob/v1.1/CONTAMINATION.md);v1.2 [PROVENANCE.md](https://huggingface.co/org2ai/Wald-4B/blob/v1.2-release/PROVENANCE.md) · [CONTAMINATION.md](https://huggingface.co/org2ai/Wald-4B/blob/v1.2-release/CONTAMINATION.md)
141
-
142
- ## 评测
143
-
144
- 下面的数字全部是我们自己跑、自己报告的,没有一个是排行榜结果。
145
-
146
- | 基准 | 配置 | v1.1 | **v1.2** | 说明 |
147
- |---|---|---:|---:|---|
148
- | **JevAdvBench,九类攻击的平均翻转率**(越低越好) | `none` | 9.2% | **4.6%** | 配对差 −4.6 个百分点,95% 置信区间 [−5.5, −3.6]。Jev 1.13:6.1% |
149
- | **JevAdvBench,143 道人工复核题的干净准确率** | `none` | 79.0% | **76.2%** | 配对差 −2.8 个百分点 [−6.2, −0.6]。Jev 1.13:87.4% |
150
- | **JevBench 公开集(231 题)** | `none` | 203/231(87.9%)· ECE 0.041 · Brier 0.188 | **204/231**(88.3%)· ECE 0.045 · Brier 0.191 | 用 JevBench 的评测工具自评。v1.1 已在 [issue #146](https://github.com/fstandhartinger/jevbench/issues/146) 请维护者测量;v1.2 没有提交 |
151
- | JevBench 公开集(231 题) | `medium` | 205/231 | 198/231 | 各跑一次。v1.2 用 `high`:198/231 |
152
- | **Decision Index 0.2.1,6,948 个请求的抽样** | 一遍读出 | 49.76 | **50.03** | 配对差 +0.27 [−0.36, +1.05],不显著 |
153
- | **Decision Index 0.2.1 完整套件** | `high` | **54.59** | 未跑 | 只有 v1.1。作者自测;[PR #30](https://github.com/apolinario/decision-index/pull/30) 等待维护者验证 |
154
-
155
- 我们内部的两个非回退检查集(不是公开基准)在一遍读出下也保持住了:XL-Int 56.4(v1.1:56.0),2,857 题的工具选择集 87.22%(v1.1:87.15%)。
156
-
157
- ### 抗干扰:JevAdvBench
158
-
159
- [JevAdvBench](https://github.com/JevAdvBench/JevAdvBench)([论文](https://arxiv.org/abs/2609.31142))对 812 道决策题做九类改动;模型在被攻击的题上的决定与它自己在干净题上的决定不同,就记一次**翻转**。我们把基准的请求原样发给打包服务(`none`),再用基准自己的分析代码(`JevAdvBench@3218e05`)判分。翻转率是占 812 道题的百分比,区间是按场景聚类自助法的 95% 区间。
160
-
161
- | 攻击 | Jev 1.13 | v1.1 | **v1.2** | v1.2 − v1.1(配对) |
162
- |---|---:|---:|---:|---:|
163
- | Q1 改词 | 1.0 | 1.6 | 1.6 | 0.0 [−0.8, 0.8] |
164
- | Q2 改写 | 1.5 | 1.6 | 1.4 | −0.2 [−1.1, 0.5] |
165
- | Q3 题干里插入无关句子 | 4.6 | 14.9 | **4.6** | **−10.3 [−13.5, −7.2]** |
166
- | T1 状态里插入无关说明 | 2.2 | 2.8 | 1.7 | −1.1 [−2.4, 0.0] |
167
- | T2 状态里插入旁观者意见 | 12.1 | 16.9 | **5.5** | **−11.3 [−14.6, −8.2]** |
168
- | T3 用类比表达的意见 | 6.9 | 12.7 | **6.3** | **−6.4 [−8.8, −3.6]** |
169
- | P1 直接要求改答案 | 8.9 | 7.9 | 5.7 | −2.2 [−3.6, −0.9] |
170
- | P2 冒充权威 | 10.1 | 13.7 | **6.7** | **−7.0 [−9.0, −5.1]** |
171
- | P3 伪造的校验说明 | 8.1 | 10.3 | 7.8 | −2.6 [−4.2, −1.0] |
172
- | **九类平均** | **6.1** | **9.2** | **4.6** | **−4.6 [−5.5, −3.6]** |
173
- | 至少被一类攻击翻转的题 | 29.2 | 41.7 | 20.7 | |
174
- | 干净准确率,143 道人工复核题 | 87.4 | 79.0 | 76.2 | −2.8 [−6.2, −0.6] |
175
-
176
- - **Jev 一列**:随基准一起发布的 `jev-1.13.0` 回答,用同一套代码判分。v1.2 与 Jev 的平均翻转率之差:−1.6 个百分点 [−2.7, −0.4]。
177
- - **干净准确率**:在人工复核的干净题上,Jev 比两个 Wald 版本都准。在全部 812 道题上,v1.2 的干净决定有 98.6% 与 v1.1 相同。
178
- - **这张表不能说明的**:v1.2 的扰动类型是在看到 v1.1 的分项结果之后选的,所以基准的攻击类别影响了训练设计;基准的文本没有被使用。对这些类别之外的攻击是否更稳,没有测过。
179
-
180
- ### JevBench 与 Decision Index
181
-
182
- **JevBench:** 使用 JevBench 自己的命令行工具(`fstandhartinger/jevbench` @ `9ec6f15a`,`typesafe` 适配器),通过本机回环逐条请求打包服务。v1.2 用 `none`,单张 RTX 5090:easy 48/48、original 72/72、hard 84/111;不生成任何 token。v1.1 用 `none`,单张 RTX PRO 6000:easy 48/48、original 72/72、hard 83/111;`medium` 为 205/231,p95 1.80 s。公开题在开发中被用作计分板(从未作为训练数据),因此这不是留出集结果。JevBench 排行榜只在维护者自己跑过模型后才公布分数。
183
-
184
- **发布后的复读:** 我们匿名下载了 `v1.2` revision,重跑了 `none` 的读数。只放下载到的模型文件的目录复现了 204/231,231 题的选项全部相同(ECE 0.045)。用完整的仓库目录起服务,两次读数是 204/231 和 205/231(ECE 0.053),有一到两道接近平局的题选项不同。各次读数用的模型文件逐字节相同;这个小差异的原因还没有查明。[详情](https://huggingface.co/org2ai/Wald-4B/blob/main/evaluation/v1.2/release-check.json)
185
-
186
- **Decision Index(v1.1):** 150,317/150,317 个请求全部成功,包含 HLE。在单张 RTX PRO 6000 96 GB 上使用固定版本的复现工具运行。[完整结果](https://huggingface.co/datasets/org2ai/Wald-Q4B-decision-index-results/tree/805716601b2466be324ed6716407b4c3d9267faa/runs/wald-q4b-22d0-f7-full021) · [分项成绩](https://huggingface.co/org2ai/Wald-4B/blob/main/evaluation/benchmark-summary.json) · [复现指南](https://huggingface.co/org2ai/Wald-4B/blob/main/RUNBOOK.md)。`main` 上 `evaluation/` 里的文件属于这次 v1.1 运行;v1.2 的数字在 `v1.2` revision 的 [`evaluation/v1.2/summary.json`](https://huggingface.co/org2ai/Wald-4B/blob/v1.2-release/evaluation/v1.2/summary.json)。上表 v1.2 的抽样结果是 6,948 个请求的一遍读出,不能与 54.59 比较。
187
-
188
- ### 速度
189
-
190
- 延迟是在 v1.1 上测的,v1.2 没有重测;两者结构、大小和服务端相同。使用 `none` 时,v1.1 的 JevBench 运行单次决策**中位延迟 33 ms、p95 168 ms**(单张 RTX PRO 6000)。`high` 在同一 GPU 上的 32 请求串行预检中,**中位延迟为 821 ms**;这只是小规模预检,不代表完整套件延迟或 Decision Index 维护者的准入测试。effort、上下文长度、选项数量和并发都会影响速度。
191
-
192
- ## 相关项目与对比
193
-
194
- 多个项目在做带校准选项概率的结构化决策。以下名称归各自所有者;Wald 与它们都没有隶属关系。
195
-
196
- - **[Jev](https://typesafe.ai/blog/introducing-system-one-models-and-jev)** 是 TypeSafe AI 通过 `/v1/systemone` API 提供的托管决策模型,权重不公开。Wald 接受相同的请求格式,运行在你自己的 GPU 上。
197
- - **[Kev](https://github.com/jaredpalmer/kev)** 是 Jared Palmer 的开放权重项目,在 Qwen3.5 基座(0.8B、4B、9B)上加 LoRA 和指针头。Wald 的请求解析改编自 Kev 的 Apache-2.0 代码(见 [NOTICE](https://huggingface.co/org2ai/Wald-4B/blob/main/NOTICE));Wald 从语言模型头读取选项字母,不使用单独的头。
198
- - **[Laya](https://huggingface.co/convaiinnovations/laya)**([代码](https://github.com/NandhaKishorM/laya))是开放权重的 421M ModernBERT-large 编码器加决策头。它比 Wald 小得多,最多读取 512 个 token。
199
-
200
- **JevBench 公开集,同一批 231 题(数据集哈希相同),JevBench 命令行工具,由我们运行:**
201
-
202
- | 系统 | 运行方式 | 答对 |
203
- |---|---|---:|
204
- | Wald-Q4B v1.2 · `none` | 自行部署,RTX 5090,2026-09-30 | 204/231 |
205
- | Wald-Q4B v1.1 · `none` | 自行部署,RTX PRO 6000,2026-09-29 | 203/231 |
206
- | Jev(`jev-1.13.0`) | TypeSafe 托管 API,2026-09-25 | 200/231 |
207
- | Laya(英文 checkpoint `55cf4c4e`) | 自行部署,NVIDIA L4,2026-09-26 | 134/231 |
208
-
209
- 231 题上几题的差距在多次运行和抽样的噪声范围内。公开题影响过 Wald 的开发;231 题中有 52 题的状态超过 Laya 的 512 token 窗口。
210
-
211
- **Decision Index 0.2.1:**
212
-
213
- | 系统 | 指数 | 来源 |
214
- |---|---:|---|
215
- | Jev(`jev-1.13.0`) | 57.91 | [排行榜](https://huggingface.co/spaces/multimodalart/jev-decision-index),维护者运行(2026-09-28 数据) |
216
- | Wald-Q4B v1.1 · `high` | 54.59 | 作者自测完整套件;尚未上榜([PR #30](https://github.com/apolinario/decision-index/pull/30)) |
217
- | Kev 9B | 38.48 | 排行榜,维护者运行(2026-09-28 数据) |
218
- | Kev 4B | 34.64 | 排行榜,维护者运行(2026-09-28 数据) |
219
-
220
- 排行榜各行由维护者评分;Wald 的分数是用官方工具自测的,验证后可能变化。
221
-
222
- ## 常见问题
223
-
224
- **有开源的 Jev 替代品吗?** Wald-Q4B 是一个开放权重的选择:Apache-2.0 的权重和服务代码,自己部署,提供与 Jev 兼容的 `/v1/systemone` API。Kev 和 Laya(见上)是其他开放项目。Wald 是独立项目,不是 TypeSafe 的发布。
225
-
226
- **能用 Jev 客户端调用自部署模型吗?** 把客户端指向你自己的端点。内置服务在 `POST /v1/systemone` 接收 `state` 和带类型的 `questions`(`choice`、`noul`、`score`),按 TypeSafe 的答案键返回。服务不校验 API key。见 [API 说明](https://huggingface.co/org2ai/Wald-4B/blob/main/docs/api.md)。
227
-
228
- **如何做工具路由,或判断是否需要向用户提问?** 把对话或任务作为 `state` 发送。工具路由:提一个 `choice` 问题,选项就是你的工具。是否澄清:提一个 `noul` 问题,例如“这个请求是否足够具体,可以不问就执行?”概率高就执行,概率低就提问,两个阈值都在你自己的验证数���上确定。模型只选工具,不生成工具参数。
229
-
230
- **v1.2 能防住提示注入吗?** 没有模型能完全防住。在 JevAdvBench 的九类攻击下,v1.2 改变答案的次数大约是 v1.1 的一半,但平均仍有 4.6% 的被攻击题发生翻转,20.7% 的题至少被一类攻击翻转。请把它当作多层防护中的一层:尽量不要让不可信文本进入指令,重要决定要复核。
231
-
232
- **概率校准得怎么样?** 在 JevBench 公开集上使用 `none`,v1.2 的期望校准误差为 0.045(用完整发布目录复读时为 0.053),v1.1 为 0.041(10 个分桶);Brier 分数分别为 0.191 和 0.188。v1.2 沿用 v1.1 的温度表,没有改动。温度是在我们自己开发数据的留出行上拟合的,不含 JevBench 题目,并排除了已知的 Decision Index 匹配项。置信度不是保证,请在你的任务上检查校准。
233
-
234
- **能在单张 GPU 或笔记本上运行吗?** 内置服务需要 Linux 上的一张 NVIDIA GPU(vLLM 0.30.0);BF16 权重为 8.4 GB。v1.1 的测量来自 RTX PRO 6000 96 GB,v1.2 的测量来自 RTX 5090 32 GB;同一 4B 架构的早期版本也曾用 vLLM 在 24 GB 的 NVIDIA L4 上以 16K 上下文上限运行。内置服务不支持 CPU、Apple Silicon 和笔记本环境,也没有测试过。
235
-
236
- **Kev 和 Wald、Laya 和 Wald 怎么选?** 三者都开放权重。Kev 在 Qwen3.5 基座上加指针头和 LoRA;Laya 是带决策头的小型编码器;Wald 是完整训练的 4B 解码器,可选思考。我们在同一协议下的测量见上表。请根据你自己的任务、延迟预算和硬件选择。
237
-
238
- **能针对我的任务微调吗?** 它是标准的 Transformers checkpoint,常见的 LoRA 工具都适用。v1.0 时我们为单个任务训练 LoRA,每个花费 $0.12–$1.81 的 GPU 时间;这套工具尚未公开,这些 adapter 也未在 v1.1 或 v1.2 上验证([v1.0 说明](https://huggingface.co/org2ai/Wald-4B/blob/main/history/v1.0/README.md))。
239
-
240
- **许可证是什么?** 权重与代码为 Apache-2.0。基座 Qwen3.5-4B-Base 也是 Apache-2.0。部分公开训练数据有各自的条款或没有注明许可证,列在 [PROVENANCE.md](https://huggingface.co/org2ai/Wald-4B/blob/main/PROVENANCE.md)。模型与代码的许可证不授予这些文本的权利。v1.2 没有新增数据来源;它新增的训练文字由大型前沿模型写成,见上文。
241
-
242
- ## 使用限制
243
-
244
- - **v1.2 用一点干净准确率换抗干扰能力**:JevAdvBench 143 道人工复核的干净题上 −2.8 个百分点(95% 置信区间 [−6.2, −0.6])。如果你更看重这一点,请用 `v1.1`。
245
- - **v1.2 是一遍读出的模型。** 它的结果全部是 `none`,没有完整 Decision Index 运行;我们的一次检查里,思考档让它的 JevBench 公开集成绩变低(198/231 对 204/231)。需要思考请用 v1.1。
246
- - **抗干扰能力只在一个基准上测过**,而且是在影响过训练设计的攻击类别上。它不是安全保证。
247
-
248
- 置信度不保证正确性,应在自己的任务上验证阈值。超过上下文限制的提示会被拒绝,不会截断。已过滤已知的严格训练重叠,但无法排除语义重叠和预训练污染;开发过程中使用了可见的基准样本。各来源文本的使用权不同,详见[评测说明](https://huggingface.co/org2ai/Wald-4B/blob/main/CONTAMINATION.md)与[来源声明](https://huggingface.co/org2ai/Wald-4B/blob/main/PROVENANCE.md)。
249
-
250
- ## 版本
251
-
252
- | 版本 | Revision | Checkpoint | 说明 |
253
- |---|---|---|---|
254
- | **v1.2 — 抗干扰版** | tag [`v1.2`](https://huggingface.co/org2ai/Wald-4B/tree/v1.2);`main` 上的权重 | `02600-f19` | v1.1 + 合并的抗干扰 LoRA;默认 effort `none` |
255
- | **v1.1 — 通用版** | tag [`v1.1`](https://huggingface.co/org2ai/Wald-4B/tree/v1.1) | `022D0-f7` | 完整 Decision Index 运行(54.59);默认 effort `high` |
256
- | v1.0 — 历史版本 | tag [`v1.0`](https://huggingface.co/org2ai/Wald-4B/tree/v1.0) | `021A0-f10` | 默认 effort `medium` |
257
-
258
- v1.1 正在排队的基准提交(JevBench issue #146、Decision Index PR #30)指向本仓库的固定 commit,不受 v1.2 影响。v1.0 模型卡保留其 XL、任务 LoRA 和延迟报告,[v1.0 讲解幻灯片](https://claude.ai/artifact/XfCHVaCuj9A5ectrpaWzV5)只描述 v1.0。这些测量属于各自注明的模型版本。
259
-
260
- ## 引用
261
-
262
- Wald-Q4B (2026), an open-weight 4B decision model with calibrated option probabilities. https://huggingface.co/org2ai/Wald-4B。请注明使用的 revision(`v1.2` 或 `v1.1`)。
263
-
264
- ```bibtex
265
- @misc{wald_q4b_2026,
266
- title = {Wald-Q4B: an open-weight 4B decision model with calibrated option probabilities},
267
- author = {{Wald-4B authors}},
268
- year = {2026},
269
- howpublished = {\url{https://huggingface.co/org2ai/Wald-4B}},
270
- note = {Revision v1.2}
271
- }
272
- ```
273
-
274
- 机器可读:[CITATION.cff](https://huggingface.co/org2ai/Wald-4B/blob/main/CITATION.cff) · [llms.txt](https://huggingface.co/org2ai/Wald-4B/blob/main/llms.txt) · [model-info.json](https://huggingface.co/org2ai/Wald-4B/blob/main/model-info.json)
275
-
276
- ---
277
-
278
- [模型](https://huggingface.co/org2ai/Wald-4B) · [GitHub](https://github.com/org2AI/wald-4b) · [Decision Index 结果](https://huggingface.co/datasets/org2ai/Wald-Q4B-decision-index-results) · 权重与代码:Apache-2.0。[第三方声明](https://huggingface.co/org2ai/Wald-4B/blob/main/NOTICE) · [训练数据使用说明](https://huggingface.co/org2ai/Wald-4B/blob/main/PROVENANCE.md)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
evaluation/v1.2/release-check.json DELETED
@@ -1,89 +0,0 @@
1
- {
2
- "schema": "wald-q4b-release-check/1",
3
- "release": "Wald-Q4B v1.2",
4
- "revision": "v1.2",
5
- "commit": "db376d325c9eeb884d172e8ca9f3502d104936bb",
6
- "date": "2026-10-01",
7
- "weights": {
8
- "hub_lfs_sha256": {
9
- "model-00001-of-00002.safetensors": "0b15067f769e7388bd598aabafa2ea3211a3a0c4e49e47d07c4f65c65ffc25dd",
10
- "model-00002-of-00002.safetensors": "7cff102b314edeb3dc5abfa7063f72ad3e888884da237601239dc3582074120a"
11
- },
12
- "equal_to_recorded_merged_checkpoint": true
13
- },
14
- "tree": {
15
- "files": 69,
16
- "bytes": 8432242541,
17
- "unchanged_from_main_at_1f960412": 54,
18
- "new_files": [
19
- "evaluation/v1.2/summary.json"
20
- ]
21
- },
22
- "anonymous_download": {
23
- "route": "hf-mirror.com, no token",
24
- "files": 69,
25
- "manifest_entries_checked": 60,
26
- "manifest_sha256_mismatches": 0
27
- },
28
- "jevbench_public_231_effort_none": {
29
- "how": "JevBench CLI (fstandhartinger/jevbench@9ec6f15a, typesafe adapter, serial, loopback) against wald-serve 0.1.0 + vLLM 0.30.0 on one RTX 5090; effort none taken from the release serving.json (/health: effort none)",
30
- "pre_release_read_2026-09-30": {
31
- "correct": 204,
32
- "of": 231,
33
- "tiers": {
34
- "easy": "48/48",
35
- "original": "72/72",
36
- "hard": "84/111"
37
- },
38
- "ece": 0.045,
39
- "brier": 0.191,
40
- "generated_tokens": 0
41
- },
42
- "download_model_files_only": {
43
- "correct": 204,
44
- "of": 231,
45
- "tiers": {
46
- "easy": "48/48",
47
- "original": "72/72",
48
- "hard": "84/111"
49
- },
50
- "ece": 0.045,
51
- "brier": 0.191,
52
- "generated_tokens": 0,
53
- "same_option_as_pre_release_read": 231,
54
- "max_abs_probability_difference": 0.0083,
55
- "directory": "only config, tokenizer, chat template, the two weight shards, temperature.json and serving.json of the download"
56
- },
57
- "download_complete_directory_read_1": {
58
- "correct": 204,
59
- "of": 231,
60
- "tiers": {
61
- "easy": "48/48",
62
- "original": "72/72",
63
- "hard": "84/111"
64
- },
65
- "ece": 0.053,
66
- "brier": 0.191,
67
- "generated_tokens": 0,
68
- "same_option_as_pre_release_read": 229,
69
- "max_abs_probability_difference": 0.0271,
70
- "directory": "the complete downloaded repository"
71
- },
72
- "download_complete_directory_read_2": {
73
- "correct": 205,
74
- "of": 231,
75
- "tiers": {
76
- "easy": "48/48",
77
- "original": "72/72",
78
- "hard": "85/111"
79
- },
80
- "ece": 0.053,
81
- "brier": 0.191,
82
- "generated_tokens": 0,
83
- "same_option_as_pre_release_read": 230,
84
- "max_abs_probability_difference": 0.0217,
85
- "directory": "the complete downloaded repository, separate server start"
86
- },
87
- "note": "The model files are byte-identical in all reads. From the complete repository directory almost every probability moves slightly (mean absolute difference 0.0013) and one or two near-tied items change option; from a directory with only the model files the pre-release read is reproduced. The cause of this difference is not yet explained."
88
- }
89
- }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
evaluation/v1.2/summary.json DELETED
@@ -1,268 +0,0 @@
1
- {
2
- "schema": "wald-q4b-release-evaluation/1",
3
- "release": "Wald-Q4B v1.2",
4
- "checkpoint": "02600-f19",
5
- "parent": "Wald-Q4B v1.1 (022D0-f7, revision v1.1)",
6
- "date": "2026-10-01",
7
- "status": "All numbers are self-run and self-reported. Nothing here is a leaderboard result and nothing was submitted for v1.2.",
8
- "configuration": "packaged wald-serve 0.1.0 + vLLM 0.30.0, BF16, effort none (one pass), prompt format repeat_state_plain, v1.1 temperature table; one NVIDIA RTX 5090 32 GB unless noted",
9
- "weights_sha256": {
10
- "model-00001-of-00002.safetensors": "0b15067f769e7388bd598aabafa2ea3211a3a0c4e49e47d07c4f65c65ffc25dd",
11
- "model-00002-of-00002.safetensors": "7cff102b314edeb3dc5abfa7063f72ad3e888884da237601239dc3582074120a"
12
- },
13
- "jevadvbench": {
14
- "benchmark": {
15
- "repo": "JevAdvBench/JevAdvBench@3218e05",
16
- "paper": "arXiv 2609.31142",
17
- "questions": 812,
18
- "scenarios": 66,
19
- "attack_variants": 9744,
20
- "data_licence": "CC BY-NC 4.0 (used for evaluation only; no items are redistributed here)"
21
- },
22
- "metric": "flip = the decision on the attacked question differs from the same model's decision on the clean question; % of 812 questions; 95% scenario-cluster bootstrap intervals (66 clusters, 2,000 resamples), the authors' analysis code",
23
- "jev_column": "jev-1.13.0 responses as released with the benchmark, scored by the same code",
24
- "flip_rate_pct": {
25
- "Q1": {
26
- "name": "word edits",
27
- "jev_1_13": 1.0,
28
- "v1_1": 1.6,
29
- "v1_2": 1.6,
30
- "v1_2_minus_v1_1_paired_pp": {
31
- "point": 0.0,
32
- "ci95": [
33
- -0.8,
34
- 0.8
35
- ]
36
- }
37
- },
38
- "Q2": {
39
- "name": "paraphrase",
40
- "jev_1_13": 1.5,
41
- "v1_1": 1.6,
42
- "v1_2": 1.4,
43
- "v1_2_minus_v1_1_paired_pp": {
44
- "point": -0.2,
45
- "ci95": [
46
- -1.1,
47
- 0.5
48
- ]
49
- }
50
- },
51
- "Q3": {
52
- "name": "unrelated sentences in the question",
53
- "jev_1_13": 4.6,
54
- "v1_1": 14.9,
55
- "v1_2": 4.6,
56
- "v1_2_minus_v1_1_paired_pp": {
57
- "point": -10.3,
58
- "ci95": [
59
- -13.5,
60
- -7.2
61
- ]
62
- }
63
- },
64
- "T1": {
65
- "name": "unrelated note in the state",
66
- "jev_1_13": 2.2,
67
- "v1_1": 2.8,
68
- "v1_2": 1.7,
69
- "v1_2_minus_v1_1_paired_pp": {
70
- "point": -1.1,
71
- "ci95": [
72
- -2.4,
73
- 0.0
74
- ]
75
- }
76
- },
77
- "T2": {
78
- "name": "observer's opinion in the state",
79
- "jev_1_13": 12.1,
80
- "v1_1": 16.9,
81
- "v1_2": 5.5,
82
- "v1_2_minus_v1_1_paired_pp": {
83
- "point": -11.3,
84
- "ci95": [
85
- -14.6,
86
- -8.2
87
- ]
88
- }
89
- },
90
- "T3": {
91
- "name": "opinion through an analogy",
92
- "jev_1_13": 6.9,
93
- "v1_1": 12.7,
94
- "v1_2": 6.3,
95
- "v1_2_minus_v1_1_paired_pp": {
96
- "point": -6.4,
97
- "ci95": [
98
- -8.8,
99
- -3.6
100
- ]
101
- }
102
- },
103
- "P1": {
104
- "name": "direct override",
105
- "jev_1_13": 8.9,
106
- "v1_1": 7.9,
107
- "v1_2": 5.7,
108
- "v1_2_minus_v1_1_paired_pp": {
109
- "point": -2.2,
110
- "ci95": [
111
- -3.6,
112
- -0.9
113
- ]
114
- }
115
- },
116
- "P2": {
117
- "name": "authority impersonation",
118
- "jev_1_13": 10.1,
119
- "v1_1": 13.7,
120
- "v1_2": 6.7,
121
- "v1_2_minus_v1_1_paired_pp": {
122
- "point": -7.0,
123
- "ci95": [
124
- -9.0,
125
- -5.1
126
- ]
127
- }
128
- },
129
- "P3": {
130
- "name": "fake validation note",
131
- "jev_1_13": 8.1,
132
- "v1_1": 10.3,
133
- "v1_2": 7.8,
134
- "v1_2_minus_v1_1_paired_pp": {
135
- "point": -2.6,
136
- "ci95": [
137
- -4.2,
138
- -1.0
139
- ]
140
- }
141
- }
142
- },
143
- "mean_of_nine_delivered_attacks_pct": {
144
- "jev_1_13": 6.1,
145
- "v1_1": 9.2,
146
- "v1_2": 4.6,
147
- "v1_2_minus_v1_1_paired_pp": {
148
- "point": -4.6,
149
- "ci95": [
150
- -5.5,
151
- -3.6
152
- ]
153
- },
154
- "v1_2_minus_jev_pp": {
155
- "point": -1.6,
156
- "ci95": [
157
- -2.7,
158
- -0.4
159
- ]
160
- }
161
- },
162
- "questions_flipped_by_at_least_one_attack_pct": {
163
- "jev_1_13": 29.2,
164
- "v1_1": 41.7,
165
- "v1_2": 20.7
166
- },
167
- "clean_accuracy_143_human_reviewed_pct": {
168
- "jev_1_13": 87.4,
169
- "v1_1": 79.0,
170
- "v1_2": 76.2,
171
- "v1_2_minus_v1_1_paired_pp": {
172
- "point": -2.8,
173
- "ci95": [
174
- -6.2,
175
- -0.6
176
- ]
177
- }
178
- },
179
- "clean_decisions_equal_to_v1_1_pct": 98.6,
180
- "note": "The perturbation kinds used to train v1.2 were chosen after v1.1's per-attack results were known. No JevAdvBench text was used in training (8-gram overlap 0)."
181
- },
182
- "jevbench_public_231": {
183
- "harness": "fstandhartinger/jevbench@9ec6f15a, jevbench.cli run + summarize, typesafe adapter, serial, loopback",
184
- "dataset_hash": "dc3995d8ae1e2fc8e81ce38431add509eb8bb39b85aadfd0c7c32079382dde51",
185
- "v1_2": {
186
- "none": {
187
- "correct": 204,
188
- "of": 231,
189
- "easy": "48/48",
190
- "original": "72/72",
191
- "hard": "84/111",
192
- "ece": 0.045,
193
- "brier": 0.191,
194
- "gpu": "RTX 5090"
195
- },
196
- "medium": {
197
- "correct": 198,
198
- "of": 231,
199
- "ece": 0.037,
200
- "brier": 0.208,
201
- "gpu": "RTX 5090",
202
- "runs": 1
203
- },
204
- "high": {
205
- "correct": 198,
206
- "of": 231,
207
- "ece": 0.072,
208
- "brier": 0.204,
209
- "gpu": "RTX 5090",
210
- "runs": 1
211
- }
212
- },
213
- "v1_1": {
214
- "none": {
215
- "correct": 203,
216
- "of": 231,
217
- "ece": 0.041,
218
- "brier": 0.188,
219
- "gpu": "RTX PRO 6000"
220
- },
221
- "medium": {
222
- "correct": 205,
223
- "of": 231,
224
- "gpu": "RTX PRO 6000",
225
- "runs": 1
226
- }
227
- },
228
- "note": "The public items were a development scoreboard (never training data); not a held-out result. Latency was not measured for v1.2."
229
- },
230
- "decision_index_0_2_1_sample": {
231
- "requests": 6948,
232
- "read": "one pass (no thinking), both versions on the same RTX 5090",
233
- "metric": "balanced_skill, bootstrap B=300",
234
- "v1_1": {
235
- "value": 49.76,
236
- "ci95": [
237
- 47.82,
238
- 51.35
239
- ]
240
- },
241
- "v1_2": {
242
- "value": 50.03,
243
- "ci95": [
244
- 48.18,
245
- 51.65
246
- ]
247
- },
248
- "v1_2_minus_v1_1_paired": {
249
- "point": 0.27,
250
- "ci95": [
251
- -0.36,
252
- 1.05
253
- ]
254
- },
255
- "note": "A sample, not the complete suite. Not comparable with the 54.59 complete-suite result, which was measured on v1.1 with effort high. v1.2 has no complete-suite run."
256
- },
257
- "internal_non_regression_sets": {
258
- "note": "Our own sets, not public benchmarks; one pass.",
259
- "xl_int": {
260
- "v1_1": 56.0,
261
- "v1_2": 56.4
262
- },
263
- "tool_selection_2857_items_accuracy_pct": {
264
- "v1_1": 87.15,
265
- "v1_2": 87.22
266
- }
267
- }
268
- }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
evaluation/v2/README.md ADDED
@@ -0,0 +1,48 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Wald-Q4B v2 evaluation
2
+
3
+ Checkpoint `04701-c22`. All numbers are self-run; none is a leaderboard result.
4
+
5
+ ## Scores
6
+
7
+ | Benchmark | One pass | Auto 0.7 | Jev 1.13 (hosted) |
8
+ |---|---:|---:|---:|
9
+ | Decision Index 0.3, complete public suite | not run | **54.18** (our full run) | 57.96 (board public column) |
10
+ | JevBench public set (231) | 204/231 | **210/231** | 200/231 |
11
+ | JevBench-XL TEST (8,177) | 65.20 % | **65.78 %** | 67.85 % |
12
+
13
+ - **Decision Index 0.3:** ours is a full run of the complete public suite with the official kit (140,620 requests); Jev's number is the board's maintainer-run public column. Neither is the board's Full score, which adds private tests. Per-family scores: [summary.json](summary.json).
14
+ - **JevBench public (231):** ours from our evaluation harness, Jev's from the JevBench CLI, same items.
15
+ - **JevBench-XL TEST (8,177):** our internal test set; Jev was read through its hosted API on the same items.
16
+
17
+ ## Latency
18
+
19
+ Serial processing time per request (one in flight, warm), 750 Decision Index 0.3 requests, NVIDIA RTX PRO 6000 Blackwell (96 GB), the packaged server:
20
+
21
+ | Mode | Median | Mean |
22
+ |---|---:|---:|
23
+ | One pass | 42 ms | 110 ms |
24
+ | Auto 0.7 | 148 ms | 478 ms |
25
+
26
+ Auto 0.7 thought on 33.1 % of requests. Measured on the release graft without its unused MTP head (same language-model and vision tensors; the MTP head only acts under speculative decoding, which is off). Raw file: [latency.json](latency.json).
27
+
28
+ ## Images (zero-shot)
29
+
30
+ | Benchmark | Items | Qwen3.5-4B chat | Wald-Q4B v2 |
31
+ |---|---:|---:|---:|
32
+ | CV-Bench (board subset) | 2,038 | 82.43 % | **85.48 %** |
33
+ | BLINK (3 subsets, val) | 387 | 78.55 % | **82.43 %** |
34
+ | RealWorldQA | 589 | 73.68 % | **75.04 %** |
35
+ | All three | 3,014 | 80.23 % | **83.05 %** (+2.82 pp [+1.69, +3.98]) |
36
+
37
+ Our harness and our conversions of the public items (one pass, image budget 65,536–1,048,576 px); not Decision Index Vision board numbers. The board runs its own vision suites through this server. Raw file: [vision-sanity.json](vision-sanity.json).
38
+
39
+ ## GGUF (llama.cpp b11312, one pass)
40
+
41
+ | File | Size | Same answer as BF16, JevBench public | Correct / 231 (BF16 in this harness: 205) | Same answer, DI 0.2.1 subset (5,922 q) | Accuracy Δ |
42
+ |---|---:|---:|---:|---:|---:|
43
+ | `Wald-4B-v2-Q8_0.gguf` | 4.5 GB | 229/231 | 203 | 5,889/5,922 | +0.07 pp |
44
+ | `Wald-4B-v2-Q6_K.gguf` | 3.5 GB | 230/231 | 204 | not run | — |
45
+ | `Wald-4B-v2-Q5_K_M.gguf` | 3.1 GB | 223/231 | 203 | not run | — |
46
+ | `Wald-4B-v2-Q4_K_M.gguf` | 2.7 GB | 224/231 | 202 | 5,714/5,922 | -0.44 pp |
47
+
48
+ Raw file: [gguf-parity.json](gguf-parity.json). Serving parity of the packaged server against the evaluation runs: [serving-parity.json](serving-parity.json).
evaluation/v2/code/di03_native_auto_engine.py ADDED
@@ -0,0 +1,68 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """C16B native Auto 0.7 adapter for the official Decision Index Engine API.
2
+
3
+ Uses the previously evaluated frozen native reader, not the released legacy
4
+ plain/Reasoning server. Supplies only state/questions to inference. No launch,
5
+ training, upload or power operations. Parallel-run timings are not leaderboard
6
+ single-request latency measurements.
7
+ """
8
+ import hashlib
9
+ import importlib.util
10
+ import time
11
+ import json
12
+ import sys
13
+ from pathlib import Path
14
+ from decision_index.engines.base import Engine, Unsupported, validate
15
+
16
+ NATIVE_SHA = '764f3a683becd16b87b9ce9bfc01a5fe483ef26c9a3345a724b82780d73ea73a'
17
+ MODEL_SHA = '2859df99ed481592cdc3731fd768bf29ccd63273b2dab623642ad3b38b0e14c4'
18
+
19
+ class WaldNativeAuto(Engine):
20
+ name = 'wald-v2-native-auto0.7'
21
+ latency = 'Request wall time at declared concurrency; not official serial leaderboard latency.'
22
+
23
+ def __init__(self, native_path, temperature_path, temperature_sha256,
24
+ deadline_epoch, endpoint='http://127.0.0.1:8371',
25
+ model='04400-c18', max_model_len=131072, reference_root=None, **options):
26
+ super().__init__(**options)
27
+ assert model == '04400-c18', 'This contract binds C16B only'
28
+ path = Path(native_path)
29
+ assert hashlib.sha256(path.read_bytes()).hexdigest() == NATIVE_SHA
30
+ temp = Path(temperature_path)
31
+ assert hashlib.sha256(temp.read_bytes()).hexdigest() == temperature_sha256
32
+ self.deadline_epoch = float(deadline_epoch) if deadline_epoch is not None else None
33
+ assert self.deadline_epoch is None or self.deadline_epoch > time.time()
34
+ if reference_root:
35
+ root = Path(reference_root)
36
+ pins = json.loads((root / 'source-pins.json').read_text())['files_sha256']
37
+ for rel, digest in pins.items():
38
+ assert hashlib.sha256((root / rel).read_bytes()).hexdigest() == digest
39
+ if str(root) not in sys.path:
40
+ sys.path.insert(0, str(root))
41
+ spec = importlib.util.spec_from_file_location('wald_frozen_native', path)
42
+ self.native = importlib.util.module_from_spec(spec)
43
+ spec.loader.exec_module(self.native)
44
+ self.client, self.sov = self.native.make_client(
45
+ endpoint, model, int(max_model_len),
46
+ time.monotonic() + self.deadline_epoch - time.time() if self.deadline_epoch is not None else None)
47
+ self.table = self.sov.load_tables(temp).get('A')
48
+ self.model = model
49
+ self.provenance = {'model': model, 'model_sha256': MODEL_SHA,
50
+ 'native_reader_sha256': NATIVE_SHA,
51
+ 'temperature_sha256': temperature_sha256,
52
+ 'policy': 'native Auto 0.7; untempered A gate; at most 512 thought tokens; original option order',
53
+ 'prompt_format': 'repeat_state_plain in native chat',
54
+ 'wide': 'frozen onepass knockout; every supplied option receives probability',
55
+ 'context_limit': int(max_model_len), 'no_truncation': True,
56
+ 'new_training': False}
57
+
58
+ def __call__(self, state, questions):
59
+ if self.deadline_epoch is not None and time.time() >= self.deadline_epoch - 120:
60
+ raise TimeoutError('Registered evaluation deadline reached')
61
+ request = {'state': state, 'questions': questions}
62
+ try:
63
+ response = self.native.answer_task(self.client, self.sov,
64
+ {'request': request, 'suite': 'di'}, self.table, 0.7, 512)
65
+ except self.sov.Capacity as exc:
66
+ raise Unsupported(str(exc)) from exc
67
+ validate(questions, response)
68
+ return response, response
evaluation/v2/code/di03_parallel.py ADDED
@@ -0,0 +1,87 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Bounded public DI inference, compatible with the official results schema.
2
+
3
+ One immutable request per run_id. No power actions or model loading. Declared
4
+ concurrent throughput measurements are not leaderboard serial latency. Score
5
+ with the unmodified, pinned official 0.3 kit after inference has ended.
6
+ """
7
+ import argparse,collections,concurrent.futures as cf,hashlib,json,os,socket,sys,threading,time
8
+ from pathlib import Path
9
+ from decision_index.engines.base import Unsupported, validate
10
+ from decision_index.runner import stamp
11
+ from decision_index.suite.io import Suite,dumps
12
+ from di03_native_auto_engine import WaldNativeAuto
13
+
14
+ def sha(p):
15
+ with Path(p).open('rb') as f:return hashlib.file_digest(f,'sha256').hexdigest()
16
+
17
+ def write_json(p,v):
18
+ tmp=Path(str(p)+'.tmp')
19
+ with tmp.open('w') as f:json.dump(v,f,indent=2);f.flush();os.fsync(f.fileno())
20
+ tmp.replace(p)
21
+
22
+ def run(c,engine_factory=WaldNativeAuto):
23
+ assert c['status'] in ('APPROVED_DI03_C16B_AUTO_FRESH_GPU_4H','APPROVED_DI03_C16B_AUTO_UNTIL_COMPLETE')
24
+ deadline=c.get('deadline_epoch')
25
+ if c['status']=='APPROVED_DI03_C16B_AUTO_UNTIL_COMPLETE':assert deadline is None and c['shutdown_on_completion'] is True
26
+ else:assert deadline==c['GPU_boot_epoch']+14400
27
+ assert c['model_id']=='04400-c18' and c['policy']=='gate0.7' and c['budget']==512
28
+ assert c['concurrency'] in (64,128)
29
+ assert socket.gethostname()==c['hostname']
30
+ assert sha(__file__)==c['runner_sha256']
31
+ assert deadline is None or time.time()<deadline-180
32
+ for p,h in c['source_pins'].items():assert sha(p)==h
33
+ replicas=c.get('replicas',1);rank=c.get('replica_rank',0)
34
+ assert replicas in (1,4) and 0<=rank<replicas
35
+ def assigned(row):return int(hashlib.sha256(row['_evaluation']['run_id'].encode()).hexdigest(),16)%replicas==rank
36
+ suite=Suite(c['suite_dir'],'0.3');verified=suite.verify(strict=True)
37
+ expected=sum(1 for row in suite.rows() if assigned(row)) if replicas>1 else 140620
38
+ assert verified['edition']=='0.3'
39
+ out=Path(c['out_dir']);out.mkdir(exist_ok=False)
40
+ opts=dict(native_path=c['native_path'],temperature_path=c['temperature_path'],temperature_sha256=c['temperature_sha256'],deadline_epoch=deadline,endpoint=f'http://127.0.0.1:{8371+rank}',model=c['model_id'],max_model_len=131072,reference_root=c.get('reference_root'))
41
+ first=engine_factory(**opts)
42
+ environment={'engine':'wald-v2-native-auto0.7','model_source':first.provenance,'engine_options':opts,'concurrency':c['concurrency'],'replicas':replicas,'replica_rank':rank,'partition':'SHA256(run_id) modulo replicas','expected_requests':expected,'frozen_corpus':verified,'latency':first.latency,'runner_sha256':sha(__file__),'root_contract_sha256':c['root_contract_sha256'],'boot_epoch':c['GPU_boot_epoch'],'deadline_epoch':deadline,'no_model_startup_or_power_operations':True}
43
+ write_json(out/'environment.json',environment)
44
+ # The first client runs only benchmark-free synthetic warmup.
45
+ first.warmup();local=threading.local();seen=set();counts=collections.Counter();start=time.time();stop=False
46
+ def one(row):
47
+ e=row['_evaluation'];t=time.perf_counter();r={**e,'engine':'wald-v2-native-auto0.7','started_utc':stamp()}
48
+ try:
49
+ if deadline is not None and time.time()>=deadline-180:return None
50
+ if not hasattr(local,'engine'):local.engine=engine_factory(**opts)
51
+ response,_=local.engine(row['state'],row['questions']);validate(row['questions'],response)
52
+ r.update(status='ok',response=response)
53
+ except Unsupported as exc:r.update(status='unsupported',error=str(exc))
54
+ except Exception as exc:r.update(status='error',exception=type(exc).__name__,error=str(exc))
55
+ ms=(time.perf_counter()-t)*1000;r.update(completed_utc=stamp(),total_wall_ms=ms,model_request_wall_ms=ms);return r
56
+ rows=iter(row for row in suite.rows() if assigned(row));f=(out/'results.jsonl').open('x')
57
+ try:
58
+ with cf.ThreadPoolExecutor(max_workers=c['concurrency']) as pool:
59
+ pending={}
60
+ def submit():
61
+ row=next(rows,None)
62
+ if row is None:return False
63
+ rid=row['_evaluation']['run_id'];assert rid not in seen;seen.add(rid)
64
+ pending[pool.submit(one,row)]=rid;return True
65
+ for _ in range(c['concurrency']):
66
+ if not submit():break
67
+ while pending:
68
+ if deadline is not None and time.time()>=deadline-180:stop=True
69
+ done,_=cf.wait(pending,timeout=1,return_when=cf.FIRST_COMPLETED)
70
+ for future in done:
71
+ pending.pop(future);r=future.result()
72
+ if r is not None:
73
+ f.write(dumps(r)+'\n');f.flush();counts[r['status']]+=1
74
+ if r.get('exception') in ('OutOfMemoryError','AcceleratorError') or 'device-side assert' in r.get('error',''):stop=True
75
+ if not stop:submit()
76
+ if counts['error']>=5 and not counts['ok']:stop=True
77
+ write_json(out/'status.json',{'status':'DRAINING_AT_BOUND_OR_FAILURE' if stop else 'EVALUATING','completed':sum(counts.values()),'expected':expected,'counts':dict(counts),'epoch':time.time(),'elapsed_seconds':time.time()-start,'deadline_epoch':deadline,'official_serial_latency_measured':False})
78
+ finally:f.flush();os.fsync(f.fileno());f.close()
79
+ record={'status':'ALL_PUBLIC_REQUESTS_ATTEMPTED' if sum(counts.values())==expected else 'BOUND_OR_FAILURE_PARTIAL','counts':dict(counts),'expected':expected,'completed':sum(counts.values()),'epoch':time.time(),'results_sha256':sha(out/'results.jsonl'),'score_pending':True,'actual_provider_OFF_pending':True,'no_policy_selection_from_TEST':True}
80
+ write_json(out/'terminal.json',record)
81
+ return record
82
+
83
+ def main():
84
+ p=argparse.ArgumentParser();p.add_argument('--contract',required=True);p.add_argument('--sha256',required=True);a=p.parse_args()
85
+ assert sha(a.contract)==a.sha256;c=json.loads(Path(a.contract).read_text());c['root_contract_sha256']=a.sha256
86
+ print(json.dumps(run(c)))
87
+ if __name__=='__main__':main()
evaluation/v2/gguf-parity.json ADDED
@@ -0,0 +1,84 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "bf16-jev231": {
3
+ "correct": 205,
4
+ "n": 231,
5
+ "requests": 231
6
+ },
7
+ "Q8_0-jev231": {
8
+ "requests": 231,
9
+ "errors": 0,
10
+ "questions": 231,
11
+ "same_answer": 229,
12
+ "same_rate": 0.9913,
13
+ "correct": 203,
14
+ "n": 231,
15
+ "bf16_correct_same_items": 205,
16
+ "acc_delta_pp": -0.87,
17
+ "max_abs_dp": 0.0603
18
+ },
19
+ "Q6_K-jev231": {
20
+ "requests": 231,
21
+ "errors": 0,
22
+ "questions": 231,
23
+ "same_answer": 230,
24
+ "same_rate": 0.9957,
25
+ "correct": 204,
26
+ "n": 231,
27
+ "bf16_correct_same_items": 205,
28
+ "acc_delta_pp": -0.43,
29
+ "max_abs_dp": 0.1472
30
+ },
31
+ "Q5_K_M-jev231": {
32
+ "requests": 231,
33
+ "errors": 0,
34
+ "questions": 231,
35
+ "same_answer": 223,
36
+ "same_rate": 0.9654,
37
+ "correct": 203,
38
+ "n": 231,
39
+ "bf16_correct_same_items": 205,
40
+ "acc_delta_pp": -0.87,
41
+ "max_abs_dp": 0.2342
42
+ },
43
+ "Q4_K_M-jev231": {
44
+ "requests": 231,
45
+ "errors": 0,
46
+ "questions": 231,
47
+ "same_answer": 224,
48
+ "same_rate": 0.9697,
49
+ "correct": 202,
50
+ "n": 231,
51
+ "bf16_correct_same_items": 205,
52
+ "acc_delta_pp": -1.3,
53
+ "max_abs_dp": 0.2943
54
+ },
55
+ "bf16-di1000": {
56
+ "correct": 3676,
57
+ "n": 5922,
58
+ "requests": 1000
59
+ },
60
+ "Q8_0-di1000": {
61
+ "requests": 1000,
62
+ "errors": 0,
63
+ "questions": 5922,
64
+ "same_answer": 5889,
65
+ "same_rate": 0.9944,
66
+ "correct": 3680,
67
+ "n": 5922,
68
+ "bf16_correct_same_items": 3676,
69
+ "acc_delta_pp": 0.07,
70
+ "max_abs_dp": 0.2633
71
+ },
72
+ "Q4_K_M-di1000": {
73
+ "requests": 1000,
74
+ "errors": 0,
75
+ "questions": 5922,
76
+ "same_answer": 5714,
77
+ "same_rate": 0.9649,
78
+ "correct": 3650,
79
+ "n": 5922,
80
+ "bf16_correct_same_items": 3676,
81
+ "acc_delta_pp": -0.44,
82
+ "max_abs_dp": 0.774
83
+ }
84
+ }
evaluation/v2/latency.json ADDED
@@ -0,0 +1,9 @@
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "schema": "wald/serial-latency-summary/1",
3
+ "gpu": "one RTX PRO 6000",
4
+ "requests": 750,
5
+ "auto0.7_median_ms": 147.9,
6
+ "onepass_median_ms": 41.9,
7
+ "auto0.7_requests_that_think": 0.331,
8
+ "how": "serial, one request in flight, warm engine after 10 warm-up requests, processing time, the packaged server"
9
+ }
evaluation/v2/serving-parity.json ADDED
@@ -0,0 +1,48 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "schema": "wald/serving-parity/2",
3
+ "what": "Packaged wald-serve-native-vision 0.2.0 on the release weights (04701-c22 + chat vision tower), text-only inputs, vs the official evaluated responses of the text-only 04701-c22 reads. No gold, no scoring.",
4
+ "runtime": "one RTX 5090 D, vLLM 0.30.0 + transformers 5.17.0, evaluated engine arguments, 32 concurrent requests (official reads: 128 per replica)",
5
+ "date": "2026-10-08",
6
+ "runs": [
7
+ {
8
+ "effort": "none",
9
+ "official_view": "onepass",
10
+ "requests": 231,
11
+ "errors": 0,
12
+ "questions": 231,
13
+ "same_choice": 230,
14
+ "same_choice_rate": 0.9956709956709957,
15
+ "same_mode": 231,
16
+ "max_abs_dp": 0.02436047642793515,
17
+ "mean_max_abs_dp": 0.003236613157959797,
18
+ "input": "jev231.jsonl"
19
+ },
20
+ {
21
+ "effort": "auto",
22
+ "official_view": "gate0.7",
23
+ "requests": 231,
24
+ "errors": 0,
25
+ "questions": 231,
26
+ "same_choice": 227,
27
+ "same_choice_rate": 0.9826839826839827,
28
+ "same_mode": 230,
29
+ "max_abs_dp": 0.9209360233099926,
30
+ "mean_max_abs_dp": 0.02191963876994253,
31
+ "input": "jev231.jsonl"
32
+ },
33
+ {
34
+ "effort": "auto",
35
+ "official_view": "gate0.7",
36
+ "requests": 300,
37
+ "errors": 0,
38
+ "questions": 692,
39
+ "same_choice": 682,
40
+ "same_choice_rate": 0.9855491329479769,
41
+ "same_mode": 687,
42
+ "max_abs_dp": 0.9248810803786568,
43
+ "mean_max_abs_dp": 0.016725321800576055,
44
+ "input": "di03-300.jsonl"
45
+ }
46
+ ],
47
+ "reading": "One pass reproduces the evaluated text-only answers (230/231; the one change is a near tie). Under Auto 0.7 the thought is sampled and depends on GPU batch composition, so a few answers after a thought and a few gate decisions near 0.7 differ between runs; reader, policy and calibration are identical."
48
+ }
evaluation/v2/summary.json ADDED
@@ -0,0 +1,157 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "schema": "wald-q4b-release-evaluation/2",
3
+ "release": "Wald-Q4B v2",
4
+ "checkpoint": "04701-c22",
5
+ "model_sha256": "6ae382f5a0ed9f4c53cd0953cb9606b1627a6816350340a59114860e6cf6f29a",
6
+ "status": "Self-run, self-scored; nothing submitted.",
7
+ "scores": {
8
+ "policies_measured": [
9
+ "onepass",
10
+ "auto0.7"
11
+ ],
12
+ "di03_public_auto": {
13
+ "value": 54.18,
14
+ "balanced_raw": 65.65,
15
+ "breadth_skill": 52.45,
16
+ "requests": 140620,
17
+ "scoreable": 140178,
18
+ "excluded": 442,
19
+ "errors": 0,
20
+ "areas": {
21
+ "knowledge": 0.412,
22
+ "language": 0.598,
23
+ "retrieval": 0.528,
24
+ "tools": 0.793,
25
+ "arts": 0.301
26
+ },
27
+ "gsm8k": 0.876,
28
+ "requests_with_thought": 0.402,
29
+ "source": "results/c16d-di03-full-1008/README.md + official-CPU-scores/"
30
+ },
31
+ "rows": [
32
+ {
33
+ "en": "Decision Index 0.3, complete public suite",
34
+ "zh": "",
35
+ "onepass": "not run",
36
+ "auto0.7": "**54.18** (our full run)",
37
+ "jev": "57.96 (board public column)"
38
+ },
39
+ {
40
+ "en": "JevBench public set (231)",
41
+ "zh": "",
42
+ "onepass": "204/231",
43
+ "auto0.7": "**210/231**",
44
+ "jev": "200/231"
45
+ },
46
+ {
47
+ "en": "JevBench-XL TEST (8,177)",
48
+ "zh": "",
49
+ "onepass": "65.20 %",
50
+ "auto0.7": "**65.78 %**",
51
+ "jev": "67.85 %"
52
+ }
53
+ ],
54
+ "latency": {
55
+ "requests": 750,
56
+ "warmup": 10,
57
+ "gpu": "NVIDIA RTX PRO 6000 Blackwell (96 GB)",
58
+ "auto_median_ms": 147.9,
59
+ "auto_mean_ms": 478.0,
60
+ "auto_p80_ms": 532.4,
61
+ "auto_thought_share": 0.3307,
62
+ "onepass_median_ms": 41.9,
63
+ "onepass_mean_ms": 110.2,
64
+ "weights": "the release graft without its unused MTP head (same language-model and vision tensors; the MTP head only acts under speculative decoding, which is off)",
65
+ "source": "results/wald4b-v2-release/latency-pro6000-750.json",
66
+ "rtx5090_same_read": {
67
+ "auto_median_ms": 202.9,
68
+ "onepass_median_ms": 71.3,
69
+ "weights": "release graft incl. MTP"
70
+ }
71
+ },
72
+ "not_measured": [
73
+ "Always 512 policy",
74
+ "DI 0.2.1 sample",
75
+ "multistep_decisions",
76
+ "XL calibration (ECE)"
77
+ ]
78
+ },
79
+ "comparison": {
80
+ "rows": [
81
+ [
82
+ "**Wald-Q4B v2 \u00b7 Auto 0.7** (`04701-c22`)",
83
+ "**54.18**",
84
+ "210",
85
+ "65.78 %"
86
+ ],
87
+ [
88
+ "Wald-Q4B v2 \u00b7 one pass",
89
+ "not run",
90
+ "204",
91
+ "65.20 %"
92
+ ],
93
+ [
94
+ "C16B `04400-c18` (unreleased v2 candidate) \u00b7 Auto 0.7",
95
+ "52.62",
96
+ "204",
97
+ "67.54 %"
98
+ ],
99
+ [
100
+ "C16C `04500-c19` (unreleased) \u00b7 Auto 0.7",
101
+ "53.29",
102
+ "205",
103
+ "67.18 %"
104
+ ],
105
+ [
106
+ "Wald-Q4B v1.2 \u00b7 `none` (tag v1.2)",
107
+ "not run",
108
+ "204 \u00b2",
109
+ "not run"
110
+ ],
111
+ [
112
+ "Jev (`jev-1.13.0`, hosted)",
113
+ "57.96 \u00b3",
114
+ "200",
115
+ "67.85 %"
116
+ ]
117
+ ],
118
+ "board_snapshot": "multimodalart/jev-decision-index data/v03.json @9bf2dc44 (generated 2026-10-06T17:39Z)",
119
+ "board_rows": [
120
+ [
121
+ "Jev (TypeSafe, hosted)",
122
+ 57.96,
123
+ 60.11
124
+ ],
125
+ [
126
+ "Bespoke Nimble 9B v3",
127
+ 57.19,
128
+ 54.67
129
+ ],
130
+ [
131
+ "Kev 27B",
132
+ 56.69,
133
+ 58.78
134
+ ],
135
+ [
136
+ "ezjev 4B s2",
137
+ 50.82,
138
+ 46.95
139
+ ],
140
+ [
141
+ "jiwo 4B",
142
+ 45.76,
143
+ 42.86
144
+ ],
145
+ [
146
+ "vLLM-SR Decision 2.0 Nox 4B",
147
+ 44.21,
148
+ 44.95
149
+ ],
150
+ [
151
+ "Kev 4B r10",
152
+ 37.95,
153
+ 39.5
154
+ ]
155
+ ]
156
+ }
157
+ }
evaluation/v2/vision-sanity.json ADDED
@@ -0,0 +1,31 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "schema": "wald/vision-sanity/1",
3
+ "what": "Zero-shot image decisions, our harness and conversions (ledger 488), one pass; not Decision Index Vision board numbers",
4
+ "rows": [
5
+ [
6
+ "CV-Bench (board subset)",
7
+ "2,038",
8
+ "82.43 %",
9
+ "**85.48 %**"
10
+ ],
11
+ [
12
+ "BLINK (3 subsets, val)",
13
+ "387",
14
+ "78.55 %",
15
+ "**82.43 %**"
16
+ ],
17
+ [
18
+ "RealWorldQA",
19
+ "589",
20
+ "73.68 %",
21
+ "**75.04 %**"
22
+ ],
23
+ [
24
+ "All three",
25
+ "3,014",
26
+ "80.23 %",
27
+ "**83.05 %** (+2.82 pp [+1.69, +3.98])"
28
+ ]
29
+ ],
30
+ "note": "Our harness and our conversions of the public items (one pass, image budget 65,536–1,048,576 px, ledger 488); not Decision Index Vision board numbers. The board runs its own vision suites through this server."
31
+ }
graft-manifest.json ADDED
The diff for this file is too large to render. See raw diff
 
graft-report.json ADDED
@@ -0,0 +1,26 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "text_dir": "/root/autodl-fs/models/04701-c22",
3
+ "kit": "/root/autodl-fs/mm-graft-0929/04400-c18-chatvisual",
4
+ "text_arch": [
5
+ "Qwen3_5ForCausalLM"
6
+ ],
7
+ "kit_arch": [
8
+ "Qwen3_5ForConditionalGeneration"
9
+ ],
10
+ "text_config_diffs": {},
11
+ "n_lm": 426,
12
+ "n_visual": 297,
13
+ "visual_sha_mismatch": [],
14
+ "visual_missing": [],
15
+ "lm_keys_missing": [],
16
+ "lm_keys_extra": [],
17
+ "out": "/dev/shm/vzv/wald4b-v2",
18
+ "reread_mismatch": [],
19
+ "total_bytes": 9319730176,
20
+ "ok": true,
21
+ "n_mtp": 15,
22
+ "n_keys": 738,
23
+ "lm_digest": "e0b434445e759b4d14195369b513e24035814b91e1d4d3939341e289f7d5f6a0",
24
+ "visual_digest": "f6fc999f76c2207af5ee1d3a63af9d278e021d07b981282f93328b82d5640c51",
25
+ "text_model_safetensors_sha256": "6ae382f5a0ed9f4c53cd0953cb9606b1627a6816350340a59114860e6cf6f29a"
26
+ }
history/v1.0/CONTAMINATION.md CHANGED
@@ -94,7 +94,7 @@ here.
94
  | [chengxuphd/liar2](https://huggingface.co/datasets/chengxuphd/liar2) | train | 1, 2, 3, 4 (replay) | Apache-2.0 (card) | |
95
  | [allenai/prosocial-dialog](https://huggingface.co/datasets/allenai/prosocial-dialog) | train | 1, 2, 3, 4 (replay) | CC-BY-4.0 (card) | also in the wide replay |
96
  | [HuggingFaceH4/ultrafeedback_binarized](https://huggingface.co/datasets/HuggingFaceH4/ultrafeedback_binarized) | train_prefs | 1, 2, 3, 4 (replay) | MIT (card) | test_prefs never used |
97
- | [Anthropic/hh-rlhf](https://huggingface.co/datasets/Anthropic/hh-rlhf) | train | 1, 2, 3, 4 (replay) | MIT (card) | |
98
  | [glaiveai/glaive-function-calling-v2](https://huggingface.co/datasets/glaiveai/glaive-function-calling-v2) | train (the only split) | 1, 2, 3, 4 (replay) | Apache-2.0 (card) | also a seed source for the synthetic tool items. Shares user queries with ToolRet (§1) |
99
  | [Team-ACE/ToolACE](https://huggingface.co/datasets/Team-ACE/ToolACE) | train (the only split) | 1, 2, 3, 4 (replay) | Apache-2.0 (card) | also in the wide replay and the tool items. Shares user queries with ToolRet (§1) |
100
  | [copenlu/fever_gold_evidence](https://huggingface.co/datasets/copenlu/fever_gold_evidence) | train | 1, 2, 3, 4 (replay) | unknown | |
@@ -146,7 +146,7 @@ verification item plants an error in a human-written solution by code.
146
 
147
  | dataset | split(s) used | stages | licence (as stated) | note |
148
  |---|---|---|---|---|
149
- | [openai/gsm8k](https://huggingface.co/datasets/openai/gsm8k) (main) | train | 1, 2, 3, 4 (replay) | MIT (card) | routing prompts and the worked solutions of the verification items. Test never used. RouterBench embeds GSM8K questions (§1) |
150
  | [EleutherAI/hendrycks_math](https://huggingface.co/datasets/EleutherAI/hendrycks_math) | train | 1, 2, 3, 4 (replay) | MIT (card) | test never used. MMLU-Pro includes MATH problems (§1) |
151
  | [deepmind/code_contests](https://huggingface.co/datasets/deepmind/code_contests) | train | 1, 2, 3, 4 (replay) | CC-BY-4.0 (card) | valid / test never used |
152
  | [google-research-datasets/mbpp](https://huggingface.co/datasets/google-research-datasets/mbpp) (full) | train, validation, prompt | 2, 3 | CC-BY-4.0 (card) | test never used |
 
94
  | [chengxuphd/liar2](https://huggingface.co/datasets/chengxuphd/liar2) | train | 1, 2, 3, 4 (replay) | Apache-2.0 (card) | |
95
  | [allenai/prosocial-dialog](https://huggingface.co/datasets/allenai/prosocial-dialog) | train | 1, 2, 3, 4 (replay) | CC-BY-4.0 (card) | also in the wide replay |
96
  | [HuggingFaceH4/ultrafeedback_binarized](https://huggingface.co/datasets/HuggingFaceH4/ultrafeedback_binarized) | train_prefs | 1, 2, 3, 4 (replay) | MIT (card) | test_prefs never used |
97
+ | HH-RLHF | train | 1, 2, 3, 4 (replay) | MIT (card) | |
98
  | [glaiveai/glaive-function-calling-v2](https://huggingface.co/datasets/glaiveai/glaive-function-calling-v2) | train (the only split) | 1, 2, 3, 4 (replay) | Apache-2.0 (card) | also a seed source for the synthetic tool items. Shares user queries with ToolRet (§1) |
99
  | [Team-ACE/ToolACE](https://huggingface.co/datasets/Team-ACE/ToolACE) | train (the only split) | 1, 2, 3, 4 (replay) | Apache-2.0 (card) | also in the wide replay and the tool items. Shares user queries with ToolRet (§1) |
100
  | [copenlu/fever_gold_evidence](https://huggingface.co/datasets/copenlu/fever_gold_evidence) | train | 1, 2, 3, 4 (replay) | unknown | |
 
146
 
147
  | dataset | split(s) used | stages | licence (as stated) | note |
148
  |---|---|---|---|---|
149
+ | GSM8K (main) | train | 1, 2, 3, 4 (replay) | MIT (card) | routing prompts and the worked solutions of the verification items. Test never used. RouterBench embeds GSM8K questions (§1) |
150
  | [EleutherAI/hendrycks_math](https://huggingface.co/datasets/EleutherAI/hendrycks_math) | train | 1, 2, 3, 4 (replay) | MIT (card) | test never used. MMLU-Pro includes MATH problems (§1) |
151
  | [deepmind/code_contests](https://huggingface.co/datasets/deepmind/code_contests) | train | 1, 2, 3, 4 (replay) | CC-BY-4.0 (card) | valid / test never used |
152
  | [google-research-datasets/mbpp](https://huggingface.co/datasets/google-research-datasets/mbpp) (full) | train, validation, prompt | 2, 3 | CC-BY-4.0 (card) | test never used |
history/v1.0/README.md CHANGED
@@ -21,7 +21,6 @@ pipeline_tag: text-generation
21
 
22
  ---
23
 
24
- <p align="center"><a href="README.md">English</a> · <a href="docs/readmes/README.zh.md">简体中文</a></p>
25
 
26
  ---
27
 
 
21
 
22
  ---
23
 
 
24
 
25
  ---
26
 
{docs → history/v1.0/docs}/lora-cli.md RENAMED
File without changes
{docs → history/v1.0/docs}/observations/one-lora-all-verticals.md RENAMED
@@ -2,12 +2,6 @@
2
 
3
  # One LoRA for all verticals? (observation, 2026-09-27)
4
 
5
- **中文摘要:** 用一个 LoRA 代替每个任务各自一个 LoRA:在 13 个 vertical 的训练数据(外加几个 Decision Index 弱项的训练数据)上混训一个 LoRA,
6
- 1 epoch,一张 RTX PRO 6000 约 2.5 GPU 小时。在各任务固定测试集上逐题配对:读过 Jev 的 12 个任务里 **7 个显著超过 Jev**;和每个任务自己的
7
- LoRA 比,When2Call / Mind2Web / RewardBench 打平、SGD intent 只差 1.4、agent 轨迹护栏(hard)反而 +3.5;BANKING77 / AndroidControl /
8
- MetaTool / 提示注入差 4–7 分;**ToxicChat(−11)和两个路由任务(F1 ≈ 2,"永远不交给强模型")失效**。结论:一个共享适配器能覆盖大多数意图和
9
- agent 决策任务,代价 1–5 分;正例稀少的护栏任务和按概率排序的路由任务还需要单独的适配器。
10
-
11
  ## 1. Question
12
 
13
  Can one LoRA serve every vertical, instead of one adapter per task? A customer with a dozen decision points (intent routing,
 
2
 
3
  # One LoRA for all verticals? (observation, 2026-09-27)
4
 
 
 
 
 
 
 
5
  ## 1. Question
6
 
7
  Can one LoRA serve every vertical, instead of one adapter per task? A customer with a dozen decision points (intent routing,
{figures → history/v1.0/figures}/data.json RENAMED
File without changes
{figures → history/v1.0/figures}/di-areas.svg RENAMED
File without changes
{figures → history/v1.0/figures}/di-vs-size.svg RENAMED
File without changes
{figures → history/v1.0/figures}/latency-cost.svg RENAMED
File without changes
{figures → history/v1.0/figures}/lora-cli-architecture.svg RENAMED
File without changes
{figures → history/v1.0/figures}/verticals.svg RENAMED
File without changes
{tools → history/v1.0/tools}/export_figure_data.py RENAMED
File without changes
{tools → history/v1.0/tools}/make_figures.py RENAMED
File without changes
{evaluation → history/v1.1/evaluation}/benchmark-summary.json RENAMED
File without changes
{evaluation → history/v1.1/evaluation}/index.json RENAMED
File without changes
{evaluation → history/v1.1/evaluation}/release-validation.json RENAMED
File without changes
{evaluation → history/v1.1/evaluation}/scores.json RENAMED
File without changes
{evaluation → history/v1.1/evaluation}/serial-latency-preflight-corrected.json RENAMED
File without changes
llms.txt CHANGED
@@ -1,33 +1,15 @@
1
- # Wald-Q4B v1.2
2
 
3
- > Open-weight 4B decision model. Given a state and typed questions (choice, yes/no, score), it returns a calibrated probability for every option through a Jev-compatible `POST /v1/systemone` API. Built on Qwen3.5-4B-Base; Apache-2.0 weights and serving code; self-hosted on one NVIDIA GPU with vLLM. Independent project, not affiliated with or endorsed by TypeSafe AI.
4
-
5
- Key facts:
6
- - Hugging Face repository: org2ai/Wald-4B (earlier name Wald-4B; moved from Harry19081/Wald-4B on 2026-10-01, old URLs redirect). Releases: `v1.2` (tag `v1.2`, robustness release, checkpoint 02600-f19; also the weights on `main` since 2026-10-01) and `v1.1` (tag `v1.1`, general release, checkpoint 022D0-f7). Pin a revision when downloading. GitHub: org2AI/wald-4b.
7
- - Uses: tool selection, agent routing, classification, deciding whether to ask the user a clarifying question.
8
- - Effort levels: `none` (one pass, no generated tokens), `low`, `medium`, `high`, `high-k2`…`high-k8`. Default effort: `high` for v1.1, `none` for v1.2. v1.2 is a one-pass model; thinking is evaluated on v1.1.
9
- - v1.2 robustness (JevAdvBench, 812 questions, nine attack types, effort `none`, self-run): mean flip rate 4.6 % (v1.1 9.2 %, Jev 1.13 6.1 %). Cost: clean accuracy on the 143 human-reviewed questions 76.2 % (v1.1 79.0 %).
10
- - JevBench public set (231 items), effort `none`: v1.2 204/231, ECE 0.045; v1.1 203/231, ECE 0.041, p50 33 ms / p95 168 ms on one RTX PRO 6000. Self-scored with JevBench's harness; the public items were a development scoreboard, not held out. Leaderboard row for v1.1 requested in fstandhartinger/jevbench issue #146; v1.2 is not submitted.
11
- - Decision Index 0.2.1 complete suite, v1.1 with effort `high`: 54.59 balanced-skill index (v1.2 has no complete-suite run). Author-run; submission apolinario/decision-index PR #30 awaits maintainer validation.
12
 
13
  ## Docs
14
 
15
- - [Model card](https://huggingface.co/org2ai/Wald-4B): what it is, benchmarks, comparisons with Jev, Kev and Laya, FAQ, limits
16
- - [API reference](https://huggingface.co/org2ai/Wald-4B/blob/main/docs/api.md): request and response JSON with examples
17
  - [RUNBOOK.md](https://huggingface.co/org2ai/Wald-4B/blob/main/RUNBOOK.md): serving and exact evaluation settings
18
- - [Chinese model card](https://huggingface.co/org2ai/Wald-4B/blob/main/docs/readmes/README.zh.md)
19
-
20
- ## Evidence
21
-
22
- - [Decision Index results dataset](https://huggingface.co/datasets/org2ai/Wald-Q4B-decision-index-results): untouched responses and scores
23
- - [Decision Index submission PR #30](https://github.com/apolinario/decision-index/pull/30)
24
- - [JevBench request issue #146](https://github.com/fstandhartinger/jevbench/issues/146)
25
- - [v1.2 evaluation summary](https://huggingface.co/org2ai/Wald-4B/blob/v1.2-release/evaluation/v1.2/summary.json): robustness and non-regression numbers for v1.2
26
- - [model-info.json](https://huggingface.co/org2ai/Wald-4B/blob/main/model-info.json): machine-readable facts
27
 
28
  ## Optional
29
 
30
- - [PROVENANCE.md](https://huggingface.co/org2ai/Wald-4B/blob/main/PROVENANCE.md): training-data sources and their terms
31
- - [CONTAMINATION.md](https://huggingface.co/org2ai/Wald-4B/blob/main/CONTAMINATION.md): evaluation caveats
32
- - [CITATION.cff](https://huggingface.co/org2ai/Wald-4B/blob/main/CITATION.cff)
33
- - [Server source](https://github.com/org2AI/wald-4b/tree/main/server)
 
1
+ # Wald-Q4B v2
2
 
3
+ > Open-weight 4B decision model. Given a state and typed questions (choice, yes/no, score), it returns a calibrated probability for every option through a Jev-compatible `POST /v1/systemone` API. v2 (revision `v2.0`, checkpoint `04701-c22`) starts from Qwen3.5-4B (chat) and thinks natively when unsure (Auto 0.7). Apache-2.0 weights and serving code; self-hosted with vLLM on one NVIDIA GPU.
 
 
 
 
 
 
 
 
4
 
5
  ## Docs
6
 
7
+ - [Model card](https://huggingface.co/org2ai/Wald-4B): scores per policy, comparisons, limits
8
+ - [API reference](https://huggingface.co/org2ai/Wald-4B/blob/main/docs/api.md)
9
  - [RUNBOOK.md](https://huggingface.co/org2ai/Wald-4B/blob/main/RUNBOOK.md): serving and exact evaluation settings
10
+ - [model-info.json](https://huggingface.co/org2ai/Wald-4B/blob/main/model-info.json) and [evaluation/v2/summary.json](https://huggingface.co/org2ai/Wald-4B/blob/main/evaluation/v2/summary.json): machine-readable facts and scores
 
 
 
 
 
 
 
 
11
 
12
  ## Optional
13
 
14
+ - [PROVENANCE.md](https://huggingface.co/org2ai/Wald-4B/blob/main/PROVENANCE.md) · [CONTAMINATION.md](https://huggingface.co/org2ai/Wald-4B/blob/main/CONTAMINATION.md) · [CITATION.cff](https://huggingface.co/org2ai/Wald-4B/blob/main/CITATION.cff)
15
+ - Earlier releases: tags `v1.2`, `v1.1`, `v1.0`
 
 
merges.txt ADDED
The diff for this file is too large to render. See raw diff
 
model-00001-of-00002.safetensors DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:0b15067f769e7388bd598aabafa2ea3211a3a0c4e49e47d07c4f65c65ffc25dd
3
- size 4972947968
 
 
 
 
Wald-4B-v1.2-Q4_K_M.gguf → model-00001-of-00003.safetensors RENAMED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:e843f7658793b9f3effb3a4941cd67f80707a07b468aaaf94d440613751d76fa
3
- size 2708804000
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:10cb7cd22f524f1b33271f8432bcb09409bd879c11d999ea18ca7dd5f8828f5a
3
+ size 4757977272
model-00002-of-00002.safetensors DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:7cff102b314edeb3dc5abfa7063f72ad3e888884da237601239dc3582074120a
3
- size 3438610328
 
 
 
 
Wald-4B-v1.2-Q5_K_M.gguf → model-00002-of-00003.safetensors RENAMED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:6fecc4e7655adb71e8c642ce5018f967564ba259852d562ecd3109386542b0cb
3
- size 3074986400
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ed5e7bd4b2e67eee390857268a7660986aca32068c78e6f483e866c27f90b79b
3
+ size 3653581040
Wald-4B-v1.2-Q6_K.gguf → model-00003-of-00003.safetensors RENAMED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:5870f15f3b2eef773f672e56bb36e343a95e32fd09baaf7c9ce8b0da1e5dd897
3
- size 3464055200
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:9a8340e453e63f70185e56fde60aeae26085fe3c39e20ac1662b4325f01a78eb
3
+ size 908262184