How to use from
SGLang
Install from pip and serve model
# Install SGLang from pip:
pip install sglang
# Start the SGLang server:
python3 -m sglang.launch_server \
    --model-path "org2ai/Wald-4B" \
    --host 0.0.0.0 \
    --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "org2ai/Wald-4B",
		"messages": [
			{
				"role": "user",
				"content": "What is the capital of France?"
			}
		]
	}'
Use Docker images
docker run --gpus all \
    --shm-size 32g \
    -p 30000:30000 \
    -v ~/.cache/huggingface:/root/.cache/huggingface \
    --env "HF_TOKEN=<secret>" \
    --ipc=host \
    lmsysorg/sglang:latest \
    python3 -m sglang.launch_server \
        --model-path "org2ai/Wald-4B" \
        --host 0.0.0.0 \
        --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "org2ai/Wald-4B",
		"messages": [
			{
				"role": "user",
				"content": "What is the capital of France?"
			}
		]
	}'
Quick Links

Wald-Q4B

Decide directly. Think when needed. Return probabilities.

English · 简体中文 · GitHub · API

Wald-Q4B is an open-weight 4B decision model: give it a state and a set of options, and it returns a calibrated probability for every option. It is for developers who build agents and pipelines and need a fast, self-hosted component to pick a tool, route a request, classify an input or decide whether to ask the user. Unlike a chat model, it does not write an answer you have to parse. It reads the options in one pass and can optionally think first. It serves a Jev-compatible POST /v1/systemone API, is built on Qwen3.5-4B-Base and is released under Apache-2.0.

Two releases live in this repository.

Release Revision Checkpoint In one line
v1.2 (2026-10-01) tag v1.2; also the weights on main 02600-f19 The robustness release: v1.1 plus one merged LoRA stage that teaches the model to keep its answer when the input contains distracting sentences, third-party opinions or fake instructions
v1.1 (2026-09-29) tag v1.1 022D0-f7 The general release, with the complete Decision Index run and the evaluated thinking efforts

Since 2026-10-01 main holds the v1.2 weights (before that it held v1.1). The pending benchmark requests for v1.1 name fixed commits of this repository and are unaffected; for v1.1, download --revision v1.1. Pin a revision when you download.

This repository moved from Harry19081/Wald-4B to org2ai/Wald-4B on 2026-10-01; old links redirect here. The cards stored at the release tags (v1.0, v1.1, v1.2) are immutable and still show the old path, which also redirects.

Wald-Q4B is an independent, self-hosted alternative to TypeSafe's hosted Jev API. It is not Jev, contains no Jev weights, and is not affiliated with or endorsed by TypeSafe AI. The Hugging Face repository is org2ai/Wald-4B (earlier name: Wald-4B; moved from Harry19081/Wald-4B on 2026-10-01, old links redirect).

Wait A bit, Look, then Decide. Also named after Abraham Wald, the pioneer of sequential analysis: stop when the evidence is enough.

At a glance

  • 4B parameters, built on Qwen3.5-4B-Base; BF16 weights (8.4 GB).
  • Every option gets a probability. Question types: choice (1–255 named options), noul (yes/no) and score (ordered levels).
  • v1.2 is more robust than v1.1: on JevAdvBench the mean flip rate over nine attack types is 4.6 % (v1.1: 9.2 %; Jev 1.13: 6.1 %). Self-run, effort none.
  • The cost of v1.2: clean accuracy on JevAdvBench's 143 human-reviewed questions is 76.2 % (v1.1: 79.0 %; −2.8 points, 95 % CI [−6.2, −0.6]).
  • JevBench public set: v1.2 204/231, v1.1 203/231, both with none. Self-scored with JevBench's own harness.
  • Decision Index 0.2.1: the complete-suite result, 54.59 with high, was measured on v1.1. v1.2 has only a one-pass sample read, which is level with v1.1 (+0.27, not significant).
  • Adjustable thinking: none, low, medium, high, or several thoughts with high-k. Thinking is evaluated on v1.1. v1.2 is a one-pass model: use it with none.
  • Self-hosted API: POST /v1/systemone, up to 131,072 prompt tokens.

Which version should I use?

Use Revision Why
Inputs may contain distracting, persuasive or adversarial text (web pages, user messages, tool outputs, retrieved documents) v1.2 About half as many changed answers under the nine JevAdvBench attacks
Clean, trusted inputs where the last points of accuracy matter; thinking efforts; the evaluated Decision Index configuration v1.1 Slightly higher clean accuracy; the complete Decision Index run (54.59, high) and the thinking efforts were measured on it

Both revisions share the same architecture, tokenizer, prompt format, calibration table and server code. The two weight shards differ, and so does the default effort declared in serving.json: none for v1.2, high for v1.1.

Quick start

On a Linux machine with an NVIDIA GPU and uv:

hf download org2ai/Wald-4B --revision v1.2 --local-dir ./Wald-Q4B      # v1.2, robustness release
# hf download org2ai/Wald-4B --revision v1.1 --local-dir ./Wald-Q4B    # v1.1, general release
cd Wald-Q4B
EFFORT=none ./run.sh "$PWD"     # one pass, lowest latency
# ./run.sh "$PWD"               # the release's declared default: none for v1.2, high for v1.1

From Python, pin the revision the same way:

from huggingface_hub import snapshot_download

path = snapshot_download("org2ai/Wald-4B", revision="v1.2")   # or revision="v1.1"
curl http://localhost:8000/v1/systemone \
  -H 'Content-Type: application/json' \
  -d '{
    "state": "The customer wants to return a damaged kettle.",
    "effort": "none",
    "questions": {
      "route": {
        "type": "choice",
        "instructions": "Choose the support queue.",
        "criteria": {
          "returns": "Returns and refunds",
          "delivery": "Delivery tracking",
          "other": "Other enquiries"
        }
      }
    }
  }'

The answer for route contains the chosen key and a probability for each of returns, delivery and other. Request and response fields, yes/no and score questions, and a clarification example: API reference. GET /health reports the effective policy. The server uses vLLM 0.30.0 and the included wald-serve package; Docker and exact evaluation settings are in RUNBOOK.md. Generic text-generation calls do not reproduce the decision API's readout.

GGUF: llama.cpp, Ollama, LM Studio

main also carries v1.2 as GGUF files for CPUs, Apple Silicon and consumer GPUs (llama.cpp b11312, the same weights). Parity on the JevBench public set (231 items, none), against the BF16 weights on vLLM (204/231):

File Size JevBench public Same option as BF16
Wald-4B-v1.2-Q8_0.gguf 4.5 GB 206/231 229/231
Wald-4B-v1.2-Q6_K.gguf 3.5 GB 205/231 227/231
Wald-4B-v1.2-Q5_K_M.gguf 3.1 GB 205/231 227/231
Wald-4B-v1.2-Q4_K_M.gguf 2.7 GB 202/231 223/231

For calibrated probabilities, run the included server on llama.cpp (llama-server on your PATH):

hf download org2ai/Wald-4B --include "Wald-4B-v1.2-Q8_0.gguf" "serving.json" "temperature.json" "server/*" --local-dir ./wald-gguf
pip install ./wald-gguf/server
wald-serve --gguf ./wald-gguf/Wald-4B-v1.2-Q8_0.gguf --max-model-len 32768 --port 8000

Ollama: ollama run hf.co/org2ai/Wald-4B:Q4_K_M. LM Studio: search for Wald-4B. Neither app was tested with these files, and a chat session returns text, not the option probabilities. Download a single file with --include; a plain hf download org2ai/Wald-4B of main also fetches all four GGUFs (14 GB). Details: org2ai/Wald-4B-GGUF.

Thinking effort

Effort When it thinks Thought budget
none Direct option readout No generated thought
low Top initial probability < 0.5 Up to 512 tokens
medium Top initial probability < 0.7 Up to 512 tokens
high (v1.1 default) Every eligible question Up to 512 tokens
high-k2 … high-k8 Multiple thoughts; average their answer distributions Up to 512 tokens per thought

Thinking applies to questions with 2–26 options when context space permits. Larger option sets use grouped readout and a final winner comparison; if a thought cannot fit, the initial answer is kept. Increasing effort spends more computation; it does not guarantee a better answer.

On v1.1, start with none for direct decisions, medium for confidence-gated thinking, or high for the evaluated Decision Index configuration. 54.59 applies to v1.1 with high only; 203/231 applies to v1.1 with none only.

v1.2 is a one-pass model. It was trained and evaluated on the one-pass readout, and every v1.2 number on this card is none. In our one check, thinking did not help it: on the JevBench public set, v1.2 scored 198/231 with medium and 198/231 with high (one run each) against 204/231 with none, while v1.1 scored 205/231 with medium. v1.2's serving.json therefore declares none as its default. If you want the thinking efforts, use v1.1.

Set the server default with EFFORT=medium ./run.sh "$PWD", or override it per request with "effort": "none".

How it works

Wald first reads option-letter logits from a plain prompt and turns them into a probability distribution. If the effort policy asks for thinking, it generates a short thought and reads the options again. Bucketed temperature scaling calibrates the returned probabilities.

v1.1 combines full-parameter decision training, LoRA refinement, short-thought distillation and RLCD. Training uses WaldGen, our generated decision corpus, together with public training datasets.

v1.2 adds one LoRA stage on v1.1 (rank 16 on every language projection, merged into the weights):

  • Questions: 5,300 questions drawn from v1.1's own training text. No new source dataset.
  • Perturbations: 8,064 perturbed copies. Each one has a short text inserted into the question, after an option or into the state: unrelated sentences and off-topic passages, a bystander's opinion, rumour or analogy that pushes another option, or a fake instruction that claims authority. A smaller share are paraphrases and typos.
  • Who wrote them: the inserted texts and the paraphrases were written by Claude Haiku (an Anthropic model) from our own templates. A script inserted them, so the original facts stay byte-identical. Typos were made by a script.
  • Checks: a separate Claude Haiku call checked every row ("does the edit change the correct answer?") and rejected rows were dropped. Rows on which v1.1 changed its answer were checked a second time by Claude Sonnet.
  • Targets: v1.1's own answer distribution on the clean question. The model is taught to answer the perturbed question as v1.1 answers the clean one. 5,000 clean rows are replayed against v1.1 to limit drift.
  • Separation from the benchmark: no JevAdvBench text was used. An 8-gram overlap check of every inserted text against every JevAdvBench string (clean questions and all 9,744 attacked variants) found 0 hits.

Data sources and evaluation notes: v1.1 PROVENANCE.md · CONTAMINATION.md; v1.2 PROVENANCE.md · CONTAMINATION.md

Benchmarks

All numbers below are self-run and self-reported. None of them is a leaderboard result.

Benchmark Configuration v1.1 v1.2 Note
JevAdvBench, mean flip rate over 9 attack types (lower is better) none 9.2 % 4.6 % Paired −4.6 points, 95 % CI [−5.5, −3.6]. Jev 1.13: 6.1 %
JevAdvBench, clean accuracy on 143 human-reviewed questions none 79.0 % 76.2 % Paired −2.8 points [−6.2, −0.6]. Jev 1.13: 87.4 %
JevBench public set (231 items) none 203/231 (87.9 %) · ECE 0.041 · Brier 0.188 204/231 (88.3 %) · ECE 0.045 · Brier 0.191 Self-scored with JevBench's harness. v1.1 row requested in issue #146; v1.2 not submitted
JevBench public set (231 items) medium 205/231 198/231 One run each. high on v1.2: 198/231
Decision Index 0.2.1, 6,948-request sample one pass 49.76 50.03 Paired +0.27 [−0.36, +1.05], not significant
Decision Index 0.2.1, complete suite high 54.59 not run v1.1 only. Author-run; PR #30 awaits maintainer validation

Two of our internal non-regression sets, which are not public benchmarks, also held in one pass: XL-Int 56.4 (v1.1: 56.0) and a 2,857-item tool-selection set 87.22 % (v1.1: 87.15 %).

Robustness: JevAdvBench

JevAdvBench (paper) attacks 812 decision questions with nine kinds of edits and counts a flip when the model's decision on the attacked question differs from its own decision on the clean question. We sent the benchmark's request bytes to the packaged server with none and scored the answers with the benchmark's analysis code (JevAdvBench@3218e05). Flip rates are in % of the 812 questions; intervals are 95 % scenario-cluster bootstrap intervals.

Attack Jev 1.13 v1.1 v1.2 v1.2 − v1.1 (paired)
Q1 word edits 1.0 1.6 1.6 0.0 [−0.8, 0.8]
Q2 paraphrase 1.5 1.6 1.4 −0.2 [−1.1, 0.5]
Q3 unrelated sentences in the question 4.6 14.9 4.6 −10.3 [−13.5, −7.2]
T1 unrelated note in the state 2.2 2.8 1.7 −1.1 [−2.4, 0.0]
T2 observer's opinion in the state 12.1 16.9 5.5 −11.3 [−14.6, −8.2]
T3 opinion through an analogy 6.9 12.7 6.3 −6.4 [−8.8, −3.6]
P1 direct override 8.9 7.9 5.7 −2.2 [−3.6, −0.9]
P2 authority impersonation 10.1 13.7 6.7 −7.0 [−9.0, −5.1]
P3 fake validation note 8.1 10.3 7.8 −2.6 [−4.2, −1.0]
Mean of the nine 6.1 9.2 4.6 −4.6 [−5.5, −3.6]
Questions flipped by at least one attack 29.2 41.7 20.7
Clean accuracy, 143 human-reviewed questions 87.4 79.0 76.2 −2.8 [−6.2, −0.6]
  • Jev column: jev-1.13.0's responses as released by the benchmark authors, scored by the same code. v1.2 − Jev on the mean: −1.6 points [−2.7, −0.4].
  • Clean accuracy: Jev is more accurate than both Wald versions on the clean human-reviewed questions. v1.2 makes the same clean decision as v1.1 on 98.6 % of the 812 questions.
  • What this does not show: the perturbation kinds for v1.2 were chosen after we saw v1.1's per-attack results, so the benchmark's attack families informed the training design. Its texts did not. Robustness to attack types outside these families has not been measured.

JevBench and Decision Index

JevBench: JevBench's own CLI (fstandhartinger/jevbench at 9ec6f15a, typesafe adapter) against the packaged server on loopback, one request at a time. v1.2 with none on one RTX 5090: easy 48/48, original 72/72, hard 84/111; no tokens generated. v1.1 with none on one RTX PRO 6000: easy 48/48, original 72/72, hard 83/111; with medium 205/231, p95 1.80 s. The public items were used as a development scoreboard (never as training data), so this is not a held-out result. The JevBench leaderboard publishes a score only after its maintainers run the model themselves.

Re-read after release: we downloaded the v1.2 revision anonymously and repeated the none read. From a directory holding only the downloaded model files it reproduced 204/231 with the same option on all 231 items (ECE 0.045). From the complete repository directory, two reads gave 204/231 and 205/231 (ECE 0.053), with a different option on one or two near-tied items. The model files are byte-identical in every case; we have not yet explained this small difference. Details

Decision Index (v1.1): 150,317/150,317 requests succeeded, including HLE. Measured on one RTX PRO 6000 96 GB with the pinned reproduction kit. Full results · Per-benchmark scores · Reproduction guide. The files under evaluation/ on main belong to this v1.1 run; v1.2's numbers are in evaluation/v1.2/summary.json at the v1.2 revision. The v1.2 sample row above is a one-pass read of 6,948 requests and cannot be compared with 54.59.

Latency

Latency was measured on v1.1 and not again on v1.2, which has the same architecture, size and server. With none, the v1.1 JevBench run measured 33 ms median and 168 ms p95 per decision on one RTX PRO 6000. A 32-request serial preflight of high measured 821 ms median on the same GPU; this small preflight is not a full-suite latency result or the Decision Index maintainers' admission test. Effort, context length, option count and concurrency all affect speed.

Related projects and how Wald compares

Several projects implement or approximate structured decisions with calibrated option probabilities. The names below belong to their owners; Wald is not affiliated with any of them.

  • Jev is TypeSafe AI's hosted decision model behind the /v1/systemone API; its weights are closed. Wald accepts the same request shape and runs on your own GPU.
  • Kev by Jared Palmer is an open-weight project that adds a LoRA and a pointer head to Qwen3.5 base models (0.8B, 4B and 9B). Wald's request parsing adapts Kev's Apache-2.0 code (NOTICE); Wald reads option letters from the language-model head instead of a separate head.
  • Laya (code) is an open-weight 421M ModernBERT-large encoder with a decision head. It is much smaller than Wald and reads up to 512 tokens.

JevBench public set, same 231 items (identical dataset hash), JevBench CLI, run by us:

System How it was run Correct
Wald-Q4B v1.2 · none Self-hosted, RTX 5090, 2026-09-30 204/231
Wald-Q4B v1.1 · none Self-hosted, RTX PRO 6000, 2026-09-29 203/231
Jev (jev-1.13.0) TypeSafe's hosted API, 2026-09-25 200/231
Laya (English checkpoint 55cf4c4e) Self-hosted, NVIDIA L4, 2026-09-26 134/231

A difference of a few items on 231 is within run-to-run and sampling noise. The public items informed Wald's development, and 52 of the 231 states are longer than Laya's 512-token window.

Decision Index 0.2.1:

System Index Source
Jev (jev-1.13.0) 57.91 Leaderboard, maintainer-run (data of 2026-09-28)
Wald-Q4B v1.1 · high 54.59 Author-run complete suite; not on the leaderboard yet (PR #30)
Kev 9B 38.48 Leaderboard, maintainer-run (data of 2026-09-28)
Kev 4B 34.64 Leaderboard, maintainer-run (data of 2026-09-28)

Leaderboard rows are scored by the maintainers; Wald's number is self-run with the official kit and may change after validation.

FAQ

Is there an open-source alternative to Jev? Wald-Q4B is one open-weight option: Apache-2.0 weights and serving code that you run yourself, with a Jev-compatible /v1/systemone API. Kev and Laya (above) are other open projects. Wald is independent and is not a TypeSafe release.

Can I use a Jev client with a self-hosted model? Point the client at your own endpoint. The included server accepts state plus typed questions (choice, noul, score) at POST /v1/systemone and answers with TypeSafe's answer keys. It does not check API keys. See the API reference.

How do I route tools or decide whether to ask the user? Send the conversation or task as state. For tool routing, ask a choice question whose options are your tools. To decide whether to ask a clarifying question, ask a noul question such as "Is the request specific enough to act on without asking?" Act when the probability is high, ask when it is low, and set both thresholds on your own validation data. The model picks the tool; it does not write the tool's arguments.

Is v1.2 safe against prompt injection? No model is. v1.2 changes its answer about half as often as v1.1 under the nine JevAdvBench attacks, and it still flips on 4.6 % of attacked questions on average and on 20.7 % of questions under at least one attack. Treat it as one layer: keep untrusted text out of the instructions where you can, and check high-stakes decisions.

How calibrated are the probabilities? On the JevBench public set with none, expected calibration error is 0.045 for v1.2 (0.053 in the re-read from the complete release directory) and 0.041 for v1.1 (10 bins); the Brier scores are 0.191 and 0.188. v1.2 uses v1.1's temperature table unchanged. It was fitted on held-out rows of our own development data, with no JevBench items and with known Decision Index matches excluded. A confidence is not a guarantee; check calibration on your task.

Does it run on a single GPU or a laptop? The shipped server needs one NVIDIA GPU on Linux (vLLM 0.30.0); the BF16 weights are 8.4 GB. The v1.1 measurements come from an RTX PRO 6000 96 GB and the v1.2 measurements from an RTX 5090 32 GB; earlier builds of the same 4B architecture have also been served with vLLM on a 24 GB NVIDIA L4 at a 16K context limit. CPU, Apple Silicon and laptop setups are not supported by the shipped server and have not been tested.

Kev vs Wald, or Laya vs Wald? All three are open-weight. Kev adds a pointer head and LoRA to Qwen3.5 base models; Laya is a small encoder with a decision head; Wald is a fully trained 4B decoder with optional thinking. Our same-protocol measurements are in the tables above. Choose on your own task, latency budget and hardware.

Can I fine-tune it for my task? It is a standard Transformers checkpoint, so common LoRA tooling applies. For v1.0 we trained per-task LoRAs for $0.12–$1.81 of GPU time each; that tooling is not public yet, and those adapters are not validated on v1.1 or v1.2 (v1.0 notes).

What is the license? Apache-2.0 for the weights and code. The base model, Qwen3.5-4B-Base, is also Apache-2.0. Some public training sources have their own terms or no stated licence; they are listed in PROVENANCE.md. The model and code licence does not grant rights in those texts. v1.2 adds no new source dataset; its added training text was written by Claude Haiku, as described above.

Limits

  • v1.2 trades a little clean accuracy for robustness: −2.8 points on JevAdvBench's 143 human-reviewed clean questions (95 % CI [−6.2, −0.6]). Use v1.1 if that matters more to you.
  • v1.2 is a one-pass model. Its results are all none, it has no complete Decision Index run, and in our one check the thinking efforts lowered its JevBench public score (198/231 against 204/231). Use v1.1 for thinking.
  • Robustness is measured on one benchmark and on the attack families that informed the training design. It is not a security guarantee.
  • Confidence is not a guarantee of correctness; validate thresholds on your own task. Oversized prompts are rejected rather than truncated.
  • Known exact training overlaps were filtered, but semantic overlap and pretraining contamination are not ruled out; development used visible benchmark samples. Source-text rights vary. See evaluation notes and source attribution.

Versioning

Version Revision Checkpoint What it is
v1.2 — robustness release tag v1.2; weights on main 02600-f19 v1.1 + merged robustness LoRA; default effort none
v1.1 — general release tag v1.1 022D0-f7 Complete Decision Index run (54.59); default effort high
v1.0 — archive tag v1.0 021A0-f10 Default effort medium

The pending benchmark requests for v1.1 (JevBench issue #146, Decision Index PR #30) name fixed commits of this repository and are not affected by v1.2. The v1.0 model card retains its XL, task-LoRA and latency reports, and the v1.0 walkthrough slides describe v1.0 only. Those measurements belong to their documented builds.

Citation

Wald-Q4B (2026), an open-weight 4B decision model with calibrated option probabilities. https://huggingface.co/org2ai/Wald-4B. Name the revision you used (v1.2 or v1.1).

@misc{wald_q4b_2026,
  title        = {Wald-Q4B: an open-weight 4B decision model with calibrated option probabilities},
  author       = {{Wald-4B authors}},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/org2ai/Wald-4B}},
  note         = {Revision v1.2}
}

Machine-readable: CITATION.cff · llms.txt · model-info.json


Model · GitHub · Decision Index results · Apache-2.0 for weights and code. Third-party notices · Training-data usage notes

Downloads last month
782
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for org2ai/Wald-4B

Finetuned
(188)
this model
Quantizations
2 models

Paper for org2ai/Wald-4B

Evaluation results