Text Generation
Transformers
Safetensors
GGUF
English
qwen3_5
image-text-to-text
decision-model
typed-decisions
calibration
calibrated-probabilities
classification
tool-selection
agent-routing
decision-index
jevbench
jev-compatible
systemone
wald
wald-q4b
qwen3.5
4b
vllm
reasoning
llama.cpp
conversational
Eval Results (legacy)
Instructions to use org2ai/Wald-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use org2ai/Wald-4B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="org2ai/Wald-4B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("org2ai/Wald-4B") model = AutoModelForMultimodalLM.from_pretrained("org2ai/Wald-4B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use org2ai/Wald-4B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "org2ai/Wald-4B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "org2ai/Wald-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/org2ai/Wald-4B
- SGLang
How to use org2ai/Wald-4B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "org2ai/Wald-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "org2ai/Wald-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "org2ai/Wald-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "org2ai/Wald-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use org2ai/Wald-4B with Docker Model Runner:
docker model run hf.co/org2ai/Wald-4B
Download evaluation/index.json from org2ai/Wald-4B: direct link, hf CLI and curl.
- Browser
- Download file 10.6 kB
-
https://huggingface.co/org2ai/Wald-4B/resolve/8bb838defd691d09bf7f6bb67499152284b9651f/evaluation/index.json
- Command line
-
hf download hf://org2ai/Wald-4B@8bb838defd691d09bf7f6bb67499152284b9651f/evaluation/index.json
-
curl -L -o index.json https://huggingface.co/org2ai/Wald-4B/resolve/8bb838defd691d09bf7f6bb67499152284b9651f/evaluation/index.json
10.6 kB
| { | |
| "edition": "0.2.1", | |
| "panel_id": "decision-index-0.2.1", | |
| "index": 54.59, | |
| "raw_index": 65.9, | |
| "scores": { | |
| "balanced_skill": 54.59, | |
| "balanced_raw": 65.9, | |
| "breadth_skill": 52.62 | |
| }, | |
| "areas": [ | |
| { | |
| "id": "knowledge", | |
| "label": "Knowledge & Reasoning", | |
| "raw": 0.539, | |
| "skill": 0.4253, | |
| "coverage": 1.0, | |
| "n": 10, | |
| "benchmarks": [ | |
| 25, | |
| 30, | |
| 31, | |
| 32, | |
| 33, | |
| 43, | |
| 44, | |
| 45, | |
| 57, | |
| 58 | |
| ] | |
| }, | |
| { | |
| "id": "language", | |
| "label": "Language Understanding", | |
| "raw": 0.7533, | |
| "skill": 0.6283, | |
| "coverage": 1.0, | |
| "n": 10, | |
| "benchmarks": [ | |
| 11, | |
| 12, | |
| 28, | |
| 29, | |
| 38, | |
| 39, | |
| 40, | |
| 41, | |
| 42, | |
| 59 | |
| ] | |
| }, | |
| { | |
| "id": "retrieval", | |
| "label": "Retrieval & Classification", | |
| "raw": 0.6454, | |
| "skill": 0.5067, | |
| "coverage": 1.0, | |
| "n": 6, | |
| "benchmarks": [ | |
| 4, | |
| 5, | |
| 36, | |
| 37, | |
| 56, | |
| 61 | |
| ] | |
| }, | |
| { | |
| "id": "tools", | |
| "label": "Tools & Automation", | |
| "raw": 0.821, | |
| "skill": 0.7949, | |
| "coverage": 1.0, | |
| "n": 5, | |
| "benchmarks": [ | |
| 1, | |
| 2, | |
| 3, | |
| 9, | |
| 62 | |
| ] | |
| }, | |
| { | |
| "id": "arts", | |
| "label": "Arts & Human Taste", | |
| "raw": 0.4569, | |
| "skill": 0.268, | |
| "coverage": 1.0, | |
| "n": 7, | |
| "benchmarks": [ | |
| 20, | |
| 21, | |
| 22, | |
| 23, | |
| 48, | |
| 50, | |
| 64 | |
| ] | |
| } | |
| ], | |
| "benchmarks": { | |
| "1": { | |
| "raw": 0.938, | |
| "skill": 0.9163, | |
| "coverage": 1.0, | |
| "random": 0.2592, | |
| "rule": "track", | |
| "in_index": true, | |
| "tracks": [] | |
| }, | |
| "2": { | |
| "raw": 0.66, | |
| "skill": 0.6073, | |
| "coverage": 1.0, | |
| "random": 0.1341, | |
| "rule": "track", | |
| "in_index": true, | |
| "tracks": [] | |
| }, | |
| "3": { | |
| "raw": 0.7874, | |
| "skill": 0.7833, | |
| "coverage": 1.0, | |
| "random": 0.0189, | |
| "rule": "chance", | |
| "in_index": true | |
| }, | |
| "4": { | |
| "raw": 0.7745, | |
| "skill": 0.7716, | |
| "coverage": 1.0, | |
| "random": 0.0127, | |
| "rule": "chance", | |
| "in_index": true | |
| }, | |
| "5": { | |
| "raw": 0.8488, | |
| "skill": 0.8479, | |
| "coverage": 1.0, | |
| "random": 0.006, | |
| "rule": "chance", | |
| "in_index": true | |
| }, | |
| "9": { | |
| "raw": 0.875, | |
| "skill": 0.875, | |
| "coverage": 1.0, | |
| "random": 0.0, | |
| "rule": "chance", | |
| "in_index": true | |
| }, | |
| "11": { | |
| "raw": 0.761, | |
| "skill": 0.6543, | |
| "coverage": 1.0, | |
| "random": 0.3085, | |
| "rule": "track", | |
| "in_index": true, | |
| "tracks": [] | |
| }, | |
| "12": { | |
| "raw": 0.6881, | |
| "skill": 0.5328, | |
| "coverage": 1.0, | |
| "random": 0.3324, | |
| "rule": "chance", | |
| "in_index": true | |
| }, | |
| "20": { | |
| "raw": 0.8322, | |
| "skill": 0.6644, | |
| "coverage": 1.0, | |
| "random": 0.5, | |
| "rule": "track", | |
| "in_index": true, | |
| "tracks": [] | |
| }, | |
| "21": { | |
| "raw": 0.6005, | |
| "skill": 0.2009, | |
| "coverage": 1.0, | |
| "random": 0.5, | |
| "rule": "track", | |
| "in_index": true, | |
| "tracks": [] | |
| }, | |
| "22": { | |
| "raw": 0.042, | |
| "skill": 0.0346, | |
| "coverage": 1.0, | |
| "random": 0.0078, | |
| "rule": "track", | |
| "in_index": true, | |
| "tracks": [] | |
| }, | |
| "23": { | |
| "raw": 0.5803, | |
| "skill": 0.1606, | |
| "coverage": 1.0, | |
| "random": 0.5, | |
| "rule": "track", | |
| "in_index": true, | |
| "tracks": [] | |
| }, | |
| "25": { | |
| "raw": 0.5102, | |
| "skill": 0.3469, | |
| "coverage": 1.0, | |
| "random": 0.25, | |
| "rule": "track", | |
| "in_index": true, | |
| "tracks": [] | |
| }, | |
| "28": { | |
| "raw": 0.7916, | |
| "skill": 0.5832, | |
| "coverage": 1.0, | |
| "random": 0.5, | |
| "rule": "chance", | |
| "in_index": true | |
| }, | |
| "29": { | |
| "raw": 0.8992, | |
| "skill": 0.8656, | |
| "coverage": 1.0, | |
| "random": 0.25, | |
| "rule": "chance", | |
| "in_index": true | |
| }, | |
| "30": { | |
| "raw": 0.9151, | |
| "skill": 0.8982, | |
| "coverage": 1.0, | |
| "random": 0.25, | |
| "rule": "track", | |
| "in_index": true, | |
| "tracks": [ | |
| { | |
| "track": "GSM8K-4choice", | |
| "score": 0.9333, | |
| "headline": false | |
| }, | |
| { | |
| "track": "GSM8K-10choice", | |
| "score": 0.8969, | |
| "headline": false | |
| } | |
| ] | |
| }, | |
| "31": { | |
| "raw": 0.1076, | |
| "skill": 0.028, | |
| "coverage": 1.0, | |
| "random": 0.0819, | |
| "rule": "track", | |
| "in_index": true, | |
| "tracks": [] | |
| }, | |
| "32": { | |
| "raw": 0.5758, | |
| "skill": 0.3256, | |
| "coverage": 1.0, | |
| "random": 0.371, | |
| "rule": "chance", | |
| "in_index": true | |
| }, | |
| "33": { | |
| "raw": 0.2503, | |
| "skill": 0.2403, | |
| "coverage": 1.0, | |
| "random": 0.0131, | |
| "rule": "chance", | |
| "in_index": true | |
| }, | |
| "36": { | |
| "raw": 0.4021, | |
| "skill": 0.3236, | |
| "coverage": 1.0, | |
| "random": 0.116, | |
| "rule": "track", | |
| "in_index": true, | |
| "tracks": [] | |
| }, | |
| "37": { | |
| "raw": 0.5261, | |
| "skill": 0.4056, | |
| "coverage": 1.0, | |
| "random": 0.2027, | |
| "rule": "track", | |
| "in_index": true, | |
| "tracks": [] | |
| }, | |
| "38": { | |
| "raw": 0.5652, | |
| "skill": 0.5513, | |
| "coverage": 1.0, | |
| "random": 0.031, | |
| "rule": "chance", | |
| "in_index": true | |
| }, | |
| "39": { | |
| "raw": 0.8918, | |
| "skill": 0.8409, | |
| "coverage": 1.0, | |
| "random": 0.3201, | |
| "rule": "chance", | |
| "in_index": true | |
| }, | |
| "40": { | |
| "raw": 0.5714, | |
| "skill": 0.4487, | |
| "coverage": 1.0, | |
| "random": 0.2227, | |
| "rule": "track", | |
| "in_index": true, | |
| "tracks": [ | |
| { | |
| "track": "A · Arabic", | |
| "score": 0.4873, | |
| "headline": false | |
| }, | |
| { | |
| "track": "A · English", | |
| "score": 0.5714, | |
| "headline": true | |
| }, | |
| { | |
| "track": "C · Arabic pairs", | |
| "score": 0.675, | |
| "headline": false | |
| }, | |
| { | |
| "track": "C · English pairs", | |
| "score": 0.905, | |
| "headline": false | |
| } | |
| ] | |
| }, | |
| "41": { | |
| "raw": 0.7784, | |
| "skill": 0.6675, | |
| "coverage": 1.0, | |
| "random": 0.3333, | |
| "rule": "track", | |
| "in_index": true, | |
| "tracks": [] | |
| }, | |
| "42": { | |
| "raw": 0.7894, | |
| "skill": 0.5903, | |
| "coverage": 1.0, | |
| "random": 0.486, | |
| "rule": "chance", | |
| "in_index": true | |
| }, | |
| "43": { | |
| "raw": 0.7825, | |
| "skill": 0.6548, | |
| "coverage": 1.0, | |
| "random": 0.3697, | |
| "rule": "track", | |
| "in_index": true, | |
| "tracks": [] | |
| }, | |
| "44": { | |
| "raw": 0.7442, | |
| "skill": 0.4884, | |
| "coverage": 1.0, | |
| "random": 0.5, | |
| "rule": "track", | |
| "in_index": true, | |
| "tracks": [] | |
| }, | |
| "45": { | |
| "raw": 0.0998, | |
| "skill": 0.0, | |
| "coverage": 1.0, | |
| "random": 0.1641, | |
| "rule": "chance", | |
| "in_index": true | |
| }, | |
| "48": { | |
| "raw": 0.1868, | |
| "skill": 0.1868, | |
| "coverage": 1.0, | |
| "random": 0.25, | |
| "rule": "vs baseline", | |
| "in_index": true | |
| }, | |
| "50": { | |
| "raw": 0.4123, | |
| "skill": 0.147, | |
| "coverage": 1.0, | |
| "random": 0.311, | |
| "rule": "track", | |
| "in_index": true, | |
| "tracks": [] | |
| }, | |
| "56": { | |
| "raw": 0.5225, | |
| "skill": 0.045, | |
| "coverage": 1.0, | |
| "random": 0.5, | |
| "rule": "chance", | |
| "in_index": true | |
| }, | |
| "57": { | |
| "raw": 0.6506, | |
| "skill": 0.607, | |
| "coverage": 1.0, | |
| "random": 0.1109, | |
| "rule": "chance", | |
| "in_index": true | |
| }, | |
| "58": { | |
| "raw": 0.7774, | |
| "skill": 0.6773, | |
| "coverage": 1.0, | |
| "random": 0.3101, | |
| "rule": "chance", | |
| "in_index": true | |
| }, | |
| "59": { | |
| "raw": 0.773, | |
| "skill": 0.5293, | |
| "coverage": 1.0, | |
| "random": 0.5177, | |
| "rule": "chance", | |
| "in_index": true | |
| }, | |
| "61": { | |
| "raw": 0.7808, | |
| "skill": 0.5616, | |
| "coverage": 1.0, | |
| "random": 0.5, | |
| "rule": "chance", | |
| "in_index": true | |
| }, | |
| "62": { | |
| "raw": 0.828, | |
| "skill": 0.7707, | |
| "coverage": 1.0, | |
| "random": 0.25, | |
| "rule": "chance", | |
| "in_index": true | |
| }, | |
| "64": { | |
| "raw": 0.5985, | |
| "skill": 0.4981, | |
| "coverage": 1.0, | |
| "random": 0.2, | |
| "rule": "chance", | |
| "in_index": true | |
| }, | |
| "6": { | |
| "raw": 0.7667, | |
| "skill": 0.453, | |
| "coverage": 1.0, | |
| "random": 0.5735, | |
| "rule": "shown, not counted", | |
| "in_index": false | |
| }, | |
| "10": { | |
| "raw": 0.0748, | |
| "skill": 0.0, | |
| "coverage": 1.0, | |
| "random": 0.399, | |
| "rule": "shown, not counted", | |
| "in_index": false | |
| }, | |
| "24": { | |
| "raw": 0.7918, | |
| "skill": 0.7224, | |
| "coverage": 1.0, | |
| "random": 0.25, | |
| "rule": "shown, not counted", | |
| "in_index": false | |
| }, | |
| "26": { | |
| "raw": 0.9827, | |
| "skill": 0.9769, | |
| "coverage": 1.0, | |
| "random": 0.2502, | |
| "rule": "shown, not counted", | |
| "in_index": false | |
| }, | |
| "27": { | |
| "raw": 0.9608, | |
| "skill": 0.9477, | |
| "coverage": 1.0, | |
| "random": 0.2502, | |
| "rule": "shown, not counted", | |
| "in_index": false | |
| }, | |
| "34": { | |
| "raw": 0.1, | |
| "skill": 0.0, | |
| "coverage": 1.0, | |
| "random": 0.1667, | |
| "rule": "shown, not counted", | |
| "in_index": false | |
| } | |
| }, | |
| "coverage": 1.0, | |
| "note": "Decision Index 0.2.1 averages 38 benchmarks in five areas. Arts & Human Taste weighs 10%; the other four share 90% in proportion to the square root of their benchmark count (knowledge 25.8%, language 25.8%, retrieval 20.0%, tools 18.3%). Inside an area, gold ★ benchmarks weigh 1.2 and the rest 1.0; the index is 100 x the weighted mean of the five areas. Each benchmark is chance-corrected first, (score - chance) / (1 - chance) clipped to 0-1, so 0 means random guessing and 100 means perfect. Every score is coverage-adjusted, so an unanswered or unsupported request counts as wrong. ForecastBench enters against its baseline: clip((0.25 - Brier) / 0.25) x coverage, so always predicting 0.5 scores zero. MMLU, ARC-Easy, ARC-Challenge, RouterBench, SGD stay on the board as non-index benchmarks. The six interactive environments are still unrun and stay out. Every entrant on the board has results on all 38 index benchmarks. Point estimates only, no uncertainty intervals yet." | |
| } | |