Text Generation
Transformers
Safetensors
GGUF
English
qwen3_5_text
decision-model
typed-decisions
calibration
calibrated-probabilities
classification
tool-selection
tool-use
agent-routing
clarification
robustness
decision-index
jevbench
jev
jev-compatible
open-jev
typesafe-compatible
systemone
kev
laya
wald
wald-q4b
qwen3.5
4b
vllm
llama.cpp
ollama
reasoning
conversational
Eval Results (legacy)
Instructions to use org2ai/Wald-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use org2ai/Wald-4B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="org2ai/Wald-4B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("org2ai/Wald-4B") model = AutoModelForCausalLM.from_pretrained("org2ai/Wald-4B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use org2ai/Wald-4B with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf org2ai/Wald-4B:Q4_K_M # Run inference directly in the terminal: llama cli -hf org2ai/Wald-4B:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf org2ai/Wald-4B:Q4_K_M # Run inference directly in the terminal: llama cli -hf org2ai/Wald-4B:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf org2ai/Wald-4B:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf org2ai/Wald-4B:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf org2ai/Wald-4B:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf org2ai/Wald-4B:Q4_K_M
Use Docker
docker model run hf.co/org2ai/Wald-4B:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use org2ai/Wald-4B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "org2ai/Wald-4B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "org2ai/Wald-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/org2ai/Wald-4B:Q4_K_M
- SGLang
How to use org2ai/Wald-4B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "org2ai/Wald-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "org2ai/Wald-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "org2ai/Wald-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "org2ai/Wald-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use org2ai/Wald-4B with Ollama:
ollama run hf.co/org2ai/Wald-4B:Q4_K_M
- Unsloth Desktop
- Pi
How to use org2ai/Wald-4B with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf org2ai/Wald-4B:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "org2ai/Wald-4B:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use org2ai/Wald-4B with Docker Model Runner:
docker model run hf.co/org2ai/Wald-4B:Q4_K_M
- Lemonade
How to use org2ai/Wald-4B with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull org2ai/Wald-4B:Q4_K_M
Run and chat with the model
lemonade run user.Wald-4B-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use org2ai/Wald-4B with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf org2ai/Wald-4B:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default org2ai/Wald-4B:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use org2ai/Wald-4B with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf org2ai/Wald-4B:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "org2ai/Wald-4B:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Card: GGUF / llama.cpp / Ollama section (EN + ZH), tags; wald-serve 0.1.1 (llama.cpp backend, vLLM path unchanged); MANIFEST with the GGUF files
9662694 verified Download MANIFEST.json from org2ai/Wald-4B: direct link, hf CLI and curl.
- Browser
- Download file 8.89 kB
-
https://huggingface.co/org2ai/Wald-4B/resolve/main/MANIFEST.json
- Command line
-
hf download hf://org2ai/Wald-4B/MANIFEST.json
-
curl -L -o MANIFEST.json https://huggingface.co/org2ai/Wald-4B/resolve/main/MANIFEST.json
8.89 kB
| { | |
| "CONTAMINATION.md": { | |
| "bytes": 2047, | |
| "sha256": "108d33ec5188b608e8b293d18ba1bb975b9849638644a682616ca8816b4e1fd1" | |
| }, | |
| "Dockerfile": { | |
| "bytes": 1201, | |
| "sha256": "b353502e02205a000a48a4ea4478328e8461ac012f246b23c7cfcba572ac69f3" | |
| }, | |
| "LICENSE": { | |
| "bytes": 11358, | |
| "sha256": "c95bae1d1ce0235ecccd3560b772ec1efb97f348a79f0fbe0a634f0c2ccefe2c" | |
| }, | |
| "NOTICE": { | |
| "bytes": 1553, | |
| "sha256": "e1f68705275e9539d5d28e0dc20eae470169af869c308a90378db3ede83dd644" | |
| }, | |
| "PROVENANCE.md": { | |
| "bytes": 2077, | |
| "sha256": "69225fb813578b0d4f9e453a9f0c140c848ec354cf6cca877bcae828c082d3d7" | |
| }, | |
| "README.md": { | |
| "bytes": 29939, | |
| "sha256": "5f02a79558567e33186c0d1e4d462b70a0fa2d960080c2af27b2e08567eb8b98" | |
| }, | |
| "RUNBOOK.md": { | |
| "bytes": 4980, | |
| "sha256": "8ff79f40f2399cb21b08f27cacb223c88f419d5d55307259fca1166cab132639" | |
| }, | |
| "chat_template.jinja": { | |
| "bytes": 7756, | |
| "sha256": "a4aee8afcf2e0711942cf848899be66016f8d14a889ff9ede07bca099c28f715" | |
| }, | |
| "config.json": { | |
| "bytes": 1979, | |
| "sha256": "63f47812d0f11118e4d252d2b3ad488707eb9287a11589f4fd382a1d31182724" | |
| }, | |
| "docs/lora-cli.md": { | |
| "bytes": 7501, | |
| "sha256": "8b4be9b9256822ff6ba545d4cac874fa405846da9d466525e958f0bbc69507d7" | |
| }, | |
| "docs/observations/one-lora-all-verticals.md": { | |
| "bytes": 9088, | |
| "sha256": "e385b92e333eaaa9395a09daf26c09420645bee242456552a55500db45919bc7" | |
| }, | |
| "docs/readmes/README.zh.md": { | |
| "bytes": 25676, | |
| "sha256": "667a7db8482decc260fdf2538de52365f3e130984a173554fb7483e1f5f02d26" | |
| }, | |
| "evaluation/benchmark-summary.json": { | |
| "bytes": 36014, | |
| "sha256": "78983f2d4f3d2c708159879ff1e66b6da7a2a10d95f732417fe8b27c39f3c1ae" | |
| }, | |
| "evaluation/index.json": { | |
| "bytes": 10559, | |
| "sha256": "6f2ef43841fc619b3ba7bdf9d4d7b5fc7d17055048dc155bd45a61f92d3dbe4b" | |
| }, | |
| "evaluation/release-validation.json": { | |
| "bytes": 413, | |
| "sha256": "163475465b5d97d1c774d284ed637352beab136fa737b5c8d464e6069d92de89" | |
| }, | |
| "evaluation/scores.json": { | |
| "bytes": 33084, | |
| "sha256": "5bccdf55495daf82396123ca1062d084e910d603536f27f4950dea037b84122e" | |
| }, | |
| "evaluation/serial-latency-preflight-corrected.json": { | |
| "bytes": 336, | |
| "sha256": "db1da0eb255ea38105a2d8e76e221e0ae50b18b6ad37737a10787240e3f8e567" | |
| }, | |
| "generation_config.json": { | |
| "bytes": 116, | |
| "sha256": "62153eb6c69f2e1f426beaa8002b7186437e949c7588167085df14e10e9c0a73" | |
| }, | |
| "history/v1.0/CONTAMINATION.md": { | |
| "bytes": 26100, | |
| "sha256": "a4f840155d12964dd9fc1dde90611c18252bb55082844145b37cd00ffe8ab090" | |
| }, | |
| "history/v1.0/README.md": { | |
| "bytes": 10639, | |
| "sha256": "f468e58384b9978060038a43434a0ed6ce2973185dcb3ff7e059fdc0314bb1d2" | |
| }, | |
| "history/v1.0/RUNBOOK.md": { | |
| "bytes": 6268, | |
| "sha256": "ab03ccc6bfbd4a9a7772e579dd1a729eec9a3ea072ca8c54ded21018c1d15ec3" | |
| }, | |
| "model-00001-of-00002.safetensors": { | |
| "bytes": 4972947968, | |
| "sha256": "0b15067f769e7388bd598aabafa2ea3211a3a0c4e49e47d07c4f65c65ffc25dd" | |
| }, | |
| "model-00002-of-00002.safetensors": { | |
| "bytes": 3438610328, | |
| "sha256": "7cff102b314edeb3dc5abfa7063f72ad3e888884da237601239dc3582074120a" | |
| }, | |
| "model-files/serving.json": { | |
| "bytes": 128, | |
| "sha256": "6b4de526ca4c939b626932dcd10d6167e21a68b1e77a090d7a5a87c7c3a353b1" | |
| }, | |
| "model.safetensors.index.json": { | |
| "bytes": 41993, | |
| "sha256": "83620b40d724f8b286580b5c7c504222b0bdbdd7dc9f068e565132409c7eec9d" | |
| }, | |
| "reference/eval/__init__.py": { | |
| "bytes": 0, | |
| "sha256": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855" | |
| }, | |
| "reference/eval/systemone_vllm.py": { | |
| "bytes": 25186, | |
| "sha256": "881195545231b7b26c7753d4a43fdce5d65a2c49b50a78aec5eee69468e59136" | |
| }, | |
| "reference/kev/__init__.py": { | |
| "bytes": 0, | |
| "sha256": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855" | |
| }, | |
| "reference/kev/api.py": { | |
| "bytes": 7086, | |
| "sha256": "23bac9af0e79702428447c62662b88910071595f6171d9e3c82d6fa35fb0dbd1" | |
| }, | |
| "reference/midtrain/__init__.py": { | |
| "bytes": 0, | |
| "sha256": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855" | |
| }, | |
| "reference/midtrain/encode.py": { | |
| "bytes": 19164, | |
| "sha256": "fb0d35c055c7052a3a3b019cd2cb4de1857ca733467680fa6182a106b045ecec" | |
| }, | |
| "reference/midtrain/letter.py": { | |
| "bytes": 12651, | |
| "sha256": "e053b6dd1790818e8473fea4ad5350c65f9de53cd59e1a750a3bdff8d5085963" | |
| }, | |
| "reference/midtrain/rawprompt.py": { | |
| "bytes": 11856, | |
| "sha256": "5abe943f8c39e84018006fcea3fd32df1df89635eba1c87639b04057654a2ffa" | |
| }, | |
| "reference/midtrain/tokens.py": { | |
| "bytes": 3488, | |
| "sha256": "7fbc482959e405dcc537bfadf0fa24df2d89bcaaab5919bcb765a70ea81e21de" | |
| }, | |
| "run.sh": { | |
| "bytes": 915, | |
| "sha256": "d5b7b3e62d43c5f0dae4779f723b1eb4408aa9c21fcbdb6ed65e1da049dd050a" | |
| }, | |
| "server/README.md": { | |
| "bytes": 1678, | |
| "sha256": "2ca738dcacd770e7c490ceeb59f2f11df7515f5bc56784eb9d476af4c7145b4d" | |
| }, | |
| "server/pyproject.toml": { | |
| "bytes": 578, | |
| "sha256": "0fd8155bfa466e0f7a95fddce986ab881cce9260187c1620884fa27e5008dfbe" | |
| }, | |
| "server/src/wald_serve/__init__.py": { | |
| "bytes": 139, | |
| "sha256": "7db760d45fae0f1d9cea6dbd3e1421a2627a92020981ae3cd56ab57181423aac" | |
| }, | |
| "server/src/wald_serve/__main__.py": { | |
| "bytes": 51, | |
| "sha256": "540fd0a5992ca535482a9d44b2e268ce3a39bb47a0576f301214afa5ef4494b9" | |
| }, | |
| "server/src/wald_serve/engine.py": { | |
| "bytes": 14814, | |
| "sha256": "40fbea84c4c5b93ef888da44508d4fec2677c58732a09d731f0bcf17c2b48dd6" | |
| }, | |
| "server/src/wald_serve/prompt.py": { | |
| "bytes": 4192, | |
| "sha256": "bb45a10734f2546eb1ea1bd7e8b5717ecd8305a4aded0b816f80cd7e85188f7a" | |
| }, | |
| "server/src/wald_serve/server.py": { | |
| "bytes": 8878, | |
| "sha256": "49fb06ee52eb2ebe7059404432386d02e049556003acb7d2f6494ba92d4b9703" | |
| }, | |
| "server/src/wald_serve/wire.py": { | |
| "bytes": 3670, | |
| "sha256": "836ce54ce93016a98f94a6beaa508cc03982e24c3d3b4d62b8ef667389cd14af" | |
| }, | |
| "server/tests/fakevllm.py": { | |
| "bytes": 4167, | |
| "sha256": "9fb222a21bc2b17221e5a1a3605973b4f71730db9bbfede439bd3388b91e2570" | |
| }, | |
| "server/tests/test_parity.py": { | |
| "bytes": 4392, | |
| "sha256": "cd1e4142090076114f2b07efce49fdc5466e302a5c79996b993f3833a3457566" | |
| }, | |
| "server/tests/test_server.py": { | |
| "bytes": 13844, | |
| "sha256": "46f2007e9743bfeb0f7f79fca093cf444a70df1bfc348229f34cd251f88a19b3" | |
| }, | |
| "serving.json": { | |
| "bytes": 128, | |
| "sha256": "6b4de526ca4c939b626932dcd10d6167e21a68b1e77a090d7a5a87c7c3a353b1" | |
| }, | |
| "temperature.json": { | |
| "bytes": 4348, | |
| "sha256": "a4e9a1983a2246de34ce4c4b48e5d501daf7c66ae4f3654a79591f0ed0059da5" | |
| }, | |
| "tokenizer.json": { | |
| "bytes": 19989509, | |
| "sha256": "bd53432f0de26d67a83b634040d4f043053da4ecd0c759e7f4b24ad4f8bb9a81" | |
| }, | |
| "tokenizer_config.json": { | |
| "bytes": 1127, | |
| "sha256": "171ecbe7ddae98d11840698f7df2b8d5b4722139db0f0620d3bbf429bd656250" | |
| }, | |
| "tools/export_contamination.py": { | |
| "bytes": 5572, | |
| "sha256": "a5723b781031c0bde60349691b82a63f150aa408afaf5621a893bde6ee82f6f8" | |
| }, | |
| "tools/export_figure_data.py": { | |
| "bytes": 10121, | |
| "sha256": "ba889a4b90d59c9487cbca34768a5ea90fafb805c5f246124d1f3475ee067a43" | |
| }, | |
| "tools/make_figures.py": { | |
| "bytes": 19067, | |
| "sha256": "207d02dea21ee9a83763490247b4aa0fe146068cf1c4922a07e4647a3fc546f8" | |
| }, | |
| "tools/prepare_model_dir.py": { | |
| "bytes": 1890, | |
| "sha256": "0c1811fbfd98fc3a8c802298de3691b5039455aceef1475b73cfced72365708e" | |
| }, | |
| "tools/secret_scan.py": { | |
| "bytes": 1698, | |
| "sha256": "ad20de47de5bbe7f72417f3d87e6f72aae7987fab913eb814b818e4a4bbc6e1e" | |
| }, | |
| "model-info.json": { | |
| "bytes": 6145, | |
| "sha256": "ab370eedeaed830d92310b952993c99ff5cef4068a95d74ea7fa9ce6efb9f06d" | |
| }, | |
| "llms.txt": { | |
| "bytes": 3316, | |
| "sha256": "4aed0ca0177819e0f6f3b04b34f8021810d83c09c0e5b9a684188fd76c638dcb" | |
| }, | |
| "CITATION.cff": { | |
| "bytes": 1063, | |
| "sha256": "75f7dd132f937aee40b8867032bb10d961bbba15fb525744c67debc272e85d1d" | |
| }, | |
| "docs/api.md": { | |
| "bytes": 6069, | |
| "sha256": "9f6a43a76824c38285856c8fb1cee632e0fa5d3c830d450fbd1120258e54997d" | |
| }, | |
| "evaluation/v1.2/summary.json": { | |
| "bytes": 6991, | |
| "sha256": "1c4231985191eb1f7916baf13dd621d3e5655f192dbad2f842f998f8ba8ffaf5" | |
| }, | |
| "evaluation/v1.2/release-check.json": { | |
| "bytes": 3036, | |
| "sha256": "0de642d7c1ddb2ececafce16c06ae2d2eb0e663ec94b4afcfab605a61aadfb70" | |
| }, | |
| "Wald-4B-v1.2-Q4_K_M.gguf": { | |
| "bytes": 2708804000, | |
| "sha256": "e843f7658793b9f3effb3a4941cd67f80707a07b468aaaf94d440613751d76fa" | |
| }, | |
| "Wald-4B-v1.2-Q5_K_M.gguf": { | |
| "bytes": 3074986400, | |
| "sha256": "6fecc4e7655adb71e8c642ce5018f967564ba259852d562ecd3109386542b0cb" | |
| }, | |
| "Wald-4B-v1.2-Q6_K.gguf": { | |
| "bytes": 3464055200, | |
| "sha256": "5870f15f3b2eef773f672e56bb36e343a95e32fd09baaf7c9ce8b0da1e5dd897" | |
| }, | |
| "Wald-4B-v1.2-Q8_0.gguf": { | |
| "bytes": 4482402720, | |
| "sha256": "09c02494ebdfb7d088a2c76726eb2b290a438dbc5658c3be16da8333fddd6038" | |
| } | |
| } | |