Instructions to use quaedra/Welp-35B-A3B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use quaedra/Welp-35B-A3B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf quaedra/Welp-35B-A3B-GGUF # Run inference directly in the terminal: llama cli -hf quaedra/Welp-35B-A3B-GGUF
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf quaedra/Welp-35B-A3B-GGUF # Run inference directly in the terminal: llama cli -hf quaedra/Welp-35B-A3B-GGUF
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf quaedra/Welp-35B-A3B-GGUF # Run inference directly in the terminal: ./llama-cli -hf quaedra/Welp-35B-A3B-GGUF
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf quaedra/Welp-35B-A3B-GGUF # Run inference directly in the terminal: ./build/bin/llama-cli -hf quaedra/Welp-35B-A3B-GGUF
Use Docker
docker model run hf.co/quaedra/Welp-35B-A3B-GGUF
- LM Studio
- Jan
- vLLM
How to use quaedra/Welp-35B-A3B-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "quaedra/Welp-35B-A3B-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "quaedra/Welp-35B-A3B-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/quaedra/Welp-35B-A3B-GGUF
- Ollama
How to use quaedra/Welp-35B-A3B-GGUF with Ollama:
ollama run hf.co/quaedra/Welp-35B-A3B-GGUF
- Unsloth Desktop
- Pi
How to use quaedra/Welp-35B-A3B-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf quaedra/Welp-35B-A3B-GGUF
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "quaedra/Welp-35B-A3B-GGUF" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use quaedra/Welp-35B-A3B-GGUF with Docker Model Runner:
docker model run hf.co/quaedra/Welp-35B-A3B-GGUF
- Lemonade
How to use quaedra/Welp-35B-A3B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull quaedra/Welp-35B-A3B-GGUF
Run and chat with the model
lemonade run user.Welp-35B-A3B-GGUF-{{QUANT_TAG}}List all available models
lemonade list
- Hermes Agent
How to use quaedra/Welp-35B-A3B-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf quaedra/Welp-35B-A3B-GGUF
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default quaedra/Welp-35B-A3B-GGUF
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use quaedra/Welp-35B-A3B-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf quaedra/Welp-35B-A3B-GGUF
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "quaedra/Welp-35B-A3B-GGUF" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Welp-35B-A3B (GGUF)
A 13.21 GB GGUF of Qwen/Qwen3.6-35B-A3B that runs on a 16 GB GPU with the full 262k context, and scores higher than Unsloth's UD-IQ3_XXS at exactly the same size.
It uses only standard llama.cpp formats, the same ones as Unsloth's UD-IQ3_XXS, so it should run wherever that file runs. Tested with llama.cpp (PrismML build b10754 and a patched build of the same).
No fine-tuning or new training data: the same model, quantized more carefully.
File
Welp-35B-A3B.gguf, 13.21 GB. SHA-256 in SHA256SUMS.
| tensors | type |
|---|---|
| routed experts gate/up | IQ2_S |
| routed experts down | IQ3_XXS (IQ4_XS in 3 layers) |
| attention, DeltaNet, shared expert, embeddings, output | Q6_K |
This is the same per-tensor type mix as Unsloth's UD-IQ3_XXS. Only the expert weights are quantized differently (see below); everything else is byte-identical to Unsloth's file.
Running
Full 262k context on a 16 GB GPU (14.8 GB used, including about 0.8 GB for the desktop):
llama-server -m Welp-35B-A3B.gguf -ngl 99 -fa on -c 262144 -ctk q4_0 -ctv q4_0 -ub 256 --jinja
For a more precise KV cache at half the context: -c 131072 -ctk q8_0 -ctv q8_0.
Results
Models that fit entirely on a 16 GB GPU, measured on one RTX 4080 Super 16 GB with the same harness, single runs.
| model | size | HumanEval | long-exact suite | context on 16 GB | decode tok/s | prefill tok/s (4k) |
|---|---|---|---|---|---|---|
| Welp-35B-A3B | 13.21 GB | 95.7% | 23/37 | 262k | 164 | 5,111 |
| Unsloth UD-IQ3_XXS | 13.21 GB | 93.9% | 21/37 | 262k | 165 | 5,165 |
| Qwen3.8-27B (Unsloth UD-Q3_K_XL) | 13.15 GB | 96.3% | 26/37 | 131k | 47 | 2,155 |
| Bonsai 2 27B | 5.95 GB | 91.5% | 15/37 | 262k | 85 | 2,278 |
- HumanEval: 164 problems, thinking off, temperature 0, max 1024 tokens.
- Long-exact suite: bonsai-ada-surgery
suite/default plan, 37 long tool-using tasks, thinking on. Bonsai 2 scores 15/37 here against 17/37 in that repo's report (RTX 4070). - Context: the largest that fits entirely on the card with a q4_0 KV cache.
- Decode: llama-bench tg128. Prefill: llama-bench pp4096.
- Perplexity is only comparable between Welp and UD-IQ3_XXS, which share the base model: wikitext-2 5.762 vs 5.873, code 1.880 vs 1.900 (context 2048, 40 chunks).
- Qwen3.8-27B is the dense model Bonsai 2 is built from. Welp solves one fewer HumanEval problem (157 vs 158) and three fewer long-exact tasks (23 vs 26), at 3.5 times its decode speed and twice its context. The HumanEval and suite differences against UD-IQ3_XXS are small enough to be run-to-run noise on their own; the perplexity gain is consistent.
HumanEval pass@1
Decode speed (tok/s, llama-bench tg128)
Prefill speed (tok/s, llama-bench pp4096)
HumanEval against GGUF size (log scale). The dashed line joins the best Qwen3.6-35B-A3B quantization at each size; Welp moves it up at 13.21 GB and ties Q4_K_M (157 of 164) at 62% of its size. The two ternary points are research builds from the Welp log and are not released.
At the full 262k context (q4_0 KV), after a 214k-token prompt: decode 107 tok/s, prefill 2,043 tok/s. A passcode hidden at 10%, 50% and 90% of that prompt was retrieved 3/3.
Speed against context
Decode tok/s (128 tokens) and prefill tok/s (next 2,048 tokens) after the given amount of context, llama-bench -d, q4_0 KV and -ub 256 for every model, RTX 4080 Super 16 GB, mean of 2 runs.
Decode speed against context
Prefill speed against context
| context | Welp-35B-A3B decode | prefill | Bonsai 2 27B decode | prefill | Qwen3.8-27B decode | prefill |
|---|---|---|---|---|---|---|
| 0 | 161.5 | 3,826 | 83.9 | 2,335 | 46.6 | 2,132 |
| 4k | 161.8 | 3,667 | 83.8 | 2,247 | 46.4 | 2,044 |
| 16k | 155.1 | 3,301 | 80.4 | 1,952 | 45.3 | 1,791 |
| 32k | 151.8 | 3,063 | 76.5 | 1,660 | 44.1 | 1,550 |
| 64k | 141.1 | 2,637 | 69.4 | 1,285 | 41.6 | 1,220 |
| 128k | 126.9 | 2,029 | 58.8 | 882 | 37.5 | 851 |
| 192k | 115.2 | 1,647 | 50.9 | 672 | does not fit | |
| 252k | 105.0 | 1,396 | 45.1 | 548 |
- Welp keeps 65% of its decode speed at 252k; at 252k it still decodes faster than either 27B model with an empty cache. Only 10 of its 40 layers use full attention (the rest are Gated DeltaNet linear attention), and 3B of its parameters are active per token.
-ub 256is the batch size Welp needs for 262k on 16 GB, so prefill here is lower than the prefill column in the table above (pp4096, default batch). UD-IQ3_XXS has the same architecture and formats as Welp and runs at the same speed.
How it was made
Each expert matrix is quantized one 256-weight block at a time with llama.cpp's own quantizer, using that expert's activation statistics as the importance matrix. After each block, its rounding error is pushed into the columns not yet quantized (GPTQ, applied block-wise so any llama.cpp format works). Layers are processed in order, each fed the output of the already quantized layers. Calibration: 256 sequences of 2048 tokens, half code, half wikitext-2 train.
Recipe and scripts: github.com/quaedra/welp.
Credits
Base model by the Qwen team. The per-tensor type mix follows Unsloth's UD-IQ3_XXS, whose non-expert tensors this file reuses. Quantization formats from llama.cpp.
License
Apache 2.0, same as the base model.
- Downloads last month
- 65
We're not able to determine the quantization variants.
Model tree for quaedra/Welp-35B-A3B-GGUF
Base model
Qwen/Qwen3.6-35B-A3B




