Instructions to use unsloth/Qwen3.5-35B-A3B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use unsloth/Qwen3.5-35B-A3B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf unsloth/Qwen3.5-35B-A3B-GGUF:UD-Q4_K_XL # Run inference directly in the terminal: llama cli -hf unsloth/Qwen3.5-35B-A3B-GGUF:UD-Q4_K_XL
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf unsloth/Qwen3.5-35B-A3B-GGUF:UD-Q4_K_XL # Run inference directly in the terminal: llama cli -hf unsloth/Qwen3.5-35B-A3B-GGUF:UD-Q4_K_XL
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf unsloth/Qwen3.5-35B-A3B-GGUF:UD-Q4_K_XL # Run inference directly in the terminal: ./llama-cli -hf unsloth/Qwen3.5-35B-A3B-GGUF:UD-Q4_K_XL
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf unsloth/Qwen3.5-35B-A3B-GGUF:UD-Q4_K_XL # Run inference directly in the terminal: ./build/bin/llama-cli -hf unsloth/Qwen3.5-35B-A3B-GGUF:UD-Q4_K_XL
Use Docker
docker model run hf.co/unsloth/Qwen3.5-35B-A3B-GGUF:UD-Q4_K_XL
- LM Studio
- Jan
- vLLM
How to use unsloth/Qwen3.5-35B-A3B-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "unsloth/Qwen3.5-35B-A3B-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "unsloth/Qwen3.5-35B-A3B-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/unsloth/Qwen3.5-35B-A3B-GGUF:UD-Q4_K_XL
- Ollama
How to use unsloth/Qwen3.5-35B-A3B-GGUF with Ollama:
ollama run hf.co/unsloth/Qwen3.5-35B-A3B-GGUF:UD-Q4_K_XL
- Unsloth Desktop
- Pi
How to use unsloth/Qwen3.5-35B-A3B-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf unsloth/Qwen3.5-35B-A3B-GGUF:UD-Q4_K_XL
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "unsloth/Qwen3.5-35B-A3B-GGUF:UD-Q4_K_XL" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use unsloth/Qwen3.5-35B-A3B-GGUF with Docker Model Runner:
docker model run hf.co/unsloth/Qwen3.5-35B-A3B-GGUF:UD-Q4_K_XL
- Lemonade
How to use unsloth/Qwen3.5-35B-A3B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull unsloth/Qwen3.5-35B-A3B-GGUF:UD-Q4_K_XL
Run and chat with the model
lemonade run user.Qwen3.5-35B-A3B-GGUF-UD-Q4_K_XL
List all available models
lemonade list
- Hermes Agent
How to use unsloth/Qwen3.5-35B-A3B-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf unsloth/Qwen3.5-35B-A3B-GGUF:UD-Q4_K_XL
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default unsloth/Qwen3.5-35B-A3B-GGUF:UD-Q4_K_XL
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use unsloth/Qwen3.5-35B-A3B-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf unsloth/Qwen3.5-35B-A3B-GGUF:UD-Q4_K_XL
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "unsloth/Qwen3.5-35B-A3B-GGUF:UD-Q4_K_XL" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
llama cpp Error: Unknown (built-in) filter 'items' for type String
srv log_server_r: done request: POST /v1/chat/completions 127.0.0.1 500
Template supports tool calls but does not natively describe tools. The fallback behaviour used may produce bad results, inspect prompt w/ --verbose & consider overriding the template.
srv operator(): got exception: {"error":{"code":500,"message":"\n------------\nWhile executing FilterExpression at line 120, column 73 in source:\n..._name, args_value in tool_call.arguments|items %}β΅ {{- '<...\n ^\nError: Unknown (built-in) filter 'items' for type String","type":"server_error"}}
srv log_server_r: done request: POST /v1/chat/completions 127.0.0.1 500
I am getting this error from presumably the prompt template in this repo
I had the exact same issue.
It was solved by updating my llama.cpp image.
Hi there, please re-download the quants and update llama.cpp image! @fullstack
This should fix it: https://github.com/ggml-org/llama.cpp/pull/19870
llama_model_load: error loading model: error loading model architecture: unknown model architecture: 'qwen35moe'
llama_model_load_from_file_impl: failed to load model
common_init_from_params: failed to load model '/LLM/Qwen3.5-35B-A3B-UD-Q4_K_XL.gguf'
srv load_model: failed to load model, '/LLM/Qwen3.5-35B-A3B-UD-Q4_K_XL.gguf'
I am using llama.cpp b8145 with Vulkan backend.
.\llama-server.exe --port 9999 --device CUDA0 -ngl 99 --temp 0.6 --min-p 0.0 --top-k 20 --top-p 0.95 --jinja -ub 2048 -b 2048 -fa on -m D:\Qwen3.5-35B-A3B-UD-Q3_K_XL.gguf -c 65536 --alias local -ctk q8_0 -ctv q8_0 -t 12 --n-cpu-moe 30 -fit off
ggml_cuda_init: found 1 CUDA devices:
Device 0: NVIDIA GeForce RTX 4070, compute capability 8.9, VMM: yes
main: n_parallel is set to auto, using n_parallel = 4 and kv_unified = true
build: 8148 (244641955) with MSVC 19.38.33135.0 for x64
system info: n_threads = 12, n_threads_batch = 12, total_threads = 16
system_info: n_threads = 12 (n_threads_batch = 12) / 16 | CUDA : ARCHS = 890 | USE_GRAPHS = 1 | PEER_MAX_BATCH_SIZE = 128 | FA_ALL_QUANTS = 1 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | AVX512 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
init: using 15 threads for HTTP server
start: binding port with default address family
main: loading model
srv load_model: loading model 'D:\Qwen3.5-35B-A3B-UD-Q3_K_XL.gguf'
llama_model_load_from_file_impl: using device CUDA0 (NVIDIA GeForce RTX 4070) (0000:01:00.0) - 11090 MiB free
gguf_init_from_file_impl: failed to read magic
llama_model_load: error loading model: llama_model_loader: failed to load model from D:\Qwen3.5-35B-A3B-UD-Q3_K_XL.gguf
llama_model_load_from_file_impl: failed to load model
common_init_from_params: failed to load model 'D:\Qwen3.5-35B-A3B-UD-Q3_K_XL.gguf'
srv load_model: failed to load model, 'D:\Qwen3.5-35B-A3B-UD-Q3_K_XL.gguf'
srv operator (): operator (): cleaning up before exit...
main: exiting due to model loading error
latest llama.cpp build
Hello, could the backslash in path being involved ? try -m D:/Qwen3.5-35B-A3B-UD-Q3_K_XL.gguf instead maybe.
Also, to be sure that other parameters does not interfer you could use the new --fit on (witch is on by default i think).
.\llama-server.exe --port 9999 --device CUDA0 --fit on --temp 0.6 --min-p 0.0 --top-k 20 --top-p 0.95 --jinja -ub 2048 -b 2048 -fa on -m D:/Qwen3.5-35B-A3B-UD-Q3_K_XL.gguf
And probably a good idea to check the downloaded model with a checksum agains sha π
Good luck
++
@CHNtentes : see https://github.com/ggml-org/llama.cpp/issues/19868
Looks like your situation could be related.
@CHNtentes : see https://github.com/ggml-org/llama.cpp/issues/19868
Looks like your situation could be related.
Thanks for your help :)
https://github.com/ggml-org/llama.cpp/pull/19870 as well.
Looks like it may be addressed as seen with release tag b8149: https://github.com/ggml-org/llama.cpp/releases/tag/b8149
Looks like your running the llama.cpp version 8148, so you might be ok if you try versions b8149 and on.
https://github.com/ggml-org/llama.cpp/pull/19870 as well.
Looks like it may be addressed as seen with release tag b8149: https://github.com/ggml-org/llama.cpp/releases/tag/b8149
Looks like your running the llama.cpp version 8148, so you might be ok if you try versions b8149 and on.
it's working normally with latest version. performance with Q3_K_XL on 4070 12G + 32G DDR5:
short prompt:
prompt eval time = 464.82 ms / 13 tokens ( 35.76 ms per token, 27.97 tokens per second)
eval time = 5883.79 ms / 367 tokens ( 16.03 ms per token, 62.37 tokens per second)
long prompt:
prompt eval time = 12036.66 ms / 20649 tokens ( 0.58 ms per token, 1715.51 tokens per second)
eval time = 40254.51 ms / 2203 tokens ( 18.27 ms per token, 54.73 tokens per second)