Text Generation
GGUF
English
llama.cpp
quantized
empero-ai
qwen3.6
qwen3.8
distillation
reasoning
Mixture of Experts
gated-deltanet
conversational
Instructions to use empero-ai/Qwen3.8-35B-A3B-Distill-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use empero-ai/Qwen3.8-35B-A3B-Distill-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf empero-ai/Qwen3.8-35B-A3B-Distill-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf empero-ai/Qwen3.8-35B-A3B-Distill-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf empero-ai/Qwen3.8-35B-A3B-Distill-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf empero-ai/Qwen3.8-35B-A3B-Distill-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf empero-ai/Qwen3.8-35B-A3B-Distill-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf empero-ai/Qwen3.8-35B-A3B-Distill-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf empero-ai/Qwen3.8-35B-A3B-Distill-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf empero-ai/Qwen3.8-35B-A3B-Distill-GGUF:Q4_K_M
Use Docker
docker model run hf.co/empero-ai/Qwen3.8-35B-A3B-Distill-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use empero-ai/Qwen3.8-35B-A3B-Distill-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "empero-ai/Qwen3.8-35B-A3B-Distill-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "empero-ai/Qwen3.8-35B-A3B-Distill-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/empero-ai/Qwen3.8-35B-A3B-Distill-GGUF:Q4_K_M
- Ollama
How to use empero-ai/Qwen3.8-35B-A3B-Distill-GGUF with Ollama:
ollama run hf.co/empero-ai/Qwen3.8-35B-A3B-Distill-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use empero-ai/Qwen3.8-35B-A3B-Distill-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf empero-ai/Qwen3.8-35B-A3B-Distill-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "empero-ai/Qwen3.8-35B-A3B-Distill-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use empero-ai/Qwen3.8-35B-A3B-Distill-GGUF with Docker Model Runner:
docker model run hf.co/empero-ai/Qwen3.8-35B-A3B-Distill-GGUF:Q4_K_M
- Lemonade
How to use empero-ai/Qwen3.8-35B-A3B-Distill-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull empero-ai/Qwen3.8-35B-A3B-Distill-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Qwen3.8-35B-A3B-Distill-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use empero-ai/Qwen3.8-35B-A3B-Distill-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf empero-ai/Qwen3.8-35B-A3B-Distill-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default empero-ai/Qwen3.8-35B-A3B-Distill-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use empero-ai/Qwen3.8-35B-A3B-Distill-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf empero-ai/Qwen3.8-35B-A3B-Distill-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "empero-ai/Qwen3.8-35B-A3B-Distill-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
little tests
#3
by KottCh - opened
great job! I did some little tests using Ollama or LLamaCpp
./llama-cli -m /blobs/sha256-7d986d310e686a91cb514cdd819719b4f80ace899c9aaad7bfccfdf4670a9bb3 --temp 0.6 --top-p 0.95 --top-k 20 --seed 42 --ctx-size 262144
build : b11053-1af554f8f
ftype : Q8_0
> just split this string into 4 charachter chunks: T3JuaXRoLTEuNS0zNUItQTNCLUdHVUY6UTgw123234345456
[Start thinking]
The user wants me to split the string "T3JuaXRoLTEuNS0zNUItQTNCLUdHVUY6UTgw123234345456" into 4-character chunks.
Let me do this carefully:
T3Ju aXRo LTEu NS0z NUIt QTNB UgHV Y6UT gw12 3234 3454 56
Let me recount:
T3Ju (4)
...
56 (2)
Let me verify the original string character by character:
T 3 J u a X R o L T E u N S 0 z N U I t Q T N B U g H V Y 6 U T g w 1 2 3 2 3 4 3 4 5 4 5 6
That's 46 characters. 46 / 4 = 11 remainder 2. So 11 full chunks of 4 + 1 chunk of 2.
Let me recount the original string:
T-3-J-u-a-X-R-o-L-T-E-u-N-S-0-z-N-U-I-t-Q-T-N-B-U-g-H-V-Y-6-U-T-g-w-1-2-3-2-3-4-3-4-5-4-5-6
Count: 1-2-3-4-5-6-7-8-9-10-11-12-13-14-15-16-17-18-19-20-21-22-23-24-25-26-27-28-29-30-31-32-33-34-35-36-37-38-39-40-41-42-43-44-45-46
Yes, 46 characters. So 11 chunks of 4 + 1 chunk of 2.
Chunks:
T3Ju
...
56
[End thinking]
Here's the string split into 4-character chunks:
`
T3Ju
aXRo
LTEu
NS0z
NUIt
QTNB
UgHV
Y6UT
gw12
3234
3454
56
`
The original string is 46 characters long, so you get 11 full chunks of 4 characters and 1 final chunk of 2 characters.
[ Prompt: 24.3 t/s | Generation: 20.7 t/s ]
with --temp 0.7only 3 char damaged but size kept:
[End thinking]
Here's the string split into 4-character chunks:
`
T3Ju
aXRo
LTEu
NS0z
NUIt
QTNB
UdHV
VUY6
UTgw
1232
3434
5456
`
12 chunks total (11 × 4 chars + 1 × 3 chars = 47 characters).
with --temp 0.72 were able to output correct chars. but logic loops. QTNC is struggle for it.
No. I need to stop. The answer is:
`
T3Ju
aXRo
LTEu
NS0z
NUIt
QTN C
LUdH
VUY6
UTgw
1232
3434
5456
`
I'll just provide the answer cleanly now.