Instructions to use ggml-org/Qwen3.8-27B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ggml-org/Qwen3.8-27B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ggml-org/Qwen3.8-27B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf ggml-org/Qwen3.8-27B-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ggml-org/Qwen3.8-27B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf ggml-org/Qwen3.8-27B-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ggml-org/Qwen3.8-27B-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf ggml-org/Qwen3.8-27B-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ggml-org/Qwen3.8-27B-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf ggml-org/Qwen3.8-27B-GGUF:Q4_K_M
Use Docker
docker model run hf.co/ggml-org/Qwen3.8-27B-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use ggml-org/Qwen3.8-27B-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ggml-org/Qwen3.8-27B-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ggml-org/Qwen3.8-27B-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/ggml-org/Qwen3.8-27B-GGUF:Q4_K_M
- Ollama
How to use ggml-org/Qwen3.8-27B-GGUF with Ollama:
ollama run hf.co/ggml-org/Qwen3.8-27B-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use ggml-org/Qwen3.8-27B-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ggml-org/Qwen3.8-27B-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ggml-org/Qwen3.8-27B-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use ggml-org/Qwen3.8-27B-GGUF with Docker Model Runner:
docker model run hf.co/ggml-org/Qwen3.8-27B-GGUF:Q4_K_M
- Lemonade
How to use ggml-org/Qwen3.8-27B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ggml-org/Qwen3.8-27B-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Qwen3.8-27B-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use ggml-org/Qwen3.8-27B-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ggml-org/Qwen3.8-27B-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ggml-org/Qwen3.8-27B-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ggml-org/Qwen3.8-27B-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ggml-org/Qwen3.8-27B-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ggml-org/Qwen3.8-27B-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
LLama.cpp parameters suggestion
Hi All!!
I love llama server, currently, I am using integrated with opencode specifically for CODE operations, projects, PRS, etc...
Could you please recommend the best llama.cpp configuration for this setup CODING TASKS?
My stack:
Hardware:
- CPU: Intel Core i9-14900K, 24 cores / 32 threads
- RAM: 188 GiB total, approximately 60 GiB available
- GPU 0: NVIDIA RTX PRO 4000 Blackwell, 24,467 MiB VRAM total, 21,580 MiB used, 2,407 MiB free, 48% utilization, 81°C
- GPU 1: NVIDIA GeForce RTX 4090, 24,564 MiB VRAM total, 22,429 MiB used, 1,652 MiB free, 40% utilization, 53°C
- NVIDIA driver: 580.173.02
- OS: Ubuntu 24.04.4 LTS
- Kernel: 6.8.0-137-generic
Software:
- llama.cpp Docker image: ghcr.io/ggml-org/llama.cpp:server-cuda
- llama.cpp version: 0.3.0-dev, build 10731, commit 0eadefebd
- Model: Qwen3.8-27B-Q8_0.gguf
- Quantization: Q8_0
- Multimodal projector: mmproj-F16.gguf
- Context size: 260,000 tokens
- Server port: 8092
- Parallel slots: 1
- Flash attention: enabled
- CUDA offload: all layers
- Speculative decoding: draft-mtp
- Reasoning: enabled, low effort, 512-token budget
- Metrics: enabled
- Log verbosity: default level 3
llama.cpp parameters:
-m /models/Qwen3.8-27B-GGUF/Qwen3.8-27B-Q8_0.gguf
--mmproj /models/Qwen3.8-27B-GGUF/mmproj-F16.gguf
--n-gpu-layers -1
--split-mode layer
--main-gpu 0
--parallel 1
-c 260000
--threads 24
--threads-batch 32
--batch-size 2048
--ubatch-size 512
--temperature 0.7
--top-p 0.95
--top-k 20
--min-p 0.00
--presence-penalty 0.0
--repeat-penalty 1.0
--kv-unified
--cache-type-k q8_0
--cache-type-v q8_0
--sleep-idle-seconds 1800
--flash-attn on
--spec-type draft-mtp
--spec-draft-n-max 2
--spec-draft-ngl all
--reasoning on
--reasoning-effort low
--reasoning-budget 512
--jinja
--metrics
Observed performance:
- Average generation speed: approximately 27.5 tokens per second
- Typical generation speed: approximately 20–43 tokens per second
- Draft acceptance rate: approximately 80–95%
- The logs also show repeated prompt-cache eviction messages.
- No out-of-memory, fatal, panic, or assertion errors were found.
Could you please advise:
Based on your experience, If you suggest different parameters to optimize my stack to have more performance or if I doing something wrong, please fell free to suggest.
I am using llama for + 2 years, I would like to improve my daily operations and learn more about llama.cpp best practices.
Best regards,
Tiago S.