Instructions to use ubergarm/Qwen3.8-27B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ubergarm/Qwen3.8-27B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ubergarm/Qwen3.8-27B-GGUF # Run inference directly in the terminal: llama cli -hf ubergarm/Qwen3.8-27B-GGUF
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ubergarm/Qwen3.8-27B-GGUF # Run inference directly in the terminal: llama cli -hf ubergarm/Qwen3.8-27B-GGUF
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ubergarm/Qwen3.8-27B-GGUF # Run inference directly in the terminal: ./llama-cli -hf ubergarm/Qwen3.8-27B-GGUF
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ubergarm/Qwen3.8-27B-GGUF # Run inference directly in the terminal: ./build/bin/llama-cli -hf ubergarm/Qwen3.8-27B-GGUF
Use Docker
docker model run hf.co/ubergarm/Qwen3.8-27B-GGUF
- LM Studio
- Jan
- vLLM
How to use ubergarm/Qwen3.8-27B-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ubergarm/Qwen3.8-27B-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ubergarm/Qwen3.8-27B-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ubergarm/Qwen3.8-27B-GGUF
- Ollama
How to use ubergarm/Qwen3.8-27B-GGUF with Ollama:
ollama run hf.co/ubergarm/Qwen3.8-27B-GGUF
- Unsloth Desktop
- Pi
How to use ubergarm/Qwen3.8-27B-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ubergarm/Qwen3.8-27B-GGUF
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ubergarm/Qwen3.8-27B-GGUF" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use ubergarm/Qwen3.8-27B-GGUF with Docker Model Runner:
docker model run hf.co/ubergarm/Qwen3.8-27B-GGUF
- Lemonade
How to use ubergarm/Qwen3.8-27B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ubergarm/Qwen3.8-27B-GGUF
Run and chat with the model
lemonade run user.Qwen3.8-27B-GGUF-{{QUANT_TAG}}List all available models
lemonade list
- Hermes Agent
How to use ubergarm/Qwen3.8-27B-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ubergarm/Qwen3.8-27B-GGUF
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ubergarm/Qwen3.8-27B-GGUF
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ubergarm/Qwen3.8-27B-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ubergarm/Qwen3.8-27B-GGUF
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ubergarm/Qwen3.8-27B-GGUF" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Looping and repeats in coding harnesses
I've tried a few different harnesses (opencode, deepseek, qwen-code) and it seems with this model it causes massive loops of nonsense after a few minutes of thinking. I dont know if this is an issue with ik_llama or the quant. I switch to Unsloth IQ4_XS with llama.cpp and the harnesses can finish the project. Hermes and other things seem to work fine with IQ4_KS
My IQ4_KS config (llama-swap with ik_llama) - using template from froggeric on both the ubergarm and unsloth models (I have also tried without the updated template with ubergarm). Also using "thinking" with medium reasoning level on both.
"Qwen3.8-27B MTP ik_llama":
cmd: >
sh -c "LD_LIBRARY_PATH=/ik-bin/bin:/usr/local/nvidia/lib:/usr/local/nvidia/lib64:/usr/local/cuda/lib64 CUDA_VISIBLE_DEVICES=0 exec /ik-bin/bin/llama-server
--port ${PORT}
--host 127.0.0.1
--model /models/qwen38/ubergarm_Qwen3.8-27B-MTP-IQ4_KS.gguf
--mmproj /models/qwen38/unsloth_mmproj-BF16.gguf
--chat-template-file /models/qwen38/froggeric_Qwen3.8_14AUG2026.jinja
-ngl 99
--ctx-size 105000
-ngld 99
-cd 4096
--spec-type mtp:n_max=2,p_min=0.0
--cache-type-k q8_0
--cache-type-v q8_0
--merge-qkv
--merge-up-gate-experts
--ctx-checkpoints 32
--parallel-tool-calls
-khad
--split-mode none
--main-gpu 0
--image-min-tokens 1024
--threads 6
--threads-batch 6
--flash-attn on
--parallel 1
--batch-size 512
--ubatch-size 512
--no-mmap
--jinja"
filters:
stripParams: "temperature, top_p, top_k, min_p, presence_penalty, repeat_penalty"
setParamsByID:
"${MODEL_ID}:thinking":
max_tokens: 32768
chat_template_kwargs:
enable_thinking: true
preserve_thinking: true
reasoning_effort: "medium"
temperature: 1.0
top_p: 0.95
top_k: 20
min_p: 0.0
presence_penalty: 0.0
repeat_penalty: 1.0
"${MODEL_ID}:thinking-coding":
chat_template_kwargs:
enable_thinking: true
preserve_thinking: true
reasoning_effort: "xhigh"
temperature: 0.4
top_p: 0.95
top_k: 20
min_p: 0.0
presence_penalty: 0.0
repeat_penalty: 1.0
"${MODEL_ID}:instruct":
max_tokens: 16384
chat_template_kwargs:
enable_thinking: false
preserve_thinking: false
temperature: 0.7
top_p: 0.8
top_k: 20
min_p: 0.0
presence_penalty: 1.5
repeat_penalty: 1.0
"${MODEL_ID}:instruct-reasoning":
max_tokens: 16384
chat_template_kwargs:
enable_thinking: false
preserve_thinking: false
temperature: 0.7
top_p: 0.8
top_k: 20
min_p: 0.0
presence_penalty: 1.5
repeat_penalty: 1.0
"${MODEL_ID}:instruct-tools":
max_tokens: 8192
chat_template_kwargs:
enable_thinking: false
preserve_thinking: false
temperature: 0.6
top_p: 0.95
top_k: 20
min_p: 0.0
presence_penalty: 0.0
repeat_penalty: 1.0
"${MODEL_ID}:thinking-tools":
max_tokens: 16384
reasoning_budget: 8192
chat_template_kwargs:
enable_thinking: true
preserve_thinking: true
reasoning_effort: "medium"
temperature: 0.6
top_p: 0.95
top_k: 20
min_p: 0.0
presence_penalty: 0.0
repeat_penalty: 1.0
"${MODEL_ID}:router":
max_tokens: 512
chat_template_kwargs:
enable_thinking: false
preserve_thinking: false
temperature: 0.1
top_p: 0.1
top_k: 10
min_p: 0.0
presence_penalty: 0.0
repeat_penalty: 1.0
Example repeat of nonsense with IQ4_KS with prompt "Create a single page html + tailwind css + javascript recreation of galaga, same to the original arcade game. include sound effects":
Hmm getting fuzzy again on details (transcribed code fragments were truncated / duplicated, probably copy glitches, etc.); as above with settled approach is a fine and I have high-confidence: use min, no issues for numbers confirmed by implementation and adjust with code layout that gets reviewed and committed: (medal at ~center; GAME OVER; small hint.) Enough already! Writing files: canvas at H:
Now finalize with these and inline implementation numbers at… 8-sec timer inline inline 68483-…, with small delay; small wait on user 7 9s later no inline issue no further confirmation required to complete layout in… proceed! I trust code inline, commit & complete at a… let me start coding file and confirm constants while final implementation… commit as: commit final (write full). Writing out carefully below inline and verify… write code to inline confirm final as follows in actual (verify: min no wait needed? I confirmed all… wait OK I can… (small confirm): h = ...min() clamp with GAP (confirmed at layout OK 5…
Proceed (min layout numbers… write… proceed; numbers to code & reconfirmed once committed? min(89…) confirm as OK for now). Let’s begin actual commit to actual layout commit and actual writing with verified: inline no (confirm)… proceeding commit commit, done inline &… proceeding final with (with verify OK in implementation), proceed writing inline inline; proceeding — numbers committed OK verified by impl and min confirm done as written? yes proceeding & committed done.) Let’s start now & actual verify constants committed — actually committing actual verify of implementation done OK verified at committed time min() numbers in written verified commit min — let me finalize in committed with verification commit commit verified numbers written, final numbers… I verified final, writing (confirmed by verified, inline layout verify final — inline implementation of the written with actual, proceed inline and confirmed layout — actual confirmed (written and implemented — done? inline and confirm OK proceed… verified: write it) & committed to (layout implementation in progress OK: actual) write out commit verified as layout confirmed and layout: done proceed actual… — no min confirmed no additional: confirmed at verified numbers min inline no inline implementation numbers final implementation no confirm & (actual write in committed): let’s stop with loop meta thinking now (inline, committed done confirmed), just implement commit write the numbers I’ll be done as code final… inline OK committed proceed no confirmed implementation layout implementation min confirm written min verified in (verify confirmed as… confirmed written and OK):
I confirm verified the write-out, let’s commit — implementation as min confirm commit and code committed numbers no: writing and inline… min final code min() final & writing as actual committed commit layout verify as OK inline verify written… 7 sec confirmed. Actual & no issue OK committed writing layout — write, actual numbers written layout committed confirmed done in write implementation and write. Write inline confirmed — code, committed — implementation write code numbers: code — numbers commit verified no, let me do
Thanks for the details, and I see your comment over at: https://huggingface.co/ubergarm/Qwen3.8-27B-GGUF/discussions/4 too.
I'm only using pi.dev harness and have only seen it loop once using the default chat template.
I'm trying to understand which combinations of harness and chat template you've tested e.g.
- opencode - looping - (which templates?)
- deepseek - looping - (which templates?)
- qwen-code - looping - (which templates?)
- hermes - works okay - (which templates?)
- pi (my testing) - works okay - default template
Were you using the baked in original Qwen official template, or the froggeric in your testing?
As mentioned in the other thread, I'll update the README with notes on how to bring your own template.
Thanks for the details, and I see your comment over at: https://huggingface.co/ubergarm/Qwen3.8-27B-GGUF/discussions/4 too.
I'm only using pi.dev harness and have only seen it loop once using the default chat template.
I'm trying to understand which combinations of harness and chat template you've tested e.g.
- opencode - looping - (which templates?)
- deepseek - looping - (which templates?)
- qwen-code - looping - (which templates?)
- hermes - works okay - (which templates?)
- pi (my testing) - works okay - default template
Were you using the baked in original Qwen official template, or the froggeric in your testing?
As mentioned in the other thread, I'll update the README with notes on how to bring your own template.
Hi Ubergarm,
I honestly didnt notice the looping since I dont really use my local AI to code much, so for my use-case your release was fine. It wasnt until I suggested your quant and model to someone and they said it looped badly then I tried for myself: https://www.reddit.com/r/LocalLLM/comments/1vtvqu9/comment/p4z6kg8/
- opencode - looping - tried with frogger and default templates
- deepseek - looping - tried with frogger
- qwen-code - looping - tried with frogger and default templates
- hermes - works okay - tried with frogger
i have experienced looping/model fixation in pi, using different chat templates, with this model and also Q6 from unsloth on same engine (ik llama) WHEN TALKING ABOUT CERTAIN TOPIC (CAD filetypes conversion).
I have tracked it down to MTP since i can continue fixated/looped pi session when turning off MTP.
happy to provide logs/more info if needed.
build: INFO [ main] build info | tid="140507548123136" timestamp=1787570968 build=4632 commit="3b81f63a"
Typical llm output:
zapisanie tylko liczy w wersji, nic co ma problem na GPL z powodu LibreDWG jako drugie narzenie dla samego odcitu plikw w 69 (polskiej):
Parsowany dwg będzie po CLI jako dwg -> DXF ACAD2018; a z tym dzxf do obo pdfa. Ale czek:
Sprawdzenie na chwylenki że PDF i SVG w odatale owa CLI jako celne — "DXF output", czy tak i PNG nie będzie się renderować; nie wiem jeszcze zez, nie wiernych jest nie; Oda converter output w tej ota nie widi nie z nie "dwi z", jest na
wywaj odfu od **ODA only CAD formarts only — ** w art: "CAD-only", ODF/PLU z OCF:
Ode jest dwu (i Dxf only. Ale tak też od pdf/png (niz na to wyraz na: 7).
- z wyjaw "w** (pdf / obo). " ** nie będzie nie pdf o pdf nie na png i "i", bo PDF: na png jest tylko (2 z w), pdf ** " nie" jest i ** o z PDF; a (n na od PNG to pdf od na dwi)
Szczerz od ODC PDF — Odc wyawia to jest pdf — pdf " z wywo** jest 5 o w w na, PDF i dwz jest: (303 w nie. z) jest — a Odf — dwc i dnf; png "nie". A więc PNG musi po "2", (53 na nie w png w), ** to owa. To na wyjavia PDF " jest pdf",
PDF; dwu o dcf ** png, PDF odf pdf pdf: w dvi nie — jest z z wy** i: jest i. i od nie o o nie i dw z png nie; na:
Df. Z OCF i w png pdf na PDF to dng — a PNG o (4: Ode)
I jeszcze jest ODE. "dvc jest — "
Ale czek na PNG na — a na z dxc; o " (n dcv): w: PDF ** dwv ** to ** od, Ocf dw z; PNG dw jest to png pdf w " w. od; jest ** PDF — i ODA. na. Z PNG, nie o; PNG i — pdf dw, dxi to PNG "w** ** ODE w o o na nie nie ** png nie 3 jest (61
jest "na.
w dvg o nie z ** — dw nie (PDF dw pdf (i (z; i " — to Ocf png nie png ** ODF o Oda jest i — dnx w PDF nie: dcv to: z z. W PNG. To (pdf — ** i to — png to.
Jako Ocf "nie w na od w" ODE dw 1 to PNG z nie PDF ** dnv ** pdf nie to o nie, PNG, PNG nie; o, " — PNG 4 "i w dw" dw na ODF i PNG dw i; i: to — nie ** PNG; a dvn; ** (PNG. Na to; (n od png ** png:
" jest PDF od o " od ODC PDF w Ocf png i, png.
— (461 (PNG) dw — o w ** jest "na png z — nie to; png i z z nie to ** w — i dwc nie nie to PDF nie " od na, jest. z, pdf jest pdf dw:
OK ** i o ODE nie dw i to, dcv pdf png " " png (7. jest, w; w to Oda i jest o ** dxf na — PNG o jest ** dvn z: PNG — a w: na pdf: dng " dw. Od to "o o — "PDF nie pdf ** PDF w to dxc 459 pdf jest w (7 na ODF pdf dw (78 w ** w (i; PNG
nie, i dcf o — i nie nie PDF to png — PDF: PNG ** — pdf; " od dxi z i ** i w dw — PNG dw o. od ODE to, png na ** na dnc nie w to dnf z z png dw " to. Na. i, PNG nie png nie ** dw (z — Odf.
Od dvn; png; 5 o: nie dwu ** OCF o png: png: png na: w PDF jest i to (1; (na ** " i Odc i w. nie — ** dcv — PNG ** pdf ** png, — w — jest nie: jest od w PDF w "w** dvg w dw pdf png, ODA na; ** — o jest dxc png i nie o jest z ** Odf;
png " o " pdf to w na Odf nie ** ODF png ** 6 w png pdf nie nie dw — dw. jest ** i ** i ODF: to PNG nie to od PDF z (dwi png. W " w o ODE "PDF od.
z w Ode o — i png i dng na (dw " ** dxf dw; dxf nie. o png pdf, PDF (6, to to jest: pdf — PNG i; i pdf nie dw "i od na "pdf na z w: to nie PDF ** png pdf: " to — na png (n (3 z dnf nie (18 od PDF — w jest png nie png nie.
Ale pdf png, dvn i jest dw jest. Z nie w — w, (z. ```
llm config
ExecStart=/work/ik_llama-server
--model /work/models/Qwen3.8-27B-GGUF/Qwen3.8-27B-MTP-IQ4_KS.gguf
--alias "Qwen3.8-27B"
-c 149504
-ctk q8_0 -ctv q8_0
-ctkd q8_0 -ctvd q8_0
--merge-qkv
-muge
-ngl 99
-t 1
-tb 1
-tm 2
--host 0.0.0.0
--port 8080
--parallel 1
--jinja
--chat-template-file /work/models/Qwen3.8-27B-GGUF/chat_template.jinja
--reasoning-format deepseek
--no-mmproj-offload
--mmproj /work/models/Qwen3.8-27B-GGUF/mmproj-Qwen3.8-27B-Q8_0.gguf
--ctx-checkpoints 16
-cram 16384
--spec-type mtp:n_max=4,p_min=0.0
Thanks for your gkubon, I wonder whether this is an issue with the MTP layer, ik_llama, or both? I didnt notice this looping during coding sessions with ubergarm's qwen3.6-27B with ik_llama and MTP, but Ive since updated ik_llama multiple times
Ahh thanks for more details and links. I'm running ik_llama.cpp@8337e4cd 2026-08-15 19:35 +0200 which has https://github.com/ikawrakow/ik_llama.cpp/pull/2322 Fix Qwen35+ MTP
So definitely update and use the latest version, might have been the issue!
It does loop from time to time, even with the latest version. But I found a temperature of 0.6 does help mitigate it by a lot!
EDIT: seems like MTP is causing this issue
yep, still broken
If you guys try unsloth quant (iq4_xs or q4) + mtp with ik_llama, does it still loop? If so, it might be an ik_llama bug
Its probably down to you quanting your kv cache