Instructions to use ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S # Run inference directly in the terminal: llama cli -hf ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S # Run inference directly in the terminal: llama cli -hf ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S # Run inference directly in the terminal: ./llama-cli -hf ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S # Run inference directly in the terminal: ./build/bin/llama-cli -hf ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S
Use Docker
docker model run hf.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S
- LM Studio
- Jan
- vLLM
How to use ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S
- Ollama
How to use ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF with Ollama:
ollama run hf.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S
- Unsloth Desktop
- Pi
How to use ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF with Docker Model Runner:
docker model run hf.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S
- Lemonade
How to use ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S
Run and chat with the model
lemonade run user.Qwen3.8-27B-GSQ-RCO-GGUF-IQ2_S
List all available models
lemonade list
- Hermes Agent
How to use ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Is this model same as Q4KM for coding related tasks?
Is this model same as Q4KM for coding related tasks?
Especially vs Unsloth quants?
Did you even take the time to read the description? There's some useful information in there.
Is this model same as Q4KM for coding related tasks?
Especially vs Unsloth quants?
just download and start using this Qwen3.8-27B-GSQ-RCO-IQ3_S-mtp.gguf and you are good to go
It seems that the model hallucinate quite quickly. I tried the biggest quant + MTP at max context. In the end, it is less useful than UD Q4M at 120k context.
Maybe I have to change some settings.
Thank you for your feedback and for your interest in our models! Would you mind sharing the run commands you used for both models along with a bit more detail about your task and the kinds of hallucinations you encountered? We may be able to suggest some settings to improve your experience and your feedback could also provide valuable insights for future releases.
It seems that the model hallucinate quite quickly. I tried the biggest quant + MTP at max context. In the end, it is less useful than UD Q4M at 120k context.
Maybe I have to change some settings.
do you use llama.cpp or something else? if you quantize the kv too much it may happen. I personally rarely get hallucinations , i have been using this model for 4 days now, so far so good, no different than UD version Q4_K_M.
Thanks, Yes I'm using llama.cpp and kv Q8_0 for model and mtp.
But I need to add that, indeed, the model is quite good and it seems that I got issues only with Cline for VSCode.
So, maybe it is because of the template ? I don't know.
models.ini:
; ============================================================
; Qwen3.8-27B-GSQ-RCO
; ============================================================
[Qwen3.8-27B-GSQ-RCO]
model = /home/gaetan/Models/Qwen3.8/27B/Qwen3.8-27B-GSQ-RCO-IQ3_S-mtp.gguf
mmproj = /home/gaetan/Models/Qwen3.8/27B/mmproj-BF16.gguf
no-mmproj-offload = true
ctx-size = 200000
np = 1
n-gpu-layers = all
split-mode = tensor
tensor-split = 1,1
kv-unified = true
flash-attn = on
ctk = q8_0
ctv = q8_0
ctkd = q8_0
ctvd = q8_0
spec-type = draft-mtp
spec-draft-n-min = 0.75
spec-draft-n-max = 2
image-min-tokens = 1024
jinja = true
reasoning-preserve = true
Personally I don't use any template other than the default.
here is my llama config:
-a qwen3.8-27b
--host 0.0.0.0 --port 8080
-ngl 99 -c 200000
--cache-type-k q8_0 --cache-type-v q4_1
--flash-attn on
-b 2048 -ub 1024
-np 1
--cache-ram 4096
--temp 0.6 --min-p 0.05 --presence-penalty 0.0 --top-p 0.95 --top-k 64
--spec-type draft-mtp
--spec-draft-n-max 3
--load-mode mlock
--metrics
--perf
--api-key-file /home/mert/.llama-api-keys
--timeout 600 \
I noticed that you don't have any temp / penalty etc. that might be the reason. " --temp 0.6 --min-p 0.05 --presence-penalty 0.0 --top-p 0.95 --top-k 64 " This is usually works for qwen3.6 and qwen3.8 for coding & agentic tasks.
So far I've only seen it hallucinate in long context text analysis. I used novels of 100k+ token as kinda needle test. I asked it questions like "how many people has A interacted with? how many opponents has B killed? ... where, when, and who are they ? list them in a table" - stuff like that. And in its thought process, you will see the model occasionally mixing up and hallucinating plots about the characters to varying degrees, even though the names and numbers come out nearly correct. Its understanding of the novel's plot is not perfect, but the summaries of the basic elements are usually almost correct.
Despite the time it takes to crunch the text, it's a fun test.
In coding, I'd have to blame this quant. The but wait maybe hold on actually avalanche is a real pain to watch.