Instructions to use ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S # Run inference directly in the terminal: llama cli -hf ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S # Run inference directly in the terminal: llama cli -hf ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S # Run inference directly in the terminal: ./llama-cli -hf ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S # Run inference directly in the terminal: ./build/bin/llama-cli -hf ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S
Use Docker
docker model run hf.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S
- LM Studio
- Jan
- vLLM
How to use ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S
- Ollama
How to use ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF with Ollama:
ollama run hf.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S
- Unsloth Desktop
- Pi
How to use ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF with Docker Model Runner:
docker model run hf.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S
- Lemonade
How to use ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S
Run and chat with the model
lemonade run user.Qwen3.8-27B-GSQ-RCO-GGUF-IQ2_S
List all available models
lemonade list
- Hermes Agent
How to use ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Miracle did not happen
Tested Qwen3.8-27B-GSQ-RCO-IQ3_S.gguf in llama-server with following arguments --threads 8 -c 131000 -np 1 -fa on --repeat-penalty 1.1 --repeat-last-n 512 --presence-penalty 0.1 --temperature 0.3 --top-p 0.9 --min-p 0.02 -ctk q8_0 -ctv q8_0 -b 2048 -ub 1024 --main-gpu 1
Got thinking loop at ~17k context at first attempt. Stopped and asked to output resulting simple game (not well known) faster. Got non-working game. So far other qwen q4 quants behaved better for me, even MoE. π
Never tested in real world use case (150k line project)
How does it behave when using the recommended settings?
We recommend using the following sets of sampling parameters for generation:
Thinking Mode: temperature=1.0, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=0.0, repetition_penalty=1.0
Instruct (or non-thinking) mode: temperature=0.7, top_p=0.80, top_k=20, min_p=0.0, presence_penalty=1.5, repetition_penalty=1.0
Please note that the support for sampling parameters varies according to inference frameworks.
I've run that exact quant for millions of tokens in Hermes @148Kcontex - never saw it even hesitate - much less loop. Even the XXS is very stable.
My first suspicion is the temp 0.3 - that is very low for qwen. If you are coding that's much too low on this model and might result in odd behaviour. I think I have seen cases before of low-temp making qwen go a bit lobotomized.
Assuming that you are coding - you should also be very careful with using repeat penalty at all. I don't think that would be a cause of the looping - but coding in general can have lots of repeating text, and sampler params should not try to block this. It's better used in creative writing / conversation ect. DRY (if available) is often a better means to the same end in those cases though.
How does it behave when using the recommended settings?
Frankly I didn't test it with recommended settings. I used the same settings I had best results with MoE (at 0.6 and more MoE variants produces too many errors for me), especially temperature. Perhaps dense model reacts to them differently
I've run that exact quant for millions of tokens in Hermes @148Kcontex - never saw it even hesitate - much less loop. Even the XXS is very stable.
How is the behavior at long contexts? For coding and tool calls?
How does it behave when using the recommended settings?
Frankly I didn't test it with recommended settings. I used the same settings I had best results with MoE (at 0.6 and more MoE variants produces too many errors for me), especially temperature. Perhaps dense model reacts to them differently
You complain about the model without using the right settings?
How does it behave when using the recommended settings?
Frankly I didn't test it with recommended settings. I used the same settings I had best results with MoE (at 0.6 and more MoE variants produces too many errors for me), especially temperature. Perhaps dense model reacts to them differently
I've never had good luck tweaking the parameters myself. Recommended settings always seemed to work the best.
Try it and see if you get the same issues. I tried IQ3_S on dual 7900XTX, but it was slower than non iMatrix quants. This is an AMD problem, according to the internet. But, the IQ3_S quant worked really well for its size, no issues on tool calling/long context work.
I've run that exact quant for millions of tokens in Hermes @148Kcontex - never saw it even hesitate - much less loop. Even the XXS is very stable.
How is the behavior at long contexts? For coding and tool calls?
I've been using it a lot recently in Hermes - for coding jobs - and so far I haven't been able to detect anything unusual. 148K context is the longest I can run, so I can't speak for really extreme context lengths - but if models start to "drift", then it's usually visible already at 16-32K. So far I have seen no evidence of any drift or loss of focus. Agent has never failed to finish a task or forgotten something important.
For tool calls - failed toolcalls are rare. Actually I have only seen it happen in "raw" terminal commands it runs from memory (no template unlike toolcalls) - and it always detects and corrects the problem if there is a command error. It's very impressive. If you give good instructions it will do the job reliably.
I think this has a lot to do with very high presicion being retained at the gated deltanet parts of layers
Both XXS and S variants are very stable. Subjectively it feels like the slightly larger "S" quant is a little more efficient at finding the solution to a problem sooner - but stability-wise they can run for hours and hours on a big task (via delegation / subagents - each get the full 148K). I customized the delegate_task tool to run sequentially since my system doesn't have the spare capacity for parallelism. But the result in total is the same = a big job autonomously run from start to end without issues.
Tested Qwen3.8-27B-GSQ-RCO-IQ3_S.gguf in llama-server with following arguments --threads 8 -c 131000 -np 1 -fa on --repeat-penalty 1.1 --repeat-last-n 512 --presence-penalty 0.1 --temperature 0.3 --top-p 0.9 --min-p 0.02 -ctk q8_0 -ctv q8_0 -b 2048 -ub 1024 --main-gpu 1
Got thinking loop at ~17k context at first attempt. Stopped and asked to output resulting simple game (not well known) faster. Got non-working game. So far other qwen q4 quants behaved better for me, even MoE. π
Never tested in real world use case (150k line project)
To explain what's probably going on: Your repeat-last-n is pretty high and your temperature is very low. Counter-intuitively, a high repeat penalty and low temperature can restrict the model's options so much that it does things like repeating the same sentence over and over. If your repeat-last-n is too high it can get into a situation where there are only a small number of viable tokens it can choose from, and by the time it gets to the end of N it has a very high chance of repeating the token the first token in N, creating a loop of whatever length your repeat-last-n is (in this case 512 tokens). It's often caused by the trained thinking interjections like "Wait, ...". If that triggers at the end of your N a loop is almost guaranteed. This is especially likely with a low temperature, because the model isn't allowed to pick less likely tokens that would break it out of the loop.
Lower temperature can help very small models (~4B size) that go off in the weeds at higher temperatures, because they often don't have great options to choose from to begin with. They get best results from just picking the most likely token in most cases. But for larger models it's the opposite - you want to give the model more room to explore other possible tokens, because it typically has a lot of good options to choose from. I wouldn't go below 0.7 for 27B, and only for coding tasks and such at that level. Qwen recommends 1.0 if you have reasoning enabled, and they enable reasoning by default.
Also, you can't directly transfer one model's settings to another architecture and expect success. You can sort of have rules of thumb, but how they were trained will determine how they will perform at different temperature settings. Even then, I think the 35B MOE recommended temperature settings from Qwen are still 0.7 for non-reasoning and 1.0 for reasoning.
Recommended settings made the miracle happen for me. On a mixed setup(rtx3090, 10G 3080, 3x12G 3060) this gave me 33 token/s while utilizing all cards.(with bf16 embeddings)