Instructions to use ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S # Run inference directly in the terminal: llama cli -hf ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S # Run inference directly in the terminal: llama cli -hf ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S # Run inference directly in the terminal: ./llama-cli -hf ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S # Run inference directly in the terminal: ./build/bin/llama-cli -hf ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S
Use Docker
docker model run hf.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S
- LM Studio
- Jan
- vLLM
How to use ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S
- Ollama
How to use ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF with Ollama:
ollama run hf.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S
- Unsloth Desktop
- Pi
How to use ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF with Docker Model Runner:
docker model run hf.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S
- Lemonade
How to use ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S
Run and chat with the model
lemonade run user.Qwen3.8-27B-GSQ-RCO-GGUF-IQ2_S
List all available models
lemonade list
- Hermes Agent
How to use ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ2_S" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Please keep the MTP/nextn tensors — and consider a ~3.25 bpw target
Thanks for this release. We benchmarked IQ3_XXS against Unsloth's UD-IQ3_S (3.44 bpw), which isn't in your comparison table — you only compare against UD-IQ2_S and UD-Q2_K_XL.
Paired runs (same llama.cpp build, same wikitext2 corpus, same seed), IQ3_XXS wins all four cells, every one outside the combined error:
| KV / ctx | UD-IQ3_S | GSQ-RCO IQ3_XXS | Δ |
|---|---|---|---|
| q4_0 / 4k | 6.4776 ± .042 | 6.3743 ± .040 | −0.103 |
| q4_0 / 16k | 6.4652 ± .044 | 6.1398 ± .039 | −0.325 |
| q4_0 / 32k | 6.8376 ± .046 | 6.6517 ± .042 | −0.186 |
| f16 / 4k | 6.4617 ± .042 | 6.3541 ± .040 | −0.108 |
Smaller file, fewer bits, better perplexity.
Two requests:
1. Keep the MTP head. IQ3_XXS ships without the multi-token-prediction tensors — we count 848 tensors and not one blk.*.nextn.eh_proj / .enorm / .hnorm / .shared_head_norm, and there's no blk.64. The base checkpoint has them and Unsloth's quants preserve them. In llama.cpp they enable --spec-type draft-mtp: speculative decoding with no external draft model. On a 16 GB card that's worth more than half a bit — it removes a separate 1.1 GB draft file. Passing them through, even at Q8_0, costs about 1 GB. It's also currently undocumented that they're missing.
2. A ~3.25 bpw build. Since RCO takes a total size budget, this should be a one-parameter change. 4 bpw is unusable for 16 GB: the file alone lands at 12.5 GiB, and with a draft model plus a 98k q4_0 KV cache we measure that ~1.2 GB over the card. 3.25 bpw is the largest target that still leaves room for both.
Also useful if cheap: the per-tensor type assignment RCO selected (JSON or table), the imatrix calibration corpus, and any long-context evaluation — your benchmarks are all short-form, and we see perplexity degrade sharply at 32k on every quant we test, which is the regime that matters for agentic coding.
Thanks for the detailed benchmarks and suggestions!
I’ve uploaded the imatrix and RCO allocations as well. MTP is on our radar too and we’d like to train a small dedicated MTP layer for our models. We’ll also do 3.25-bit and 3.5-bit models. Thanks for the suggestion!
Thanks for shipping the imatrix and the RCO allocations so quickly — both answered things we couldn't get at from the GGUF alone, and both fed into what follows.
A correction on our side first
We said 848 tensors. The real count is 851, and your allocation file says 851 too — it matches our copy of the GGUF tensor-for-tensor, zero differences in either direction. Our number came from a bad hand-count, apologies. The conclusion is unchanged, and now it comes from your file rather than our reading.
The MTP head is already gone before quantization
This is the one that may save you work. Your imatrix covers blocks 0..63 and contains zero nextn tensors. So the head wasn't dropped by the quantizer — it was already absent when you collected the statistics. That points at the conversion step, not the quantization config.
It matters because you mentioned wanting to train a small dedicated MTP layer. You may not need to. Unsloth's Qwen3.8-27B-UD-IQ3_S.gguf has 866 tensors to your 851, and the difference is exactly the 15 tensors of blk.64:
blk.64.attn_norm.weight blk.64.attn_q.weight
blk.64.attn_q_norm.weight blk.64.attn_k.weight
blk.64.attn_k_norm.weight blk.64.attn_v.weight
blk.64.attn_output.weight blk.64.post_attention_norm.weight
blk.64.ffn_gate.weight blk.64.ffn_up.weight
blk.64.ffn_down.weight blk.64.nextn.eh_proj.weight
blk.64.nextn.enorm.weight blk.64.nextn.hnorm.weight
blk.64.nextn.shared_head_norm.weight
A full transformer block plus the four nextn-specific tensors, already present in the base checkpoint. If your HF-to-GGUF path preserved it, --spec-type draft-mtp works for the cost of passing tensors through — no training run. Worth checking before spending the compute.
A third column for the table we sent you
We've since run the same paired setup on Unsloth's UD-IQ4_XS (4.24 bpw), same base model, same binary, same corpus, same --seed 1234. Adding it to the numbers from our earlier post:
| ctx (KV) | UD-IQ3_S (3.44) | GSQ-RCO IQ3_XXS (3.00) | UD-IQ4_XS (4.24) |
|---|---|---|---|
| 4096 (q4_0) | 6.4776 ±.042 | 6.3743 ±.040 | 6.3882 ±.041 |
| 16384 (q4_0) | 6.4652 ±.044 | 6.1398 ±.039 | 6.0625 ±.039 |
| 32768 (q4_0) | 6.8376 ±.046 | 6.6517 ±.042 | 6.4783 ±.041 |
| 4096 (f16) | 6.4617 ±.042 | 6.3541 ±.040 | 6.3642 ±.041 |
Against 1.24 bpw more, your IQ3_XXS ties at 4k in both KV precisions (0.2x the combined error, in both directions) and loses only at 16k and 32k (1.4x and 2.9x). So the gap that a full extra bit buys is specifically a long-context gap — everything else it buys is nothing.
The part we didn't expect: your imatrix is 1000 chunks × 4096 tokens, so the calibration never saw past 4k, and the model still holds level with a 4.24 bpw quant there and beats a 3.44 bpw one everywhere. Whatever RCO is doing generalises well past its calibration length.
A guess at the mechanism from your allocation file — hypothesis, not something we measured. The 16 full-attention blocks (3, 7, 11 … 63) average 3.18 bpw, against 2.96 for the 48 DeltaNet blocks and 2.88 for the FFN. The budget went where long-range information lives. Two exceptions stood out: blk.3.attn_k and blk.11.attn_v at IQ1_S. That's where we'd look first if it ever does degrade in long context.
Still interested in any long-context evaluation from your side — all three quants above get worse from 16k to 32k, and that's the regime we actually work in.
3.25 vs 3.5, now with a measurement instead of arithmetic
Good news that both are coming. Our earlier "4 bpw is unusable on 16 GB" was a calculation; we've now measured it on that same 4.24 bpw file, and it's worse than we said. On a 16376 MiB card with a 98k q4_0 KV cache:
| window | tok/s | peak | |
|---|---|---|---|
| UD-IQ3_S (3.44) + native MTP | 98304 | 80.0 | 14404 |
| UD-IQ4_XS (4.24) + native MTP | 98304 | — | OOM |
| UD-IQ4_XS (4.24), no speculation | 98304 | 38.2 | 15398 |
At 4.24 bpw the 98k cache only fits if you give up speculative decoding entirely — a 2.1x clock penalty for 0.8 bpw. Its ceiling with speculation is between 72k and 80k.
Scaling from your measured 14402 MiB peak (IQ3_XXS + 1.1 GB external draft + 98k q4_0 KV), 3.25 bpw lands near 15200 and 3.5 near 16000 of 16376. So 3.25 works comfortably and 3.5 is on the edge with an external draft.
Which loops back to the MTP point: if the head survives conversion, the external draft goes away and 3.5 becomes comfortable. On a 16 GB card that conversion fix is worth more than the extra half bit.
One small thing
The README still doesn't mention that the MTP tensors are absent, and doesn't document the imatrix or the allocation files you just added.
Hi guys this is truely a phenominal model - running on rtx 5080 16GB VRAM and I can do in llama.cpp:
GGML_CUDA_DISABLE_GRAPHS=1 ./llama.cpp/build/bin/llama-server -hf ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ3_XXS --alias local-model --no-mmproj -np 1 -fa on -ctk q4_0 -ctv q4_0 -ngl 66 -fit off -b 2048 -ub 256 -t 14 -tb 28 -rea on --reasoning-budget 4096 --reasoning-budget-message "Reasoning budget reached. Finish your analysis and provide the complete final answer." --no-reasoning-preserve --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --presence-penalty 0.0 --frequency-penalty 0.0 --repeat-penalty 1.0 --cache-ram 8192 --host 0.0.0.0 --port 8080 -lv 4 -c 262144
at full context 262k, q4 kv cache and (prompt prefill) prompt eval is 487.37 tokens per second and (decode) eval time is 62.97 tokens per second
and also :
GGML_CUDA_DISABLE_GRAPHS=1 ./llama.cpp/build/bin/llama-server -hf ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF:IQ3_XXS --alias local-model --no-mmproj -np 1 -fa on -ngl 66 -fit off -b 2048 -ub 128 -t 14 -tb 28 -rea on --reasoning-budget 4096 --reasoning-budget-message "Reasoning budget reached. Finish your analysis and provide the complete final answer." --no-reasoning-preserve --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --presence-penalty 0.0 --frequency-penalty 0.0 --repeat-penalty 1.0 --cache-ram 8192 --host 0.0.0.0 --port 8080 -lv 4 -ctk q8_0 -ctv q4_0 -c 196608
at 192k, q8 kv cache and (prompt prefill) prompt eval is 362.41 tokens per second and (decode) eval time is 63.01 tokens per second
I agree, this model deserves more attention... I've tested IQ3_XXS with MTP headers merged from the Unsloth UD_IQ3_XXS and the speed and quality was amazing with just a RTX 5070Ti