Is this model same as Q4KM for coding related tasks?

#22
by soyalemujica - opened

Is this model same as Q4KM for coding related tasks?
Especially vs Unsloth quants?

Did you even take the time to read the description? There's some useful information in there.

Is this model same as Q4KM for coding related tasks?
Especially vs Unsloth quants?

just download and start using this Qwen3.8-27B-GSQ-RCO-IQ3_S-mtp.gguf and you are good to go

It seems that the model hallucinate quite quickly. I tried the biggest quant + MTP at max context. In the end, it is less useful than UD Q4M at 120k context.
Maybe I have to change some settings.

IST Austria Distributed Algorithms and Systems Lab org

Thank you for your feedback and for your interest in our models! Would you mind sharing the run commands you used for both models along with a bit more detail about your task and the kinds of hallucinations you encountered? We may be able to suggest some settings to improve your experience and your feedback could also provide valuable insights for future releases.

It seems that the model hallucinate quite quickly. I tried the biggest quant + MTP at max context. In the end, it is less useful than UD Q4M at 120k context.
Maybe I have to change some settings.

do you use llama.cpp or something else? if you quantize the kv too much it may happen. I personally rarely get hallucinations , i have been using this model for 4 days now, so far so good, no different than UD version Q4_K_M.

Thanks, Yes I'm using llama.cpp and kv Q8_0 for model and mtp.

But I need to add that, indeed, the model is quite good and it seems that I got issues only with Cline for VSCode.
So, maybe it is because of the template ? I don't know.
models.ini:

; ============================================================
; Qwen3.8-27B-GSQ-RCO
; ============================================================

[Qwen3.8-27B-GSQ-RCO]

model = /home/gaetan/Models/Qwen3.8/27B/Qwen3.8-27B-GSQ-RCO-IQ3_S-mtp.gguf
mmproj = /home/gaetan/Models/Qwen3.8/27B/mmproj-BF16.gguf

no-mmproj-offload = true

ctx-size = 200000

np = 1

n-gpu-layers = all
split-mode = tensor
tensor-split = 1,1

kv-unified = true

flash-attn = on

ctk = q8_0
ctv = q8_0
ctkd = q8_0
ctvd = q8_0

spec-type = draft-mtp
spec-draft-n-min = 0.75
spec-draft-n-max = 2

image-min-tokens = 1024

jinja = true
reasoning-preserve = true

Personally I don't use any template other than the default.
here is my llama config:

-a qwen3.8-27b
--host 0.0.0.0 --port 8080
-ngl 99 -c 200000
--cache-type-k q8_0 --cache-type-v q4_1
--flash-attn on
-b 2048 -ub 1024
-np 1
--cache-ram 4096
--temp 0.6 --min-p 0.05 --presence-penalty 0.0 --top-p 0.95 --top-k 64
--spec-type draft-mtp
--spec-draft-n-max 3
--load-mode mlock
--metrics
--perf
--api-key-file /home/mert/.llama-api-keys
--timeout 600 \

I noticed that you don't have any temp / penalty etc. that might be the reason. " --temp 0.6 --min-p 0.05 --presence-penalty 0.0 --top-p 0.95 --top-k 64 " This is usually works for qwen3.6 and qwen3.8 for coding & agentic tasks.

So far I've only seen it hallucinate in long context text analysis. I used novels of 100k+ token as kinda needle test. I asked it questions like "how many people has A interacted with? how many opponents has B killed? ... where, when, and who are they ? list them in a table" - stuff like that. And in its thought process, you will see the model occasionally mixing up and hallucinating plots about the characters to varying degrees, even though the names and numbers come out nearly correct. Its understanding of the novel's plot is not perfect, but the summaries of the basic elements are usually almost correct.
Despite the time it takes to crunch the text, it's a fun test.
In coding, I'd have to blame this quant. The but wait maybe hold on actually avalanche is a real pain to watch.

Sign up or log in to comment