Garbled response from Q8_0

#3
by cnsiva - opened

I've tested Qwythos-9B-v2-Q8_0.gguf and I'm getting garbled responses. My GPU is an RTX 4070 Ti Super. Here is my configuration. The same configuration works with Q6_K

llama-server
--model ~/models/empero-ai/Qwythos-9B-v2-GGUF/Qwythos-9B-v2-Q8_0.gguf
--alias qwythos-9b-pro
--path ~/github-sources/llama.cpp/build/tools/ui/dist
--mmproj ~/models/empero-ai/Qwythos-9B-v2-GGUF/mmproj-Qwythos-9B-v2-BF16.gguf
--n-gpu-layers 999
--n-cpu-moe 0
--cache-type-k q8_0
--cache-type-v q8_0
--parallel 4
--cache-ram 24576
--cache-reuse 256
--kv-unified
--cache-idle-slots
--context-shift
--slot-save-path ~/.cache/llama-slots/qwythos-9b
--cont-batching
--threads 8
--threads-batch 8
--threads-http 4
--prio 2
--prio-batch 1
--numa isolate
--flash-attn on
--ctx-size 262144
--batch-size 4096
--ubatch-size 512
--reasoning auto
--reasoning-format auto
--reasoning-budget -1
--reasoning-budget-message " [Logic Finalized] "
--jinja
--temp 0.6
--top-p 0.95
--top-k 20
--host 0.0.0.0
--port 8080
--log-disable
--metrics

Try without kv cache quanization. q35 architecture is very sensitive to it.

It works fine with the MTP version.

Hey @cnsiva

We have tested the model locally and did not find any issues with the Q8 variant.

We would recommend you to remove the cache quantization and the reasoning budget

The MTP variant works seamlessly with Q8_0. However, I’ve noticed that the file/picture input doesn’t function correctly ( —mmproj ~/models/empero-ai/Qwythos-9B-v2-GGUF/mmproj-Qwythos-9B-v2-BF16.gguf).
On the other hand, it performs exceptionally well without the mmproj, achieving a remarkable speed of 120+ t/s, which is truly impressive.

Here is my llama.cpp service file for this model without ( —mmproj ~/models/empero-ai/Qwythos-9B-v2-GGUF/mmproj-Qwythos-9B-v2-BF16.gguf).

[Unit]
Description=Llama.cpp Server - Config: qwythos-9b.conf
After=network.target nss-lookup.target
Wants=nvidia-suspend.service nvidia-hibernate.service nvidia-resume.service

[Service]
Type=simple
User=siva
CPUAffinity=0-7,16-23
LimitMEMLOCK=infinity

Environment=GGML_CUDA_REGISTER_HOST=1

ExecStart=~/github-sources/llama.cpp/build/bin/llama-server
--model ~/models/empero-ai/Qwythos-9B-v2-GGUF/Qwythos-9B-v2-MTP-Q6_K.gguf
--alias qwythos-9b-pro
--path ~/github-sources/llama.cpp/build/tools/ui/dist
--n-gpu-layers 999
--n-cpu-moe 0
--cache-type-k q8_0
--cache-type-v q8_0
--parallel 6
--cache-ram 24576
--cache-reuse 256
--kv-unified
--cache-idle-slots
--slot-save-path /home/siva/.cache/llama-slots/qwythos-9b
--cont-batching
--threads 8
--threads-batch 8
--threads-http 4
--prio 2
--prio-batch 1
--numa isolate
--flash-attn on
--ctx-size 262144
--batch-size 4096
--ubatch-size 512
--reasoning auto
--reasoning-format auto
--reasoning-budget -1
--reasoning-budget-message "[Logic Finalized]"
--jinja
--temp 0.6
--top-p 0.95
--top-k 20
--host 0.0.0.0
--port 8080
--log-disable
--spec-type draft-mtp
--spec-draft-n-max 3
--image-min-tokens 1024
--metrics

Restart=always
RestartSec=5
StandardOutput=journal
StandardError=journal
SyslogIdentifier=llama-server

[Install]
WantedBy=multi-user.target

Sign up or log in to comment