Qwen3.5-122B-A10B + MTP on AMD Strix Halo (27.8 tok/s @ ~23k context)

#1
by rauko - opened

I've been benchmarking Qwen3.5-122B-A10B-Opus-Reasoning-Q4_K_XL + MTP on an AMD Ryzen AI MAX+ 395 (Strix Halo, 128 GB LPDDR5X, ~256 GB/s UMA) using the latest llama.cpp (llama-server build 10205 with ROCm).

After testing multiple combinations my overall best results:

llama-server \
-m Qwen3.5-122B-A10B-Opus-Reasoning-Q4_K_XL.gguf \
--model-draft mtp-opus-q4kxl-draft.gguf \
--spec-type draft-mtp \
--spec-draft-n-max 2 \
--spec-draft-ngl all \
-ngl 999 \
-fa on \
-lm none \
-c 262144 \
--batch-size 8192 \
--ubatch-size 4096 \
--threads 16 \
--threads-batch 16 \
--parallel 1 \
--cache-type-k q8_0 \
--cache-type-v q8_0 \
--cache-ram 32768 \
--jinja \
--chat-template-file chat_template.jinja

Results:

I'd be interested to compare results with other Strix Halo owners, especially if anyone has found additional optimizations for ROCm or llama.cpp.

679 tok/s prefill · 37 tok/s decode - 64k context using Vulkan, Strix Halo 128GB (framework desktop board)

/home/sixvolts/llama-b10066-vulkan/llama-b10066/llama-server
-m /home/sixvolts/models/qwen-3.5-122B-opus-mtp/Qwen3.5-122B-A10B-Opus-Reasoning-Q4_K_XL.gguf
-md /home/sixvolts/models/qwen-3.5-122B-opus-mtp/mtp-opus-q4kxl-draft.gguf
--spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-n-min 1
-ngl 99 -fa on --no-mmap
--parallel 6 -c 393216 --no-kv-unified
-ctk q8_0 -ctv q8_0 -ctkd q8_0 -ctvd q8_0
-t 12 -b 2048 -ub 512
--jinja
--host 0.0.0.0 --port 8080
--alias qwen3.5-122b-opus

Sign up or log in to comment