1.40 to .48 token

#2
by BrokeAmerican - opened

I was hoping to go faster. But my 72gb vram and 64gb system ram went slower with mtp. Unless I did something wrong. The smaller models nearly doubled though line qwen 27b and 122b. I was wishing I could run this model.

I was hoping to go faster. But my 72gb vram and 64gb system ram went slower with mtp. Unless I did something wrong. The smaller models nearly doubled though line qwen 27b and 122b. I was wishing I could run this model.

MTP currently increases memory consumption. If a model previously barely fit in memory, then with MTP enabled, due to insufficient RAM it may start offloading to SSD, for example, resulting in performance degradation.
In my case, MTP improved token generation (TG) speed from 12 tokens/sec to 15 tokens/sec using draft=3. However, prompt processing (PP) speed dropped by almost half—from 150 tokens/sec down to 80 tokens/sec.
That's why I currently run two versions of the model: one with MTP and one without. For long prompts, I use the version without MTP (at least until they figure out how to optimize PP speed). For short prompts where the model needs to do extensive reasoning/generation, I switch to the MTP-enabled version.

before the merge I used https://github.com/ikawrakow/ik_llama.cpp

with setting like:

@echo off
"E:\llama_ai\llama_ikawrakowik-b1111-bin-cuda-13.1\llama-server.exe" ^
  -m "E:\llama_ai\models\Qwen3.5-397B-A17B\UD-IQ3_XSS\Qwen3.5-397B-A17B-UD-IQ3_XXS-00001-of-00004.gguf" ^
  --alias "Qwen3.5-397B-A17B-GGUF:UD-IQ3_XXS" ^
  --gpu-layers 999 ^
  -ot "\.([0-9]|[1-9][0-9]|[0-9][0-9][0-9])\.ffn_(gate|up|down)_exps.=CPU" ^
  --flash-attn on ^
  --no-mmap ^
-cuda fusion=1,offload-batch-size=8,mmq-id-size=1000 ^
  --cache-type-k q8_0 ^
  --cache-type-v q8_0 ^
  --cache-ram 0 ^
  --ctx-size 100536 ^
-ub 8192 ^
-b 8192 ^
 -mqkv  ^
 -muge ^
  --context-shift auto ^
  --slot-save-path "E:\llama_ai\kv_cache\Qwen3.5-397B-A17B" ^
--reasoning on ^
  --threads 16 ^
  --ctx-checkpoints-interval 20192 ^
  --ctx-checkpoints 8 ^
  --recurrent-ckpt-mode gpu-fallback ^
  --parallel 1 ^
  --host 0.0.0.0 ^
  --port 11434 ^
  --seed 3407 ^
  --temp 1.0 ^
  --top-p 0.9 ^
  --min-p 0.01 ^
  --top-k 40 ^
  --jinja

pause

now with the Details
llama + spec: MTP Support (#22673 https://github.com/ggml-org/llama.cpp/pull/22673)
I am confused .. should I use ikawrakow rather to benefit from better PP speed ? I am using a notebook 5090m with 192gb ram (4000M memory though...), and cannot PIN the memory .. the ik_llama helped to boost PP speed like x3 .. but I think this was about mmq having the most effect ..

any advice on the llama.cpp main ? does it support also -cuda fusion options or which should I use best with this unsloth uploads?

Sign up or log in to comment