Report: 56 t/s on RTX 4090D (48GB VRAM) with UD-Q6_K_XL

#25
by SlavikF - opened

compare results below to

System:

  • Nvidia RTX 4090D 48GB VRAM
  • Intel Xeon W5-3425 with 12 cores
  • DDR5-4800 RAM

Speed:

  • PP: start with 2200 t/s on small context, goes under 1600 t/s on long context
  • TG: 60 t/s
prompt eval time =   22s / 40564 tokens ( 1808.13 tokens per second)
       eval time =   42s /  2489 tokens (   59.15 tokens per second)

My docker compose:

services:
  llama-router:
    image: ghcr.io/ggml-org/llama.cpp:server-cuda12-b9209
    container_name: router
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              capabilities: [gpu]
    ports:
      - "8080:8080"
    volumes:
      - /home/slavik/.cache/huggingface/hub:/root/.cache/huggingface/hub:ro
      - ./models.ini:/app/models.ini:ro
    entrypoint: ["./llama-server"]
    command: >
      --models-max 1
      --models-preset ./models.ini
      --host 0.0.0.0  --port 8080

my INI file:

version = 1

[unsloth/Qwen3.6-27B-MTP-GGUF:Q6_K_XL]
ctx-size=262144
temp=0.6
top-p=0.95
top-k=20
min-p=0.00
alias=local-vl-qwen27B
spec-type=draft-mtp
spec-draft-n-max=4

using nvtop I see 46.8 GB of VRAM used.

I also ran additional testing to see how the value of spec-draft-n-max affects the speed and VRAM.

Constants:

  • unsloth/Qwen3.6-27B-MTP-GGUF:Q6_K_XL
  • context=196k
  • same query with ~40k tokens prompt

Results:

spec-draft-n-max T/G (t/s) VRAM (GB) PP
no MTP 31.15 38.32 2052
2 56.76 40.88 ~1800
4 60.45 42.10 ~1800
6 59.86 43.26 ~1800

You got the 4090 48gb version! So cool. I get similar speeds. I use a 4090 and 3090 before 27.5 tokens 200k ctx. Around 50 with mtp 4-6. Sometimes 40-66 tokens depending on typical tasks. I bet that 48gb allows you to do great video creations.

@SlavikF Thanks for your report. https://huggingface.co/spaces/oobabooga/accurate-gguf-vram-calculator shows Estimated memory usage: 65432 MiB for full fp16, 262144, context. ?

b7a387788e3f1d4bea1a618bc1d4250faa8647cb
llama-server -hf unsloth/Qwen3.6-27B-MTP-GGUF:UD-Q4_K_XL --alias "qwen3.6:27b" --mlock --no-mmap -ngl all --keep 24576 -cram 16384 --context-shift -c 262144 -ctk q8_0 -ctv q8_0 -np 1 --spec-type draft-mtp --spec-draft-n-max 2 -fa on --host 0.0.0.0 --port 11434

Sign up or log in to comment