LLama.cpp parameters suggestion

#5
by tstello - opened

Hi All!!

I love llama server, currently, I am using integrated with opencode specifically for CODE operations, projects, PRS, etc...

Could you please recommend the best llama.cpp configuration for this setup CODING TASKS?

My stack:
Hardware:

  • CPU: Intel Core i9-14900K, 24 cores / 32 threads
  • RAM: 188 GiB total, approximately 60 GiB available
  • GPU 0: NVIDIA RTX PRO 4000 Blackwell, 24,467 MiB VRAM total, 21,580 MiB used, 2,407 MiB free, 48% utilization, 81°C
  • GPU 1: NVIDIA GeForce RTX 4090, 24,564 MiB VRAM total, 22,429 MiB used, 1,652 MiB free, 40% utilization, 53°C
  • NVIDIA driver: 580.173.02
  • OS: Ubuntu 24.04.4 LTS
  • Kernel: 6.8.0-137-generic

Software:

  • llama.cpp Docker image: ghcr.io/ggml-org/llama.cpp:server-cuda
  • llama.cpp version: 0.3.0-dev, build 10731, commit 0eadefebd
  • Model: Qwen3.8-27B-Q8_0.gguf
  • Quantization: Q8_0
  • Multimodal projector: mmproj-F16.gguf
  • Context size: 260,000 tokens
  • Server port: 8092
  • Parallel slots: 1
  • Flash attention: enabled
  • CUDA offload: all layers
  • Speculative decoding: draft-mtp
  • Reasoning: enabled, low effort, 512-token budget
  • Metrics: enabled
  • Log verbosity: default level 3

llama.cpp parameters:

-m /models/Qwen3.8-27B-GGUF/Qwen3.8-27B-Q8_0.gguf
--mmproj /models/Qwen3.8-27B-GGUF/mmproj-F16.gguf
--n-gpu-layers -1
--split-mode layer
--main-gpu 0
--parallel 1
-c 260000
--threads 24
--threads-batch 32
--batch-size 2048
--ubatch-size 512
--temperature 0.7
--top-p 0.95
--top-k 20
--min-p 0.00
--presence-penalty 0.0
--repeat-penalty 1.0
--kv-unified
--cache-type-k q8_0
--cache-type-v q8_0
--sleep-idle-seconds 1800
--flash-attn on
--spec-type draft-mtp
--spec-draft-n-max 2
--spec-draft-ngl all
--reasoning on
--reasoning-effort low
--reasoning-budget 512
--jinja
--metrics

Observed performance:

  • Average generation speed: approximately 27.5 tokens per second
  • Typical generation speed: approximately 20–43 tokens per second
  • Draft acceptance rate: approximately 80–95%
  • The logs also show repeated prompt-cache eviction messages.
  • No out-of-memory, fatal, panic, or assertion errors were found.

Could you please advise:
Based on your experience, If you suggest different parameters to optimize my stack to have more performance or if I doing something wrong, please fell free to suggest.

I am using llama for + 2 years, I would like to improve my daily operations and learn more about llama.cpp best practices.

Best regards,
Tiago S.

Sign up or log in to comment