DGX Spark llama.cpp serving script

#3
by donghwan-shin - opened

Since some of the flags used in README.md are deprecated, I am using the following updated version on my DGX Spark (GB10), in case someone wants to use the same:

#!/bin/bash

llama-server \
  -m /home/donghwan/.cache/huggingface/hub/models--s-batman--Ornith-1.0-35B-NVFP4-MTP-GGUF/snapshots/8113f1b7635c01fab34c7afd8fcbd3ac2962b582/ornith-1.0-35b-NVFP4-MTP.gguf \
  --alias ornith35b \
  --host 0.0.0.0 --port 8080 --metrics \
  -t 20 --no-warmup --load-mode mlock \
  -fa on -ctk q8_0 -ctv q8_0 \
  -b 2048 -ub 2048 -c 1024000 -np 4 -ngl 99 \
  --chat-template-file /home/donghwan/models/chat_template.jinja \
  --spec-type draft-mtp --spec-draft-n-max 3 \
  --temp 0.9 --top-p 0.95 --top-k 20 --min-p 0.01 --repeat-penalty 1.0

Note that I am using 4 parallel slots, with a 256k context length each.

Sign up or log in to comment