Instructions to use unsloth/DeepSeek-V4-Flash-Vision-Exp-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use unsloth/DeepSeek-V4-Flash-Vision-Exp-GGUF with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("unsloth/DeepSeek-V4-Flash-Vision-Exp-GGUF", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use unsloth/DeepSeek-V4-Flash-Vision-Exp-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf unsloth/DeepSeek-V4-Flash-Vision-Exp-GGUF:UD-Q4_K_XL # Run inference directly in the terminal: llama cli -hf unsloth/DeepSeek-V4-Flash-Vision-Exp-GGUF:UD-Q4_K_XL
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf unsloth/DeepSeek-V4-Flash-Vision-Exp-GGUF:UD-Q4_K_XL # Run inference directly in the terminal: llama cli -hf unsloth/DeepSeek-V4-Flash-Vision-Exp-GGUF:UD-Q4_K_XL
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf unsloth/DeepSeek-V4-Flash-Vision-Exp-GGUF:UD-Q4_K_XL # Run inference directly in the terminal: ./llama-cli -hf unsloth/DeepSeek-V4-Flash-Vision-Exp-GGUF:UD-Q4_K_XL
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf unsloth/DeepSeek-V4-Flash-Vision-Exp-GGUF:UD-Q4_K_XL # Run inference directly in the terminal: ./build/bin/llama-cli -hf unsloth/DeepSeek-V4-Flash-Vision-Exp-GGUF:UD-Q4_K_XL
Use Docker
docker model run hf.co/unsloth/DeepSeek-V4-Flash-Vision-Exp-GGUF:UD-Q4_K_XL
- LM Studio
- Jan
- Ollama
How to use unsloth/DeepSeek-V4-Flash-Vision-Exp-GGUF with Ollama:
ollama run hf.co/unsloth/DeepSeek-V4-Flash-Vision-Exp-GGUF:UD-Q4_K_XL
- Unsloth Desktop
- Pi
How to use unsloth/DeepSeek-V4-Flash-Vision-Exp-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf unsloth/DeepSeek-V4-Flash-Vision-Exp-GGUF:UD-Q4_K_XL
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "unsloth/DeepSeek-V4-Flash-Vision-Exp-GGUF:UD-Q4_K_XL" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use unsloth/DeepSeek-V4-Flash-Vision-Exp-GGUF with Docker Model Runner:
docker model run hf.co/unsloth/DeepSeek-V4-Flash-Vision-Exp-GGUF:UD-Q4_K_XL
- Lemonade
How to use unsloth/DeepSeek-V4-Flash-Vision-Exp-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull unsloth/DeepSeek-V4-Flash-Vision-Exp-GGUF:UD-Q4_K_XL
Run and chat with the model
lemonade run user.DeepSeek-V4-Flash-Vision-Exp-GGUF-UD-Q4_K_XL
List all available models
lemonade list
- Hermes Agent
How to use unsloth/DeepSeek-V4-Flash-Vision-Exp-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf unsloth/DeepSeek-V4-Flash-Vision-Exp-GGUF:UD-Q4_K_XL
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default unsloth/DeepSeek-V4-Flash-Vision-Exp-GGUF:UD-Q4_K_XL
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use unsloth/DeepSeek-V4-Flash-Vision-Exp-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf unsloth/DeepSeek-V4-Flash-Vision-Exp-GGUF:UD-Q4_K_XL
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "unsloth/DeepSeek-V4-Flash-Vision-Exp-GGUF:UD-Q4_K_XL" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Does the current model have dspark support?
DeepSeek-V4-Flash-0731-GGUF contains a "dspark" model, but the current model does not exist. May I ask if it is supported?
It works for me - just use the original dspark gguf that was provided with the non-vision unsloth model.
I was able to
app-b10796-mix-659e406-windows-x64-cuda13-portable\llama-server.exe (switches not specific to the execution)
-hf unsloth/DeepSeek-V4-Flash-Vision-Exp-GGUF:UD-Q8_K_XL --ctx-size 262144 --temp 1.0 --top-p 0.95 --reasoning on --cache-type-k q8_0 --cache-type-v q8_0 --fit on --spec-type draft-dspark --spec-draft-n-max 4 -hfd unsloth/DeepSeek-V4-Flash-0731-GGUF:Q8_0 -lv 4 --seed 7417 -np 1
and it pulled the dspark from the other repo like I was hoping...
but this ran at 15 token/s whereas without dspark i get ~30 token/s
so idk what is going on there
Unstudio log says: "DSpark requested but no matching dspark-*.gguf sidecar was found; loading without speculative decoding."
But there is no sidecard, might try the sidecar from the non-vision model, but that's shaky at best. Looks like dspark is baked/fused into the .gguf, but doesn't work with unsloth studio/llama.cpp at all, and is only for vLLM and Sglang running on nvidia.
Per my Qwen:
The Vision-Exp model does have DSpark — but it's fused into the checkpoint, not a separate sidecar file. Unlike the 0731 base (where unsloth ships a standalone dspark-DeepSeek-V4-Flash-0731-BF16.gguf
Why you can't just switch it on: the runtimes that read that fused module are:
vLLM: --speculative-config '{"method":"dspark","model":"deepseek-ai/DeepSeek-V4-Flash-Vision-Exp","num_speculative_tokens":3,...}'
SGLang: --speculative-algorithm DSPARK (no draft path)
…both of which, for the vision model specifically, are pre-release and NVIDIA-only.
I'm guessing someone is working on this, but it might be nice to note in the model card. Especially since it's only noted in the log, not in the UI other than slower performance compared to the original non-vision gguf.
oh what a bummer, lack of this kills it for me.