Instructions to use deepseek-ai/DeepSeek-V4-Flash-0731 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use deepseek-ai/DeepSeek-V4-Flash-0731 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="deepseek-ai/DeepSeek-V4-Flash-0731")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("deepseek-ai/DeepSeek-V4-Flash-0731") model = AutoModelForCausalLM.from_pretrained("deepseek-ai/DeepSeek-V4-Flash-0731", device_map="auto") - Inference
- HuggingChat
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use deepseek-ai/DeepSeek-V4-Flash-0731 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "deepseek-ai/DeepSeek-V4-Flash-0731" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "deepseek-ai/DeepSeek-V4-Flash-0731", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/deepseek-ai/DeepSeek-V4-Flash-0731
- SGLang
How to use deepseek-ai/DeepSeek-V4-Flash-0731 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "deepseek-ai/DeepSeek-V4-Flash-0731" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "deepseek-ai/DeepSeek-V4-Flash-0731", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "deepseek-ai/DeepSeek-V4-Flash-0731" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "deepseek-ai/DeepSeek-V4-Flash-0731", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use deepseek-ai/DeepSeek-V4-Flash-0731 with Docker Model Runner:
docker model run hf.co/deepseek-ai/DeepSeek-V4-Flash-0731
Serving recipe: full 3x1M context on 2x NVIDIA DGX Spark (GB10), 3.39M-token KV pool, 653 Tok/s Peak Decode (list content @ c=16)
For anyone deploying this checkpoint on DGX Spark class hardware: we published a complete serving recipe for 2 units at TP=2, running the full 1,048,576-token context with a 3.39M-token KV pool (3.23x concurrency at 1M) and up to 653 tok/s decode (list content, c=16; ~310 tok/s mixed at c=12).
https://github.com/dkmode22/DeepSeek-V4-Flash-0731-3.23x1M-context-653-toks-2x-DGX-Spark-GB10
Includes the runtime patches (since upstreamed into vLLM 0.26), launch scripts, a bench kit, and the failure modes we hit. One thing worth knowing before you burn a day on it: this checkpoint is much more sensitive to the DSpark shared-expert loader bug than the preview was (acceptance collapses to ~0.14 unpatched), so do not serve it without tonyd2wild's loader fix (included in the repo, with measurements). The recipe as a whole builds on his DSpark serving work for DGX Spark: https://github.com/tonyd2wild/DeepSeek-v4-Flash-0731-DSpark-1M-NVFP4-KV-2x-DGX-Spark
Two things I wonder about this:
- What is average single user (c=1) decode performance at what specific prompt length?
- I see that you are using nvfp4_ds_mla for KV cache, does it keep stable over long context with this model?
For comparison:
https://forums.developer.nvidia.com/t/deepseek-v4-flash-0731-dspark-1m-nvfp4-kv-2x-dgx-spark/378824
Both good questions. We required a maintenance window to verigy as the lane serves 11 production agents. The second one made us go and measure something we'd been asserting rather than proving, so thanks for that. The repo is updated with the full results.
Does nvfp4_ds_mla stay stable over long context? In our testing yes, out to about 994K tokens. We ran a needle ladder on the exact published config, with nothing changed and no relaunch, at 297,404 / 695,018 / 991,668 and 994,397 tokens. Each run had two needles, one at roughly 50% depth and one at 99%, because a single deep needle can be satisfied by a model that only attended to the tail.
All four passed. The two 1M runs were about 19 minutes apart under different background load and agreed on TTFT to 0.19 s (0.02%), so we're reading this as deterministic rather than a lucky pass. Prefill degrades smoothly at roughly -19% per depth doubling, with no cliff.
One caveat worth stating up front: needle retrieval shows the KV holds retrievable information at depth. It does not show that generation quality is unchanged at depth, and we haven't measured that.
c=1 decode, and at what prompt length. Our previously quoted ~70 tok/s was measured against a ~40-token prompt, so it was effectively a zero-context number and didn't answer your question. Measured properly, per content kind, median of 3, with the engine idle-verified before and after every cell:
| Prompt tokens | list | prose |
|---|---|---|
| ~1,000 | 83.55 tok/s (accept 0.97) | 45.25 tok/s (accept 0.43) |
| ~131,072 | 81.19 tok/s (accept 0.96) | 42.14 tok/s (accept 0.40) |
| ~1,000,000 | 75.58 tok/s (accept 0.96) | 36.51 tok/s (accept 0.41) |
list runs 83.55 -> 75.58 tok/s from ~1K to ~1M. That's -9.5% and strictly monotonic; prose is -19.3%. The more useful finding is that speculative-decode acceptance does not degrade with depth (list holds 0.973 -> 0.959 at ~1M), so the falloff looks like attention cost over the resident KV rather than the drafter weakening.
The repo now carries the full matrix (4 kinds x 5 depths x 3 repeats, with acceptance per cell) and both harnesses. That includes a gate fix that makes the long-context probe usable on a long-lived engine. The obvious "is there room in the KV pool" check reads kv_cache_usage_perc, which counts evictable prefix-cache blocks as used, so it refuses every deep request on an engine that would serve them perfectly well.
https://github.com/dkmode22/DeepSeek-V4-Flash-0731-3.23x1M-context-653-toks-2x-DGX-Spark-GB10
I am a bit confused. In the header you mention 3x1M concurrency while in the config it shows max concurrency at 24 and 1M context window. This should not be possible from kv-cache perspective. For now on 2 DGX Sparks cluster I have not managed to go beyond 12 concurrency with any context window as compute and stability is the limit. This is with standard ray setup. If I go with D12x container setup, I managed to increase to 30 unstable and 24 stable concurrency. What is your experience?
We did have stability issues initially which was the reason for the patches in the repo. We are now at 4.5x 1M with the 11 agents with VLLM configured for c=24 and it has been stable. The primary agent lanes are capped at 256K context, however their subagents get full 1M lanes when needed and KV cache sits happily at ~95%. By far the best model we have run to date for agentic workloads.
One caveat worth stating up front: needle retrieval shows the KV holds retrievable information at depth. It does not show that generation quality is unchanged at depth, and we haven't measured that.
Could you please elaborate what you mean or what you measured exactly? Have you quantized needle finding degredation? Because as far as I can read it, you were simply measuring performance only.
Per the repo: 994,397-token prompt with mid- AND deep-needle retrieval verified on this config; monotonic ladder 297,404 / 695,018 / 991,668 / 994,397 tokens, 4/4 PASS, two 1M runs agreeing on TTFT to 0.19 s (0.02%).
I pulled the environment down to capture those details and update the repo when you asked your initial question on this thread. As it has been stable with no issues for over 2 weeks decoding ~30M tokens per day at depth, I won't be pulling it down for any further benchmarking.
Thanks for sharing!
I just asked because I was exploring the impact of different KV cache configuration knobs on a model's output precision. I found e.g. that a simple one-needle search is too simple to tell. Here are some KV cache precision benchmark results and take-aways, performed with Qwen3.8-27B though: