smalinin commited on
Commit
6eefbdc
·
verified ·
1 Parent(s): 44479c4

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +32 -39
README.md CHANGED
@@ -10,6 +10,7 @@ tags:
10
  - deepseek-v4.1
11
  - llama.cpp
12
  quantized_by: vcruz305
 
13
  ---
14
 
15
  # DeepSeek-V4.1-Flash GGUF
@@ -22,49 +23,41 @@ tables, hyper-connections and sparse attention. It is not V4-Flash-0731.
22
  ## Recipe
23
 
24
  How to build the engine, serve it, and the gotchas, plus the current status:
25
- [vcruz305/DeepSeek-V4.1-Flash-GGUF-DGX-Spark-recipe](https://github.com/vcruz305/DeepSeek-V4.1-Flash-GGUF-DGX-Spark-recipe)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
26
 
27
- ## Status
28
-
29
- **These files do not run on upstream llama.cpp yet.** Conversion works and is open as
30
- [ggml-org/llama.cpp#28696](https://github.com/ggml-org/llama.cpp/pull/28696). The runtime is in
31
- progress on the `runtime/deepseek41` branch of
32
- [vcruz305/llama.cpp](https://github.com/vcruz305/llama.cpp): the loader, the engram tables and the
33
- hyper-connections work and are verified against the reference implementation, and the sparse
34
- attention is the remaining piece.
35
 
36
- Weights land here as each rung finishes. Anything converted before 2026-09-10 carries
37
- `general.architecture = deepseek4` and is being redone as `deepseek41`.
38
-
39
- The architecture string is `deepseek41`, following llama.cpp's habit of dropping the `_v`
40
- (`deepseek_v2` became `deepseek2`, `deepseek_v3.2` became `deepseek32`).
41
 
42
- **2026-09-11 fix:** the 4 Engram KV keys (`head_count`, `key_length`, `max_ngram_size`,
43
- `layer_ids`) were written with a hardcoded `deepseek4.engram.*` prefix instead of resolving
44
- `{arch}.engram.*` like every other arch-scoped key in the file. `general.architecture` and all
45
- 38 other arch-scoped keys were already correct (`deepseek41.*`); only these 4 were wrong, which
46
- would have made the `runtime/deepseek41` loader fail to find Engram config on an otherwise
47
- loadable file. Fixed in place via a KV-only rewrite (tensor data untouched, verified
48
- byte-identical by SHA-256) on all five quant rungs' first shard, where GGUF split files store
49
- metadata. Confirmed live: all five now read `deepseek41.engram.*`.
50
 
51
  ## Files
52
 
53
- Ladder in order: **Q2_K, Q3_K_M, Q4_K_M**. Measured tensor payload:
54
-
55
- | File | Quant | Bytes | GiB |
56
- | --- | --- | ---: | ---: |
57
- | `DeepSeek-V4.1-Flash-Q2_K.gguf` | Q2_K | 264,514,761,248 | 246.3 |
58
- | `DeepSeek-V4.1-Flash-Q3_K_M.gguf` | Q3_K_M | 347,270,954,112 | 323.4 |
59
- | `DeepSeek-V4.1-Flash-Q4_K_M.gguf` | Q4_K_M | pending | |
60
-
61
- Split into parts, since each exceeds the Hub's single file limit.
62
-
63
- Q5_K_M is skipped unless asked for. The routed experts arrive as MXFP4 at 4.25 bpw, so higher rungs
64
- move parts of the mixture *up* rather than down: Q3_K_M already lands at 0.684 of the Q8_0 staging
65
- file, and Q5_K_M would be close enough to Q8_0 to be poor value.
66
-
67
- Most of the file is the two engram tables, roughly 196.6B parameters between them. They follow the
68
- rung, 99,611 to 40,284 MiB each between q8_0 and q3_K.
69
 
70
- Apache/MIT from upstream. Credit: DeepSeek. GGUF pack: Victor Cruz (`vcruz305`).
 
10
  - deepseek-v4.1
11
  - llama.cpp
12
  quantized_by: vcruz305
13
+ quality recovered: smalinin
14
  ---
15
 
16
  # DeepSeek-V4.1-Flash GGUF
 
23
  ## Recipe
24
 
25
  How to build the engine, serve it, and the gotchas, plus the current status:
26
+ [https://github.com/smalinin/llama.cpp/tree/my_build_deepseek41](https://github.com/smalinin/llama.cpp/tree/my_build_deepseek41)
27
+
28
+ ```
29
+ ./llama-server \
30
+ --model ./DeepSeek-V4.1-Flash-Q2_K-00001-of-00007.gguf \
31
+ --host 127.0.0.1 \
32
+ --port 8080 \
33
+ --ctx-size 64000 \
34
+ --batch-size 2048 \
35
+ --ubatch-size 256 \
36
+ --parallel 1 \
37
+ --n-gpu-layers auto \
38
+ --fit on \
39
+ --fit-ctx 64000 \
40
+ --fit-target 2048 \
41
+ --load-mode mmap \
42
+ --lazy-mode auto \
43
+ --flash-attn on \
44
+ --no-warmup \
45
+ --no-context-shift \
46
+ --jinja \
47
+ --chat-template-file ./models/templates/deepseek-ai-DeepSeek-V4.1.jinja \
48
+ --chat-template-kwargs '{"reasoning_effort":80,"enable_thinking":true}' \
49
+ --reasoning-format deepseek \
50
+ --no-reasoning-preserve \
51
+ --no-prefill-assistant
52
+
53
+ ```
54
 
 
 
 
 
 
 
 
 
55
 
56
+ ## Status
 
 
 
 
57
 
58
+ **These files do not run on upstream llama.cpp yet.**
 
 
 
 
 
 
 
59
 
60
  ## Files
61
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
62
 
63
+ Apache/MIT from upstream.