danielhanchen commited on
Commit
1f1e542
·
verified ·
1 Parent(s): 365d657

Rename MTP GGUFs

Browse files
.gitattributes CHANGED
@@ -46,3 +46,7 @@ MTP/gemma-4-31B-it-Q8_0-MTP.gguf filter=lfs diff=lfs merge=lfs -text
46
  MTP/gemma-4-31B-it-BF16-MTP.gguf filter=lfs diff=lfs merge=lfs -text
47
  MTP/gemma-4-31B-it-F16-MTP.gguf filter=lfs diff=lfs merge=lfs -text
48
  MTP/gemma-4-31B-it-Q4_0-MTP.gguf filter=lfs diff=lfs merge=lfs -text
 
 
 
 
 
46
  MTP/gemma-4-31B-it-BF16-MTP.gguf filter=lfs diff=lfs merge=lfs -text
47
  MTP/gemma-4-31B-it-F16-MTP.gguf filter=lfs diff=lfs merge=lfs -text
48
  MTP/gemma-4-31B-it-Q4_0-MTP.gguf filter=lfs diff=lfs merge=lfs -text
49
+ MTP/mtp-gemma-4-31B-it-BF16.gguf filter=lfs diff=lfs merge=lfs -text
50
+ MTP/mtp-gemma-4-31B-it-F16.gguf filter=lfs diff=lfs merge=lfs -text
51
+ MTP/mtp-gemma-4-31B-it-Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
52
+ MTP/mtp-gemma-4-31B-it-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
MTP/README.md CHANGED
@@ -11,10 +11,10 @@ MTP was merged into llama.cpp on 2026-06-07 (PR ggml-org/llama.cpp#23398). You n
11
  The recommended drafter is a **smart Q4_0**: the native 4-bit QAT drafter (about 97% of its weights are byte-exact on the int4 grid), near-lossless versus higher precision while roughly half the size. It sits at the repo root as `mtp-gemma-4-31B-it.gguf` so `-hf` finds it automatically, and the same file plus higher-precision drafters are in `MTP/`:
12
 
13
  - `mtp-gemma-4-31B-it.gguf` (repo root, smart Q4_0, recommended; used by `-hf`)
14
- - `MTP/gemma-4-31B-it-Q4_0-MTP.gguf` (same smart Q4_0)
15
- - `MTP/gemma-4-31B-it-Q8_0-MTP.gguf`
16
- - `MTP/gemma-4-31B-it-BF16-MTP.gguf`
17
- - `MTP/gemma-4-31B-it-F16-MTP.gguf`
18
 
19
  ## Build llama.cpp
20
 
@@ -46,11 +46,11 @@ Use this to choose a precision or point at a local file.
46
 
47
  ```bash
48
  hf download unsloth/gemma-4-31B-it-qat-GGUF gemma-4-31B-it-qat-UD-Q4_K_XL.gguf --local-dir .
49
- hf download unsloth/gemma-4-31B-it-qat-GGUF MTP/gemma-4-31B-it-Q8_0-MTP.gguf --local-dir .
50
 
51
  ./build/bin/llama-server \
52
  -m gemma-4-31B-it-qat-UD-Q4_K_XL.gguf \
53
- --model-draft MTP/gemma-4-31B-it-Q8_0-MTP.gguf \
54
  --spec-type draft-mtp --spec-draft-n-max 4 \
55
  -ngl 999 -fa on
56
  ```
 
11
  The recommended drafter is a **smart Q4_0**: the native 4-bit QAT drafter (about 97% of its weights are byte-exact on the int4 grid), near-lossless versus higher precision while roughly half the size. It sits at the repo root as `mtp-gemma-4-31B-it.gguf` so `-hf` finds it automatically, and the same file plus higher-precision drafters are in `MTP/`:
12
 
13
  - `mtp-gemma-4-31B-it.gguf` (repo root, smart Q4_0, recommended; used by `-hf`)
14
+ - `MTP/mtp-gemma-4-31B-it-Q4_0.gguf` (same smart Q4_0)
15
+ - `MTP/mtp-gemma-4-31B-it-Q8_0.gguf`
16
+ - `MTP/mtp-gemma-4-31B-it-BF16.gguf`
17
+ - `MTP/mtp-gemma-4-31B-it-F16.gguf`
18
 
19
  ## Build llama.cpp
20
 
 
46
 
47
  ```bash
48
  hf download unsloth/gemma-4-31B-it-qat-GGUF gemma-4-31B-it-qat-UD-Q4_K_XL.gguf --local-dir .
49
+ hf download unsloth/gemma-4-31B-it-qat-GGUF MTP/mtp-gemma-4-31B-it-Q8_0.gguf --local-dir .
50
 
51
  ./build/bin/llama-server \
52
  -m gemma-4-31B-it-qat-UD-Q4_K_XL.gguf \
53
+ --model-draft MTP/mtp-gemma-4-31B-it-Q8_0.gguf \
54
  --spec-type draft-mtp --spec-draft-n-max 4 \
55
  -ngl 999 -fa on
56
  ```
MTP/{gemma-4-31B-it-BF16-MTP.gguf → mtp-gemma-4-31B-it-BF16.gguf} RENAMED
File without changes
MTP/{gemma-4-31B-it-F16-MTP.gguf → mtp-gemma-4-31B-it-F16.gguf} RENAMED
File without changes
MTP/{gemma-4-31B-it-Q4_0-MTP.gguf → mtp-gemma-4-31B-it-Q4_0.gguf} RENAMED
File without changes
MTP/{gemma-4-31B-it-Q8_0-MTP.gguf → mtp-gemma-4-31B-it-Q8_0.gguf} RENAMED
File without changes