.gitattributes CHANGED
@@ -64,14 +64,3 @@ BF16/gemma-4-26B-A4B-it-BF16-00002-of-00002.gguf filter=lfs diff=lfs merge=lfs -
64
  gemma-4-26B-A4B-it-UD-IQ4_S.gguf filter=lfs diff=lfs merge=lfs -text
65
  gemma-4-26B-A4B-it-UD-IQ4_NL_XL.gguf filter=lfs diff=lfs merge=lfs -text
66
  gemma-4-26B-A4B-it-UD-Q4_K_S_XS.gguf filter=lfs diff=lfs merge=lfs -text
67
- MTP/gemma-4-26B-A4B-it-MTP-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
68
- MTP/gemma-4-26B-A4B-it-MTP-BF16.gguf filter=lfs diff=lfs merge=lfs -text
69
- MTP/gemma-4-26B-A4B-it-MTP-F16.gguf filter=lfs diff=lfs merge=lfs -text
70
- mtp-gemma-4-26B-A4B-it-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
71
- MTP/mtp-gemma-4-26B-A4B-it-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
72
- MTP/mtp-gemma-4-26B-A4B-it-BF16.gguf filter=lfs diff=lfs merge=lfs -text
73
- MTP/mtp-gemma-4-26B-A4B-it-F16.gguf filter=lfs diff=lfs merge=lfs -text
74
- mtp-gemma-4-26B-A4B-it.gguf filter=lfs diff=lfs merge=lfs -text
75
- MTP/gemma-4-26B-A4B-it-Q8_0-MTP.gguf filter=lfs diff=lfs merge=lfs -text
76
- MTP/gemma-4-26B-A4B-it-BF16-MTP.gguf filter=lfs diff=lfs merge=lfs -text
77
- MTP/gemma-4-26B-A4B-it-F16-MTP.gguf filter=lfs diff=lfs merge=lfs -text
 
64
  gemma-4-26B-A4B-it-UD-IQ4_S.gguf filter=lfs diff=lfs merge=lfs -text
65
  gemma-4-26B-A4B-it-UD-IQ4_NL_XL.gguf filter=lfs diff=lfs merge=lfs -text
66
  gemma-4-26B-A4B-it-UD-Q4_K_S_XS.gguf filter=lfs diff=lfs merge=lfs -text
 
 
 
 
 
 
 
 
 
 
 
BF16/gemma-4-26B-A4B-it-BF16-00001-of-00002.gguf CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:56fb42adfb78ab308cd81675bc0f6ce4d4f3d9cccfbff4e9b053713eed976374
3
- size 49923215552
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:39822285661ea71080dc15832e6603c06ce5c19613e20d59ebe108de1d0b2c62
3
+ size 49923213088
MTP/README.md DELETED
@@ -1,56 +0,0 @@
1
- # Gemma 4 26B-A4B MTP drafter
2
-
3
- Multi-Token Prediction (MTP) drafter for `unsloth/gemma-4-26B-A4B-it-GGUF`. It runs as a speculative draft model that shares the target's KV cache and speeds up text generation.
4
-
5
- Verified on a single B200 with the `gemma-4-26B-A4B-it-UD-Q4_K_M.gguf` target: 159 tok/s without MTP, 257 tok/s with MTP, 0.72 draft acceptance.
6
-
7
- MTP was merged into llama.cpp on 2026-06-07 (PR ggml-org/llama.cpp#23398). You need a llama.cpp build from after that date. Older builds cannot load these (arch `gemma4-assistant`).
8
-
9
- ## Files
10
-
11
- For `-hf` auto-discovery a Q8_0 drafter sits at the repo root as `mtp-gemma-4-26B-A4B-it.gguf`. The same three precisions live in `MTP/`:
12
-
13
- - `mtp-gemma-4-26B-A4B-it-Q8_0.gguf` (smallest, recommended; mirrored at the repo root as `mtp-gemma-4-26B-A4B-it.gguf`)
14
- - `mtp-gemma-4-26B-A4B-it-BF16.gguf`
15
- - `mtp-gemma-4-26B-A4B-it-F16.gguf`
16
-
17
- ## Build llama.cpp
18
-
19
- ```bash
20
- git clone https://github.com/ggml-org/llama.cpp
21
- cd llama.cpp
22
-
23
- # CUDA build. Set the arch for your GPU: 89 (RTX 4090), 90 (H100), 100 (B200).
24
- cmake -B build -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=90
25
- cmake --build build --config Release -j --target llama-server
26
- ```
27
-
28
- ## Run, the easy way
29
-
30
- A recent llama.cpp finds the drafter automatically from the root `mtp-` file, so `-hf` is all you need. No `--model-draft`.
31
-
32
- ```bash
33
- ./build/bin/llama-server \
34
- -hf unsloth/gemma-4-26B-A4B-it-GGUF:UD-Q4_K_M \
35
- --spec-type draft-mtp --spec-draft-n-max 4 \
36
- -ngl 999 -fa on
37
- ```
38
-
39
- If your build is too old to auto-discover the sibling, use the explicit form below.
40
-
41
- ## Run with an explicit drafter
42
-
43
- Use this to choose a precision or point at a local file.
44
-
45
- ```bash
46
- hf download unsloth/gemma-4-26B-A4B-it-GGUF gemma-4-26B-A4B-it-UD-Q4_K_M.gguf --local-dir .
47
- hf download unsloth/gemma-4-26B-A4B-it-GGUF MTP/mtp-gemma-4-26B-A4B-it-Q8_0.gguf --local-dir .
48
-
49
- ./build/bin/llama-server \
50
- -m gemma-4-26B-A4B-it-UD-Q4_K_M.gguf \
51
- --model-draft MTP/mtp-gemma-4-26B-A4B-it-Q8_0.gguf \
52
- --spec-type draft-mtp --spec-draft-n-max 4 \
53
- -ngl 999 -fa on
54
- ```
55
-
56
- Multi GPU: add `--spec-draft-device CUDA0 -sm layer`. The drafter pairs with any quant of the 26B-A4B. Quantized KV cache works.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
MTP/mtp-gemma-4-26B-A4B-it-F16.gguf DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:36bf2f6710cf06ff1a5d026cca49ba88bd3451d8588a484ca5a79c5e30aa45f2
3
- size 855228576
 
 
 
 
MTP/mtp-gemma-4-26B-A4B-it-Q8_0.gguf DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:6326fb9f5e487aa8dcdd313a091e3c67724cb2a666ec3b7d2895b5b26d93ed1b
3
- size 461766816
 
 
 
 
README.md CHANGED
@@ -27,7 +27,6 @@ tags:
27
  </div>
28
 
29
  <ul style="margin: 0;">
30
- <li><b>Jun 9 Update:</b> Added MTP support. See our <a href="https://unsloth.ai/docs/models/mtp">MTP Guide</a>.</li>
31
  <li><b>Apr 11 Update:</b> Re-download for Google's latest chat template and llama.cpp fixes.</li>
32
  <li>Gemma 4 can now be run and fine-tuned in <a href="https://unsloth.ai/docs/new/studio">Unsloth Studio</a>. <a href="https://unsloth.ai/docs/models/gemma-4">Read our guide</a>.</li>
33
  <li>See all versions of Gemma 4 (GGUF, 16-bit etc.) <a href="https://huggingface.co/collections/unsloth/gemma-4">in our collection</a>.</li>
 
27
  </div>
28
 
29
  <ul style="margin: 0;">
 
30
  <li><b>Apr 11 Update:</b> Re-download for Google's latest chat template and llama.cpp fixes.</li>
31
  <li>Gemma 4 can now be run and fine-tuned in <a href="https://unsloth.ai/docs/new/studio">Unsloth Studio</a>. <a href="https://unsloth.ai/docs/models/gemma-4">Read our guide</a>.</li>
32
  <li>See all versions of Gemma 4 (GGUF, 16-bit etc.) <a href="https://huggingface.co/collections/unsloth/gemma-4">in our collection</a>.</li>
gemma-4-26B-A4B-it-MXFP4_MOE.gguf CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:a8908250f2a72a5824382d488158f12b65effa24e6c6b1244e5b4818ac0a1459
3
- size 16551048928
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ba25642c352fda86a38ae47114b9603e8ec1fb527e0a2ddf458001fcc53aa866
3
+ size 16630345024
gemma-4-26B-A4B-it-Q8_0.gguf CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:5f7cbd0f4564e84342fc34321a09acb54b1a3da9215124e5bf444baa6dda152c
3
- size 26859861728
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:af5f307a0d72b99be5fa547573ea161dafc7164467f33601598a6e64a43f6d80
3
+ size 26859859264
gemma-4-26B-A4B-it-UD-IQ2_M.gguf CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:3fe245f4436ab44c103291222e81c514e53ca382871b555df24fea11e36b3c3a
3
- size 10014755296
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:394f5074fc5d07af610c5bec113ac2979de542d80da20e290ce70bf0572ef7b0
3
+ size 9974943040
gemma-4-26B-A4B-it-UD-IQ2_XXS.gguf CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:52f98e6fc6df62438dff8f57ac049f60b4c36acf93786b6aeeedc129acc02343
3
- size 9922480608
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:bd1866a6ae011422534d53f78a259d3f4ffd24f02e112e2c2e8dbd49aafa32a9
3
+ size 9882668352
gemma-4-26B-A4B-it-UD-IQ3_S.gguf CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:878be93f9c238ea853b3fd1eb602637ce3cf1cddea56dc345d9a7bf2d6093e29
3
- size 11289671136
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:2f8dc9c2c7756e115f10f38983f5369df036ac24a82c09c668014bea249b1b44
3
+ size 11219406656
gemma-4-26B-A4B-it-UD-IQ3_XXS.gguf CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:3b7aec683e176e1d957d007e0fb590611d82aec4db358d99b7f9714b6bc7bfc2
3
- size 11416548832
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:3e54bc50c8a25d4e761abaeba13205a7956b3d9d1f7b10482adc6156c5c5276b
3
+ size 11219406656
gemma-4-26B-A4B-it-UD-IQ4_NL.gguf CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:eeb867f279ea5a3d52a0dc15fe8ada677b3328a328530957f3f9a5da93cb10b8
3
- size 13613037280
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:0fa0c968cda1c34f95a53dad4ebfc1a7c12628ca38185f352dbe043436064343
3
+ size 13418753344
mtp-gemma-4-26B-A4B-it.gguf → gemma-4-26B-A4B-it-UD-IQ4_NL_XL.gguf RENAMED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:6326fb9f5e487aa8dcdd313a091e3c67724cb2a666ec3b7d2895b5b26d93ed1b
3
- size 461766816
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c1ac61e3af04b40c7688dfe0686c54d68ff5ed26178a56f0b18c8a0e2b94db83
3
+ size 14556687680
gemma-4-26B-A4B-it-UD-IQ4_XS.gguf CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:babd1e389d386352f71600765d37390f7dc993fbfad6725caccf996ffe34aecf
3
- size 13597177568
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ab05e3504fae607c205dae0117d58e3fddf0e88f968a0fbf307437c68398bad9
3
+ size 13418753344
gemma-4-26B-A4B-it-UD-Q2_K_XL.gguf CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:2a1d26dfe6ea00a467940a5728316af6edb366bbdba950d65b85d232392fb658
3
- size 10546934240
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:1e3a85588478801c24759be68f139ae6dff9a1f092ff06ede0f4b9a4ad5a5cb3
3
+ size 10545963072
gemma-4-26B-A4B-it-UD-Q3_K_M.gguf CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:811882f6df81c8ea204a5d382f196d6bd601245091786912d4c70df41cdfa052
3
- size 12728497888
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:6296b39d581668c4b654617d6d49c9edca15635fb803c3c9f6d813771df85d1e
3
+ size 12526284096
MTP/mtp-gemma-4-26B-A4B-it-BF16.gguf → gemma-4-26B-A4B-it-UD-Q3_K_S.gguf RENAMED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:01cdef467d12d51474b821542eb113cd8ac6ac3649b2a9f5f0a6f3a2eacd59a0
3
- size 855228576
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:04a284239b76d92e473d4923cc6ee10fe406a8a91042d4810af16905300a4e28
3
+ size 12526284096
gemma-4-26B-A4B-it-UD-Q3_K_XL.gguf CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:90a918830420e6a36e01c0d219e563f1d0ca7f223dffd90ad4e02ef4f3253fde
3
- size 12907280096
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:5798a39ad0b64bb054c8ef6c0885d089885a948a82da21424219b6935c18b9f6
3
+ size 12875558208
gemma-4-26B-A4B-it-UD-Q4_K_M.gguf CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:f2c28b3dc4776931ac6f879e11f203dec637ea0f14267a86ec8f6165f63f293f
3
- size 16947541728
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:b8707e57f676d8dd1b80f623b45200cc92e6966b0e95275e606f412095a49fde
3
+ size 16868240704
gemma-4-26B-A4B-it-UD-Q4_K_S.gguf CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:675de4b188a3072c9da8260fe2eddfdb7d03b16f615b6619021989d929a51b2b
3
- size 16487610080
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c0d205a16df9359b8d97e8c53824a0c4200d3d8e2a0c7629065f15ca78b708e7
3
+ size 16392449344
gemma-4-26B-A4B-it-UD-Q4_K_XL.gguf CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:ef728c8e0c337fd1067b947af006e38a9ef2419e56feced4fd29b4bf0636e30c
3
- size 17010980576
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:87121467dee16160e5678be823187079784ccda8f20d3ad2f77a77739581a148
3
+ size 17090276672
gemma-4-26B-A4B-it-UD-Q5_K_M.gguf CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:769d386a69d43782321c1bad04d41d29a2e84b2c06e6a277cd99fd6265ec0e80
3
- size 21150365408
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:f19acf0c0329932bc8e66ee904398140569658ccb65075182f198b3b07edcf7f
3
+ size 21150362944
gemma-4-26B-A4B-it-UD-Q5_K_S.gguf CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:ce87883f935900407ce0cb4b1fc4d423b6af8c2970392387cc4bc0ae8e572825
3
- size 18850707168
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:8645e8668f98d13d1d9bb58e2d3672fb4debb1ead401ca3adc12a6ddb442426f
3
+ size 18771406144
gemma-4-26B-A4B-it-UD-Q5_K_XL.gguf CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:d7553b64566bcb0b8b7d712e12e80213f3f02e756ab17099628925ce2e39e597
3
- size 21217769184
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ffb15f755bcda2dced5524490ce8a4106e557879b3e6f1a1d0d225016205bd64
3
+ size 21217766720
gemma-4-26B-A4B-it-UD-Q6_K.gguf CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:76bf3eee49cca5b076b34a4b575ecd505dabc470512b46f1207b4f9da08385f0
3
- size 23172478688
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:59ac620521cd8519ad4720f09ea58bce4b9f0c35684f5c062666cd6562c82f31
3
+ size 23172476224
gemma-4-26B-A4B-it-UD-Q6_K_XL.gguf CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:b01ee10a1423c17f9c4384f1fc569726b8782c5403557ff138ceb9468ca49d6b
3
- size 23295391456
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:8a8697b42b455ace3853bd5eca725a8a071757f3ae0678cab932918211e394db
3
+ size 23295388992
gemma-4-26B-A4B-it-UD-Q8_K_XL.gguf CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:d4bf9791d727d7b88aeea89aba309c68086a4d51cf337047c4e51dde7e243058
3
- size 27636232928
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:9c961effefe478b33e4f4f63ba225aa51b3ab11aa3d68fe2ad34ce4f082a241b
3
+ size 27866185024
imatrix_unsloth.gguf_file CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:8065033ecc77d819b6bca2fe1ae5e0b0e63b2ab2cb1221663a7d7265ee7f28e9
3
  size 56941536
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:6186d77efe82b850b1c52dbd408f722e8bcad0258485191d23655bdb14356f10
3
  size 56941536
mmproj-BF16.gguf CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:41926ed5f1403cf5add23b0684992805ea6f97253096132e769e65646b8cef9d
3
- size 1194828256
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:fc2ebf4c44528daa2cea7b39891712847ca5e4f87dcf578054a06c46bfe6da27
3
+ size 1194828384
mmproj-F16.gguf CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:418a6d8723067cd712235facbbc5cba6c8fbbd413fc1292d2aace5a027d5a42f
3
- size 1193058784
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:90a77a586ccdbabe2b5d28d9da0b61271bdb426c04f723bd01786992f00d317e
3
+ size 1193058912
mmproj-F32.gguf CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:ec31640a1f68fd7883e3ef7ef1afdc98d8b42867ff49ea16649a000447bcf163
3
- size 2291200480
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ce12ca17e5f479ff292cd66817960e4d3f1b09671f744e415c98a55d7725c9ed
3
+ size 2291200608