ubergarm commited on
Commit
3f0415f
·
1 Parent(s): ef6d1c1

Add IQ4_KT for full GPU offload

Browse files
Files changed (1) hide show
  1. README.md +45 -0
README.md CHANGED
@@ -195,6 +195,51 @@ custom=$(
195
 
196
  </details>
197
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
198
 
199
  ## `IQ3_K` 14.509 GiB (4.082 BPW)
200
  Final estimate: PPL = 7.4360 +/- 0.05162
 
195
 
196
  </details>
197
 
198
+ ## `IQ4_KT` 14.438 GiB (4.062 BPW)
199
+ Final estimate: PPL = 7.5020 +/- 0.05230
200
+
201
+ Mostly pure IQ4_KT meant for full GPU offload similar to [turboderp-org/exllamav3](https://github.com/turboderp-org/exllamav3) [check out ArtusDev's HuggingFace Page](https://huggingface.co/ArtusDev) for someh excellent EXL3 quants!
202
+
203
+ <details>
204
+
205
+ <summary>👈 Secret Recipe</summary>
206
+
207
+ ```bash
208
+ #!/usr/bin/env bash
209
+
210
+ custom="
211
+ # 48 Repeating Layers [0-47]
212
+
213
+ # Attention
214
+ blk\..*\.attn_q.*=iq4_kt
215
+ blk\..*\.attn_k.*=iq4_kt
216
+ blk\..*\.attn_v.*=iq4_kt
217
+ blk\..*\.attn_output.*=iq4_kt
218
+
219
+ # Routed Experts
220
+ blk\..*\.ffn_down_exps\.weight=iq4_kt
221
+ blk\..*\.ffn_(gate|up)_exps\.weight=iq4_kt
222
+
223
+ # Non-Repeating Layers
224
+ token_embd\.weight=iq4_kt
225
+ output\.weight=iq6_k
226
+ "
227
+
228
+ custom=$(
229
+ echo "$custom" | grep -v '^#' | \
230
+ sed -Ez 's:\n+:,:g;s:,$::;s:^,::'
231
+ )
232
+
233
+ ./build/bin/llama-quantize \
234
+ --custom-q "$custom" \
235
+ --imatrix /mnt/raid/models/ubergarm/Qwen3-30B-A3B-Thinking-2507-GGUF/imatrix-eaddario-combined-all-medium-Qwen3-30B-A3B-Thinking-2507-BF16.dat \
236
+ /mnt/raid/models/ubergarm/Qwen3-30B-A3B-Thinking-2507-GGUF/Qwen3-30B-A3B-Thinking-2507-BF16-00001-of-00002.gguf \
237
+ /mnt/raid/models/ubergarm/Qwen3-30B-A3B-Thinking-2507-GGUF/Qwen3-30B-A3B-Thinking-2507-IQ4_KT.gguf \
238
+ IQ4_KT \
239
+ 192
240
+ ```
241
+
242
+ </details>
243
 
244
  ## `IQ3_K` 14.509 GiB (4.082 BPW)
245
  Final estimate: PPL = 7.4360 +/- 0.05162