UD-TQ2_0 loads fully on 8x RTX PRO 6000 but logits are corrupted

#18
by bperin42 - opened

I have unsloth/Kimi-K3-GGUF UD-TQ2_0 fully loading and serving through a custom llama.cpp build on 8× RTX PRO 6000 (768 GB VRAM).

Runtime:

Unsloth Kimi-K3 llama.cpp PR #48 commit 768d2a481a99cb75ec9a03b95dadbd35e7acf496
CUDA 12.4.1
-ngl 99 -sm layer -c 131072
mmproj-BF16.gguf

Model loads successfully:

2.78T parameters
~551 GB model
all 93 layers offloaded
tokenizer/detokenizer verified correct

The GGUF contains tensor type 64. I currently map type 64 to GGML_TYPE_TQ2_0 and use a 46-byte block size so offsets align and the model loads.

However inference logits are clearly corrupted. /completion repeatedly predicts token ID 30 (?), with top logits concentrated around token IDs 0..30.

Tokenizer round-trip is correct:
Hello world -> [19180,2695] -> Hello world

So the remaining issue appears to be the actual UD-TQ2_0 type-64 dequantization/kernel semantics rather than parsing, tokenizer, or serving.

Question: Is there an existing implementation/spec for the type-64 UD-TQ2_0 block layout/dequantization used by this Kimi-K3 GGUF? Has anyone gotten this exact quant producing correct logits in llama.cpp yet?

Sign up or log in to comment