Insanely good 2-bit quant

#40
by cogitech - opened

I decided to try UD-Q2_K_XL on my RTX3060 12GB with 64k KV cache (at Q4_0) just for the hell of it. I am running it via llama.cpp's llama-server which is serving Hermes Agent. It is blowing my mind! This 2-bit model successfully navigated a gauntlet of "AI traps" that usually break much larger models. Here is the summary of its wins:

The Einstein Riddle (Zebra Puzzle): Proved complex constraint satisfaction by correctly mapping 15+ variables (nationalities, houses, pets) without losing track of the "bottleneck" clues.

The Car Wash Challenge: Demonstrated spatial awareness and object permanence by recognizing that walking to a car wash doesn't move the car.

The 30 Shirts Task: Showed an understanding of parallel vs. serial processing, correctly identifying that more items don't increase time if the heat source (the sun) is constant.

Sally’s Sisters: Solved a Theory of Mind trap by accurately calculating family relationships and recognizing that Sally herself is one of the sisters.

The Killer in the Room: Handled state-change logic and self-identity, realizing that the user’s actions (killing) added a new killer to the total count.

The Boiling Water/Freezing Room: Applied thermodynamics and temporal reasoning, correctly predicting that 10 minutes is insufficient for thermal equilibrium.

This is currently the best way to use 12GB of VRAM, IMHO. Although, I do realize for pure coding prowess, Qwen3.6 is probably better. That isn't my primary use case, so this 2-bit quant of Gemma4 wins, and it does it all so eloquently that I actually enjoy chatting with it.

Sign up or log in to comment