Surprisingly good PTQ1_0 quantization – and ~39.5 tok/s on an RTX 4070

#27
by TheWegemann - opened

I have to give some credit here because I went into this test rather skeptical.

I'm generally not a fan of aggressively quantizing large models just to make them fit. In my experience there is usually a point where the quality loss becomes obvious, even if benchmarks still look surprisingly good. Personally, I'm much more interested in models that are trained with ternary/very-low-bit weights from the beginning – BitNet-style approaches and similar architectures – rather than taking a conventional model afterwards and squeezing it as hard as possible.

So I honestly didn't expect much from a 27B model compressed into a 5.95 GB PTQ1_0 GGUF.

But after actually using Ternary Bonsai 2, I have to say: this is a seriously impressive quantization.

I'm running the PTQ1_0 version fully on an RTX 4070 12 GB. In a real conversation at roughly 11.6K context, generation speed is extremely stable at around 39.4–39.5 tokens/sec. More importantly, so far I don't see the kind of obvious degradation I normally associate with extremely aggressive quantization.

The model still reasons coherently, follows the conversation over multiple turns and can correct its own mistakes. One particularly interesting test was when it initially confused ternary weights with "3-bit" weights. After I pasted the relevant part of the model card, it correctly recognized the mistake, explained the distinction between ternary {-1, 0, +1} weights and 3-bit quantization, and reasoned through the approximate bits-per-weight calculation. The resulting English response was surprisingly strong and coherent.

One caveat: the German is bad. Sometimes hilariously bad. There are malformed sentences, invented/incorrect word constructions and phrasing that no native German speaker would use.

However, I explicitly do not blame PTQ1_0 for this.

I've tested Qwen models in substantially less aggressive / higher-quality quantizations before and have repeatedly seen the same problem. Qwen can produce German that looks superficially fluent but becomes strangely constructed, semantically awkward or simply broken once you actually read it as a native speaker. The behavior I'm seeing in Bonsai 2 is very familiar to me from the underlying Qwen family.

In English, the difference is striking. The model suddenly sounds much more coherent and natural, and its reasoning is considerably more convincing.

So while I still prefer the idea of training genuinely ternary models from scratch rather than forcing conventional models down to extremely low bitrates afterwards, PTQ1_0 has changed my opinion somewhat about how far post-training ternary quantization can be pushed.

I expected obvious damage at this size.

So far, I'm not seeing it.

5.95 GB for a 27B model, ~39.5 tok/s on a normal RTX 4070, with reasoning quality that still appears remarkably intact is genuinely impressive work.

I'm going to continue testing it, especially tool calling / agentic behavior and vision. Those will probably tell me much more than another benchmark table will. :)

Sign up or log in to comment