Really useful quantified small models PQ2_0. Have good intelligence and pretty fast. (With evalscope testing result)

#31
by zhijin123 - opened

I tested the PQ2_0 with PrismML-Eng/llama.cpp "prism-b10709-9a9394a" version. Test is ran on my 5090 laptop with speed ~7x t/s.
It only took ~11.9G VRAM with kv-cache q8_0 128k context. It passes my intellengence test scripts and the evalscope tool_bench score is pretty good (0.4200).
Ternary-Bonsai-2-27B-PQ2_0 with reasoning-effort "medium"
image
Ref. - Unsloth Qwen3.8-27b-UD_Q4_K_XL with reasoning-effort "xhigh", evalscope tool_bench score (0.4706)
image
Reminder and Conclusion:
I tested Ternary-Bonsai-2-27B-PQ2_0 with both reasoning-effort xhigh and medium. When testing with "xhigh", PQ2_0 caused thinking loop and made testing not finish. After switching to "medium", it ran stable and complete full testing. So suggest set the reasoning-effort to "meiudm".
Next:
Will test real coding and compare with Unsloth Qwen3.8-27b-UD-Q4_K_XL.

zhijin123 changed discussion title from Really useful quantified small models PQ2_0. Have good intelligence and pretty fast. to Really useful quantified small models PQ2_0. Have good intelligence and pretty fast. (With evalscope testing result)

Sign up or log in to comment