Does it support 4-bit?

#1
by yichenyanyu - opened

Precision policy: Q8_0 and higher are supported. 4-bit and 5-bit quantizations (legacy and k-quant) are rejected at load time with an actionable error, because this graph's kernels are not validated below Q8_0 and otherwise decode to empty text.
Here is a 4-bit version supported in llama.cpp:https://huggingface.co/FIT17/Confucius4-R2T2-GGUF
I'd like to know if this is also supported in audio.cpp

其实也可以支持Q4的,只不过我没有测试效果,而且当时磁盘空间不太多,不想电脑上再多一份权重了; 代码已经写好了,用coding agent改一下应该就可以支持Q4了

As of https://github.com/0xShug0/audio.cpp/issues/678 more quants works now in audio.cpp.

I have been trying it out and it's great so far. The reason why i wanted something like r2t2 to work streaming is to get a bigger model to transcribe while i am talking and basically free. The q8 version is 1/2 real time and this like 1/3 real time so folks with a less powerful mac will definitely benefit.

I am implementing a https://www.typewhisper.com/ plugin and once my audio.cpp pr (actually partially davidxifeng's work) is merged will add both https://huggingface.co/davidxifeng/Confucius4-R2T2-gguf q16 q8 and https://huggingface.co/Nairod785/Confucius4-R2T2-Q4_K_M-GGUF q4 models.

Cheers Florian

Sign up or log in to comment