Please Q5-K-M GGUF.

#2
by sunnyboxs - opened

Please provide a Q5-K-M (or smaller) quantized GGUF version so that the model can be run on systems with limited VRAM.
Thanks!

IQ4_XS is smaller than Q5?

I would rather they make an IQ4_NL (1G bigger) for Metal because PPL/KLD, prefill/decode worth the 1G extra over the RSS shrink. :-)

OrcaRouter org

the current model is mixed-precision and equivalent to ~3 bit already in terms of footprint, just with significantly better capacity retention

Man I want a higher parameter version maybe 70.

Sign up or log in to comment