Very interesting work, very useful

#1
by waynemchuck - opened

Could you possibly make more quants, like bartowski's q4_0, q4_1? My GPUs are best to run these types of quants. Thanks anyway!

Great work!
I tested 4 versions of similar of this size (mradermacher, unsloth, AnonimousA and this one), but only this one was able to maintain its original communication ability in other languages, such as Hungarian.
The other small quants started acting very strangely, even at 80GB+ in size.

Could you possibly make more quants, like bartowski's q4_0, q4_1? My GPUs are best to run these types of quants. Thanks anyway!

Yeah, let me see what i can figure out.

Great work!
I tested 4 versions of similar of this size (mradermacher, unsloth, AnonimousA and this one), but only this one was able to maintain its original communication ability in other languages, such as Hungarian.
The other small quants started acting very strangely, even at 80GB+ in size.

i second this even on Q3, Q4_K_M or _K_XL will be great for this, smaller size with same intelligence level and not a REAP

Could you possibly make more quants, like bartowski's q4_0, q4_1? My GPUs are best to run these types of quants. Thanks anyway!

May i ask what your hardware is? The answer may matter for compatibility. These quants are a bit older and harder to generate.

Can you do Q4_K_XL? as they are similar with Q3_K_XL?

Could you possibly make more quants, like bartowski's q4_0, q4_1? My GPUs are best to run these types of quants. Thanks anyway!

May i ask what your hardware is? The answer may matter for compatibility. These quants are a bit older and harder to generate.

Thanks a lot for asking. My GPUs are AMD Instinct Mi50, which are also quite old but still working fine with vanilla llama.cpp. That is why they work best with the old quant types. I also believe there many other old GPUs working best with q4_0, q4_1, such as Tesla p100, p40.

In unsloth studio, it only uses half of my memory while running agentic works.
21 tok/sec, bc I'm using it on a very low-power settings. Is there a way to speed this up without using more power?
image
Is teher a way for

In unsloth studio, it only uses half of my memory while running agentic works.
21 tok/sec, bc I'm using it on a very low-power settings. Is there a way to speed this up without using more power?
image
Is teher a way for

In addition to the flag --override-tensor "per_layer_token_embd.weight=CPU" , you might want to override some experts weights to the cpu's ram. I don't remember the command, but you can search for it. Probably it would improve the performance. On the other hand, if your ram is ddr4, I guess there is not much room for improvement since more than 20 tokens/second is already quite good. If ddr5, then the performance may be better.

Could you possibly make more quants, like bartowski's q4_0, q4_1? My GPUs are best to run these types of quants. Thanks anyway!

May i ask what your hardware is? The answer may matter for compatibility. These quants are a bit older and harder to generate.

Thanks a lot for asking. My GPUs are AMD Instinct Mi50, which are also quite old but still working fine with vanilla llama.cpp. That is why they work best with the old quant types. I also believe there many other old GPUs working best with q4_0, q4_1, such as Tesla p100, p40.

It's up. Note that this quant is 75.4 GB — bigger than the Q3_K_XL, since the n-gram table is what makes this model small and Q4_0 stores everything else at more bits. I couldn't test it here effectively (only 48 GB RAM), so if you get it running let me know whether it loads and what tok/s you see on the MI50s.

Could you possibly make more quants, like bartowski's q4_0, q4_1? My GPUs are best to run these types of quants. Thanks anyway!

May i ask what your hardware is? The answer may matter for compatibility. These quants are a bit older and harder to generate.

Thanks a lot for asking. My GPUs are AMD Instinct Mi50, which are also quite old but still working fine with vanilla llama.cpp. That is why they work best with the old quant types. I also believe there many other old GPUs working best with q4_0, q4_1, such as Tesla p100, p40.

It's up. Note that this quant is 75.4 GB — bigger than the Q3_K_XL, since the n-gram table is what makes this model small and Q4_0 stores everything else at more bits. I couldn't test it here effectively (only 48 GB RAM), so if you get it running let me know whether it loads and what tok/s you see on the MI50s.

Thank you so much. I tested it, the speed is identical to the standard q4_0 which I have to offload to the ram. The intelligence seems to be affected a bit. Still, thanks again!

Sign up or log in to comment