Why no NVFP4 GGUF?

#6
by mattzink - opened

llama.cpp and derivatives have supported it for a long time?

Unsloth AI org

We are investigating!!!

https://huggingface.co/Avifenesh/Qwen3.8-27B-NVFP4-MTP-GGUF
in the meanwhile
Some bit exact and benchmarks and you'll be able to use https://github.com/avifenesh/memra if it fits your hardware.

I tried the Avifenesh and it was really slow on llama.cpp and don't want to try memra thanks, there's this Felipebburk profile, I used his and worked fine, but the speed was almost the same of his NVFP4 compared to Unsloth at IQ4_NL with two 5060ti for a total of 32GB of RAM on split-mode tensor, with Unsloth sligthly faster but his making the GPUs work less.

Also his, doesn't have the mmproj

I would love to have a GGUF NVFP4 from Unsloth directly.

I built a set of hybrid GGUFs from the sources of this repository, if it helps anyone:

There's a linked benchmark perform by the community too.

Thank you esatapedico, I can confirm your NVFP4 GGUF works.
It would be great if it came directly from unsloth though!
So please unsloth team include GGUF with your NVFP4 releases!? ๐Ÿ™

Sign up or log in to comment