Instructions to use unsloth/Qwen3.8-27B-NVFP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Local Apps Settings
- Unsloth Desktop
Why no NVFP4 GGUF?
llama.cpp and derivatives have supported it for a long time?
We are investigating!!!
https://huggingface.co/Avifenesh/Qwen3.8-27B-NVFP4-MTP-GGUF
in the meanwhile
Some bit exact and benchmarks and you'll be able to use https://github.com/avifenesh/memra if it fits your hardware.
I tried the Avifenesh and it was really slow on llama.cpp and don't want to try memra thanks, there's this Felipebburk profile, I used his and worked fine, but the speed was almost the same of his NVFP4 compared to Unsloth at IQ4_NL with two 5060ti for a total of 32GB of RAM on split-mode tensor, with Unsloth sligthly faster but his making the GPUs work less.
Also his, doesn't have the mmproj
I would love to have a GGUF NVFP4 from Unsloth directly.
I built a set of hybrid GGUFs from the sources of this repository, if it helps anyone:
- https://huggingface.co/esatapedico/Qwen3.8-27B-NVFP4-MTP-GGUF
- https://huggingface.co/esatapedico/Qwen3.8-27B-NVFP4-BUDGET-GGUF
There's a linked benchmark perform by the community too.
Thank you esatapedico, I can confirm your NVFP4 GGUF works.
It would be great if it came directly from unsloth though!
So please unsloth team include GGUF with your NVFP4 releases!? ๐