Llama.cpp NVFP4 support

#1
by jpsequeira - opened

Are these good to use as they are?
Are you just looking for testers?

Owner

@jpsequeira
It's too early to test, for now nvfp ggufs are worse than q4 or mxfp quants due to lack of some cuda parts in llama.cpp - some tests here (for 122b model): https://github.com/ggml-org/llama.cpp/pull/21095#issuecomment-4147740036
Waiting for proper support too, hope we don't neet to requantize again.

Thanks for the explanation.

jpsequeira changed discussion status to closed

Sign up or log in to comment