Qwen3.8-Flash-Next W4A16 NVFP4 GGUF

GGUF conversion of axiomofmind/Qwen3.8-Flash-Next-W4A16-NVFP4, based on Qwen/Qwen3.8-Flash-Next.

The routed-expert weights use W4A16 NVFP4. Attention, shared experts, routers, embeddings, PLE, and other retained tensors remain in BF16 or F32.

Files

File Description Size
Qwen3.8-Flash-Next-W4A16-NVFP4-BF16attn-PLE-noMTP.gguf Main text model 180.4 GB
mmproj-Qwen3.8-Flash-Next-BF16.gguf BF16 vision projector 907.5 MB

The vision projector is required for image inputs. It is not required for text-only use.

This GGUF does not include MTP weights.

Requirements

A llama.cpp build with Qwen4Exp and NVFP4 GGUF support is required.

For multimodal use, load the main model together with the included mmproj file.

License

This model is distributed under the Qwen Community License 1.0. Refer to the official model card for architecture details, usage guidance, and limitations.

Downloads last month
1,455
GGUF
Model size
177B params
Architecture
qwen4exp
Hardware compatibility
Log In to add your hardware

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for axiomofmind/Qwen3.8-Flash-Next-W4A16-NVFP4-GGUF

Quantized
(1)
this model

Collection including axiomofmind/Qwen3.8-Flash-Next-W4A16-NVFP4-GGUF