NVFP4 model ?

#29
by Vazox - opened

Hi, Can we get nvfp4 version, Thanks.

will be uploaded in 1 hours.

added.

hey, could we get the NVFP4 safetensor please (no gguf, .safetensors)? thank you so much.

hey, could we get the NVFP4 safetensor please (no gguf, .safetensors)? thank you so much.

added.

Thank you very much. Do you think having the text encoder in NVFP4 would be better ?

@EyeBlinder probably not worth much. The two text encoder files in this repo are byte-identical to Comfy-Org's stock Qwen3-VL-8B (same LFS sha256), so the uncensoring lives in the diffusion model. The encoder also runs only once per prompt, so a smaller one saves memory, not speed. The existing options are bf16 at 17.5 GB, int8_convrot at 9.35 GB, and Comfy-Org's qwen3vl_8b_w4a8 at 6.31 GB. An NVFP4 build would land around 5–5.5 GB, only about 1 GB below w4a8. With the 4.20 GB NVFP4 DiT and the 0.68 GB VAE, w4a8 already keeps the weights near 11 GB.

Disclosure: I made a free calculator for this model's DiT/encoder/VAE combinations, if you want to check your GPU: https://modelvram.com/qwen-image-2-1-vram-calculator/

got it! thank you for the explanation.

Sign up or log in to comment