Text Generation
Transformers
Safetensors
PyTorch
nemotron_h
nvidia
nemotron-3
latent-moe
mtp
conversational
custom_code
8-bit precision
modelopt

Model size is halved for NVFP4

#32
by YouNeedCryDear - opened
NVIDIA org

Is it correct to show model size as half for this NVFP4 quantized model? In the description it is clearly says the model is total 120B but the metadata on the right shows only 67B size. Note that for FP8 and FP16 variants, they are the same size.

Sign up or log in to comment