SM120 support

#1
by InformaticsSolutions - opened

Hi,
thank you so much for creating and sharing this quant. I am trying to load it on 2xRTX6000 Pro (SM120) and i'm getting this error:

RuntimeError: dispatch_scaled_mm, /home/user1/vllm/csrc/libtorch_stable/quantization/w8a8/cutlass/c3x/scaled_mm_helper.hpp:34, Int8 not supported on SM120. Use FP8 quantization instead, or run on older arch (SM < 100)

I am confused about this Int8 not being supported on SM120: this quant https://huggingface.co/Minachist/Qwen3.6-27B-INT8-AutoRound/tree/W8A16-GS128 works great. Is there something about the quark quant making it incompatible? Or am i doing something wrong?
My setup:

vllm version 0.23.0
cuda  13.2, V13.2.78
NVIDIA driver 595.58.03 

thanks a lot!

SM120 does have INT8 capabilities in general, but vLLM currently does not support the W8A8 INT8 scaled_mm kernel path on SM120.
In vLLM's SM120 scaled_mm implementation, the FP8 kernel is registered, but the INT8 dispatch is explicitly set to nullptr, so W8A8 INT8 checkpoints fail with that error.
Weight-only INT8 models such as W8A16/AutoRound use a different kernel path, which is why they can still work on the same GPU.

InformaticsSolutions changed discussion status to closed

Sign up or log in to comment