RTX 5090 performence
#4
by jasionkajakub - opened
IF you won't flag mtp and vision thower model weight is only 19.58 GiB and you coud achive 2.11x context 256k KV cashe with 0.9 gpu memory utilization. Prefil with 16k batch size for 64k prompt is 10500 tps for 128k is 13350tps and 256k giving 8680 tps prefill. inference is stable 220-240 tps. MTP is not recomended.
ps. Tkanks NVIDIA <3 we are waiting for 3.6 27b but not like gemma 4 31it taking 32GB.