nice quantization, you can sqweeze more kv .
two 3090 here, i got your t/s numbers, a bit more, im using some genesis patches and some own patches, im about 90t/s average on long runs , but got it to 350k kv or more, i think i can get it to 450k but im having issues when it builds the triton kernels, also some issues with expandable_segments on pytorch i think i have wrong but , keep on it u can sqweeze a bit more.
why i want more kv? with about 600k my harness works perfectly, and thats what im aiming .
Absolutely there is further optimizations to be had when using genesis patches. https://github.com/noonghunna/club-3090 is a great resource that I had used in the past and it will most likely be updated for Qwen3.8-27B in 24 hours.
Mind posting your config and I'll see if I can update the README to include the Genesis patches? I assume you referred to the above repo for that idea.
I assume you are not using fp8 kv cache?