fits perfect on 4 x A6000

#1
by WiSFoR - opened

With turbo quant enabled, barely makes to 262144 max context length.
MTP support could fit too (0.96 gpu) but resulted to gibberish output when enabled.
This could be the best model to run on ampere gpu since deepseek support on it has no promise.
Thank you for this model 😃

Sign up or log in to comment