Vision tower stripped to save space. Fits in 48gb VRAM with 16k context, kv quantized to 6,6
The KL is kind of high, so I verified it against
another quant, and it measured in line.
2.50bpw exl3 quantization of
Mistral-Medium-3.5-128B TEXT ONLY via
exllamav3.
repo generated automatically with
ezexl3.