Dolphin Mistral 24B Venice Edition 1.2 GGUF

GGUF quantizations of dphn/Dolphin-Mistral-24B-Venice-Edition, built from version 1.2, with vision adapters and an importance matrix. All the Dolphin-Mistral-24B-Venice-Edition GGUFs I could find on Hugging Face were converted from an earlier release, so I built these from version 1.2, introduced in 337ce042026e: "Updated to version 1.2 - vision + 131k context + improved tool calling".

Quality vs size

Every quant was measured by KL-divergence against the BF16 weights on held-out wikitext-2 (128 chunks).

image/png

Quant Size bpw Mean KLD vs BF16 PPL ratio Notes
Q3_K_M 10.69 GB 3.89 0.049230 1.0585
IQ4_XS 11.88 GB 4.33 0.020138 1.0244 Best pick under 12 GB.
Q4_K_M 13.35 GB 4.87 0.016559 1.0221 Good default choice.
Q5_K_M 15.61 GB 5.69 0.003989 1.0049
Q6_K 18.02 GB 6.57 0.001576 1.0018
Q8_0 23.33 GB 8.50 0.000188 1.0004 Effectively lossless.

BF16 (43.92 GB) is the reference and is included for anyone wanting to re-quantize without re-downloading the safetensors.

Downloads last month
3,693
GGUF
Model size
24B params
Architecture
mistral3
Hardware compatibility
Log In to add your hardware

3-bit

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for douganger/Dolphin-Mistral-24B-Venice-Edition-1.2-GGUF