Model

This repo contains specialized MoE-quants for MiniMax-M3 with the Indexer Tensors preserved at FP32. A BF16 MMPROJ file for image vision input has also been provided.

The text model files should load on mainline as this PR got merged. The MMPROJ requires this one, which is a superset of the first PR.

Quant Size Mixture PPL 1-(Mean PPL(Q)/PPL(base)) KLD
Q4_K_M 246.11 GiB (4.96 BPW) Q8_0 / Q4_K / Q4_K / Q5_K 5.287961 Β± 0.035451 +1.9814% 0.069890 Β± 0.000978

kld_graph ppl_graph

Downloads last month
210
GGUF
Model size
426B params
Architecture
minimax-m3
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for Serpen/Minimax-M3-MSA-GGUF

Quantized
(58)
this model