--- base_model: - Motif-Technologies/Motif-3 --- Requires [this PR](https://github.com/ggml-org/llama.cpp/pull/26298) to run. This repo contains specialized MoE-quants for Motif-Technologies/Motif-3. The idea being that given the huge size of the FFN tensors compared to the rest of the tensors in the model, it should be possible to achieve a better quality while keeping the overall size of the entire model smaller compared to a similar naive quantization. To that end, the quantization type default is kept in high quality and the FFN UP + FFN GATE tensors are quanted down along with the FFN DOWN tensors. | Quant | Size | Mixture | PPL | 1-(Mean PPL(Q)/PPL(base)) | KLD | | :----- | :-------------------- | :------------------------------- | :------------------- | :------------------------ | :------------------ | | Q8_0 | 314.95 GiB (8.60 BPW) | Q8_0 | 35.367303 ± 0.451587 | +2.8870% | 0.010134 ± 0.000061 | | Q5_K_M | 220.07 GiB (6.01 BPW) | Q8_0 / Q5_K / Q5_K / Q6_K | 35.336996 ± 0.450859 | +2.7989% | 0.015087 ± 0.000088 | | Q4_K_M | 183.47 GiB (5.01 BPW) | Q8_0 / Q4_K / Q4_K / Q5_K | 35.635889 ± 0.455428 | +3.6684% | 0.026133 ± 0.000135 | | IQ4_XS | 143.12 GiB (3.91 BPW) | Q8_0 / IQ3_S / IQ3_S / IQ4_XS | 36.272353 ± 0.463193 | +5.5199% | 0.058303 ± 0.000287 | | IQ3_M | 123.79 GiB (3.38 BPW) | Q6_K / IQ3_XXS / IQ3_XXS / IQ3_S | 38.364498 ± 0.492873 | +11.6062% | 0.103530 ± 0.000507 | | IQ3_S | 111.84 GiB (3.05 BPW) | Q6_K / IQ2_S / IQ2_S / IQ3_S | 39.318385 ± 0.502525 | +14.3811% | 0.157063 ± 0.000752 | | IQ2_S | 101.38 GiB (2.77 BPW) | Q6_K / IQ2_XS / IQ2_XS / IQ3_XXS | 41.701314 ± 0.536307 | +21.3133% | 0.222056 ± 0.001044 | ![kld_graph](kld_data/01_kld_vs_filesize.png "Chart showing Pareto KLD analysis of quants") ![ppl_graph](kld_data/02_ppl_vs_filesize.png "Chart showing Pareto PPL analysis of quants")