How did you quantize without imatrix?

#27
by Wladastic - opened

I think I saw an imatrix file in this repo yesterday, you have removed it from commit history it seems.

In another discussion you meantioned that it is still in progress.
Does that mean that the Quants besides Q4 and Q8 XL are going to be replaced?

Sadly I do not have enough ram to fit the original ~160GB, so my imatrix with 200k token mix of my oneshot benchmarks, openclaw and hermes prompts are calculating since yesterday and I would like to compare our outputs there.
I only made one GGUF that is inspired by unsloth and APEX a bit, its below the IQ1XXS in size but still very smart, I am trying to find a way to keep a higher precision for the most used experts, so am currently rewriting the llamacpp quantization script...

Sign up or log in to comment