awesome, some questions

#1
by ragallo - opened

First I love your work, I recognize the name and I go ther ealong with a few others.

I have the M3U 256 gig too which is near what you have but architecturally the same more or less, so it helps me get a preview of what to expect + your configurations and instructions with oMLX.

The questions I have are the following, hoping you may answer:

  1. Is there a real difference between q4 and q6 or q8/9? As I understand it, q4 is worse on smaller models, potentially more so with MLX (though that may be outdated due to MLX advancements). Q4 is better on larger models, perhaps 70b+? We can run models easily at Q8, but is Q6 the real sweet spot for larger ram + bandwidth users? I've also read that on paper a larger quant should be better but can be worse even.

  2. Do you plan on establishing at least oMLX benchmarks? It is a lot to ask, so I'll be running on my own but the biggest data you provide that is most helpful is the highlevel such as decode speed - its rarely shared but THANK YOU for the effort here alone.

  3. Will you make an oq4/6/8 version - you've uplaoded these, ty!

  4. Would you recommend setting KV cache to Q8 in oMLX?

  5. Are there any configurations for MTP, DSPARK, etc.? Sorry, I'm taking a quick rbeak from work so haven't been able to review in fine details. I noticed you just upload MLX q6 and q8 - cool!

Please always include great instructions and configs for both CLI and in-app settings of the oMLx app - great work.

TensorFold org

Thanks, really appreciate it.

Yes, there is a real difference between Q4, Q6, and Q8, but it varies a lot by model.
On 256 GB systems, I generally like Q6 as the sweet spot. Q4 is great when memory or speed matters, while Q8 is useful when you want to minimize quantization loss.

Yes, I want to add more oMLX benchmarks.
Decode speed, memory use, context, and exact settings are probably the most useful numbers to share.

Yep, I’ve started uploading oQ4/Q6/Q8 where it makes sense.
With 256 GB, I’d generally start with Q8 KV cache. I’d only go lower if long context or concurrency makes memory an issue.
I’m also looking at MTP/DSpark configs. oMLX is changing quickly, so I’d rather test them properly before recommending specific settings.

And agreed, I’ll keep including both CLI commands and oMLX app settings with uploads.

for me the bottleneck is prefill speed and memory, not cache - the cache itself is much smaller than what it needs for prefill (not that it makes much sense to me tbh)
TQ slows down the model a bit, so whether you need it is up to you. If you run a single agent that mostly runs serialized requests or chats then there's not much benefit in enabling TQ. If you run many different workloads with numerous varying prefixes then TQ will help and 8bit will have absolutely negligible impact but save 50% of memory (but only for cache, not prefill)

oQ4/oQ4e works quite well for dense model like 27B, but avoid it for MoE (35B-A3B etc.) - there is degrades a lot and the lowest quant I'd consider is oQ6

MTP/DFlash is up to how it benchmarks on your machine, in my experience MTP is simply better, faster, with lower memory overhead. DFlash is sometimes faster for short prompts but you lose MTP - I'd rather keep MTP for everything

Sign up or log in to comment