Can the 51B n-gram table be kept on NVMe/SSD?

#44
by NamerPRO - opened

Does this Qwen3.8-Flash-Next UD-Q4_K_XL GGUF support the same n-gram SSD offloading approach as AtomicChat's version - i.e. keeping the 51B n-gram embedding table in a separate GGUF file, memory-mapped on NVMe, with only the required rows loaded on demand?

If so, is there already a supported way to run it with 64 GB RAM?

https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF/discussions/35#6a9168ae2cc54f730719e80c

I used this guy's weights where he ripped out the ngram table into its own gguf shard, now ive 20 tps on 64GB of DDR4 + RTX 4090 running a 151GB model

https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF/discussions/35#6a9168ae2cc54f730719e80c

I used this guy's weights where he ripped out the ngram table into its own gguf shard, now ive 20 tps on 64GB of DDR4 + RTX 4090 running a 151GB model

Thank you very much. I will defenitly try it!

I think key to maxing token rate when you offload is to offload in blocks intelligently so that the data most accessed lands on your fastest storage, and the least accessed on your slowest. Just offloading it all to SSD is a terrible idea.

There are solid examples on reddit.

That said. If you are not on Unified memory, you should probably be running 3.8 27B anyway.

I think key to maxing token rate when you offload is to offload in blocks intelligently so that the data most accessed lands on your fastest storage, and the least accessed on your slowest. Just offloading it all to SSD is a terrible idea.

There are solid examples on reddit.

That said. If you are not on Unified memory, you should probably be running 3.8 27B anyway.

as long as you leave some RAM leftover, mmap should naturally handle that?

I think key to maxing token rate when you offload is to offload in blocks intelligently so that the data most accessed lands on your fastest storage, and the least accessed on your slowest. Just offloading it all to SSD is a terrible idea.

There are solid examples on reddit.

That said. If you are not on Unified memory, you should probably be running 3.8 27B anyway.

as long as you leave some RAM leftover, mmap should naturally handle that?

I honestly don't know. But the testing I have done on a Threadripper Pro 7965, 128GB DDR5-6000, with 2 x A5000 24GB Nvlink gives me around 70 tokens/sec at Q4, while optimized ram-offload 3.8 Flash-next maybe 15 or so.

I think key to maxing token rate when you offload is to offload in blocks intelligently so that the data most accessed lands on your fastest storage, and the least accessed on your slowest. Just offloading it all to SSD is a terrible idea.

There are solid examples on reddit.

That said. If you are not on Unified memory, you should probably be running 3.8 27B anyway.

as long as you leave some RAM leftover, mmap should naturally handle that?

I honestly don't know. But the testing I have done on a Threadripper Pro 7965, 128GB DDR5-6000, with 2 x A5000 24GB Nvlink gives me around 70 tokens/sec at Q4, while optimized ram-offload 3.8 Flash-next maybe 15 or so.

Isn’t that because of different reasons and not just n-gram becoming bottleneck ? E.g. with your high bandwidth 4-8 channel DDR5 the expert layers are streamed to VRAM , so no wonder it’s fast as long as it fits etc. Perhaps the n-gram ram-offload is not that optimized , or at least not for your setup.

Sign up or log in to comment