n-gram from SSD: how exactly?

#75
by dilavni - opened

There are anecdotes of people doing this, but I have no idea how. I have unified memory system so n-gram from RAM is the same as from VRAM. udiq4xss with 256k bf16 ctx is ~105gb RAM/VRAM with llama.cpp's own ram usage on top. I want to load n-gram from SSD and go for a higher quant.

use -lm mmap --lazy-mode on.

On unified memory this is worth doing. The n-gram table is per_layer_token_embd, 28.8 GB (26.8 GiB) in every UD quant. Going by the GGUF sizes on this repo, lazy-reading it takes UD-IQ4_XS from 93.7 GB resident down to about 64.9 GB, and UD-Q4_K_XL from 111.3 GB down to about 82.5 GB. Your ~11 GB of KV cache and overhead at 256K comes on top of that, so Q4_K_XL would land around 94 GB, below what you use now.

On current mainline the flag is --lazy-mode on (-lzm on); -lm mmap isn't required, because the lazy tensor is mmapped either way. Pass on explicitly rather than relying on the default: since llama.cpp#28326, auto switches to off when a device reports no mmap support, which covers many iGPU/unified setups. Check the load log for tensor per_layer_token_embd (size = ... MiB) lazy read enabled. RAM will still rise as rows get touched, but that is page cache, which the OS can drop under pressure.

On unified memory this is worth doing. The n-gram table is per_layer_token_embd, 28.8 GB (26.8 GiB) in every UD quant. Going by the GGUF sizes on this repo, lazy-reading it takes UD-IQ4_XS from 93.7 GB resident down to about 64.9 GB, and UD-Q4_K_XL from 111.3 GB down to about 82.5 GB. Your ~11 GB of KV cache and overhead at 256K comes on top of that, so Q4_K_XL would land around 94 GB, below what you use now.

On current mainline the flag is --lazy-mode on (-lzm on); -lm mmap isn't required, because the lazy tensor is mmapped either way. Pass on explicitly rather than relying on the default: since llama.cpp#28326, auto switches to off when a device reports no mmap support, which covers many iGPU/unified setups. Check the load log for tensor per_layer_token_embd (size = ... MiB) lazy read enabled. RAM will still rise as rows get touched, but that is page cache, which the OS can drop under pressure.

this works. thanks a bunch

Sign up or log in to comment