KV Cache Precision Benchmarks

#182
by Rieker - opened

I used Qwen3.8-27B to test and measure how different KV cache configurations affect quality.

Comparing the following configuration knobs:

  • KV cache data type (BF16, FP8)
  • TurboQuant KV cache compression presets (k8v4, 4bit, k3v4, 3bit)
  • Context scaling by using YaRN (native 256k, extended 512k, 1M tokens max. context)
  • Weights data type (BF16, FP8)

Setup

  • Model: Qwen3.8-27B (BF16 and FP8)
  • Inference Engine: vLLM nightly snapshot
  • Benchmark: OpenAI MRCR v2 (post-12/5/2025 bugfix), 2/4/8-needle
  • Protocol: 15 samples/run (5 per needle bucket n=2/4/8), temp=0, seed=42, concurrency=1, max-tokens=2048, thinking off. Score < 0.5 = hard fail.
  • Sample sets: in-range ~212–227k tokens (all configs below unless noted); out-of-range ~525–547k (section 4 only).
  • Hardware: single DGX Spark

Decisions

I picked OpenAI's MRCR benchmark* for this task, as I quickly realized that a primitive one-needle search, buried in a large repitition of the same sentence was too simple for even aggressive KV cache configurations to fail.

The MRCR benchmark OTOH challenges the LLM much more by using multi-needle search, packed into sentences of high similarity. So it makes the results much more sensitive to KV value precision.

Basically this benchmark tests for information retrieval only though, i.e. whether the right needle was found in a large context, and whether the nearby context of that needle was returned by the LLM exactly verbatim in its response.

What this benchmark does not tell is the impact on complex tasks like agentic software engineering. I can imagine that those complex tasks are much more sensitive to V-precision than this benchmark does, which is probably more K-precision sensitive.

NOTE: I restricted the benchmark to only 15 samples per configuration, so there is some noticeable noise in the results. I had to restrict it simply because of the very long prefill time on my single DGX Spark machine. Even with 15 samples it took several days to run all these benchmarks, so I had to make a tradeoff.

1. KV dtype × Max. Context (in-range samples)

Achieved score as primary value, hard-fails as secondary information in round brackets:

KV dtype 256k (no YaRN) 512k (YaRN f=2) 1M (YaRN f=4)
BF16 0.833 (3/15) 0.741 (5/15) 0.742 (5/15)
FP8 0.717 (5/15) 0.737 (5/15) 0.742 (5/15)
TQ-k8v4 0.833 (3/15) 0.833 (3/15) 0.741 (5/15)
TQ-4bit_nc 0.777 (4/15) 0.833 (3/15) 0.692 (6/15)
TQ-k3v4_nc 0.777 (4/15) 0.686 (5/15) 0.686 (6/15)
TQ-3bit_nc 0.673 (6/15) 0.777 (4/15) 0.638 (7/15)
  • The scores show that extending max. context by using YaRN is not free.
  • While BF16 KV cache dtype shines at native context, it appears to fall down to FP8 level on YaRN extended context.
  • FP8 remains quite stable over all context lengths.
  • TQ-k8v4 appearing to be better than FP8 is most probably just noise due to the low amount of samples used. The key data type of this TurboQuant preset is plain FP8 type (no K-compression, only V-compression).

2. Per-needle results (KV dtypes × Max. Context)

KV dtype Max. Context n2 n4 n8
BF16 256k (no YaRN) 0.998 (0/5) 0.832 (1/5) 0.669 (2/5)
BF16 512k (YaRN f=2) 0.998 (0/5) 0.832 (1/5) 0.394 (4/5)
BF16 1M (YaRN f=4) 0.998 (0/5) 0.704 (2/5) 0.523 (3/5)
FP8 256k (no YaRN) 0.998 (0/5) 0.798 (1/5) 0.355 (4/5)
FP8 512k (YaRN f=2) 0.998 (0/5) 0.832 (1/5) 0.380 (4/5)
FP8 1M (YaRN f=4) 0.998 (0/5) 0.704 (2/5) 0.523 (3/5)
TQ-k8v4 256k (no YaRN) 0.998 (0/5) 0.832 (1/5) 0.669 (2/5)
TQ-k8v4 512k (YaRN f=2) 0.998 (0/5) 0.832 (1/5) 0.669 (2/5)
TQ-k8v4 1M (YaRN f=4) 0.998 (0/5) 0.555 (3/5) 0.669 (2/5)
TQ-3bit_nc 256k (no YaRN) 0.844 (1/5) 0.821 (1/5) 0.355 (4/5)
TQ-3bit_nc 512k (YaRN f=2) 0.998 (0/5) 0.812 (1/5) 0.523 (3/5)
TQ-3bit_nc 1M (YaRN f=4) 0.997 (0/5) 0.538 (3/5) 0.380 (4/5)
  • Here you can see that needle count matters, almost no configuration failed on only 2 needles.
  • Only the most aggressive TurboQuant presets failed on small needle count.

3. TurboQuant compression vs. quality (no-YaRN)

Note that I have not measured the perplexity (PPL) values in the following table, they were taken directly from the comments attached to vLLM's TQ presets.

preset compression PPL score fails
TQ-k8v4 2.6× +1.17% 0.833 3/15
TQ-4bit_nc 3.8× +2.71% 0.777 4/15
TQ-k3v4_nc 3.5× +10.63% 0.777 4/15
TQ-3bit_nc 4.9× +20.59% 0.673 6/15
  • TQ-k8v4 appears to be almost free, while providing good compression ratio.
  • TQ-4bit_nc appears to be a good trade-off between compression and quality.
  • The other two TurboQuant presets are probably too aggressive.

4. Out-of-range ~547k samples (YaRN f=4, 1M max. context)

While the previous benchmark results were all taken by using the same samples <256k tokens, I also ran some benchmarks on samples that were about ~547k tokens in length, so true long context prompts. However I had no more time for an extensive exploration.

KV dtype score fails
BF16 0.789 4/15
FP8 0.696 6/15
  • While BF16 fell off before on smaller prompts on YaRN extended context, on true large prompts like here though it keeps its dominance over FP8.
  • Maybe both suffer more hard fails on true long prompts than on short prompts, however hard to compare, as these are completely different sample sets than before.

Weight quantizations

KV BF16 weights FP8 weights
BF16KV 0.827 (3/15) 0.833 (3/15)
FP8KV 0.782 (4/15) 0.717 (5/15)

I have only performed few benchmarks on this aspect, but they suggested that BF16 vs. FP8 weights had no real impact on the benchmarks' quality results. They were either bit-identical or within noise range. Therefore I used FP8 weights exclusively for all previous benchmarks instead.

Sign up or log in to comment