DeCoPrune: Efficient KV-Cache Pruning for Autoregressive Video Diffusion via Denoising Consistency
Abstract
Autoregressive video diffusion supports streaming generation and interactive control, but its KV cache grows continuously with the generated history. Existing compression strategies either discard history using fixed windows or select tokens through local attention and similarity signals, which do not directly measure whether the current chunk contributes information beyond the retained context. We introduce DeCoPrune, a training-free method that treats cache compression as a denoising-consistency problem. We find empirically that denoising difficulty provides a useful proxy for a token's value in long-term retention: tokens with larger step-to-final discrepancies tend to carry visual evidence that is less predictable from the retained context. DeCoPrune measures each current-chunk token's denoising difficulty using the discrepancy between its intermediate clean prediction and final denoised value, retaining high-discrepancy tokens in the long-term cache while pruning those with low discrepancy. To evaluate information retention, we introduce CMBench, comprising 58 approximately one-minute generated or real-world context episodes and 116 Reappear or Revisit continuation tasks that require recalling specific previously observed objects or scenes. Experiments with LingBot World v2 show that DeCoPrune preserves near-FullKV long-range recall while pruning over 85% of historical KV tokens and accelerating continuation generation by over 4times, substantially outperforming the evaluated compression baselines at comparable budgets. These results indicate that denoising consistency can serve as a model-intrinsic signal for retaining long-range information while reducing autoregressive inference cost. Our project homepage is https://decoprune.github.io. The code is available at https://github.com/DeCoPrune/CMBench, and the benchmark at https://huggingface.co/datasets/Aoraku/CMBench.
Community
🎬 Can a video model remember what it saw a minute ago?
AR video diffusion enables streaming generation, but its KV cache grows with the video. Compressing that cache can leave the motion smooth while earlier objects and scenes are forgotten.
To retain the right memories, we turn to denoising itself.
We find empirically that denoising difficulty is an effective proxy for a token’s long-term retention value: tokens with larger discrepancies between their intermediate clean predictions and final denoised outputs tend to carry visual information that is difficult to predict from the existing context alone. Conversely, content that the context already makes predictable tends to approach the final output at an intermediate denoising step.
This observation leads to DeCoPrune (Denoising Consistency Prune): retain high-discrepancy tokens and prune low-discrepancy ones. The method is training-free and reuses predictions already available along the same denoising trajectory for online scoring, with no extra model forward passes. The generation process itself provides a signal for what to remember.
To evaluate whether those memories are preserved, we introduce CMBench: 58 approximately one-minute synthetic and real-world videos, with 116 Reappear and Revisit tasks. Can the model bring back the same object or return to an earlier scene? These tasks directly test long-range visual recall.
On LingBot World v2, DeCoPrune reduces cumulative historical KV tokens by 85.43% and accelerates continuation generation by 4.14×, while maintaining long-range recall close to FullKV.
Get this paper in your agent:
hf papers read 2609.39096 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 1
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper