首页 > AI前沿 > EpiKV: Epiphany-Aware KV Cache Eviction Without the Attention Matrix

EpiKV: Epiphany-Aware KV Cache Eviction Without the Attention Matrix

arXiv自然语言 2026-10-01 12:00 8 阅读 查看原文

Reasoning models can generate chains of thought tens of thousands of tokens long, making the key--value (KV) cache that holds them a major bottleneck for inference throughput.

Existing eviction policies for long reasoning traces typically rank cached tokens using attention weights, requiring access to the attention matrix and making them incompatible with fast inference kernels.

In this work

We study the limits of such policies under tight cache budgets. Surprisingly, we find that under the strongest of them the generations that finish are wrong about as often as without eviction; most of the accuracy loss comes from generations that enter loops and run until the length limit, and retaining more tokens according to a fixed importance score exacerbates this behavior.

What stops the looping is keeping the tokens the model's recent queries point to, and the forward pass the model already runs reveals them without the attention matrix.

Motivated by this observation, we introduce epiphany-aware KV cache eviction EpiKV, which combines hidden-state shifts with the model's recent query--key relevance to rank cached tokens without materializing the attention matrix.

On multiple benchmarks, EpiKV matches or outperforms the strongest attention-based eviction baselines while running directly in vLLM with unmodified attention kernels.