首页 > AI前沿 > AvoKV-E: Payload-Aware KV Cache Eviction for Long Reasoning

AvoKV-E: Payload-Aware KV Cache Eviction for Long Reasoning

arXiv机器学习 2026-10-02 16:37 5 阅读 查看原文

Long-output reasoning shifts the KV-cache bottleneck from the fixed prompt to the generated trace.

Existing reasoning-cache eviction methods largely treat cached entries as routing objects, estimating whether an old key will still be read, will recur, or can be replaced.

This routing-only view overlooks two effects: low-attention entries can carry large value payloads whose removal changes future predictions, and newly generated states can appear stale before later queries have had a chance to read them.

We introduce AvoKV-E

AvoKV-E, a training-free eviction policy that first delays eligibility for recent states and then ranks eligible entries using candidate-normalized read pressure, key redundancy, and value-payload potential.

According to empirical evaluation across different models and datasets, AvoKV-E matches or exceeds redundancy-aware, recurrence-based, and thought-adaptive eviction baselines at matched active-KV budgets, with its largest gains in the tightest-cache regime.

Component and counterfactual analyses

Component and counterfactual analyses further connect these gains to delayed observation, payload-aware scoring, redundancy, and scale-robust normalization.

Together, the results show that long-reasoning KV eviction should preserve not only keys that are likely to be read, but also the value payloads that sustain the reasoning trajectory.