首页 > AI前沿 > Rehearse Everything, Remember Nothing: Attic-KV Rehearses What Will Be Read

Rehearse Everything, Remember Nothing: Attic-KV Rehearses What Will Be Read

arXiv自然语言 2026-10-08 23:22 4 阅读 查看原文

Many key-value (KV) caches are compressed before anyone knows what will be asked of them: a document cached for retrieval, a prompt prefix shared across requests, the memory of a long conversation.

The prevailing approach scores KV entries by rehearsal: the model rereads the context and keeps the entries it attends to, assuming that the more completely a cache rehearses its context, the better it remembers it.

We show that under tight budgets this assumption backfires: rehearse everything, remember nothing.

At a 3% keep ratio, rereading the whole context keeps 31.5 of 96.5 points on RULER, and on LongBench's natural-text tasks it falls below methods that rehearse nothing at all.

The cause is that a cache keeps what it rehearses: rereading spreads the budget across the whole context, so the answer's own entries survive at little more than chance.

Like a student before an exam, a cache remembers more by testing itself than by rereading.

Two principles follow: rehearse what will be read, and rehearse as much as there is.

We instantiate them as Attic-KV (Attic for short), a training-free rehearsal in which the model quizzes itself with question-answer pairs that quote the context, alongside anchor tokens in a content-adaptive amount.

Changing only the rehearsal lifts three hosts that score it in three different ways: Attic alone is the best training-free method in all eight settings we test on RULER and LongBench's natural-text tasks, and plugged into the gradient-based KVgrad and the trained RestoreKV+, it raises them by up to 17.1 and 28.1 points.

Its advantage grows as the budget shrinks, reaching 41.9 points over full rereading at a 3% keep ratio, and it compresses faster than rereading the whole context.