首页 > AI前沿 > Dual-QK: Sharp Queries and Flat Keys for Prunable 2-bit KV Caches

Dual-QK: Sharp Queries and Flat Keys for Prunable 2-bit KV Caches

arXiv机器学习 2026-10-07 18:50 3 阅读 查看原文

Long inputs and extended generation increase the storage and access costs of the key-value (KV) cache.

Low-bit quantization reduces storage and memory traffic, while query-channel pruning can further reduce key-cache reads.

Rotation-based quantization redistributes the energy of key outliers across channels.

To maintain computational invariance, the same orthogonal transform must be applied to queries, preserving query-key dot products.

However, this rotation can disperse query energy, weakening the separation between a few large components to retain and many small ones to prune.

We introduce Dual-QK, which uses paired non-orthogonal query and key transforms to address this conflict.

Using calibrated query and key statistics, Dual-QK combines partial key whitening with a query-aligned basis to balance key scales for INT2 quantization and concentrate query energy for dynamic channel pruning.

Channel-0 protection and bucket-relative RoPE support low-bit accuracy over long contexts.

Experiments on four models across five generative benchmarks and long-context retrieval tasks show improved accuracy over OSCAR on most tasks at 40% query-channel sparsity.

At a 128K context, Dual-QK provides $6.8\times$ KV-cache compression and an estimated $8.3\times$ reduction in KV read volume relative to unpruned BF16.

Under the evaluated configurations, our SGLang implementation achieves up to $3.75\times$ the decoding throughput of unpruned BF16.