首页 > AI前沿 > Shared Stopping Decisions Change Answers in HQQ Cache Quantization

Shared Stopping Decisions Change Answers in HQQ Cache Quantization

arXiv自然语言 2026-10-05 20:47 5 阅读 查看原文

Language-model systems batch questions for throughput, but unrelated questions should not change a target's answer when its input and numerical execution are fixed.

We study compression of the key and value cache, which stores attention representations reused during generation.

Request-local Groups and Transformer's HQQ Backend

With request-local groups, Transformers' Half-Quadratic Quantization (HQQ) backend updates compression parameters separately but uses a shared average error to decide when all updates stop.

Replacing only the question batched with the target changes four-bit HQQ answers in 170/384 test comparisons across two models.

Replaying the other execution's update counts reproduces its complete answer and cache fingerprints in every changed pair, in both directions.

FP32 Stopping Mean and Cache Differences

Computing the stopping mean in FP32 reduces cache differences but leaves answer changes.

Native HQQ also changes confirmed numerical correctness in eight arithmetic pairs.

Companion Dependence and Request-local Stopping

Fixed iterations and request-local stopping remove observed companion dependence under matched controls.

Request-local stopping remains sensitive to synthetic padding changes at the tensor level.

Fixing the original iteration budget removes this decision path without tuning.

Repair and Quality Advantage

Neither repair has an established quality advantage, and natural rebatching still changes answers.

Request-independence Audits

Request-independence audits must cover stopping decisions as well as quantization groups.