首页 > AI前沿 > Replay the Curvature: Accurate and Scalable NVFP4 Quantization for Large Language Model Inference

Replay the Curvature: Accurate and Scalable NVFP4 Quantization for Large Language Model Inference

arXiv自然语言 2026-09-29 11:55 7 阅读 查看原文

Large language models make weight storage and memory traffic major inference costs, motivating low-precision formats that represent each weight with only a few bits.

Schur Replay, A SCALE-SELECTION ALGORITHM THAT REPRODUCES THE GPTQ UPDATES CAUSED BY EACH BLOCK SCALE AND SCORES THE RESULTING BLOCK ERROR AFTER ACCOUNTING FOR COMPENSATION FROM UNQUANTIZED COLUMNS.