首页 > AI前沿 > LampAttention: Look-Ahead Mixed-Precision FlashAttention for Dedicated Accelerators

LampAttention: Look-Ahead Mixed-Precision FlashAttention for Dedicated Accelerators

arXiv机器学习 2026-09-30 17:18 9 阅读 查看原文

While most attention logits can be computed in low precision without degrading numerical stability, current attention kernels fail to exploit this phenomenon.

We introduce a novel hardware-algorithm co-design in the form of mixed-precision FlashAttention.

Our method accumulates key-query products and evaluates their exponentials in 8-bit formats, then adaptively identifies sensitive sub-blocks and recomputes them in 16-bit formats.

We propose the specifications for a dedicated accelerator capable of executing this pipeline efficiently.

Simulated experiments with Qwen3 and Gemma 3 show that rerouting a selective minority of sub-blocks to high precision is sufficient to recover the baseline model performance.