首页 > AI前沿 > Switching Linear Attention

Switching Linear Attention

arXiv机器学习 2026-09-30 13:37 8 阅读 查看原文

Designing expressive sequence layers with efficient inference remains a central challenge in modern machine learning.

Standard softmax attention achieves excellent sequence modeling performance through rich nonlinear token interactions, but it requires a key-value cache that grows linearly with sequence length, limiting its scalability.

Linear attention enables efficient recurrent computation with a constant memory footprint, yet its reduced expressivity often yields inferior modeling performance.

We introduce Switching Linear Attention (SwiLA), a novel sequence layer that bridges this gap by enhancing representational capacity while retaining the fixed-size recurrent state of linear attention.

We derive the SwiLA recurrence from the test-time regression framework, casting the state update rule as online expectation-maximization in a mixture of linear regressions model.

At test time, each output dimension dynamically selects among multiple linear attention components based on the input.

Across associative recall, in-context language learning, and language modeling benchmarks, SwiLA shows strong performance and narrows the gap to softmax attention, even surpassing it in several settings.