首页 > AI前沿 > MoRE: Scaling mixture of experts with hardware-aware low-rank routing

MoRE: Scaling mixture of experts with hardware-aware low-rank routing

arXiv自然语言 2026-10-01 12:00 8 阅读 查看原文

Mixture-of-Experts (MoE) layers are central to frontier language models, and recent architectures push toward more and smaller experts.

In this regime, the standard linear router becomes a bottleneck: with $M$ experts and hidden dimension $h$, its per-token cost $\Theta(Mh)$ dominates the MoE layer once $M$ is large.

We introduce MoRE (Mixture of Rank-reduced-routed Experts), which factorizes the router weight matrix at rank $r$ and reduces the routing cost to $O((h + M)r)$.

We prove that rank logarithmic in $M$ suffices for routing expressivity when the number of active experts is fixed, and is necessary up to precision factors.

We also prove that logarithmic rank preserves load balance in a Gaussian memorization model, and training on a synthetic phonebook task shows that low rank does not hurt memorization.

At matched active FLOPs, the factorization allows a factor of $\Theta(h/r)$ more experts.

To realize this gain in wall-clock time, we design a fused Triton kernel at inference that avoids expensive memory operations on HBM.

Empirically, MoRE improves memorization on the phonebook task and performance on knowledge-intensive Q\&A benchmarks after pretraining, while matching reasoning ability.

Code available at https://github.com/Matheart/MoRE_code.