首页 > AI前沿 > The Row Normalization Puzzle in Muon

The Row Normalization Puzzle in Muon

arXiv机器学习 2026-09-30 14:51 7 阅读 查看原文

This paper examines how row-wise renormalization affects Muon, focusing on the gap between NorMuon's worst-case guarantees and its practical performance (Li et al.).

Despite its growing adoption and promising performance in large language model (LLM) pretraining, NorMuon's worst-case guarantees remain poorly understood.

One fundamental question is: Does row normalization yield provable convergence gains, potentially through its interaction with approximate polar computation and exponential moving-average momentum?

Our results show that row normalization introduces a dimension-dependent factor in the worst-case iteration complexity under the operator-norm geometry, which persists even with exact polar computation and any fixed momentum parameters.

Indeed, we establish an algorithm-dependent lower bound and a matching upper bound in deterministic settings, and extend our upper bound analysis to stochastic settings.

Both upper-bound analyses allow approximate polar computation.

Experiments show that NorMuon is slower than Muon on synthetic problems inspired by our worst-case construction, yet outperforms Muon in LLM pretraining.

These findings sharpen the puzzle of why row normalization helps in practice and complement the recent findings of Dewulf et al.