首页 > AI前沿 > Convergence of Practical Muon with Finite Newton-Schulz Iterations and Nesterov Momentum

Convergence of Practical Muon with Finite Newton-Schulz Iterations and Nesterov Momentum

arXiv机器学习 2026-09-30 20:19 4 阅读 查看原文

Practical Muon maintains momentum and performs a small, fixed number of Newton--Schulz iterations separately for each parameter matrix, often with a Nesterov correction.

We analyze these layer-wise finite-step updates jointly on a coupled nonconvex objective, rather than replacing them by exact polar factors or one global orthogonalization.

Under gradient-dependent $(\mathcal L_0,\mathcal L_1,q)$-smoothness and conditionally unbiased stochastic gradients with bounded layer-wise variance, we establish an $\mathcal O(T^{-1/4})$ bound on the expected average Frobenius gradient norm.

The analysis retains the Nesterov recursion and requires neither bounded stochastic gradients, symmetric noise, nor a uniform positive lower bound on the nonzero output singular values.

Its constants contain no explicit matrix-dimension or rank factors when the number of blocks and problem constants are fixed.

The proof follows a descent inequality and a decomposition of the momentum tracking error into initialization, noise, and drift.

For the original five-step quintic, we verify the required scalar-map bounds analytically; the result also allows step-dependent coefficients satisfying the same bounds.

A complementary nuclear-norm result quantifies rank dependence under a stronger spectral condition.

The vanishing rate uses coupled learning-rate and momentum schedules, including the standard single-coefficient Nesterov rule.