Practical Muon maintains momentum and performs a small, fixed number of Newton--Schulz iterations separately for each parameter matrix, often with a Nesterov correction.
We analyze these layer-wise finite-step updates jointly on a coupled nonconvex objective, rather than replacing them by exact polar factors or one global orthogonalization.
Under gradient-dependent $(\mathcal L_0,\mathcal L_1,q)$-smoothness and conditionally unbiased stochastic gradients with bounded layer-wise variance, we establish an $\mathcal O(T^{-1/4})$ bound on the expected average Frobenius gradient norm.
The analysis retains the Nesterov recursion and requires neither bounded stochastic gradients, symmetric noise, nor a uniform positive lower bound on the nonzero output singular values.
Its constants contain no explicit matrix-dimension or rank factors when the number of blocks and problem constants are fixed.
The proof follows a descent inequality and a decomposition of the momentum tracking error into initialization, noise, and drift.
For the original five-step quintic, we verify the required scalar-map bounds analytically; the result also allows step-dependent coefficients satisfying the same bounds.
A complementary nuclear-norm result quantifies rank dependence under a stronger spectral condition.
The vanishing rate uses coupled learning-rate and momentum schedules, including the standard single-coefficient Nesterov rule.