首页 > AI前沿 > Inner Momentum for Differentially Private Muon

Inner Momentum for Differentially Private Muon

arXiv机器学习 2026-10-02 11:10 5 阅读 查看原文

Differentially private training clips each per-example gradient before adding noise.

This clipping is radial for each example, yet unequal clipping factors can distort the relative singular-vector geometry of their average.

Muon is particularly exposed to this effect, since its update is an approximate polar factor UV^T that depends only on the singular vectors that clipping can shift.

To curb this degradation, we propose averaging each sampled example's Muon gradient over the current model and a short history of recent models before clipping.

The clipped batch matrix then separates into a common rescaling and a covariance residual R between sampled gradients and clipping values, with ||R||_F ≤ sigma_lambda sigma_G, bounding the clipping-induced distortion directly.

We further show that a finite Newton-Schulz iteration preserves the polar factor of its input under these spectral conditions, confirming that our correction survives orthogonalization.

In private GPT-2 fine-tuning on E2E and DART at epsilon in {1, 2, 4, 8}, DP-Muon-IM improves BLEU and ROUGE-L over DP-Muon in every seed-matched comparison,

and non-private diagnostics show 2-4% lower pre-noise polar error.