
Free Heavy-Tailed Lunch for Muon: A Theoretical Justification of Empirical Success
Researchers have closed a theoretical gap around Muon and Scion, non-Euclidean optimizers that have outperformed standard methods in Transformer training without clear justification. The work proves these matrix-based approaches achieve better sample complexity in heavy-tailed gradient regimes, where real-world noise patterns deviate from Gaussian assumptions. This validates an emerging class of optimization techniques gaining traction in large-scale model training, potentially reshaping how practitioners choose solvers for production workloads.62


























