Seminar

Implicit Biases of Large Learning Rates in Machine Learning

June 20, 2025 · Molei Tao · Georgia Institute of Technology
Implicit Biases of Large Learning Rates in Machine Learning

This talk discusses some nontrivial but often pleasant effects of large learning rates, which are commonly used in machine-learning practice for improved empirical performance but defy traditional theoretical analyses. I first quantify how large learning rates can help gradient descent in multiscale landscapes escape local minima via chaotic dynamics — an alternative to the commonly known escape mechanism due to stochastic-gradient noise. I then show how large learning rates provably bias toward flatter minimizers, unifying and explaining several recently observed phenomena (Edge of Stability, loss catapulting, and balancing) as algorithmic implicit biases of large learning rates. These results are enabled by a new global convergence result of gradient descent for certain nonconvex functions without Lipschitz gradient.

Speaker

Molei Tao is a full professor in the School of Mathematics at Georgia Tech, working on the mathematical foundations of machine learning. He received his B.S. from Tsinghua University and Ph.D. from Caltech, was a Courant Instructor at NYU, serves as an Area Chair for NeurIPS, ICLR and ICML, and received the NSF CAREER Award and the AISTATS best-paper award, among other recognitions. He is a founding PI of the Georgia Tech AI4Science Institute.