Understanding Optimization in Deep Learning with Central Flows

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cohen, Jeremy M., Damian, Alex, Talwalkar, Ameet, Kolter, J. Zico, Lee, Jason D.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908557257998336
author Cohen, Jeremy M.
Damian, Alex
Talwalkar, Ameet
Kolter, J. Zico
Lee, Jason D.
author_facet Cohen, Jeremy M.
Damian, Alex
Talwalkar, Ameet
Kolter, J. Zico
Lee, Jason D.
contents Traditional theories of optimization cannot describe the dynamics of optimization in deep learning, even in the simple setting of deterministic training. The challenge is that optimizers typically operate in a complex, oscillatory regime called the "edge of stability." In this paper, we develop theory that can describe the dynamics of optimization in this regime. Our key insight is that while the *exact* trajectory of an oscillatory optimizer may be challenging to analyze, the *time-averaged* (i.e. smoothed) trajectory is often much more tractable. To analyze an optimizer, we derive a differential equation called a "central flow" that characterizes this time-averaged trajectory. We empirically show that these central flows can predict long-term optimization trajectories for generic neural networks with a high degree of numerical accuracy. By interpreting these central flows, we are able to understand how gradient descent makes progress even as the loss sometimes goes up; how adaptive optimizers "adapt" to the local loss landscape; and how adaptive optimizers implicitly navigate towards regions where they can take larger steps. Our results suggest that central flows can be a valuable theoretical tool for reasoning about optimization in deep learning.
format Preprint
id arxiv_https___arxiv_org_abs_2410_24206
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Understanding Optimization in Deep Learning with Central Flows
Cohen, Jeremy M.
Damian, Alex
Talwalkar, Ameet
Kolter, J. Zico
Lee, Jason D.
Machine Learning
Artificial Intelligence
Optimization and Control
Traditional theories of optimization cannot describe the dynamics of optimization in deep learning, even in the simple setting of deterministic training. The challenge is that optimizers typically operate in a complex, oscillatory regime called the "edge of stability." In this paper, we develop theory that can describe the dynamics of optimization in this regime. Our key insight is that while the *exact* trajectory of an oscillatory optimizer may be challenging to analyze, the *time-averaged* (i.e. smoothed) trajectory is often much more tractable. To analyze an optimizer, we derive a differential equation called a "central flow" that characterizes this time-averaged trajectory. We empirically show that these central flows can predict long-term optimization trajectories for generic neural networks with a high degree of numerical accuracy. By interpreting these central flows, we are able to understand how gradient descent makes progress even as the loss sometimes goes up; how adaptive optimizers "adapt" to the local loss landscape; and how adaptive optimizers implicitly navigate towards regions where they can take larger steps. Our results suggest that central flows can be a valuable theoretical tool for reasoning about optimization in deep learning.
title Understanding Optimization in Deep Learning with Central Flows
topic Machine Learning
Artificial Intelligence
Optimization and Control
url https://arxiv.org/abs/2410.24206