Optimization Insights into Deep Diagonal Linear Networks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Labarrière, Hippolyte, Molinari, Cesare, Rosasco, Lorenzo, Vega, Cristian, Villa, Silvia
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911383253155840
author Labarrière, Hippolyte
Molinari, Cesare
Rosasco, Lorenzo
Vega, Cristian
Villa, Silvia
author_facet Labarrière, Hippolyte
Molinari, Cesare
Rosasco, Lorenzo
Vega, Cristian
Villa, Silvia
contents Gradient-based methods successfully train highly overparameterized models in practice, even though the associated optimization problems are markedly nonconvex. Understanding the mechanisms that make such methods effective has become a central problem in modern optimization. To investigate this question in a tractable setting, we study Deep Diagonal Linear Networks. These are multilayer architectures with a reparameterization that preserves convexity in the effective parameter, while inducing a nontrivial geometry in the optimization landscape. Under mild initialization conditions, we show that gradient flow on the layer parameters induces a mirror-flow dynamic in the effective parameter space. This structural insight yields explicit convergence guarantees, including exponential decay of the loss under a Polyak-Lojasiewicz condition, and clarifies how the parametrization and initialization scale govern the training speed. Overall, our results demonstrate that deep diagonal over parameterizations, despite their apparent complexity, can endow standard gradient methods with well-behaved and interpretable optimization dynamics.
format Preprint
id arxiv_https___arxiv_org_abs_2412_16765
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Optimization Insights into Deep Diagonal Linear Networks
Labarrière, Hippolyte
Molinari, Cesare
Rosasco, Lorenzo
Vega, Cristian
Villa, Silvia
Machine Learning
Optimization and Control
Gradient-based methods successfully train highly overparameterized models in practice, even though the associated optimization problems are markedly nonconvex. Understanding the mechanisms that make such methods effective has become a central problem in modern optimization. To investigate this question in a tractable setting, we study Deep Diagonal Linear Networks. These are multilayer architectures with a reparameterization that preserves convexity in the effective parameter, while inducing a nontrivial geometry in the optimization landscape. Under mild initialization conditions, we show that gradient flow on the layer parameters induces a mirror-flow dynamic in the effective parameter space. This structural insight yields explicit convergence guarantees, including exponential decay of the loss under a Polyak-Lojasiewicz condition, and clarifies how the parametrization and initialization scale govern the training speed. Overall, our results demonstrate that deep diagonal over parameterizations, despite their apparent complexity, can endow standard gradient methods with well-behaved and interpretable optimization dynamics.
title Optimization Insights into Deep Diagonal Linear Networks
topic Machine Learning
Optimization and Control
url https://arxiv.org/abs/2412.16765