Diagonal Linear Networks and the Lasso Regularization Path

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
1. Verfasser: Berthier, Raphaël
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918394660388864
author Berthier, Raphaël
author_facet Berthier, Raphaël
contents Diagonal linear networks are neural networks with linear activation and diagonal weight matrices. Their theoretical interest is that their implicit regularization can be rigorously analyzed: from a small initialization, the training of diagonal linear networks converges to the linear predictor with minimal 1-norm among minimizers of the training loss. In this paper, we deepen this analysis showing that the full training trajectory of diagonal linear networks is closely related to the lasso regularization path. In this connection, the training time plays the role of an inverse regularization parameter. Both rigorous results and simulations are provided to illustrate this conclusion. Under a monotonicity assumption on the lasso regularization path, the connection is exact while in the general case, we show an approximate connection.
format Preprint
id arxiv_https___arxiv_org_abs_2509_18766
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Diagonal Linear Networks and the Lasso Regularization Path
Berthier, Raphaël
Machine Learning
Optimization and Control
62J07, 68T07
G.3
Diagonal linear networks are neural networks with linear activation and diagonal weight matrices. Their theoretical interest is that their implicit regularization can be rigorously analyzed: from a small initialization, the training of diagonal linear networks converges to the linear predictor with minimal 1-norm among minimizers of the training loss. In this paper, we deepen this analysis showing that the full training trajectory of diagonal linear networks is closely related to the lasso regularization path. In this connection, the training time plays the role of an inverse regularization parameter. Both rigorous results and simulations are provided to illustrate this conclusion. Under a monotonicity assumption on the lasso regularization path, the connection is exact while in the general case, we show an approximate connection.
title Diagonal Linear Networks and the Lasso Regularization Path
topic Machine Learning
Optimization and Control
62J07, 68T07
G.3
url https://arxiv.org/abs/2509.18766