Near-optimal estimates for the $\ell^p$-Lipschitz constants of deep random ReLU neural networks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dirksen, Sjoerd, Finke, Patrick, Geuchen, Paul, Stöger, Dominik, Voigtlaender, Felix
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908419651272704
author Dirksen, Sjoerd
Finke, Patrick
Geuchen, Paul
Stöger, Dominik
Voigtlaender, Felix
author_facet Dirksen, Sjoerd
Finke, Patrick
Geuchen, Paul
Stöger, Dominik
Voigtlaender, Felix
contents This paper studies the $\ell^p$-Lipschitz constants of ReLU neural networks $Φ: \mathbb{R}^d \to \mathbb{R}$ with random parameters for $p \in [1,\infty]$. The distribution of the weights follows a variant of the He initialization and the biases are drawn from symmetric distributions. We derive high probability upper and lower bounds for wide networks that differ at most by a factor that is logarithmic in the network's width and linear in its depth. In the special case of shallow networks, we obtain matching bounds. Remarkably, the behavior of the $\ell^p$-Lipschitz constant varies significantly between the regimes $ p \in [1,2) $ and $ p \in [2,\infty] $. For $p \in [2,\infty]$, the $\ell^p$-Lipschitz constant behaves similarly to $\Vert g\Vert_{p'}$, where $g \in \mathbb{R}^d$ is a $d$-dimensional standard Gaussian vector and $1/p + 1/p' = 1$. In contrast, for $p \in [1,2)$, the $\ell^p$-Lipschitz constant aligns more closely to $\Vert g \Vert_{2}$.
format Preprint
id arxiv_https___arxiv_org_abs_2506_19695
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Near-optimal estimates for the $\ell^p$-Lipschitz constants of deep random ReLU neural networks
Dirksen, Sjoerd
Finke, Patrick
Geuchen, Paul
Stöger, Dominik
Voigtlaender, Felix
Machine Learning
Probability
68T07, 26A16, 60B20, 60G15
This paper studies the $\ell^p$-Lipschitz constants of ReLU neural networks $Φ: \mathbb{R}^d \to \mathbb{R}$ with random parameters for $p \in [1,\infty]$. The distribution of the weights follows a variant of the He initialization and the biases are drawn from symmetric distributions. We derive high probability upper and lower bounds for wide networks that differ at most by a factor that is logarithmic in the network's width and linear in its depth. In the special case of shallow networks, we obtain matching bounds. Remarkably, the behavior of the $\ell^p$-Lipschitz constant varies significantly between the regimes $ p \in [1,2) $ and $ p \in [2,\infty] $. For $p \in [2,\infty]$, the $\ell^p$-Lipschitz constant behaves similarly to $\Vert g\Vert_{p'}$, where $g \in \mathbb{R}^d$ is a $d$-dimensional standard Gaussian vector and $1/p + 1/p' = 1$. In contrast, for $p \in [1,2)$, the $\ell^p$-Lipschitz constant aligns more closely to $\Vert g \Vert_{2}$.
title Near-optimal estimates for the $\ell^p$-Lipschitz constants of deep random ReLU neural networks
topic Machine Learning
Probability
68T07, 26A16, 60B20, 60G15
url https://arxiv.org/abs/2506.19695