The Riemannian Geometry Associated to Gradient Flows of Linear Convolutional Networks

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Achour, El Mehdi, Kohn, Kathlén, Rauhut, Holger
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910103816372224
author Achour, El Mehdi
Kohn, Kathlén
Rauhut, Holger
author_facet Achour, El Mehdi
Kohn, Kathlén
Rauhut, Holger
contents We study geometric properties of the gradient flow for learning deep linear convolutional networks. For linear fully connected networks, it has been shown recently that the corresponding gradient flow on parameter space can be written as a Riemannian gradient flow on function space (i.e., on the product of weight matrices) if the initialization satisfies a so-called balancedness condition. We establish that the gradient flow on parameter space for learning linear convolutional networks can be written as a Riemannian gradient flow on function space regardless of the initialization. This result holds for $D$-dimensional convolutions with $D \geq 2$, and for $D =1$ it holds if all so-called strides of the convolutions are greater than one. The corresponding Riemannian metric depends on the initialization.
format Preprint
id arxiv_https___arxiv_org_abs_2507_06367
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Riemannian Geometry Associated to Gradient Flows of Linear Convolutional Networks
Achour, El Mehdi
Kohn, Kathlén
Rauhut, Holger
Machine Learning
Algebraic Geometry
We study geometric properties of the gradient flow for learning deep linear convolutional networks. For linear fully connected networks, it has been shown recently that the corresponding gradient flow on parameter space can be written as a Riemannian gradient flow on function space (i.e., on the product of weight matrices) if the initialization satisfies a so-called balancedness condition. We establish that the gradient flow on parameter space for learning linear convolutional networks can be written as a Riemannian gradient flow on function space regardless of the initialization. This result holds for $D$-dimensional convolutions with $D \geq 2$, and for $D =1$ it holds if all so-called strides of the convolutions are greater than one. The corresponding Riemannian metric depends on the initialization.
title The Riemannian Geometry Associated to Gradient Flows of Linear Convolutional Networks
topic Machine Learning
Algebraic Geometry
url https://arxiv.org/abs/2507.06367