Geometric Inductive Biases of Deep Networks: The Role of Data and Architecture

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Movahedi, Sajad, Orvieto, Antonio, Moosavi-Dezfooli, Seyed-Mohsen
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909613663715328
author Movahedi, Sajad
Orvieto, Antonio
Moosavi-Dezfooli, Seyed-Mohsen
author_facet Movahedi, Sajad
Orvieto, Antonio
Moosavi-Dezfooli, Seyed-Mohsen
contents In this paper, we propose the $\textit{geometric invariance hypothesis (GIH)}$, which argues that the input space curvature of a neural network remains invariant under transformation in certain architecture-dependent directions during training. We investigate a simple, non-linear binary classification problem residing on a plane in a high dimensional space and observe that$\unicode{x2014}$unlike MLPs$\unicode{x2014}$ResNets fail to generalize depending on the orientation of the plane. Motivated by this example, we define a neural network's $\textbf{average geometry}$ and $\textbf{average geometry evolution}$ as compact $\textit{architecture-dependent}$ summaries of the model's input-output geometry and its evolution during training. By investigating the average geometry evolution at initialization, we discover that the geometry of a neural network evolves according to the data covariance projected onto its average geometry. This means that the geometry only changes in a subset of the input space when the average geometry is low-rank, such as in ResNets. This causes an architecture-dependent invariance property in the input space curvature, which we dub GIH. Finally, we present extensive experimental results to observe the consequences of GIH and how it relates to generalization in neural networks.
format Preprint
id arxiv_https___arxiv_org_abs_2410_12025
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Geometric Inductive Biases of Deep Networks: The Role of Data and Architecture
Movahedi, Sajad
Orvieto, Antonio
Moosavi-Dezfooli, Seyed-Mohsen
Machine Learning
In this paper, we propose the $\textit{geometric invariance hypothesis (GIH)}$, which argues that the input space curvature of a neural network remains invariant under transformation in certain architecture-dependent directions during training. We investigate a simple, non-linear binary classification problem residing on a plane in a high dimensional space and observe that$\unicode{x2014}$unlike MLPs$\unicode{x2014}$ResNets fail to generalize depending on the orientation of the plane. Motivated by this example, we define a neural network's $\textbf{average geometry}$ and $\textbf{average geometry evolution}$ as compact $\textit{architecture-dependent}$ summaries of the model's input-output geometry and its evolution during training. By investigating the average geometry evolution at initialization, we discover that the geometry of a neural network evolves according to the data covariance projected onto its average geometry. This means that the geometry only changes in a subset of the input space when the average geometry is low-rank, such as in ResNets. This causes an architecture-dependent invariance property in the input space curvature, which we dub GIH. Finally, we present extensive experimental results to observe the consequences of GIH and how it relates to generalization in neural networks.
title Geometric Inductive Biases of Deep Networks: The Role of Data and Architecture
topic Machine Learning
url https://arxiv.org/abs/2410.12025