ResNets Are Deeper Than You Think

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mehmeti-Göpel, Christian H. X. Ali, Wand, Michael
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912436058062848
author Mehmeti-Göpel, Christian H. X. Ali
Wand, Michael
author_facet Mehmeti-Göpel, Christian H. X. Ali
Wand, Michael
contents Residual connections remain ubiquitous in modern neural network architectures nearly a decade after their introduction. Their widespread adoption is often credited to their dramatically improved trainability: residual networks train faster, more stably, and achieve higher accuracy than their feedforward counterparts. While numerous techniques, ranging from improved initialization to advanced learning rate schedules, have been proposed to close the performance gap between residual and feedforward networks, this gap has persisted. In this work, we propose an alternative explanation: residual networks do not merely reparameterize feedforward networks, but instead inhabit a different function space. We design a controlled post-training comparison to isolate generalization performance from trainability; we find that variable-depth architectures, similar to ResNets, consistently outperform fixed-depth networks, even when optimization is unlikely to make a difference. These results suggest that residual connections confer performance advantages beyond optimization, pointing instead to a deeper inductive bias aligned with the structure of natural data.
format Preprint
id arxiv_https___arxiv_org_abs_2506_14386
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ResNets Are Deeper Than You Think
Mehmeti-Göpel, Christian H. X. Ali
Wand, Michael
Machine Learning
Artificial Intelligence
Residual connections remain ubiquitous in modern neural network architectures nearly a decade after their introduction. Their widespread adoption is often credited to their dramatically improved trainability: residual networks train faster, more stably, and achieve higher accuracy than their feedforward counterparts. While numerous techniques, ranging from improved initialization to advanced learning rate schedules, have been proposed to close the performance gap between residual and feedforward networks, this gap has persisted. In this work, we propose an alternative explanation: residual networks do not merely reparameterize feedforward networks, but instead inhabit a different function space. We design a controlled post-training comparison to isolate generalization performance from trainability; we find that variable-depth architectures, similar to ResNets, consistently outperform fixed-depth networks, even when optimization is unlikely to make a difference. These results suggest that residual connections confer performance advantages beyond optimization, pointing instead to a deeper inductive bias aligned with the structure of natural data.
title ResNets Are Deeper Than You Think
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2506.14386