The Resurrection of the ReLU

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Horuz, Coşku Can, Kasenbacher, Geoffrey, Higuchi, Saya, Kairat, Sebastian, Stoltz, Jendrik, Pesl, Moritz, Moser, Bernhard A., Linse, Christoph, Martinetz, Thomas, Otte, Sebastian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918037203976192
author Horuz, Coşku Can
Kasenbacher, Geoffrey
Higuchi, Saya
Kairat, Sebastian
Stoltz, Jendrik
Pesl, Moritz
Moser, Bernhard A.
Linse, Christoph
Martinetz, Thomas
Otte, Sebastian
author_facet Horuz, Coşku Can
Kasenbacher, Geoffrey
Higuchi, Saya
Kairat, Sebastian
Stoltz, Jendrik
Pesl, Moritz
Moser, Bernhard A.
Linse, Christoph
Martinetz, Thomas
Otte, Sebastian
contents Modeling sophisticated activation functions within deep learning architectures has evolved into a distinct research direction. Functions such as GELU, SELU, and SiLU offer smooth gradients and improved convergence properties, making them popular choices in state-of-the-art models. Despite this trend, the classical ReLU remains appealing due to its simplicity, inherent sparsity, and other advantageous topological characteristics. However, ReLU units are prone to becoming irreversibly inactive - a phenomenon known as the dying ReLU problem - which limits their overall effectiveness. In this work, we introduce surrogate gradient learning for ReLU (SUGAR) as a novel, plug-and-play regularizer for deep architectures. SUGAR preserves the standard ReLU function during the forward pass but replaces its derivative in the backward pass with a smooth surrogate that avoids zeroing out gradients. We demonstrate that SUGAR, when paired with a well-chosen surrogate function, substantially enhances generalization performance over convolutional network architectures such as VGG-16 and ResNet-18, providing sparser activations while effectively resurrecting dead ReLUs. Moreover, we show that even in modern architectures like Conv2NeXt and Swin Transformer - which typically employ GELU - substituting these with SUGAR yields competitive and even slightly superior performance. These findings challenge the prevailing notion that advanced activation functions are necessary for optimal performance. Instead, they suggest that the conventional ReLU, particularly with appropriate gradient handling, can serve as a strong, versatile revived classic across a broad range of deep learning vision models.
format Preprint
id arxiv_https___arxiv_org_abs_2505_22074
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Resurrection of the ReLU
Horuz, Coşku Can
Kasenbacher, Geoffrey
Higuchi, Saya
Kairat, Sebastian
Stoltz, Jendrik
Pesl, Moritz
Moser, Bernhard A.
Linse, Christoph
Martinetz, Thomas
Otte, Sebastian
Machine Learning
Artificial Intelligence
Modeling sophisticated activation functions within deep learning architectures has evolved into a distinct research direction. Functions such as GELU, SELU, and SiLU offer smooth gradients and improved convergence properties, making them popular choices in state-of-the-art models. Despite this trend, the classical ReLU remains appealing due to its simplicity, inherent sparsity, and other advantageous topological characteristics. However, ReLU units are prone to becoming irreversibly inactive - a phenomenon known as the dying ReLU problem - which limits their overall effectiveness. In this work, we introduce surrogate gradient learning for ReLU (SUGAR) as a novel, plug-and-play regularizer for deep architectures. SUGAR preserves the standard ReLU function during the forward pass but replaces its derivative in the backward pass with a smooth surrogate that avoids zeroing out gradients. We demonstrate that SUGAR, when paired with a well-chosen surrogate function, substantially enhances generalization performance over convolutional network architectures such as VGG-16 and ResNet-18, providing sparser activations while effectively resurrecting dead ReLUs. Moreover, we show that even in modern architectures like Conv2NeXt and Swin Transformer - which typically employ GELU - substituting these with SUGAR yields competitive and even slightly superior performance. These findings challenge the prevailing notion that advanced activation functions are necessary for optimal performance. Instead, they suggest that the conventional ReLU, particularly with appropriate gradient handling, can serve as a strong, versatile revived classic across a broad range of deep learning vision models.
title The Resurrection of the ReLU
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2505.22074