$λ$-GELU: Learning Gating Hardness for Controlled ReLU-ization in Deep Networks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pérez-Corral, Cristian, Fernández-Hernández, Alberto, Mestre, Jose I., Dolz, Manuel F., Quintana-Ortí, Enrique S.
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914441564520448
author Pérez-Corral, Cristian
Fernández-Hernández, Alberto
Mestre, Jose I.
Dolz, Manuel F.
Quintana-Ortí, Enrique S.
author_facet Pérez-Corral, Cristian
Fernández-Hernández, Alberto
Mestre, Jose I.
Dolz, Manuel F.
Quintana-Ortí, Enrique S.
contents Gaussian Error Linear Unit (GELU) is a widely used smooth alternative to Rectifier Linear Unit (ReLU), yet many deployment, compression, and analysis toolchains are most naturally expressed for piecewise-linear (ReLU-type) networks. We study a hardness-parameterized formulation of GELU, f(x;λ)=xΦ(λ x), where Φ is the Gaussian CDF and λ \in [1, infty) controls gate sharpness, with the goal of turning smooth gated training into a controlled path toward ReLU-compatible models. Learning λ is non-trivial: naive updates yield unstable dynamics and effective gradient attenuation, so we introduce a constrained reparameterization and an optimizer-aware update scheme. Empirically, across a diverse set of model--dataset pairs spanning MLPs, CNNs, and Transformers, we observe structured layerwise hardness profiles and assess their robustness under different initializations. We further study a deterministic ReLU-ization strategy in which the learned gates are progressively hardened toward a principled target, enabling a post-training substitution of λ-GELU by ReLU with reduced disruption. Overall, λ-GELU provides a minimal and interpretable knob to profile and control gating hardness, bridging smooth training with ReLU-centric downstream pipelines.
format Preprint
id arxiv_https___arxiv_org_abs_2603_21991
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle $λ$-GELU: Learning Gating Hardness for Controlled ReLU-ization in Deep Networks
Pérez-Corral, Cristian
Fernández-Hernández, Alberto
Mestre, Jose I.
Dolz, Manuel F.
Quintana-Ortí, Enrique S.
Machine Learning
Artificial Intelligence
Gaussian Error Linear Unit (GELU) is a widely used smooth alternative to Rectifier Linear Unit (ReLU), yet many deployment, compression, and analysis toolchains are most naturally expressed for piecewise-linear (ReLU-type) networks. We study a hardness-parameterized formulation of GELU, f(x;λ)=xΦ(λ x), where Φ is the Gaussian CDF and λ \in [1, infty) controls gate sharpness, with the goal of turning smooth gated training into a controlled path toward ReLU-compatible models. Learning λ is non-trivial: naive updates yield unstable dynamics and effective gradient attenuation, so we introduce a constrained reparameterization and an optimizer-aware update scheme. Empirically, across a diverse set of model--dataset pairs spanning MLPs, CNNs, and Transformers, we observe structured layerwise hardness profiles and assess their robustness under different initializations. We further study a deterministic ReLU-ization strategy in which the learned gates are progressively hardened toward a principled target, enabling a post-training substitution of λ-GELU by ReLU with reduced disruption. Overall, λ-GELU provides a minimal and interpretable knob to profile and control gating hardness, bridging smooth training with ReLU-centric downstream pipelines.
title $λ$-GELU: Learning Gating Hardness for Controlled ReLU-ization in Deep Networks
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2603.21991