$λ$-GELU: Learning Gating Hardness for Controlled ReLU-ization in Deep Networks
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914441564520448 |
|---|---|
| author | Pérez-Corral, Cristian Fernández-Hernández, Alberto Mestre, Jose I. Dolz, Manuel F. Quintana-Ortí, Enrique S. |
| author_facet | Pérez-Corral, Cristian Fernández-Hernández, Alberto Mestre, Jose I. Dolz, Manuel F. Quintana-Ortí, Enrique S. |
| contents | Gaussian Error Linear Unit (GELU) is a widely used smooth alternative to Rectifier Linear Unit (ReLU), yet many deployment, compression, and analysis toolchains are most naturally expressed for piecewise-linear (ReLU-type) networks. We study a hardness-parameterized formulation of GELU, f(x;λ)=xΦ(λ x), where Φ is the Gaussian CDF and λ \in [1, infty) controls gate sharpness, with the goal of turning smooth gated training into a controlled path toward ReLU-compatible models. Learning λ is non-trivial: naive updates yield unstable dynamics and effective gradient attenuation, so we introduce a constrained reparameterization and an optimizer-aware update scheme.
Empirically, across a diverse set of model--dataset pairs spanning MLPs, CNNs, and Transformers, we observe structured layerwise hardness profiles and assess their robustness under different initializations. We further study a deterministic ReLU-ization strategy in which the learned gates are progressively hardened toward a principled target, enabling a post-training substitution of λ-GELU by ReLU with reduced disruption. Overall, λ-GELU provides a minimal and interpretable knob to profile and control gating hardness, bridging smooth training with ReLU-centric downstream pipelines. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_21991 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | $λ$-GELU: Learning Gating Hardness for Controlled ReLU-ization in Deep Networks Pérez-Corral, Cristian Fernández-Hernández, Alberto Mestre, Jose I. Dolz, Manuel F. Quintana-Ortí, Enrique S. Machine Learning Artificial Intelligence Gaussian Error Linear Unit (GELU) is a widely used smooth alternative to Rectifier Linear Unit (ReLU), yet many deployment, compression, and analysis toolchains are most naturally expressed for piecewise-linear (ReLU-type) networks. We study a hardness-parameterized formulation of GELU, f(x;λ)=xΦ(λ x), where Φ is the Gaussian CDF and λ \in [1, infty) controls gate sharpness, with the goal of turning smooth gated training into a controlled path toward ReLU-compatible models. Learning λ is non-trivial: naive updates yield unstable dynamics and effective gradient attenuation, so we introduce a constrained reparameterization and an optimizer-aware update scheme. Empirically, across a diverse set of model--dataset pairs spanning MLPs, CNNs, and Transformers, we observe structured layerwise hardness profiles and assess their robustness under different initializations. We further study a deterministic ReLU-ization strategy in which the learned gates are progressively hardened toward a principled target, enabling a post-training substitution of λ-GELU by ReLU with reduced disruption. Overall, λ-GELU provides a minimal and interpretable knob to profile and control gating hardness, bridging smooth training with ReLU-centric downstream pipelines. |
| title | $λ$-GELU: Learning Gating Hardness for Controlled ReLU-ization in Deep Networks |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2603.21991 |