ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Zhengyan, Song, Yixin, Yu, Guanghui, Han, Xu, Lin, Yankai, Xiao, Chaojun, Song, Chenyang, Liu, Zhiyuan, Mi, Zeyu, Sun, Maosong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
by: Xiao, Chaojun, et al.
Published: (2024)
by: Xiao, Chaojun, et al.
Published: (2024)
Exploring the Benefit of Activation Sparsity in Pre-training
by: Zhang, Zhengyan, et al.
Published: (2024)
by: Zhang, Zhengyan, et al.
Published: (2024)
Sparsing Law: Towards Large Language Models with Greater Activation Sparsity
by: Luo, Yuqi, et al.
Published: (2024)
by: Luo, Yuqi, et al.
Published: (2024)
ReCA: A Parametric ReLU Composite Activation Function
by: Chidiac, John, et al.
Published: (2025)
by: Chidiac, John, et al.
Published: (2025)
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity
by: Song, Chenyang, et al.
Published: (2025)
by: Song, Chenyang, et al.
Published: (2025)
Turbo Sparse: Achieving LLM SOTA Performance with Minimal Activated Parameters
by: Song, Yixin, et al.
Published: (2024)
by: Song, Yixin, et al.
Published: (2024)
Discrete Functional Geometry of ReLU Networks via ReLU Transition Graphs
by: Dhayalkar, Sahil Rajesh
Published: (2025)
by: Dhayalkar, Sahil Rajesh
Published: (2025)
Efficient Quantum Circuits for Machine Learning Activation Functions including Constant T-depth ReLU
by: Zi, Wei, et al.
Published: (2024)
by: Zi, Wei, et al.
Published: (2024)
Brownian ReLU(Br-ReLU): A New Activation Function for a Long-Short Term Memory (LSTM) Network
by: Awiakye-Marfo, George, et al.
Published: (2026)
by: Awiakye-Marfo, George, et al.
Published: (2026)
Deep Network Approximation: Beyond ReLU to Diverse Activation Functions
by: Zhang, Shijun, et al.
Published: (2023)
by: Zhang, Shijun, et al.
Published: (2023)
N-ReLU: Zero-Mean Stochastic Extension of ReLU
by: Manik, Md Motaleb Hossen, et al.
Published: (2025)
by: Manik, Md Motaleb Hossen, et al.
Published: (2025)
The Geometry of ReLU Networks through the ReLU Transition Graph
by: Dhayalkar, Sahil Rajesh
Published: (2025)
by: Dhayalkar, Sahil Rajesh
Published: (2025)
ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language Models
by: Song, Chenyang, et al.
Published: (2024)
by: Song, Chenyang, et al.
Published: (2024)
A Significantly Better Class of Activation Functions Than ReLU Like Activation Functions
by: Noel, Mathew Mithra, et al.
Published: (2024)
by: Noel, Mathew Mithra, et al.
Published: (2024)
The Resurrection of the ReLU
by: Horuz, Coşku Can, et al.
Published: (2025)
by: Horuz, Coşku Can, et al.
Published: (2025)
The Optimal Condition Number for ReLU Function
by: Xia, Yu, et al.
Published: (2025)
by: Xia, Yu, et al.
Published: (2025)
DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices
by: Song, Chenyang, et al.
Published: (2026)
by: Song, Chenyang, et al.
Published: (2026)
Representation Learning for Natural Language Processing
by: Liu, Zhiyuan, et al.
Published: (2020)
by: Liu, Zhiyuan, et al.
Published: (2020)
How Does the ReLU Activation Affect the Implicit Bias of Gradient Descent on High-dimensional Neural Network Regression?
by: Lai, Kuo-Wei, et al.
Published: (2026)
by: Lai, Kuo-Wei, et al.
Published: (2026)
Activation-Descent Regularization for Input Optimization of ReLU Networks
by: Yu, Hongzhan, et al.
Published: (2024)
by: Yu, Hongzhan, et al.
Published: (2024)
Topological Signatures of ReLU Neural Network Activation Patterns
by: Bosca, Vicente, et al.
Published: (2025)
by: Bosca, Vicente, et al.
Published: (2025)
Is ReLU Adversarially Robust?
by: Sooksatra, Korn, et al.
Published: (2024)
by: Sooksatra, Korn, et al.
Published: (2024)
Sobolev Approximation of Deep ReLU Networks in Log-Barron Space
by: Song, Changhoon, et al.
Published: (2026)
by: Song, Changhoon, et al.
Published: (2026)
Functional dimension of feedforward ReLU neural networks
by: Grigsby, J. Elisenda, et al.
Published: (2022)
by: Grigsby, J. Elisenda, et al.
Published: (2022)
Sign-Aware Gated Sparse Autoencoders: Modeling Anticorrelated Features with Bi-Jump-ReLU Activations
by: Wieciech, Bartosz, et al.
Published: (2026)
by: Wieciech, Bartosz, et al.
Published: (2026)
Variator: Accelerating Pre-trained Models with Plug-and-Play Compression Modules
by: Xiao, Chaojun, et al.
Published: (2023)
by: Xiao, Chaojun, et al.
Published: (2023)
SEEV: Synthesis with Efficient Exact Verification for ReLU Neural Barrier Functions
by: Zhang, Hongchao, et al.
Published: (2024)
by: Zhang, Hongchao, et al.
Published: (2024)
Better NTK Conditioning: A Free Lunch from (ReLU) Nonlinear Activation in Wide Neural Networks
by: Liu, Chaoyue, et al.
Published: (2023)
by: Liu, Chaoyue, et al.
Published: (2023)
SurvReLU: Inherently Interpretable Survival Analysis via Deep ReLU Networks
by: Sun, Xiaotong, et al.
Published: (2024)
by: Sun, Xiaotong, et al.
Published: (2024)
Neural Characteristic Activation Analysis and Geometric Parameterization for ReLU Networks
by: Chen, Wenlin, et al.
Published: (2023)
by: Chen, Wenlin, et al.
Published: (2023)
Agnostic Learning of Arbitrary ReLU Activation under Gaussian Marginals
by: Guo, Anxin, et al.
Published: (2024)
by: Guo, Anxin, et al.
Published: (2024)
Uncovering Layer-Dependent Activation Sparsity Patterns in ReLU Transformers
by: Wild, Cody, et al.
Published: (2024)
by: Wild, Cody, et al.
Published: (2024)
Continuous Specialization Transition in the Soft Committee Machine with ReLU Activation
by: Afanah, Assem, et al.
Published: (2026)
by: Afanah, Assem, et al.
Published: (2026)
Agnostic Learning of General ReLU Activation Using Gradient Descent
by: Awasthi, Pranjal, et al.
Published: (2022)
by: Awasthi, Pranjal, et al.
Published: (2022)
Newton-Direction-Based ReLU-Thresholding Methods for Nonnegative Sparse Signal Recovery
by: Bian, Ning, et al.
Published: (2026)
by: Bian, Ning, et al.
Published: (2026)
Zorro: A Flexible and Differentiable Parametric Family of Activation Functions That Extends ReLU and GELU
by: Roodschild, Matias, et al.
Published: (2024)
by: Roodschild, Matias, et al.
Published: (2024)
Injectivity capacity of ReLU gates
by: Stojnic, Mihailo
Published: (2024)
by: Stojnic, Mihailo
Published: (2024)
ReLU Networks as Random Functions: Their Distribution in Probability Space
by: Chaudhari, Shreyas, et al.
Published: (2025)
by: Chaudhari, Shreyas, et al.
Published: (2025)
Replacing K-infinity Function with Leaky ReLU in Barrier Function Design: A Union of Invariant Sets Approach for ReLU-Based Dynamical Systems
by: Samanipour, Pouya, et al.
Published: (2025)
by: Samanipour, Pouya, et al.
Published: (2025)
Component-based Sketching for Deep ReLU Nets
by: Wang, Di, et al.
Published: (2024)
by: Wang, Di, et al.
Published: (2024)
Similar Items
-
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
by: Xiao, Chaojun, et al.
Published: (2024) -
Exploring the Benefit of Activation Sparsity in Pre-training
by: Zhang, Zhengyan, et al.
Published: (2024) -
Sparsing Law: Towards Large Language Models with Greater Activation Sparsity
by: Luo, Yuqi, et al.
Published: (2024) -
ReCA: A Parametric ReLU Composite Activation Function
by: Chidiac, John, et al.
Published: (2025) -
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity
by: Song, Chenyang, et al.
Published: (2025)