To 2:4 Sparsity and Beyond: Neuron-level Activation Function to Accelerate LLM Pre-Training
Fuente:
arXiv
Saved in:
| Main Authors: | Madhyastha, Meghana, Haziza, Daniel, Cai, Jesse, Ardalani, Newsha, Bu, Zhiqi, Wu, Carole-Jean |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Accelerating Transformer Inference and Training with 2:4 Activation Sparsity
by: Haziza, Daniel, et al.
Published: (2025)
by: Haziza, Daniel, et al.
Published: (2025)
Composer: A Search Framework for Hybrid Neural Architecture Design
by: Acun, Bilge, et al.
Published: (2025)
by: Acun, Bilge, et al.
Published: (2025)
Masked Matrix Multiplication for Emergent Sparsity
by: Wheatman, Brian, et al.
Published: (2024)
by: Wheatman, Brian, et al.
Published: (2024)
MAD Max Beyond Single-Node: Enabling Large Machine Learning Model Acceleration on Distributed Systems
by: Hsia, Samuel, et al.
Published: (2023)
by: Hsia, Samuel, et al.
Published: (2023)
Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use Instead
by: Kang, Feiyang, et al.
Published: (2025)
by: Kang, Feiyang, et al.
Published: (2025)
Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls
by: Kang, Feiyang, et al.
Published: (2025)
by: Kang, Feiyang, et al.
Published: (2025)
CATransformers: Carbon Aware Transformers Through Joint Model-Hardware Optimization
by: Wang, Irene, et al.
Published: (2025)
by: Wang, Irene, et al.
Published: (2025)
Towards Decentralized and Sustainable Foundation Model Training with the Edge
by: Xue, Leyang, et al.
Published: (2025)
by: Xue, Leyang, et al.
Published: (2025)
Accelerating Transformer Pre-training with 2:4 Sparsity
by: Hu, Yuezhou, et al.
Published: (2024)
by: Hu, Yuezhou, et al.
Published: (2024)
On Harnessing Idle Compute at the Edge for Foundation Model Training
by: Xue, Leyang, et al.
Published: (2025)
by: Xue, Leyang, et al.
Published: (2025)
Annotating the Pangenome Reveals the Diversity in the Genetic Basis for Metabolic Enzymes
by: Ardalani, Omid
Published: (2025)
by: Ardalani, Omid
Published: (2025)
Sieve: Multimodal Dataset Pruning Using Image Captioning Models
by: Mahmoud, Anas, et al.
Published: (2023)
by: Mahmoud, Anas, et al.
Published: (2023)
ELAS: Efficient Pre-Training of Low-Rank Large Language Models via 2:4 Activation Sparsity
by: Li, Jiaxi, et al.
Published: (2026)
by: Li, Jiaxi, et al.
Published: (2026)
Training-Free Activation Sparsity in Large Language Models
by: Liu, James, et al.
Published: (2024)
by: Liu, James, et al.
Published: (2024)
Exploring the Benefit of Activation Sparsity in Pre-training
by: Zhang, Zhengyan, et al.
Published: (2024)
by: Zhang, Zhengyan, et al.
Published: (2024)
Scaling depth capacity via zero/one-layer model expansion
by: Bu, Zhiqi
Published: (2025)
by: Bu, Zhiqi
Published: (2025)
A Cognitively Grounded Bayesian Framework for Misinformation Susceptibility
by: Madhyastha, Pranava
Published: (2026)
by: Madhyastha, Pranava
Published: (2026)
Case report of high origin of radial, ulnar, and profunda brachii arteries, its clinical implications and review of the literature
by: Sampath Madhyastha
Published: (2009)
by: Sampath Madhyastha
Published: (2009)
SparseInfer: Training-free Prediction of Activation Sparsity for Fast LLM Inference
by: Shin, Jiho, et al.
Published: (2024)
by: Shin, Jiho, et al.
Published: (2024)
Text Quality-Based Pruning for Efficient Training of Language Models
by: Sharma, Vasu, et al.
Published: (2024)
by: Sharma, Vasu, et al.
Published: (2024)
Improving Model Fusion by Training-time Neuron Alignment with Fixed Neuron Anchors
by: Li, Zexi, et al.
Published: (2024)
by: Li, Zexi, et al.
Published: (2024)
Resting Neurons, Active Insights: Robustifying Activation Sparsity in LLMs via Spontaneity
by: Xu, Haotian, et al.
Published: (2025)
by: Xu, Haotian, et al.
Published: (2025)
Sparsity-Accelerated Training for Large Language Models
by: Ma, Da, et al.
Published: (2024)
by: Ma, Da, et al.
Published: (2024)
Decomposing the Neurons: Activation Sparsity via Mixture of Experts for Continual Test Time Adaptation
by: Zhang, Rongyu, et al.
Published: (2024)
by: Zhang, Rongyu, et al.
Published: (2024)
Joint Training Across Multiple Activation Sparsity Regimes
by: Wang, Haotian
Published: (2026)
by: Wang, Haotian
Published: (2026)
Post-Training Statistical Calibration for Higher Activation Sparsity
by: Chua, Vui Seng, et al.
Published: (2024)
by: Chua, Vui Seng, et al.
Published: (2024)
Machine learning methods for finite population parameter estimation in survey sampling
by: Dagdoug, Mehdi, et al.
Published: (2026)
by: Dagdoug, Mehdi, et al.
Published: (2026)
Beyond Public Access in LLM Pre-Training Data
by: Rosenblat, Sruly, et al.
Published: (2025)
by: Rosenblat, Sruly, et al.
Published: (2025)
Wasserstein Distances, Neuronal Entanglement, and Sparsity
by: Sawmya, Shashata, et al.
Published: (2024)
by: Sawmya, Shashata, et al.
Published: (2024)
Noise or Nuance: An Investigation Into Useful Information and Filtering For LLM Driven AKBC
by: Clay, Alex, et al.
Published: (2025)
by: Clay, Alex, et al.
Published: (2025)
PowLU: An Activation Function for Stable Pre-Training of LLMs
by: Jiang, Peijie, et al.
Published: (2026)
by: Jiang, Peijie, et al.
Published: (2026)
Pre-training Differentially Private Models with Limited Public Data
by: Bu, Zhiqi, et al.
Published: (2024)
by: Bu, Zhiqi, et al.
Published: (2024)
LLM-Assisted Visual Analytics: Opportunities and Challenges
by: Hutchinson, Maeve, et al.
Published: (2024)
by: Hutchinson, Maeve, et al.
Published: (2024)
Adaptive parameter-efficient fine-tuning via Hessian-informed subset selection
by: Xu, Shiyun, et al.
Published: (2025)
by: Xu, Shiyun, et al.
Published: (2025)
Gradient descent with generalized Newton's method
by: Bu, Zhiqi, et al.
Published: (2024)
by: Bu, Zhiqi, et al.
Published: (2024)
Towards hyperparameter-free optimization with differential privacy
by: Bu, Zhiqi, et al.
Published: (2025)
by: Bu, Zhiqi, et al.
Published: (2025)
Weight Block Sparsity: Training, Compilation, and AI Engine Accelerators
by: D'Alberto, Paolo, et al.
Published: (2024)
by: D'Alberto, Paolo, et al.
Published: (2024)
Accelerating Training with Neuron Interaction and Nowcasting Networks
by: Knyazev, Boris, et al.
Published: (2024)
by: Knyazev, Boris, et al.
Published: (2024)
Variable Selection for Linear Regression Imputation in Surveys
by: An, Ziming, et al.
Published: (2026)
by: An, Ziming, et al.
Published: (2026)
Working Memory Constraints Scaffold Learning in Transformers under Data Scarcity
by: Madhyastha, Pranava, et al.
Published: (2026)
by: Madhyastha, Pranava, et al.
Published: (2026)
Similar Items
-
Accelerating Transformer Inference and Training with 2:4 Activation Sparsity
by: Haziza, Daniel, et al.
Published: (2025) -
Composer: A Search Framework for Hybrid Neural Architecture Design
by: Acun, Bilge, et al.
Published: (2025) -
Masked Matrix Multiplication for Emergent Sparsity
by: Wheatman, Brian, et al.
Published: (2024) -
MAD Max Beyond Single-Node: Enabling Large Machine Learning Model Acceleration on Distributed Systems
by: Hsia, Samuel, et al.
Published: (2023) -
Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use Instead
by: Kang, Feiyang, et al.
Published: (2025)