Accelerating Transformer Inference and Training with 2:4 Activation Sparsity
Fuente:
arXiv
Guardado en:
| Autores principales: | Haziza, Daniel, Chou, Timothy, Choudhary, Dhruv, Wehrstedt, Luca, Massa, Francisco, Yu, Jiecao, Jeong, Geonhwa, Rao, Supriya, Labatut, Patrick, Cai, Jesse |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
To 2:4 Sparsity and Beyond: Neuron-level Activation Function to Accelerate LLM Pre-Training
por: Madhyastha, Meghana, et al.
Publicado: (2026)
por: Madhyastha, Meghana, et al.
Publicado: (2026)
Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns
por: Bambhaniya, Abhimanyu, et al.
Publicado: (2026)
por: Bambhaniya, Abhimanyu, et al.
Publicado: (2026)
Inducing Semi-Structured Sparsity by Masking for Efficient Model Inference in Convolutional Networks
por: Danhofer, David A.
Publicado: (2024)
por: Danhofer, David A.
Publicado: (2024)
Sparsing Law: Towards Large Language Models with Greater Activation Sparsity
por: Luo, Yuqi, et al.
Publicado: (2024)
por: Luo, Yuqi, et al.
Publicado: (2024)
MotherNet: Fast Training and Inference via Hyper-Network Transformers
por: Müller, Andreas, et al.
Publicado: (2023)
por: Müller, Andreas, et al.
Publicado: (2023)
Mean-Pooled Cosine Similarity is Not Length-Invariant: Theory and Cross-Domain Evidence for a Length-Invariant Alternative
por: Mitra, Sibayan, et al.
Publicado: (2026)
por: Mitra, Sibayan, et al.
Publicado: (2026)
Emotion-Inspired Learning Signals (EILS): A Homeostatic Framework for Adaptive Autonomous Agents
por: Tiwari, Dhruv
Publicado: (2025)
por: Tiwari, Dhruv
Publicado: (2025)
BlindSight: Harnessing Sparsity for Efficient Vision-Language Models
por: Srikrishnan, Tharun Adithya, et al.
Publicado: (2025)
por: Srikrishnan, Tharun Adithya, et al.
Publicado: (2025)
ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language Models
por: Song, Chenyang, et al.
Publicado: (2024)
por: Song, Chenyang, et al.
Publicado: (2024)
Steering Conceptual Bias via Transformer Latent-Subspace Activation
por: Sharma, Vansh, et al.
Publicado: (2025)
por: Sharma, Vansh, et al.
Publicado: (2025)
The Data Efficiency Frontier of Financial Foundation Models: Scaling Laws from Continued Pretraining
por: Ponnock, Jesse
Publicado: (2025)
por: Ponnock, Jesse
Publicado: (2025)
Inference-Time Machine Unlearning via Gated Activation Redirection
por: Turani, Vinícius Conte, et al.
Publicado: (2026)
por: Turani, Vinícius Conte, et al.
Publicado: (2026)
Nonparametric Partial Disentanglement via Mechanism Sparsity: Sparse Actions, Interventions and Sparse Temporal Dependencies
por: Lachapelle, Sébastien, et al.
Publicado: (2024)
por: Lachapelle, Sébastien, et al.
Publicado: (2024)
Optimizing Inference in Transformer-Based Models: A Multi-Method Benchmark
por: Ho, Siu Hang, et al.
Publicado: (2025)
por: Ho, Siu Hang, et al.
Publicado: (2025)
CrossVLA: Cross-Paradigm Post-Training and Inference Optimization for Vision-Language-Action Models
por: Liu, Zhi
Publicado: (2026)
por: Liu, Zhi
Publicado: (2026)
Generative AI for Strategic Plan Development
por: Ponnock, Jesse
Publicado: (2025)
por: Ponnock, Jesse
Publicado: (2025)
GradES: Significantly Faster Training in Transformers with Gradient-Based Early Stopping
por: Wen, Qifu, et al.
Publicado: (2025)
por: Wen, Qifu, et al.
Publicado: (2025)
Pointing-Guided Target Estimation via Transformer-Based Attention
por: Müller, Luca, et al.
Publicado: (2025)
por: Müller, Luca, et al.
Publicado: (2025)
Painless Activation Steering: An Automated, Lightweight Approach for Post-Training Large Language Models
por: Cui, Sasha, et al.
Publicado: (2025)
por: Cui, Sasha, et al.
Publicado: (2025)
Near-Oracle KV Selection via Pre-hoc Sparsity for Long-Context Inference
por: Gao, Yifei, et al.
Publicado: (2026)
por: Gao, Yifei, et al.
Publicado: (2026)
Once-For-All: A Train-Once and Select-Anytime Framework for Multimodal Instruction Tuning
por: Dong, Mingkang, et al.
Publicado: (2026)
por: Dong, Mingkang, et al.
Publicado: (2026)
The Effect of Mobility Trajectory Sparsity on Epidemic Modeling Outcomes
por: Delussu, Federico, et al.
Publicado: (2026)
por: Delussu, Federico, et al.
Publicado: (2026)
CASE: Contrastive Activation for Saliency Estimation
por: Williamson, Dane, et al.
Publicado: (2025)
por: Williamson, Dane, et al.
Publicado: (2025)
VQ4DiT: Efficient Post-Training Vector Quantization for Diffusion Transformers
por: Deng, Juncan, et al.
Publicado: (2024)
por: Deng, Juncan, et al.
Publicado: (2024)
Controlling Language and Diffusion Models by Transporting Activations
por: Rodriguez, Pau, et al.
Publicado: (2024)
por: Rodriguez, Pau, et al.
Publicado: (2024)
Product-of-Experts Training Reduces Dataset Artifacts in Natural Language Inference
por: Mathew, Aby Mammen
Publicado: (2026)
por: Mathew, Aby Mammen
Publicado: (2026)
Efficient Reasoning via Thought-Training and Thought-Free Inference
por: Wu, Canhui, et al.
Publicado: (2025)
por: Wu, Canhui, et al.
Publicado: (2025)
PolyGLU: State-Conditional Activation Routing in Transformer Feed-Forward Networks
por: Medeiros, Daniel Nobrega
Publicado: (2026)
por: Medeiros, Daniel Nobrega
Publicado: (2026)
ReBoot: Encrypted Training of Deep Neural Networks with CKKS Bootstrapping
por: Pirillo, Alberto, et al.
Publicado: (2025)
por: Pirillo, Alberto, et al.
Publicado: (2025)
${\tt KRAFT}$: Sampling-Based Kinodynamic Replanning and Feedback Control over Approximate, Identified Models of Vehicular Systems
por: Sivaramakrishnan, Aravind, et al.
Publicado: (2024)
por: Sivaramakrishnan, Aravind, et al.
Publicado: (2024)
The Efficacy of Semantics-Preserving Transformations in Self-Supervised Learning for Medical Ultrasound
por: VanBerlo, Blake, et al.
Publicado: (2025)
por: VanBerlo, Blake, et al.
Publicado: (2025)
Exploring the Relationship: Transformative Adaptive Activation Functions in Comparison to Other Activation Functions
por: Kunc, Vladimír
Publicado: (2024)
por: Kunc, Vladimír
Publicado: (2024)
FAME: Feature Activation Map Explanation on Image Classification and Face Recognition
por: Zhang, Xinyi, et al.
Publicado: (2026)
por: Zhang, Xinyi, et al.
Publicado: (2026)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
por: Raoufi, Behnam, et al.
Publicado: (2025)
por: Raoufi, Behnam, et al.
Publicado: (2025)
ForAug: Recombining Foregrounds and Backgrounds to Improve Vision Transformer Training with Bias Mitigation
por: Nauen, Tobias Christian, et al.
Publicado: (2025)
por: Nauen, Tobias Christian, et al.
Publicado: (2025)
Closed-Loop Neural Activation Control in Vision-Language-Action Models
por: Babu, Abhijith, et al.
Publicado: (2026)
por: Babu, Abhijith, et al.
Publicado: (2026)
TAGC: Optimizing Gradient Communication in Distributed Transformer Training
por: Polyakov, Igor, et al.
Publicado: (2025)
por: Polyakov, Igor, et al.
Publicado: (2025)
PowLU: An Activation Function for Stable Pre-Training of LLMs
por: Jiang, Peijie, et al.
Publicado: (2026)
por: Jiang, Peijie, et al.
Publicado: (2026)
Quality Versus Sparsity in Image Recovery by Dictionary Learning Using Iterative Shrinkage
por: Khoshghiaferezaee, Mohammadsadegh, et al.
Publicado: (2025)
por: Khoshghiaferezaee, Mohammadsadegh, et al.
Publicado: (2025)
RPRA: Predicting an LLM-Judge for Efficient but Performant Inference
por: Ashley, Dylan R., et al.
Publicado: (2026)
por: Ashley, Dylan R., et al.
Publicado: (2026)
Ejemplares similares
-
To 2:4 Sparsity and Beyond: Neuron-level Activation Function to Accelerate LLM Pre-Training
por: Madhyastha, Meghana, et al.
Publicado: (2026) -
Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns
por: Bambhaniya, Abhimanyu, et al.
Publicado: (2026) -
Inducing Semi-Structured Sparsity by Masking for Efficient Model Inference in Convolutional Networks
por: Danhofer, David A.
Publicado: (2024) -
Sparsing Law: Towards Large Language Models with Greater Activation Sparsity
por: Luo, Yuqi, et al.
Publicado: (2024) -
MotherNet: Fast Training and Inference via Hyper-Network Transformers
por: Müller, Andreas, et al.
Publicado: (2023)