Spark Transformer: Reactivating Sparsity in FFN and Attention
Fuente:
arXiv
Guardado en:
| Autores principales: | You, Chong, Wu, Kan, Jia, Zhipeng, Chen, Lin, Bhojanapalli, Srinadh, Guo, Jiaxian, Evci, Utku, Wassenberg, Jan, Netrapalli, Praneeth, Willcock, Jeremiah J., Subramanian, Suvinay, Chern, Felix, Andreev, Alek, Pathak, Shreya, Yu, Felix, Jain, Prateek, Culler, David E., Levy, Henry M., Kumar, Sanjiv |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
HiRE: High Recall Approximate Top-$k$ Estimation for Efficient LLM Inference
por: L, Yashas Samaga B, et al.
Publicado: (2024)
por: L, Yashas Samaga B, et al.
Publicado: (2024)
Progressive Gradient Flow for Robust N:M Sparsity Training in Transformers
por: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Publicado: (2024)
por: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Publicado: (2024)
Scalable In-context Ranking with Generative Models
por: Gupta, Nilesh, et al.
Publicado: (2025)
por: Gupta, Nilesh, et al.
Publicado: (2025)
On student-teacher deviations in distillation: does it pay to disobey?
por: Nagarajan, Vaishnavh, et al.
Publicado: (2023)
por: Nagarajan, Vaishnavh, et al.
Publicado: (2023)
Mimetic Initialization Helps State Space Models Learn to Recall
por: Trockman, Asher, et al.
Publicado: (2024)
por: Trockman, Asher, et al.
Publicado: (2024)
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws
por: Jin, Tian, et al.
Publicado: (2025)
por: Jin, Tian, et al.
Publicado: (2025)
Dynamic Sparse Training with Structured Sparsity
por: Lasby, Mike, et al.
Publicado: (2023)
por: Lasby, Mike, et al.
Publicado: (2023)
A model of errors in transformers
por: Raju, Suvrat, et al.
Publicado: (2026)
por: Raju, Suvrat, et al.
Publicado: (2026)
The Feature Speed Formula: a flexible approach to scale hyper-parameters of deep neural networks
por: Chizat, Lénaïc, et al.
Publicado: (2023)
por: Chizat, Lénaïc, et al.
Publicado: (2023)
Compression Scaling Laws:Unifying Sparsity and Quantization
por: Frantar, Elias, et al.
Publicado: (2025)
por: Frantar, Elias, et al.
Publicado: (2025)
Tandem Transformers for Inference Efficient LLMs
por: S, Aishwarya P, et al.
Publicado: (2024)
por: S, Aishwarya P, et al.
Publicado: (2024)
Dual-Encoders for Extreme Multi-Label Classification
por: Gupta, Nilesh, et al.
Publicado: (2023)
por: Gupta, Nilesh, et al.
Publicado: (2023)
Arithmetic Transformers Can Length-Generalize in Both Operand Length and Count
por: Cho, Hanseul, et al.
Publicado: (2024)
por: Cho, Hanseul, et al.
Publicado: (2024)
A Faster Generalized Two-Stage Approximate Top-K
por: Samaga, Yashas, et al.
Publicado: (2025)
por: Samaga, Yashas, et al.
Publicado: (2025)
Polnische Außen- und Sicherheitspolitik 2005-2015 und die Strategie der begrenzten Unabhängigkeit
por: Wassenberg, Florian
Publicado: (2022)
por: Wassenberg, Florian
Publicado: (2022)
Autoregressive Ranking: Bridging the Gap Between Dual and Cross Encoders
por: Rozonoyer, Benjamin, et al.
Publicado: (2026)
por: Rozonoyer, Benjamin, et al.
Publicado: (2026)
Fast Forward: Accelerating LLM Prefill with Predictive FFN Sparsity
por: Gautam, Aayush, et al.
Publicado: (2026)
por: Gautam, Aayush, et al.
Publicado: (2026)
Characterizing VLA Models: Identifying the Action Generation Bottleneck for Edge AI Architectures
por: Vishwanathan, Manoj, et al.
Publicado: (2026)
por: Vishwanathan, Manoj, et al.
Publicado: (2026)
Efficient Language Model Architectures for Differentially Private Federated Learning
por: Ro, Jae Hun, et al.
Publicado: (2024)
por: Ro, Jae Hun, et al.
Publicado: (2024)
Position Coupling: Improving Length Generalization of Arithmetic Transformers Using Task Structure
por: Cho, Hanseul, et al.
Publicado: (2024)
por: Cho, Hanseul, et al.
Publicado: (2024)
Second Order Methods for Bandit Optimization and Control
por: Suggala, Arun, et al.
Publicado: (2024)
por: Suggala, Arun, et al.
Publicado: (2024)
Functional Interpolation for Relative Positions Improves Long Context Transformers
por: Li, Shanda, et al.
Publicado: (2023)
por: Li, Shanda, et al.
Publicado: (2023)
FG-Attn: Leveraging Fine-Grained Sparsity In Diffusion Transformers
por: Durvasula, Sankeerth, et al.
Publicado: (2025)
por: Durvasula, Sankeerth, et al.
Publicado: (2025)
How Far Can Disaggregation Go? A Design-Space Exploration of Attention-FFN Disaggregation for Efficient MoE LLM Serving
por: Wu, Hanjiang, et al.
Publicado: (2026)
por: Wu, Hanjiang, et al.
Publicado: (2026)
Inferring Asteroseismic Parameters from Short Observations Using Deep Learning: Application to TESS and K2 Red Giants
por: Ghanghas, Nipun, et al.
Publicado: (2026)
por: Ghanghas, Nipun, et al.
Publicado: (2026)
Sparsity Moves Computation: How FFN Architecture Reshapes Attention in Small Transformers
por: Smithline, Gabriel, et al.
Publicado: (2026)
por: Smithline, Gabriel, et al.
Publicado: (2026)
Learning Fine-grained Parameter Sharing via Sparse Tensor Decomposition
por: Üyük, Cem, et al.
Publicado: (2024)
por: Üyük, Cem, et al.
Publicado: (2024)
Theoretical modelling of the translation process
por: Jeremiah Felix Nwachukwu
Publicado: (2024)
por: Jeremiah Felix Nwachukwu
Publicado: (2024)
Effective Interplay between Sparsity and Quantization: From Theory to Practice
por: Harma, Simla Burcu, et al.
Publicado: (2024)
por: Harma, Simla Burcu, et al.
Publicado: (2024)
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity
por: Song, Chenyang, et al.
Publicado: (2025)
por: Song, Chenyang, et al.
Publicado: (2025)
AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding
por: Li, Handong, et al.
Publicado: (2026)
por: Li, Handong, et al.
Publicado: (2026)
Sparsity-Driven Parallel Imaging Consistency for Improved Self-Supervised MRI Reconstruction
por: Alçalar, Yaşar Utku, et al.
Publicado: (2025)
por: Alçalar, Yaşar Utku, et al.
Publicado: (2025)
Inertial migration of slender prolate and thin oblate spheroids in plane Poiseuille flow
por: Anand, Prateek, et al.
Publicado: (2025)
por: Anand, Prateek, et al.
Publicado: (2025)
Explainable Artificial Intelligence Credit Risk Assessment using Machine Learning
por: Shreya, et al.
Publicado: (2025)
por: Shreya, et al.
Publicado: (2025)
Scalable Machine Learning Training Infrastructure for Online Ads Recommendation and Auction Scoring Modeling at Google
por: Kurian, George, et al.
Publicado: (2025)
por: Kurian, George, et al.
Publicado: (2025)
Towards Optimal Adapter Placement for Efficient Transfer Learning
por: Nowak, Aleksandra I., et al.
Publicado: (2024)
por: Nowak, Aleksandra I., et al.
Publicado: (2024)
RAGO: Systematic Performance Optimization for Retrieval-Augmented Generation Serving
por: Jiang, Wenqi, et al.
Publicado: (2025)
por: Jiang, Wenqi, et al.
Publicado: (2025)
Compressing Many-Shots in In-Context Learning
por: Khatri, Devvrit, et al.
Publicado: (2025)
por: Khatri, Devvrit, et al.
Publicado: (2025)
Enabling Dynamic Sparsity in Quantized LLM Inference
por: Wang, Rongxiang, et al.
Publicado: (2025)
por: Wang, Rongxiang, et al.
Publicado: (2025)
vmintf/WsFFN: WsFFN Pre Alpha v0.0.2a-e Experimental Version Release
por: 민성 Skystarry
Publicado: (2025)
por: 민성 Skystarry
Publicado: (2025)
Ejemplares similares
-
HiRE: High Recall Approximate Top-$k$ Estimation for Efficient LLM Inference
por: L, Yashas Samaga B, et al.
Publicado: (2024) -
Progressive Gradient Flow for Robust N:M Sparsity Training in Transformers
por: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Publicado: (2024) -
Scalable In-context Ranking with Generative Models
por: Gupta, Nilesh, et al.
Publicado: (2025) -
On student-teacher deviations in distillation: does it pay to disobey?
por: Nagarajan, Vaishnavh, et al.
Publicado: (2023) -
Mimetic Initialization Helps State Space Models Learn to Recall
por: Trockman, Asher, et al.
Publicado: (2024)