Toward Efficient Permutation for Hierarchical N:M Sparsity on GPUs
Fuente:
arXiv
Salvato in:
| Autori principali: | Yu, Seungmin, Yi, Xiaodie, Lee, Hayun, Shin, Dongkun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Realizing Unaligned Block-wise Pruning for DNN Acceleration on Mobile Devices
di: Lee, Hayun, et al.
Pubblicazione: (2024)
di: Lee, Hayun, et al.
Pubblicazione: (2024)
PermLLM: Learnable Channel Permutation for N:M Sparse Large Language Models
di: Zou, Lancheng, et al.
Pubblicazione: (2025)
di: Zou, Lancheng, et al.
Pubblicazione: (2025)
Minimum Variance Unbiased N:M Sparsity for the Neural Gradients
di: Chmiel, Brian, et al.
Pubblicazione: (2022)
di: Chmiel, Brian, et al.
Pubblicazione: (2022)
Amber Pruner: Leveraging N:M Activation Sparsity for Efficient Prefill in Large Language Models
di: An, Tai, et al.
Pubblicazione: (2025)
di: An, Tai, et al.
Pubblicazione: (2025)
E-Sparse: Boosting the Large Language Model Inference through Entropy-based N:M Sparsity
di: Li, Yun, et al.
Pubblicazione: (2023)
di: Li, Yun, et al.
Pubblicazione: (2023)
Motivating Next-Gen Accelerators with Flexible (N:M) Activation Sparsity via Benchmarking Lightweight Post-Training Sparsification Approaches
di: Alanova, Shirin, et al.
Pubblicazione: (2025)
di: Alanova, Shirin, et al.
Pubblicazione: (2025)
TSENOR: Highly-Efficient Algorithm for Finding Transposable N:M Sparse Masks
di: Meng, Xiang, et al.
Pubblicazione: (2025)
di: Meng, Xiang, et al.
Pubblicazione: (2025)
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator
di: Ramachandran, Akshat, et al.
Pubblicazione: (2025)
di: Ramachandran, Akshat, et al.
Pubblicazione: (2025)
Beyond 2:4: exploring V:N:M sparsity for efficient transformer inference on GPUs
di: Zhao, Kang, et al.
Pubblicazione: (2024)
di: Zhao, Kang, et al.
Pubblicazione: (2024)
DistrAttention: An Efficient and Flexible Self-Attention Mechanism on Modern GPUs
di: Jin, Haolin, et al.
Pubblicazione: (2025)
di: Jin, Haolin, et al.
Pubblicazione: (2025)
Memory-Efficient Acceleration of Block Low-Rank Foundation Models on Resource Constrained GPUs
di: Abillama, Pierre, et al.
Pubblicazione: (2025)
di: Abillama, Pierre, et al.
Pubblicazione: (2025)
Sparsity and Out-of-Distribution Generalization
di: Aaronson, Scott, et al.
Pubblicazione: (2026)
di: Aaronson, Scott, et al.
Pubblicazione: (2026)
Towards the Connection between Activation Sparsity and Flat Minima
di: Peng, Ze, et al.
Pubblicazione: (2026)
di: Peng, Ze, et al.
Pubblicazione: (2026)
An Efficient Training Algorithm for Models with Block-wise Sparsity
di: Zhu, Ding, et al.
Pubblicazione: (2025)
di: Zhu, Ding, et al.
Pubblicazione: (2025)
The Unseen Frontier: Pushing the Limits of LLM Sparsity with Surrogate-Free ADMM
di: Lee, Kwanhee, et al.
Pubblicazione: (2025)
di: Lee, Kwanhee, et al.
Pubblicazione: (2025)
RaaS: Reasoning-Aware Attention Sparsity for Efficient LLM Reasoning
di: Hu, Junhao, et al.
Pubblicazione: (2025)
di: Hu, Junhao, et al.
Pubblicazione: (2025)
MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs
di: Sun, Yan, et al.
Pubblicazione: (2025)
di: Sun, Yan, et al.
Pubblicazione: (2025)
From Patterns to Predictions: A Shapelet-Based Framework for Directional Forecasting in Noisy Financial Markets
di: Kim, Juwon, et al.
Pubblicazione: (2025)
di: Kim, Juwon, et al.
Pubblicazione: (2025)
EcoSpa: Efficient Transformer Training with Coupled Sparsity
di: Xiao, Jinqi, et al.
Pubblicazione: (2025)
di: Xiao, Jinqi, et al.
Pubblicazione: (2025)
Progressive Gradient Flow for Robust N:M Sparsity Training in Transformers
di: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Pubblicazione: (2024)
di: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Pubblicazione: (2024)
Beyond One-Way Pruning: Bidirectional Pruning-Regrowth for Extreme Accuracy-Sparsity Tradeoff
di: Liu, Junchen, et al.
Pubblicazione: (2025)
di: Liu, Junchen, et al.
Pubblicazione: (2025)
S$^{2}$FT: Efficient, Scalable and Generalizable LLM Fine-tuning by Structured Sparsity
di: Yang, Xinyu, et al.
Pubblicazione: (2024)
di: Yang, Xinyu, et al.
Pubblicazione: (2024)
Uncovering Intra-expert Activation Sparsity for Efficient Mixture-of-Expert Model Execution
di: Park, Jongseok, et al.
Pubblicazione: (2026)
di: Park, Jongseok, et al.
Pubblicazione: (2026)
Memory-Efficient LLM Training with Dynamic Sparsity: From Stability to Practical Scaling
di: Xiao, Qiao, et al.
Pubblicazione: (2026)
di: Xiao, Qiao, et al.
Pubblicazione: (2026)
LiNR: Model Based Neural Retrieval on GPUs at LinkedIn
di: Borisyuk, Fedor, et al.
Pubblicazione: (2024)
di: Borisyuk, Fedor, et al.
Pubblicazione: (2024)
Improving Decision Sparsity
di: Sun, Yiyang, et al.
Pubblicazione: (2024)
di: Sun, Yiyang, et al.
Pubblicazione: (2024)
Homeostasis and Sparsity in Transformer
di: Kotyuzanskiy, Leonid, et al.
Pubblicazione: (2024)
di: Kotyuzanskiy, Leonid, et al.
Pubblicazione: (2024)
Weight Concentration Regularization for Improving Pruning Robustness Under High Sparsity
di: Yun, Vincent-Daniel, et al.
Pubblicazione: (2025)
di: Yun, Vincent-Daniel, et al.
Pubblicazione: (2025)
Reinforcement Learning With Sparse-Executing Actions via Sparsity Regularization
di: Pang, Jing-Cheng, et al.
Pubblicazione: (2021)
di: Pang, Jing-Cheng, et al.
Pubblicazione: (2021)
Polar Sparsity: High Throughput Batched LLM Inferencing with Scalable Contextual Sparsity
di: Shrestha, Susav, et al.
Pubblicazione: (2025)
di: Shrestha, Susav, et al.
Pubblicazione: (2025)
Fast Forward: Accelerating LLM Prefill with Predictive FFN Sparsity
di: Gautam, Aayush, et al.
Pubblicazione: (2026)
di: Gautam, Aayush, et al.
Pubblicazione: (2026)
Monte Carlo Permutation Search
di: Cazenave, Tristan
Pubblicazione: (2025)
di: Cazenave, Tristan
Pubblicazione: (2025)
Hierarchical Balance Packing: Towards Efficient Supervised Fine-tuning for Long-Context LLM
di: Yao, Yongqiang, et al.
Pubblicazione: (2025)
di: Yao, Yongqiang, et al.
Pubblicazione: (2025)
Sparsity-Aware Low-Rank Representation for Efficient Fine-Tuning of Large Language Models
di: Zhang, Longteng, et al.
Pubblicazione: (2026)
di: Zhang, Longteng, et al.
Pubblicazione: (2026)
Sparsity and Superposition in Mixture of Experts
di: Chaudhari, Marmik, et al.
Pubblicazione: (2025)
di: Chaudhari, Marmik, et al.
Pubblicazione: (2025)
Electromagnetic Simulations of Antennas on GPUs for Machine Learning Applications
di: Temiz, Murat, et al.
Pubblicazione: (2025)
di: Temiz, Murat, et al.
Pubblicazione: (2025)
Toward Data Efficient Model Merging between Different Datasets without Performance Degradation
di: Yamada, Masanori, et al.
Pubblicazione: (2023)
di: Yamada, Masanori, et al.
Pubblicazione: (2023)
mLoRA: Fine-Tuning LoRA Adapters via Highly-Efficient Pipeline Parallelism in Multiple GPUs
di: Ye, Zhengmao, et al.
Pubblicazione: (2023)
di: Ye, Zhengmao, et al.
Pubblicazione: (2023)
BESA: Pruning Large Language Models with Blockwise Parameter-Efficient Sparsity Allocation
di: Xu, Peng, et al.
Pubblicazione: (2024)
di: Xu, Peng, et al.
Pubblicazione: (2024)
Permutation Equivariant Model-based Offline Reinforcement Learning for Auto-bidding
di: Mou, Zhiyu, et al.
Pubblicazione: (2025)
di: Mou, Zhiyu, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Realizing Unaligned Block-wise Pruning for DNN Acceleration on Mobile Devices
di: Lee, Hayun, et al.
Pubblicazione: (2024) -
PermLLM: Learnable Channel Permutation for N:M Sparse Large Language Models
di: Zou, Lancheng, et al.
Pubblicazione: (2025) -
Minimum Variance Unbiased N:M Sparsity for the Neural Gradients
di: Chmiel, Brian, et al.
Pubblicazione: (2022) -
Amber Pruner: Leveraging N:M Activation Sparsity for Efficient Prefill in Large Language Models
di: An, Tai, et al.
Pubblicazione: (2025) -
E-Sparse: Boosting the Large Language Model Inference through Entropy-based N:M Sparsity
di: Li, Yun, et al.
Pubblicazione: (2023)