ELAS: Efficient Pre-Training of Low-Rank Large Language Models via 2:4 Activation Sparsity
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Jiaxi, Yin, Lu, Shen, Li, Xu, Jinjin, Liu, Yuhui, Wang, Wenwu, Liu, Shiwei, Wang, Xilu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LOST: Low-rank and Sparse Pre-training for Large Language Models
von: Li, Jiaxi, et al.
Veröffentlicht: (2025)
von: Li, Jiaxi, et al.
Veröffentlicht: (2025)
OWLed: Outlier-weighed Layerwise Pruning for Efficient Autonomous Driving Framework
von: Li, Jiaxi, et al.
Veröffentlicht: (2024)
von: Li, Jiaxi, et al.
Veröffentlicht: (2024)
ActTail: Global Activation Sparsity in Large Language Models
von: Hou, Wenwen, et al.
Veröffentlicht: (2026)
von: Hou, Wenwen, et al.
Veröffentlicht: (2026)
CoLA: Compute-Efficient Pre-Training of LLMs via Low-Rank Activation
von: Liu, Ziyue, et al.
Veröffentlicht: (2025)
von: Liu, Ziyue, et al.
Veröffentlicht: (2025)
R-Sparse: Rank-Aware Activation Sparsity for Efficient LLM Inference
von: Zhang, Zhenyu, et al.
Veröffentlicht: (2025)
von: Zhang, Zhenyu, et al.
Veröffentlicht: (2025)
Full-Rank No More: Low-Rank Weight Training for Modern Speech Recognition Models
von: Fernandez-Lopez, Adriana, et al.
Veröffentlicht: (2024)
von: Fernandez-Lopez, Adriana, et al.
Veröffentlicht: (2024)
SPP: Sparsity-Preserved Parameter-Efficient Fine-Tuning for Large Language Models
von: Lu, Xudong, et al.
Veröffentlicht: (2024)
von: Lu, Xudong, et al.
Veröffentlicht: (2024)
Training-Free Activation Sparsity in Large Language Models
von: Liu, James, et al.
Veröffentlicht: (2024)
von: Liu, James, et al.
Veröffentlicht: (2024)
Sparsity-Accelerated Training for Large Language Models
von: Ma, Da, et al.
Veröffentlicht: (2024)
von: Ma, Da, et al.
Veröffentlicht: (2024)
Sparsity-Aware Low-Rank Representation for Efficient Fine-Tuning of Large Language Models
von: Zhang, Longteng, et al.
Veröffentlicht: (2026)
von: Zhang, Longteng, et al.
Veröffentlicht: (2026)
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN
von: Li, Pengxiang, et al.
Veröffentlicht: (2024)
von: Li, Pengxiang, et al.
Veröffentlicht: (2024)
Comet: Accelerating Private Inference for Large Language Model by Predicting Activation Sparsity
von: Yan, Guang, et al.
Veröffentlicht: (2025)
von: Yan, Guang, et al.
Veröffentlicht: (2025)
From Low Rank Gradient Subspace Stabilization to Low-Rank Weights: Observations, Theories, and Applications
von: Jaiswal, Ajay, et al.
Veröffentlicht: (2024)
von: Jaiswal, Ajay, et al.
Veröffentlicht: (2024)
SoLA: Leveraging Soft Activation Sparsity and Low-Rank Decomposition for Large Language Model Compression
von: Huang, Xinhao, et al.
Veröffentlicht: (2026)
von: Huang, Xinhao, et al.
Veröffentlicht: (2026)
Beyond Higher Rank: Token-wise Input-Output Projections for Efficient Low-Rank Adaptation
von: Li, Shiwei, et al.
Veröffentlicht: (2025)
von: Li, Shiwei, et al.
Veröffentlicht: (2025)
CLIP Brings Better Features to Visual Aesthetics Learners
von: Xu, Liwu, et al.
Veröffentlicht: (2023)
von: Xu, Liwu, et al.
Veröffentlicht: (2023)
TSegAgent: Zero-Shot Tooth Segmentation via Geometry-Aware Vision-Language Agents
von: Zhuang, Shaojie, et al.
Veröffentlicht: (2026)
von: Zhuang, Shaojie, et al.
Veröffentlicht: (2026)
Junk DNA Hypothesis: Pruning Small Pre-Trained Weights Irreversibly and Monotonically Impairs "Difficult" Downstream Tasks in LLMs
von: Yin, Lu, et al.
Veröffentlicht: (2023)
von: Yin, Lu, et al.
Veröffentlicht: (2023)
Exploring the Benefit of Activation Sparsity in Pre-training
von: Zhang, Zhengyan, et al.
Veröffentlicht: (2024)
von: Zhang, Zhengyan, et al.
Veröffentlicht: (2024)
Q-GaLore: Quantized GaLore with INT4 Projection and Layer-Adaptive Low-Rank Gradients
von: Zhang, Zhenyu, et al.
Veröffentlicht: (2024)
von: Zhang, Zhenyu, et al.
Veröffentlicht: (2024)
GPTailor: Large Language Model Pruning Through Layer Cutting and Stitching
von: Su, Guinan, et al.
Veröffentlicht: (2025)
von: Su, Guinan, et al.
Veröffentlicht: (2025)
Amber Pruner: Leveraging N:M Activation Sparsity for Efficient Prefill in Large Language Models
von: An, Tai, et al.
Veröffentlicht: (2025)
von: An, Tai, et al.
Veröffentlicht: (2025)
Astra: Activation-Space Tail-Eigenvector Low-Rank Adaptation of Large Language Models
von: Liu, Kainan, et al.
Veröffentlicht: (2026)
von: Liu, Kainan, et al.
Veröffentlicht: (2026)
The Panaceas for Improving Low-Rank Decomposition in Communication-Efficient Federated Learning
von: Li, Shiwei, et al.
Veröffentlicht: (2025)
von: Li, Shiwei, et al.
Veröffentlicht: (2025)
Data Efficient Adaptation in Large Language Models via Continuous Low-Rank Fine-Tuning
von: Han, Xiao, et al.
Veröffentlicht: (2025)
von: Han, Xiao, et al.
Veröffentlicht: (2025)
ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language Models
von: Song, Chenyang, et al.
Veröffentlicht: (2024)
von: Song, Chenyang, et al.
Veröffentlicht: (2024)
InfoMerge: Information-aware Token Compression for Efficient Video Large Language Models
von: Liu, Xinxin, et al.
Veröffentlicht: (2026)
von: Liu, Xinxin, et al.
Veröffentlicht: (2026)
Learning Dynamics in Continual Pre-Training for Large Language Models
von: Wang, Xingjin, et al.
Veröffentlicht: (2025)
von: Wang, Xingjin, et al.
Veröffentlicht: (2025)
Joint Training Across Multiple Activation Sparsity Regimes
von: Wang, Haotian
Veröffentlicht: (2026)
von: Wang, Haotian
Veröffentlicht: (2026)
Collaborative Low-Rank Adaptation for Pre-Trained Vision Transformers
von: Liu, Zheng, et al.
Veröffentlicht: (2025)
von: Liu, Zheng, et al.
Veröffentlicht: (2025)
To 2:4 Sparsity and Beyond: Neuron-level Activation Function to Accelerate LLM Pre-Training
von: Madhyastha, Meghana, et al.
Veröffentlicht: (2026)
von: Madhyastha, Meghana, et al.
Veröffentlicht: (2026)
Sparsity Induction for Accurate Post-Training Pruning of Large Language Models
von: Jiang, Minhao, et al.
Veröffentlicht: (2026)
von: Jiang, Minhao, et al.
Veröffentlicht: (2026)
SymbolicLight V1: Spike-Gated Dual-Path Language Modeling with High Activation Sparsity and Sub-Billion-Scale Pre-Training Evidence
von: Liu, Ting
Veröffentlicht: (2026)
von: Liu, Ting
Veröffentlicht: (2026)
WiSparse: Boosting LLM Inference Efficiency with Weight-Aware Mixed Activation Sparsity
von: Chen, Lei, et al.
Veröffentlicht: (2026)
von: Chen, Lei, et al.
Veröffentlicht: (2026)
DL-QAT: Weight-Decomposed Low-Rank Quantization-Aware Training for Large Language Models
von: Ke, Wenjin, et al.
Veröffentlicht: (2025)
von: Ke, Wenjin, et al.
Veröffentlicht: (2025)
Pivoting Factorization: A Compact Meta Low-Rank Representation of Sparsity for Efficient Inference in Large Language Models
von: Zhao, Jialin, et al.
Veröffentlicht: (2025)
von: Zhao, Jialin, et al.
Veröffentlicht: (2025)
The Curse of Depth in Large Language Models
von: Sun, Wenfang, et al.
Veröffentlicht: (2025)
von: Sun, Wenfang, et al.
Veröffentlicht: (2025)
Sparsing Law: Towards Large Language Models with Greater Activation Sparsity
von: Luo, Yuqi, et al.
Veröffentlicht: (2024)
von: Luo, Yuqi, et al.
Veröffentlicht: (2024)
DuoGPT: Training-free Dual Sparsity through Activation-aware Pruning in LLMs
von: Yin, Ruokai, et al.
Veröffentlicht: (2025)
von: Yin, Ruokai, et al.
Veröffentlicht: (2025)
Training-Free Bayesianization for Low-Rank Adapters of Large Language Models
von: Shi, Haizhou, et al.
Veröffentlicht: (2024)
von: Shi, Haizhou, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
LOST: Low-rank and Sparse Pre-training for Large Language Models
von: Li, Jiaxi, et al.
Veröffentlicht: (2025) -
OWLed: Outlier-weighed Layerwise Pruning for Efficient Autonomous Driving Framework
von: Li, Jiaxi, et al.
Veröffentlicht: (2024) -
ActTail: Global Activation Sparsity in Large Language Models
von: Hou, Wenwen, et al.
Veröffentlicht: (2026) -
CoLA: Compute-Efficient Pre-Training of LLMs via Low-Rank Activation
von: Liu, Ziyue, et al.
Veröffentlicht: (2025) -
R-Sparse: Rank-Aware Activation Sparsity for Efficient LLM Inference
von: Zhang, Zhenyu, et al.
Veröffentlicht: (2025)