Improving MLLM Training Efficiency via Stage-Aware Sparsity
Fuente:
arXiv
Saved in:
| Main Authors: | Shi, Kean, Chen, Liang, Zhao, Haozhe, Chang, Baobao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
WiSparse: Boosting LLM Inference Efficiency with Weight-Aware Mixed Activation Sparsity
by: Chen, Lei, et al.
Published: (2026)
by: Chen, Lei, et al.
Published: (2026)
Improving MoE Compute Efficiency by Composing Weight and Data Sparsity
by: Kilian, Maciej, et al.
Published: (2026)
by: Kilian, Maciej, et al.
Published: (2026)
Improving Decision Sparsity
by: Sun, Yiyang, et al.
Published: (2024)
by: Sun, Yiyang, et al.
Published: (2024)
ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching
by: Zhao, Youpeng, et al.
Published: (2024)
by: Zhao, Youpeng, et al.
Published: (2024)
Sparsity-Aware Evolution for Model Merging
by: Zhang, Huan, et al.
Published: (2026)
by: Zhang, Huan, et al.
Published: (2026)
ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training
by: Liang, Yu, et al.
Published: (2026)
by: Liang, Yu, et al.
Published: (2026)
Maximum Redundancy Pruning: A Principle-Driven Layerwise Sparsity Allocation for LLMs
by: Gao, Chang, et al.
Published: (2025)
by: Gao, Chang, et al.
Published: (2025)
RaaS: Reasoning-Aware Attention Sparsity for Efficient LLM Reasoning
by: Hu, Junhao, et al.
Published: (2025)
by: Hu, Junhao, et al.
Published: (2025)
Scaling Attention via Feature Sparsity
by: Xie, Yan, et al.
Published: (2026)
by: Xie, Yan, et al.
Published: (2026)
When Does Value-Aware KV Eviction Help? A Fixed-Contract Diagnostic for Non-Monotone Cache Compression
by: Zhang, Ruijie, et al.
Published: (2026)
by: Zhang, Ruijie, et al.
Published: (2026)
On the Interplay Between Sparsity and Training in Deep Reinforcement Learning
by: Davelouis, Fatima, et al.
Published: (2025)
by: Davelouis, Fatima, et al.
Published: (2025)
An Efficient Training Algorithm for Models with Block-wise Sparsity
by: Zhu, Ding, et al.
Published: (2025)
by: Zhu, Ding, et al.
Published: (2025)
Joint Training Across Multiple Activation Sparsity Regimes
by: Wang, Haotian
Published: (2026)
by: Wang, Haotian
Published: (2026)
Post-Training Statistical Calibration for Higher Activation Sparsity
by: Chua, Vui Seng, et al.
Published: (2024)
by: Chua, Vui Seng, et al.
Published: (2024)
DeepSpeed Data Efficiency: Improving Deep Learning Model Quality and Training Efficiency via Efficient Data Sampling and Routing
by: Li, Conglong, et al.
Published: (2022)
by: Li, Conglong, et al.
Published: (2022)
Introspective X Training: Feedback Conditioning Improves Scaling Across all LLM Training Stages
by: Cui, Brandon, et al.
Published: (2026)
by: Cui, Brandon, et al.
Published: (2026)
The Wolf Within: Covert Injection of Malice into MLLM Societies via an MLLM Operative
by: Tan, Zhen, et al.
Published: (2024)
by: Tan, Zhen, et al.
Published: (2024)
Sparsity-Aware Low-Rank Representation for Efficient Fine-Tuning of Large Language Models
by: Zhang, Longteng, et al.
Published: (2026)
by: Zhang, Longteng, et al.
Published: (2026)
LOTUS: Improving Transformer Efficiency with Sparsity Pruning and Data Lottery Tickets
by: Upadhyay, Ojasw
Published: (2024)
by: Upadhyay, Ojasw
Published: (2024)
PLUM: Improving Inference Efficiency By Leveraging Repetition-Sparsity Trade-Off
by: Kuhar, Sachit, et al.
Published: (2023)
by: Kuhar, Sachit, et al.
Published: (2023)
Android Coach: Improve Online Agentic Training Efficiency with Single State Multiple Actions
by: Gan, Guo, et al.
Published: (2026)
by: Gan, Guo, et al.
Published: (2026)
StableQAT: Stable Quantization-Aware Training at Ultra-Low Bitwidths
by: Chen, Tianyi, et al.
Published: (2026)
by: Chen, Tianyi, et al.
Published: (2026)
Efficient LLM Reasoning via Variational Posterior Guidance with Efficiency Awareness
by: Chen, Zizhao, et al.
Published: (2026)
by: Chen, Zizhao, et al.
Published: (2026)
EcoSpa: Efficient Transformer Training with Coupled Sparsity
by: Xiao, Jinqi, et al.
Published: (2025)
by: Xiao, Jinqi, et al.
Published: (2025)
Activation Sparsity Opportunities for Compressing General Large Language Models
by: Dhar, Nobel, et al.
Published: (2024)
by: Dhar, Nobel, et al.
Published: (2024)
Improving the Robustness of Distantly-Supervised Named Entity Recognition via Uncertainty-Aware Teacher Learning and Student-Student Collaborative Learning
by: Si, Shuzheng, et al.
Published: (2023)
by: Si, Shuzheng, et al.
Published: (2023)
ELAS: Efficient Pre-Training of Low-Rank Large Language Models via 2:4 Activation Sparsity
by: Li, Jiaxi, et al.
Published: (2026)
by: Li, Jiaxi, et al.
Published: (2026)
Weight Concentration Regularization for Improving Pruning Robustness Under High Sparsity
by: Yun, Vincent-Daniel, et al.
Published: (2025)
by: Yun, Vincent-Daniel, et al.
Published: (2025)
Concept Influence: Leveraging Interpretability to Improve Performance and Efficiency in Training Data Attribution
by: Kowal, Matthew, et al.
Published: (2026)
by: Kowal, Matthew, et al.
Published: (2026)
AMSbench: A Comprehensive Benchmark for Evaluating MLLM Capabilities in AMS Circuits
by: Shi, Yichen, et al.
Published: (2025)
by: Shi, Yichen, et al.
Published: (2025)
Memory-Efficient LLM Training with Dynamic Sparsity: From Stability to Practical Scaling
by: Xiao, Qiao, et al.
Published: (2026)
by: Xiao, Qiao, et al.
Published: (2026)
Post-Training Sparse Attention with Double Sparsity
by: Yang, Shuo, et al.
Published: (2024)
by: Yang, Shuo, et al.
Published: (2024)
FDRMFL:Multi-modal Federated Feature Extraction Model Based on Information Maximization and Contrastive Learning
by: Wu, Haozhe
Published: (2025)
by: Wu, Haozhe
Published: (2025)
ContextRL: Enhancing MLLM's Knowledge Discovery Efficiency with Context-Augmented RL
by: Lu, Xingyu, et al.
Published: (2026)
by: Lu, Xingyu, et al.
Published: (2026)
ES-Merging: Biological MLLM Merging via Embedding Space Signals
by: Lee, Wonbin, et al.
Published: (2026)
by: Lee, Wonbin, et al.
Published: (2026)
Sparse-Reg: Improving Sample Complexity in Offline Reinforcement Learning using Sparsity
by: Arnob, Samin Yeasar, et al.
Published: (2025)
by: Arnob, Samin Yeasar, et al.
Published: (2025)
Strategies for Improving Communication Efficiency in Distributed and Federated Learning: Compression, Local Training, and Personalization
by: Yi, Kai
Published: (2025)
by: Yi, Kai
Published: (2025)
Towards the Connection between Activation Sparsity and Flat Minima
by: Peng, Ze, et al.
Published: (2026)
by: Peng, Ze, et al.
Published: (2026)
MLLM Is a Strong Reranker: Advancing Multimodal Retrieval-augmented Generation via Knowledge-enhanced Reranking and Noise-injected Training
by: Chen, Zhanpeng, et al.
Published: (2024)
by: Chen, Zhanpeng, et al.
Published: (2024)
Train One Sparse Autoencoder Across Multiple Sparsity Budgets to Preserve Interpretability and Accuracy
by: Balagansky, Nikita, et al.
Published: (2025)
by: Balagansky, Nikita, et al.
Published: (2025)
Similar Items
-
WiSparse: Boosting LLM Inference Efficiency with Weight-Aware Mixed Activation Sparsity
by: Chen, Lei, et al.
Published: (2026) -
Improving MoE Compute Efficiency by Composing Weight and Data Sparsity
by: Kilian, Maciej, et al.
Published: (2026) -
Improving Decision Sparsity
by: Sun, Yiyang, et al.
Published: (2024) -
ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching
by: Zhao, Youpeng, et al.
Published: (2024) -
Sparsity-Aware Evolution for Model Merging
by: Zhang, Huan, et al.
Published: (2026)