Leave it to the Specialist: Repair Sparse LLMs with Sparse Fine-Tuning via Sparsity Evolution
Fuente:
arXiv
Saved in:
| Main Authors: | Xiao, Qiao, Ansell, Alan, Wu, Boqian, Yin, Lu, Pechenizkiy, Mykola, Liu, Shiwei, Mocanu, Decebal Constantin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
E2ENet: Dynamic Sparse Feature Fusion for Accurate and Efficient 3D Medical Image Segmentation
by: Wu, Boqian, et al.
Published: (2023)
by: Wu, Boqian, et al.
Published: (2023)
Are Sparse Neural Networks Better Hard Sample Learners?
by: Xiao, Qiao, et al.
Published: (2024)
by: Xiao, Qiao, et al.
Published: (2024)
Addressing the Collaboration Dilemma in Low-Data Federated Learning via Transient Sparsity
by: Xiao, Qiao, et al.
Published: (2025)
by: Xiao, Qiao, et al.
Published: (2025)
Dynamic Sparse Training versus Dense Training: The Unexpected Winner in Image Corruption Robustness
by: Wu, Boqian, et al.
Published: (2024)
by: Wu, Boqian, et al.
Published: (2024)
NeuroTrails: Training with Dynamic Sparse Heads as the Key to Effective Ensembling
by: Grooten, Bram, et al.
Published: (2025)
by: Grooten, Bram, et al.
Published: (2025)
When Data Is Scarce: Scaling Sparse Language Models with Repeated Training
by: Wu, Boqian, et al.
Published: (2026)
by: Wu, Boqian, et al.
Published: (2026)
Adaptive Sparsity Level during Training for Efficient Time Series Forecasting with Transformers
by: Atashgahi, Zahra, et al.
Published: (2023)
by: Atashgahi, Zahra, et al.
Published: (2023)
Memory-Efficient LLM Training with Dynamic Sparsity: From Stability to Practical Scaling
by: Xiao, Qiao, et al.
Published: (2026)
by: Xiao, Qiao, et al.
Published: (2026)
Dynamic Data Pruning for Automatic Speech Recognition
by: Xiao, Qiao, et al.
Published: (2024)
by: Xiao, Qiao, et al.
Published: (2024)
Unveiling the Power of Sparse Neural Networks for Feature Selection
by: Atashgahi, Zahra, et al.
Published: (2024)
by: Atashgahi, Zahra, et al.
Published: (2024)
Deep Ensembling with No Overhead for either Training or Testing: The All-Round Blessings of Dynamic Sparsity
by: Liu, Shiwei, et al.
Published: (2021)
by: Liu, Shiwei, et al.
Published: (2021)
Sparse-to-Sparse Training of Diffusion Models
by: Oliveira, Inês Cardoso, et al.
Published: (2025)
by: Oliveira, Inês Cardoso, et al.
Published: (2025)
Boosting Robustness in Preference-Based Reinforcement Learning with Dynamic Sparsity
by: Muslimani, Calarina, et al.
Published: (2024)
by: Muslimani, Calarina, et al.
Published: (2024)
You Can Have Better Graph Neural Networks by Not Training Weights at All: Finding Untrained GNNs Tickets
by: Huang, Tianjin, et al.
Published: (2022)
by: Huang, Tianjin, et al.
Published: (2022)
Scaling Sparse Fine-Tuning to Large Language Models
by: Ansell, Alan, et al.
Published: (2024)
by: Ansell, Alan, et al.
Published: (2024)
Nerva: a Truly Sparse Implementation of Neural Networks
by: Wesselink, Wieger, et al.
Published: (2024)
by: Wesselink, Wieger, et al.
Published: (2024)
Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity
by: Yin, Lu, et al.
Published: (2023)
by: Yin, Lu, et al.
Published: (2023)
Batch Matrix-form Equations and Implementation of Multilayer Perceptrons
by: Wesselink, Wieger, et al.
Published: (2025)
by: Wesselink, Wieger, et al.
Published: (2025)
SparseLoRA: Accelerating LLM Fine-Tuning with Contextual Sparsity
by: Khaki, Samir, et al.
Published: (2025)
by: Khaki, Samir, et al.
Published: (2025)
Dynamic Sparse No Training: Training-Free Fine-tuning for Sparse LLMs
by: Zhang, Yuxin, et al.
Published: (2023)
by: Zhang, Yuxin, et al.
Published: (2023)
(PASS) Visual Prompt Locates Good Structure Sparsity through a Recurrent HyperNetwork
by: Huang, Tianjin, et al.
Published: (2024)
by: Huang, Tianjin, et al.
Published: (2024)
EBFT: Effective and Block-Wise Fine-Tuning for Sparse LLMs
by: Guo, Song, et al.
Published: (2024)
by: Guo, Song, et al.
Published: (2024)
Self-Regulated Neurogenesis for Online Data-Incremental Learning
by: Yildirim, Murat Onur, et al.
Published: (2024)
by: Yildirim, Murat Onur, et al.
Published: (2024)
Enhancing Adversarial Training via Reweighting Optimization Trajectory
by: Huang, Tianjin, et al.
Published: (2023)
by: Huang, Tianjin, et al.
Published: (2023)
Robust Active Learning (RoAL): Countering Dynamic Adversaries in Active Learning with Elastic Weight Consolidation
by: Fajri, Ricky Maulana, et al.
Published: (2024)
by: Fajri, Ricky Maulana, et al.
Published: (2024)
Zeroth-Order Fine-Tuning of LLMs with Extreme Sparsity
by: Guo, Wentao, et al.
Published: (2024)
by: Guo, Wentao, et al.
Published: (2024)
LiMTR: Time Series Motion Prediction for Diverse Road Users through Multimodal Feature Integration
by: Oerlemans, Camiel, et al.
Published: (2024)
by: Oerlemans, Camiel, et al.
Published: (2024)
MSRS: Training Multimodal Speech Recognition Models from Scratch with Sparse Mask Optimization
by: Fernandez-Lopez, Adriana, et al.
Published: (2024)
by: Fernandez-Lopez, Adriana, et al.
Published: (2024)
SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse-Linear Attention
by: Zhang, Jintao, et al.
Published: (2025)
by: Zhang, Jintao, et al.
Published: (2025)
Reinforcement Learning With Sparse-Executing Actions via Sparsity Regularization
by: Pang, Jing-Cheng, et al.
Published: (2021)
by: Pang, Jing-Cheng, et al.
Published: (2021)
ProxSparse: Regularized Learning of Semi-Structured Sparsity Masks for Pretrained LLMs
by: Liu, Hongyi, et al.
Published: (2025)
by: Liu, Hongyi, et al.
Published: (2025)
SparseST: Exploiting Data Sparsity in Spatiotemporal Modeling and Prediction
by: Wu, Junfeng, et al.
Published: (2025)
by: Wu, Junfeng, et al.
Published: (2025)
HASARD: A Benchmark for Vision-Based Safe Reinforcement Learning in Embodied Agents
by: Tomilin, Tristan, et al.
Published: (2025)
by: Tomilin, Tristan, et al.
Published: (2025)
Everyone deserves their voice to be heard: Analyzing Predictive Gender Bias in ASR Models Applied to Dutch Speech Data
by: Raes, Rik, et al.
Published: (2024)
by: Raes, Rik, et al.
Published: (2024)
Efficient Exploration in Average-Reward Constrained Reinforcement Learning: Achieving Near-Optimal Regret With Posterior Sampling
by: Provodin, Danil, et al.
Published: (2024)
by: Provodin, Danil, et al.
Published: (2024)
FairSNA: Algorithmic Fairness in Social Network Analysis
by: Saxena, Akrati, et al.
Published: (2022)
by: Saxena, Akrati, et al.
Published: (2022)
Are There Exceptions to Goodhart's Law? On the Moral Justification of Fairness-Aware Machine Learning
by: Weerts, Hilde, et al.
Published: (2022)
by: Weerts, Hilde, et al.
Published: (2022)
Sparsity via Sparse Group $k$-max Regularization
by: Tao, Qinghua, et al.
Published: (2024)
by: Tao, Qinghua, et al.
Published: (2024)
Sparse Fine-Tuning of Transformers for Generative Tasks
by: Chen, Wei, et al.
Published: (2025)
by: Chen, Wei, et al.
Published: (2025)
Visual Prompting Upgrades Neural Network Sparsification: A Data-Model Perspective
by: Jin, Can, et al.
Published: (2023)
by: Jin, Can, et al.
Published: (2023)
Similar Items
-
E2ENet: Dynamic Sparse Feature Fusion for Accurate and Efficient 3D Medical Image Segmentation
by: Wu, Boqian, et al.
Published: (2023) -
Are Sparse Neural Networks Better Hard Sample Learners?
by: Xiao, Qiao, et al.
Published: (2024) -
Addressing the Collaboration Dilemma in Low-Data Federated Learning via Transient Sparsity
by: Xiao, Qiao, et al.
Published: (2025) -
Dynamic Sparse Training versus Dense Training: The Unexpected Winner in Image Corruption Robustness
by: Wu, Boqian, et al.
Published: (2024) -
NeuroTrails: Training with Dynamic Sparse Heads as the Key to Effective Ensembling
by: Grooten, Bram, et al.
Published: (2025)