Data-Free Pruning of Self-Attention Layers in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Saikumar, Dhananjay, Varghese, Blesson |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Signal Collapse in One-Shot Pruning: When Sparse Models Fail to Distinguish Neural Representations
by: Saikumar, Dhananjay, et al.
Published: (2025)
by: Saikumar, Dhananjay, et al.
Published: (2025)
DRIVE: Dual Gradient-Based Rapid Iterative Pruning
by: Saikumar, Dhananjay, et al.
Published: (2024)
by: Saikumar, Dhananjay, et al.
Published: (2024)
NeuroFlux: Memory-Efficient CNN Training Using Adaptive Local Learning
by: Saikumar, Dhananjay, et al.
Published: (2024)
by: Saikumar, Dhananjay, et al.
Published: (2024)
Mosaic: Composite Projection Pruning for Resource-efficient LLMs
by: Eccles, Bailey J., et al.
Published: (2025)
by: Eccles, Bailey J., et al.
Published: (2025)
Rapid Deployment of DNNs for Edge Computing via Structured Pruning at Initialization
by: Eccles, Bailey J., et al.
Published: (2024)
by: Eccles, Bailey J., et al.
Published: (2024)
DNNShifter: An Efficient DNN Pruning System for Edge Computing
by: Eccles, Bailey J., et al.
Published: (2023)
by: Eccles, Bailey J., et al.
Published: (2023)
Ghosted Layers: Unconstrained Activation Alignment for Recovering Layer-Pruned LLMs
by: Yun, Vincent-Daniel, et al.
Published: (2026)
by: Yun, Vincent-Daniel, et al.
Published: (2026)
When Fewer Layers Break More Chains: Layer Pruning Harms Test-Time Scaling in LLMs
by: Wang, Keyu, et al.
Published: (2025)
by: Wang, Keyu, et al.
Published: (2025)
Post-Pruning Accuracy Recovery via Data-Free Knowledge Distillation
by: Tripurwar, Chinmay, et al.
Published: (2025)
by: Tripurwar, Chinmay, et al.
Published: (2025)
SwiftPrune: Hessian-Free Weight Pruning for Large Language Models
by: Kang, Yuhan, et al.
Published: (2025)
by: Kang, Yuhan, et al.
Published: (2025)
Layer Collapse Can be Induced by Unstructured Pruning
by: Liao, Zhu, et al.
Published: (2024)
by: Liao, Zhu, et al.
Published: (2024)
PrunePEFT: Iterative Hybrid Pruning for Parameter-Efficient Fine-tuning of LLMs
by: Yu, Tongzhou, et al.
Published: (2025)
by: Yu, Tongzhou, et al.
Published: (2025)
A Little Human Data Goes A Long Way
by: Ashok, Dhananjay, et al.
Published: (2024)
by: Ashok, Dhananjay, et al.
Published: (2024)
Dispatch-Aware Ragged Attention for Pruned Vision Transformers
by: Abdellatif, Seifeldin, et al.
Published: (2026)
by: Abdellatif, Seifeldin, et al.
Published: (2026)
Safe Pruning LoRA: Robust Distance-Guided Pruning for Safety Alignment in Adaptation of LLMs
by: Ao, Shuang, et al.
Published: (2025)
by: Ao, Shuang, et al.
Published: (2025)
CLOVER: Cross-Layer Orthogonal Vectors Pruning and Fine-Tuning
by: Meng, Fanxu, et al.
Published: (2024)
by: Meng, Fanxu, et al.
Published: (2024)
On the Limits of Layer Pruning for Generative Reasoning in Large Language Models
by: Shrestha, Safal, et al.
Published: (2026)
by: Shrestha, Safal, et al.
Published: (2026)
PERP: Rethinking the Prune-Retrain Paradigm in the Era of LLMs
by: Zimmer, Max, et al.
Published: (2023)
by: Zimmer, Max, et al.
Published: (2023)
Multi-Layer Attention-Based Explainability via Transformers for Tabular Data
by: Gavito, Andrea Treviño, et al.
Published: (2023)
by: Gavito, Andrea Treviño, et al.
Published: (2023)
The Structural Scalpel: Automated Contiguous Layer Pruning for Large Language Models
by: Lu, Yao, et al.
Published: (2025)
by: Lu, Yao, et al.
Published: (2025)
TopoPrune: Robust Data Pruning via Unified Latent Space Topology
by: Roy, Arjun, et al.
Published: (2026)
by: Roy, Arjun, et al.
Published: (2026)
Mixture of Layers with Hybrid Attention
by: Ternovtsii, Ivan, et al.
Published: (2026)
by: Ternovtsii, Ivan, et al.
Published: (2026)
Hierarchical Attention-based Graph Neural Network with Relevance-driven Pruning
by: Kum, Seungwoo
Published: (2026)
by: Kum, Seungwoo
Published: (2026)
Severing Spurious Correlations with Data Pruning
by: Mulchandani, Varun, et al.
Published: (2025)
by: Mulchandani, Varun, et al.
Published: (2025)
EMP: Enhance Memory in Data Pruning
by: Xiao, Jinying, et al.
Published: (2024)
by: Xiao, Jinying, et al.
Published: (2024)
MaskPrune: Mask-based LLM Pruning for Layer-wise Uniform Structures
by: Qin, Jiayu, et al.
Published: (2025)
by: Qin, Jiayu, et al.
Published: (2025)
Adaptive Layer Selection for Layer-Wise Token Pruning in LLM Inference
by: Taniguchi, Rei, et al.
Published: (2026)
by: Taniguchi, Rei, et al.
Published: (2026)
A Free Lunch in LLM Compression: Revisiting Retraining after Pruning
by: Wagner, Moritz, et al.
Published: (2025)
by: Wagner, Moritz, et al.
Published: (2025)
Probe Pruning: Accelerating LLMs through Dynamic Pruning via Model-Probing
by: Le, Qi, et al.
Published: (2025)
by: Le, Qi, et al.
Published: (2025)
IDAP++: Advancing Divergence-Based Pruning via Filter-Level and Layer-Level Optimization
by: Samarin, Aleksei, et al.
Published: (2025)
by: Samarin, Aleksei, et al.
Published: (2025)
A Generic Layer Pruning Method for Signal Modulation Recognition Deep Learning Models
by: Lu, Yao, et al.
Published: (2024)
by: Lu, Yao, et al.
Published: (2024)
UniPruning: Unifying Local Metric and Global Feedback for Scalable Sparse LLMs
by: Ding, Yizhuo, et al.
Published: (2025)
by: Ding, Yizhuo, et al.
Published: (2025)
Maximum Redundancy Pruning: A Principle-Driven Layerwise Sparsity Allocation for LLMs
by: Gao, Chang, et al.
Published: (2025)
by: Gao, Chang, et al.
Published: (2025)
Geometric Median (GM) Matching for Robust Data Pruning
by: Acharya, Anish, et al.
Published: (2024)
by: Acharya, Anish, et al.
Published: (2024)
Dynamic Vocabulary Pruning in Early-Exit LLMs
by: Vincenti, Jort, et al.
Published: (2024)
by: Vincenti, Jort, et al.
Published: (2024)
DriftGuard: Mitigating Asynchronous Data Drift in Federated Learning
by: Han, Yizhou, et al.
Published: (2026)
by: Han, Yizhou, et al.
Published: (2026)
FAIR-Pruner: A Flexible Framework for Automatic Layer-Wise Pruning via Tolerance of Difference
by: Lin, Chenqing, et al.
Published: (2025)
by: Lin, Chenqing, et al.
Published: (2025)
When Attention Collapses: How Degenerate Layers in LLMs Enable Smaller, Stronger Models
by: Sanyal, Sunny, et al.
Published: (2024)
by: Sanyal, Sunny, et al.
Published: (2024)
IntraSlice: Towards High-Performance Structural Pruning with Block-Intra PCA for LLMs
by: Li, Meng, et al.
Published: (2026)
by: Li, Meng, et al.
Published: (2026)
Language Models Can Predict Their Own Behavior
by: Ashok, Dhananjay, et al.
Published: (2025)
by: Ashok, Dhananjay, et al.
Published: (2025)
Similar Items
-
Signal Collapse in One-Shot Pruning: When Sparse Models Fail to Distinguish Neural Representations
by: Saikumar, Dhananjay, et al.
Published: (2025) -
DRIVE: Dual Gradient-Based Rapid Iterative Pruning
by: Saikumar, Dhananjay, et al.
Published: (2024) -
NeuroFlux: Memory-Efficient CNN Training Using Adaptive Local Learning
by: Saikumar, Dhananjay, et al.
Published: (2024) -
Mosaic: Composite Projection Pruning for Resource-efficient LLMs
by: Eccles, Bailey J., et al.
Published: (2025) -
Rapid Deployment of DNNs for Edge Computing via Structured Pruning at Initialization
by: Eccles, Bailey J., et al.
Published: (2024)