Mosaic: Composite Projection Pruning for Resource-efficient LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Eccles, Bailey J., Wong, Leon, Varghese, Blesson |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rapid Deployment of DNNs for Edge Computing via Structured Pruning at Initialization
by: Eccles, Bailey J., et al.
Published: (2024)
by: Eccles, Bailey J., et al.
Published: (2024)
Data-Free Pruning of Self-Attention Layers in LLMs
by: Saikumar, Dhananjay, et al.
Published: (2025)
by: Saikumar, Dhananjay, et al.
Published: (2025)
DNNShifter: An Efficient DNN Pruning System for Edge Computing
by: Eccles, Bailey J., et al.
Published: (2023)
by: Eccles, Bailey J., et al.
Published: (2023)
Signal Collapse in One-Shot Pruning: When Sparse Models Fail to Distinguish Neural Representations
by: Saikumar, Dhananjay, et al.
Published: (2025)
by: Saikumar, Dhananjay, et al.
Published: (2025)
FedOptima: Optimizing Resource Utilization in Federated Learning
by: Zhang, Zihan, et al.
Published: (2025)
by: Zhang, Zihan, et al.
Published: (2025)
Mosaic Pruning: A Hierarchical Framework for Generalizable Pruning of Mixture-of-Experts Models
by: Hu, Wentao, et al.
Published: (2025)
by: Hu, Wentao, et al.
Published: (2025)
DRIVE: Dual Gradient-Based Rapid Iterative Pruning
by: Saikumar, Dhananjay, et al.
Published: (2024)
by: Saikumar, Dhananjay, et al.
Published: (2024)
Ampere: Communication-Efficient and High-Accuracy Split Federated Learning
by: Zhang, Zihan, et al.
Published: (2025)
by: Zhang, Zihan, et al.
Published: (2025)
Memory Mosaics
by: Zhang, Jianyu, et al.
Published: (2024)
by: Zhang, Jianyu, et al.
Published: (2024)
MobiEdit: Resource-efficient Knowledge Editing for Personalized On-device LLMs
by: Lu, Zhenyan, et al.
Published: (2025)
by: Lu, Zhenyan, et al.
Published: (2025)
PrunePEFT: Iterative Hybrid Pruning for Parameter-Efficient Fine-tuning of LLMs
by: Yu, Tongzhou, et al.
Published: (2025)
by: Yu, Tongzhou, et al.
Published: (2025)
Safe Pruning LoRA: Robust Distance-Guided Pruning for Safety Alignment in Adaptation of LLMs
by: Ao, Shuang, et al.
Published: (2025)
by: Ao, Shuang, et al.
Published: (2025)
PERP: Rethinking the Prune-Retrain Paradigm in the Era of LLMs
by: Zimmer, Max, et al.
Published: (2023)
by: Zimmer, Max, et al.
Published: (2023)
Resource-Constrained Affect Modelling via Variance Regularisation Pruning
by: Pinitas, Kosmas, et al.
Published: (2026)
by: Pinitas, Kosmas, et al.
Published: (2026)
UnIT: Scalable Unstructured Inference-Time Pruning for MAC-efficient Neural Inference on MCUs
by: Neth, Ashe, et al.
Published: (2025)
by: Neth, Ashe, et al.
Published: (2025)
Resource-Aware Neural Network Pruning Using Graph-based Reinforcement Learning
by: Balemans, Dieter, et al.
Published: (2025)
by: Balemans, Dieter, et al.
Published: (2025)
Probe Pruning: Accelerating LLMs through Dynamic Pruning via Model-Probing
by: Le, Qi, et al.
Published: (2025)
by: Le, Qi, et al.
Published: (2025)
NeuroFlux: Memory-Efficient CNN Training Using Adaptive Local Learning
by: Saikumar, Dhananjay, et al.
Published: (2024)
by: Saikumar, Dhananjay, et al.
Published: (2024)
UniPruning: Unifying Local Metric and Global Feedback for Scalable Sparse LLMs
by: Ding, Yizhuo, et al.
Published: (2025)
by: Ding, Yizhuo, et al.
Published: (2025)
Maximum Redundancy Pruning: A Principle-Driven Layerwise Sparsity Allocation for LLMs
by: Gao, Chang, et al.
Published: (2025)
by: Gao, Chang, et al.
Published: (2025)
Mosaic of Modalities: A Comprehensive Benchmark for Multimodal Graph Learning
by: Zhu, Jing, et al.
Published: (2024)
by: Zhu, Jing, et al.
Published: (2024)
Dynamic Vocabulary Pruning in Early-Exit LLMs
by: Vincenti, Jort, et al.
Published: (2024)
by: Vincenti, Jort, et al.
Published: (2024)
IntraSlice: Towards High-Performance Structural Pruning with Block-Intra PCA for LLMs
by: Li, Meng, et al.
Published: (2026)
by: Li, Meng, et al.
Published: (2026)
Ghosted Layers: Unconstrained Activation Alignment for Recovering Layer-Pruned LLMs
by: Yun, Vincent-Daniel, et al.
Published: (2026)
by: Yun, Vincent-Daniel, et al.
Published: (2026)
AudioMosaic: Contrastive Masked Audio Representation Learning
by: Huang, Hanxun, et al.
Published: (2026)
by: Huang, Hanxun, et al.
Published: (2026)
DONOD: Efficient and Generalizable Instruction Fine-Tuning for LLMs via Model-Intrinsic Dataset Pruning
by: Hu, Jucheng, et al.
Published: (2025)
by: Hu, Jucheng, et al.
Published: (2025)
SLoPe: Double-Pruned Sparse Plus Lazy Low-Rank Adapter Pretraining of LLMs
by: Mozaffari, Mohammad, et al.
Published: (2024)
by: Mozaffari, Mohammad, et al.
Published: (2024)
Two-Stage Regularization-Based Structured Pruning for LLMs
by: Feng, Mingkuan, et al.
Published: (2025)
by: Feng, Mingkuan, et al.
Published: (2025)
Resource-efficient Layer-wise Federated Self-supervised Learning
by: Tun, Ye Lin, et al.
Published: (2024)
by: Tun, Ye Lin, et al.
Published: (2024)
When Fewer Layers Break More Chains: Layer Pruning Harms Test-Time Scaling in LLMs
by: Wang, Keyu, et al.
Published: (2025)
by: Wang, Keyu, et al.
Published: (2025)
OPTIMA: Optimal One-shot Pruning for LLMs via Quadratic Programming Reconstruction
by: Mozaffari, Mohammad, et al.
Published: (2025)
by: Mozaffari, Mohammad, et al.
Published: (2025)
Frugal Machine Learning for Energy-efficient, and Resource-aware Artificial Intelligence
by: Violos, John, et al.
Published: (2025)
by: Violos, John, et al.
Published: (2025)
Sparsest Models Elude Pruning: An Exposé of Pruning's Current Capabilities
by: Zhang, Stephen, et al.
Published: (2024)
by: Zhang, Stephen, et al.
Published: (2024)
REAM: Merging Improves Pruning of Experts in LLMs
by: Jha, Saurav, et al.
Published: (2026)
by: Jha, Saurav, et al.
Published: (2026)
SwiftPrune: Hessian-Free Weight Pruning for Large Language Models
by: Kang, Yuhan, et al.
Published: (2025)
by: Kang, Yuhan, et al.
Published: (2025)
IPPRO: Importance-based Pruning with PRojective Offset for Magnitude-indifferent Structural Pruning
by: Jung, Jaeheun, et al.
Published: (2025)
by: Jung, Jaeheun, et al.
Published: (2025)
TopoPrune: Robust Data Pruning via Unified Latent Space Topology
by: Roy, Arjun, et al.
Published: (2026)
by: Roy, Arjun, et al.
Published: (2026)
Junk DNA Hypothesis: Pruning Small Pre-Trained Weights Irreversibly and Monotonically Impairs "Difficult" Downstream Tasks in LLMs
by: Yin, Lu, et al.
Published: (2023)
by: Yin, Lu, et al.
Published: (2023)
2SSP: A Two-Stage Framework for Structured Pruning of LLMs
by: Sandri, Fabrizio, et al.
Published: (2025)
by: Sandri, Fabrizio, et al.
Published: (2025)
ASFL: An Adaptive Model Splitting and Resource Allocation Framework for Split Federated Learning
by: Meng, Chuiyang, et al.
Published: (2026)
by: Meng, Chuiyang, et al.
Published: (2026)
Similar Items
-
Rapid Deployment of DNNs for Edge Computing via Structured Pruning at Initialization
by: Eccles, Bailey J., et al.
Published: (2024) -
Data-Free Pruning of Self-Attention Layers in LLMs
by: Saikumar, Dhananjay, et al.
Published: (2025) -
DNNShifter: An Efficient DNN Pruning System for Edge Computing
by: Eccles, Bailey J., et al.
Published: (2023) -
Signal Collapse in One-Shot Pruning: When Sparse Models Fail to Distinguish Neural Representations
by: Saikumar, Dhananjay, et al.
Published: (2025) -
FedOptima: Optimizing Resource Utilization in Federated Learning
by: Zhang, Zihan, et al.
Published: (2025)