One Shot vs. Iterative: Rethinking Pruning Strategies for Model Compression
Fuente:
arXiv
Guardado en:
| Autores principales: | Janusz, Mikołaj, Wojnar, Tomasz, Li, Yawei, Benini, Luca, Adamczewski, Kamil |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Shapley Pruning for Neural Network Compression
por: Adamczewski, Kamil, et al.
Publicado: (2024)
por: Adamczewski, Kamil, et al.
Publicado: (2024)
LUNA: Efficient and Topology-Agnostic Foundation Model for EEG Signal Analysis
por: Döner, Berkay, et al.
Publicado: (2025)
por: Döner, Berkay, et al.
Publicado: (2025)
RDKV: Rate-Distortion Bit Allocation for Joint Eviction and Quantization of the KV Cache
por: Zhang, Junkai, et al.
Publicado: (2026)
por: Zhang, Junkai, et al.
Publicado: (2026)
FEMBA: Efficient and Scalable EEG Analysis with a Bidirectional Mamba Foundation Model
por: Tegon, Anna, et al.
Publicado: (2025)
por: Tegon, Anna, et al.
Publicado: (2025)
OMENN: One Matrix to Explain Neural Networks
por: Wróbel, Adam, et al.
Publicado: (2024)
por: Wróbel, Adam, et al.
Publicado: (2024)
Projected Compression: Trainable Projection for Efficient Transformer Compression
por: Stefaniak, Maciej, et al.
Publicado: (2025)
por: Stefaniak, Maciej, et al.
Publicado: (2025)
PhysioWave: A Multi-Scale Wavelet-Transformer for Physiological Signal Representation
por: Chen, Yanlong, et al.
Publicado: (2025)
por: Chen, Yanlong, et al.
Publicado: (2025)
Finetuning and Quantization of EEG-Based Foundational BioSignal Models on ECG and PPG Data for Blood Pressure Estimation
por: Tóth, Bálint, et al.
Publicado: (2025)
por: Tóth, Bálint, et al.
Publicado: (2025)
Attention Sinks as Internal Signals for Hallucination Detection in Large Language Models
por: Binkowski, Jakub, et al.
Publicado: (2026)
por: Binkowski, Jakub, et al.
Publicado: (2026)
PrunePEFT: Iterative Hybrid Pruning for Parameter-Efficient Fine-tuning of LLMs
por: Yu, Tongzhou, et al.
Publicado: (2025)
por: Yu, Tongzhou, et al.
Publicado: (2025)
Scaling Laws for Fine-Grained Mixture of Experts
por: Krajewski, Jakub, et al.
Publicado: (2024)
por: Krajewski, Jakub, et al.
Publicado: (2024)
TinyMyo: a Tiny Foundation Model for Flexible EMG Signal Processing at the Edge
por: Fasulo, Matteo, et al.
Publicado: (2025)
por: Fasulo, Matteo, et al.
Publicado: (2025)
PERP: Rethinking the Prune-Retrain Paradigm in the Era of LLMs
por: Zimmer, Max, et al.
Publicado: (2023)
por: Zimmer, Max, et al.
Publicado: (2023)
Rethinking Pruning for Backdoor Mitigation: An Optimization Perspective
por: Li, Nan, et al.
Publicado: (2024)
por: Li, Nan, et al.
Publicado: (2024)
Structured vs. Unstructured Pruning: An Exponential Gap
por: Ferre', Davide, et al.
Publicado: (2026)
por: Ferre', Davide, et al.
Publicado: (2026)
Sparse Weight Averaging with Multiple Particles for Iterative Magnitude Pruning
por: Choi, Moonseok, et al.
Publicado: (2023)
por: Choi, Moonseok, et al.
Publicado: (2023)
Chasing COMET: Leveraging Minimum Bayes Risk Decoding for Self-Improving Machine Translation
por: Guttmann, Kamil, et al.
Publicado: (2024)
por: Guttmann, Kamil, et al.
Publicado: (2024)
REAP the Experts: Why Pruning Prevails for One-Shot MoE compression
por: Lasby, Mike, et al.
Publicado: (2025)
por: Lasby, Mike, et al.
Publicado: (2025)
Locality-Aware Redundancy Pruning for LLM Depth Compression
por: Yun, Vincent-Daniel, et al.
Publicado: (2026)
por: Yun, Vincent-Daniel, et al.
Publicado: (2026)
Beyond One-Way Pruning: Bidirectional Pruning-Regrowth for Extreme Accuracy-Sparsity Tradeoff
por: Liu, Junchen, et al.
Publicado: (2025)
por: Liu, Junchen, et al.
Publicado: (2025)
Vanishing Contributions: A Unified Framework for Smooth and Iterative Model Compression
por: Nikiforos, Lorenzo, et al.
Publicado: (2025)
por: Nikiforos, Lorenzo, et al.
Publicado: (2025)
One Self-Configurable Model to Solve Many Abstract Visual Reasoning Problems
por: Małkiński, Mikołaj, et al.
Publicado: (2023)
por: Małkiński, Mikołaj, et al.
Publicado: (2023)
Signal Collapse in One-Shot Pruning: When Sparse Models Fail to Distinguish Neural Representations
por: Saikumar, Dhananjay, et al.
Publicado: (2025)
por: Saikumar, Dhananjay, et al.
Publicado: (2025)
FedMap: Iterative Magnitude-Based Pruning for Communication-Efficient Federated Learning
por: Herzog, Alexander, et al.
Publicado: (2024)
por: Herzog, Alexander, et al.
Publicado: (2024)
Compressing Many-Shots in In-Context Learning
por: Khatri, Devvrit, et al.
Publicado: (2025)
por: Khatri, Devvrit, et al.
Publicado: (2025)
Rethinking Key-Value Cache Compression Techniques for Large Language Model Serving
por: Gao, Wei, et al.
Publicado: (2025)
por: Gao, Wei, et al.
Publicado: (2025)
SymDrift: One-Shot Generative Modeling under Symmetries
por: Darouich, Samir, et al.
Publicado: (2026)
por: Darouich, Samir, et al.
Publicado: (2026)
A Free Lunch in LLM Compression: Revisiting Retraining after Pruning
por: Wagner, Moritz, et al.
Publicado: (2025)
por: Wagner, Moritz, et al.
Publicado: (2025)
RAP: KV-Cache Compression via RoPE-Aligned Pruning
por: Xin, Jihao, et al.
Publicado: (2026)
por: Xin, Jihao, et al.
Publicado: (2026)
Is Meta-Learning Out? Rethinking Unsupervised Few-Shot Classification with Limited Entropy
por: Guan, Yunchuan, et al.
Publicado: (2025)
por: Guan, Yunchuan, et al.
Publicado: (2025)
One-Shot Clustering for Federated Learning
por: Zuziak, Maciej Krzysztof, et al.
Publicado: (2025)
por: Zuziak, Maciej Krzysztof, et al.
Publicado: (2025)
Is One Score Enough? Rethinking the Evaluation of Sequentially Evolving LLM Memory
por: Dong, Songwei, et al.
Publicado: (2026)
por: Dong, Songwei, et al.
Publicado: (2026)
Rethinking the Harmonic Loss via Non-Euclidean Distance Layers
por: Miller-Golub, Maxwell, et al.
Publicado: (2026)
por: Miller-Golub, Maxwell, et al.
Publicado: (2026)
CAOS: Conformal Aggregation of One-Shot Predictors
por: Waldron, Maja
Publicado: (2026)
por: Waldron, Maja
Publicado: (2026)
Semi-Supervised One-Shot Imitation Learning
por: Wu, Philipp, et al.
Publicado: (2024)
por: Wu, Philipp, et al.
Publicado: (2024)
Real vs. Semi-Simulated: Rethinking Evaluation for Treatment Effect Estimation
por: Panagopoulos, George
Publicado: (2026)
por: Panagopoulos, George
Publicado: (2026)
SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator
por: Chen, Guoxuan, et al.
Publicado: (2024)
por: Chen, Guoxuan, et al.
Publicado: (2024)
Sparsest Models Elude Pruning: An Exposé of Pruning's Current Capabilities
por: Zhang, Stephen, et al.
Publicado: (2024)
por: Zhang, Stephen, et al.
Publicado: (2024)
Exploring Federated Pruning for Large Language Models
por: Guo, Pengxin, et al.
Publicado: (2025)
por: Guo, Pengxin, et al.
Publicado: (2025)
Towards a Larger Model via One-Shot Federated Learning on Heterogeneous Client Models
por: Ye, Wenxuan, et al.
Publicado: (2025)
por: Ye, Wenxuan, et al.
Publicado: (2025)
Ejemplares similares
-
Shapley Pruning for Neural Network Compression
por: Adamczewski, Kamil, et al.
Publicado: (2024) -
LUNA: Efficient and Topology-Agnostic Foundation Model for EEG Signal Analysis
por: Döner, Berkay, et al.
Publicado: (2025) -
RDKV: Rate-Distortion Bit Allocation for Joint Eviction and Quantization of the KV Cache
por: Zhang, Junkai, et al.
Publicado: (2026) -
FEMBA: Efficient and Scalable EEG Analysis with a Bidirectional Mamba Foundation Model
por: Tegon, Anna, et al.
Publicado: (2025) -
OMENN: One Matrix to Explain Neural Networks
por: Wróbel, Adam, et al.
Publicado: (2024)