Sparse Model Soups: A Recipe for Improved Pruning via Model Averaging
Fuente:
arXiv
Saved in:
| Main Authors: | Zimmer, Max, Spiegel, Christoph, Pokutta, Sebastian |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PERP: Rethinking the Prune-Retrain Paradigm in the Era of LLMs
by: Zimmer, Max, et al.
Published: (2023)
by: Zimmer, Max, et al.
Published: (2023)
SparseSwaps: Tractable LLM Pruning Mask Refinement at Scale
by: Zimmer, Max, et al.
Published: (2025)
by: Zimmer, Max, et al.
Published: (2025)
A Free Lunch in LLM Compression: Revisiting Retraining after Pruning
by: Wagner, Moritz, et al.
Published: (2025)
by: Wagner, Moritz, et al.
Published: (2025)
Neural Sum-of-Squares: Certifying the Nonnegativity of Polynomials with Transformers
by: Pelleriti, Nico, et al.
Published: (2025)
by: Pelleriti, Nico, et al.
Published: (2025)
The Agentic Researcher: A Practical Guide to AI-Assisted Research in Mathematics and Machine Learning
by: Zimmer, Max, et al.
Published: (2026)
by: Zimmer, Max, et al.
Published: (2026)
Compression-aware Training of Neural Networks using Frank-Wolfe
by: Zimmer, Max, et al.
Published: (2022)
by: Zimmer, Max, et al.
Published: (2022)
On the Byzantine-Resilience of Distillation-Based Federated Learning
by: Roux, Christophe, et al.
Published: (2024)
by: Roux, Christophe, et al.
Published: (2024)
From Associations to Activations: Comparing Behavioral and Hidden-State Semantic Geometry in LLMs
by: Schiekiera, Louis, et al.
Published: (2026)
by: Schiekiera, Louis, et al.
Published: (2026)
Sparse Weight Averaging with Multiple Particles for Iterative Magnitude Pruning
by: Choi, Moonseok, et al.
Published: (2023)
by: Choi, Moonseok, et al.
Published: (2023)
Neural Discovery in Mathematics: Do Machines Dream of Colored Planes?
by: Mundinger, Konrad, et al.
Published: (2025)
by: Mundinger, Konrad, et al.
Published: (2025)
Don't Be Greedy, Just Relax! Pruning LLMs via Frank-Wolfe
by: Roux, Christophe, et al.
Published: (2025)
by: Roux, Christophe, et al.
Published: (2025)
Bone Soups: A Seek-and-Soup Model Merging Approach for Controllable Multi-Objective Generation
by: Xie, Guofu, et al.
Published: (2025)
by: Xie, Guofu, et al.
Published: (2025)
SAFE: Finding Sparse and Flat Minima to Improve Pruning
by: Lee, Dongyeop, et al.
Published: (2025)
by: Lee, Dongyeop, et al.
Published: (2025)
Interpretability Guarantees with Merlin-Arthur Classifiers
by: Wäldchen, Stephan, et al.
Published: (2022)
by: Wäldchen, Stephan, et al.
Published: (2022)
ECG-Soup: Harnessing Multi-Layer Synergy for ECG Foundation Models
by: Nguyen, Phu X., et al.
Published: (2025)
by: Nguyen, Phu X., et al.
Published: (2025)
The Right to be Forgotten in Pruning: Unveil Machine Unlearning on Sparse Models
by: Xiao, Yang, et al.
Published: (2025)
by: Xiao, Yang, et al.
Published: (2025)
Capturing Temporal Dynamics in Large-Scale Canopy Tree Height Estimation
by: Pauls, Jan, et al.
Published: (2025)
by: Pauls, Jan, et al.
Published: (2025)
VC-Soup: Value-Consistency Guided Multi-Value Alignment for Large Language Models
by: Xu, Hefei, et al.
Published: (2026)
by: Xu, Hefei, et al.
Published: (2026)
Understanding and Improving Model Averaging in Federated Learning on Heterogeneous Data
by: Zhou, Tailin, et al.
Published: (2023)
by: Zhou, Tailin, et al.
Published: (2023)
A Training Data Recipe to Accelerate A* Search with Language Models
by: Gupta, Devaansh, et al.
Published: (2024)
by: Gupta, Devaansh, et al.
Published: (2024)
State Soup: In-Context Skill Learning, Retrieval and Mixing
by: Pióro, Maciej, et al.
Published: (2024)
by: Pióro, Maciej, et al.
Published: (2024)
Quamba: A Post-Training Quantization Recipe for Selective State Space Models
by: Chiang, Hung-Yueh, et al.
Published: (2024)
by: Chiang, Hung-Yueh, et al.
Published: (2024)
Approximating Latent Manifolds in Neural Networks via Vanishing Ideals
by: Pelleriti, Nico, et al.
Published: (2025)
by: Pelleriti, Nico, et al.
Published: (2025)
Resource-Constrained Affect Modelling via Variance Regularisation Pruning
by: Pinitas, Kosmas, et al.
Published: (2026)
by: Pinitas, Kosmas, et al.
Published: (2026)
Sparsest Models Elude Pruning: An Exposé of Pruning's Current Capabilities
by: Zhang, Stephen, et al.
Published: (2024)
by: Zhang, Stephen, et al.
Published: (2024)
Mosaic Pruning: A Hierarchical Framework for Generalizable Pruning of Mixture-of-Experts Models
by: Hu, Wentao, et al.
Published: (2025)
by: Hu, Wentao, et al.
Published: (2025)
Neural Parameter Regression for Explicit Representations of PDE Solution Operators
by: Mundinger, Konrad, et al.
Published: (2024)
by: Mundinger, Konrad, et al.
Published: (2024)
Soup to go: mitigating forgetting during continual learning with model averaging
by: Kleiman, Anat, et al.
Published: (2025)
by: Kleiman, Anat, et al.
Published: (2025)
SparseJEPA: Sparse Representation Learning of Joint Embedding Predictive Architectures
by: Hartman, Max, et al.
Published: (2025)
by: Hartman, Max, et al.
Published: (2025)
SwiftPrune: Hessian-Free Weight Pruning for Large Language Models
by: Kang, Yuhan, et al.
Published: (2025)
by: Kang, Yuhan, et al.
Published: (2025)
Diffusion Soup: Model Merging for Text-to-Image Diffusion Models
by: Biggs, Benjamin, et al.
Published: (2024)
by: Biggs, Benjamin, et al.
Published: (2024)
Estimating Canopy Height at Scale
by: Pauls, Jan, et al.
Published: (2024)
by: Pauls, Jan, et al.
Published: (2024)
CAFP: A Post-Processing Framework for Group Fairness via Counterfactual Model Averaging
by: Arévalo, Irina, et al.
Published: (2026)
by: Arévalo, Irina, et al.
Published: (2026)
Probe Pruning: Accelerating LLMs through Dynamic Pruning via Model-Probing
by: Le, Qi, et al.
Published: (2025)
by: Le, Qi, et al.
Published: (2025)
Numerical Pruning for Efficient Autoregressive Models
by: Shen, Xuan, et al.
Published: (2024)
by: Shen, Xuan, et al.
Published: (2024)
Symmetric Pruning of Large Language Models
by: Yi, Kai, et al.
Published: (2025)
by: Yi, Kai, et al.
Published: (2025)
Revisiting Weight Averaging for Model Merging
by: Choi, Jiho, et al.
Published: (2024)
by: Choi, Jiho, et al.
Published: (2024)
OATS: Outlier-Aware Pruning Through Sparse and Low Rank Decomposition
by: Zhang, Stephen, et al.
Published: (2024)
by: Zhang, Stephen, et al.
Published: (2024)
Improving Robustness In Sparse Autoencoders via Masked Regularization
by: Narayanaswamy, Vivek, et al.
Published: (2026)
by: Narayanaswamy, Vivek, et al.
Published: (2026)
Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
by: Kantamneni, Subhash, et al.
Published: (2025)
by: Kantamneni, Subhash, et al.
Published: (2025)
Similar Items
-
PERP: Rethinking the Prune-Retrain Paradigm in the Era of LLMs
by: Zimmer, Max, et al.
Published: (2023) -
SparseSwaps: Tractable LLM Pruning Mask Refinement at Scale
by: Zimmer, Max, et al.
Published: (2025) -
A Free Lunch in LLM Compression: Revisiting Retraining after Pruning
by: Wagner, Moritz, et al.
Published: (2025) -
Neural Sum-of-Squares: Certifying the Nonnegativity of Polynomials with Transformers
by: Pelleriti, Nico, et al.
Published: (2025) -
The Agentic Researcher: A Practical Guide to AI-Assisted Research in Mathematics and Machine Learning
by: Zimmer, Max, et al.
Published: (2026)