SAFE: Finding Sparse and Flat Minima to Improve Pruning
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Dongyeop, Lee, Kwanhee, Chung, Jinseok, Lee, Namhoon |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SASSHA: Sharpness-aware Adaptive Second-order Optimization with Stable Hessian Approximation
by: Shin, Dahun, et al.
Published: (2025)
by: Shin, Dahun, et al.
Published: (2025)
The Unseen Frontier: Pushing the Limits of LLM Sparsity with Surrogate-Free ADMM
by: Lee, Kwanhee, et al.
Published: (2025)
by: Lee, Kwanhee, et al.
Published: (2025)
Zeroth-Order Optimization Finds Flat Minima
by: Zhang, Liang, et al.
Published: (2025)
by: Zhang, Liang, et al.
Published: (2025)
Are Flat Minima an Illusion?
by: Bennett, Michael Timothy
Published: (2026)
by: Bennett, Michael Timothy
Published: (2026)
Sparse Weight Averaging with Multiple Particles for Iterative Magnitude Pruning
by: Choi, Moonseok, et al.
Published: (2023)
by: Choi, Moonseok, et al.
Published: (2023)
Towards the Connection between Activation Sparsity and Flat Minima
by: Peng, Ze, et al.
Published: (2026)
by: Peng, Ze, et al.
Published: (2026)
IPPRO: Importance-based Pruning with PRojective Offset for Magnitude-indifferent Structural Pruning
by: Jung, Jaeheun, et al.
Published: (2025)
by: Jung, Jaeheun, et al.
Published: (2025)
DP-FedPGN: Finding Global Flat Minima for Differentially Private Federated Learning via Penalizing Gradient Norm
by: Liu, Junkang, et al.
Published: (2025)
by: Liu, Junkang, et al.
Published: (2025)
A Flat Minima Perspective on Understanding Augmentations and Model Robustness
by: Yoo, Weebum, et al.
Published: (2025)
by: Yoo, Weebum, et al.
Published: (2025)
An Analysis of Concept Bottleneck Models: Measuring, Understanding, and Mitigating the Impact of Noisy Annotations
by: Park, Seonghwan, et al.
Published: (2025)
by: Park, Seonghwan, et al.
Published: (2025)
Weight Concentration Regularization for Improving Pruning Robustness Under High Sparsity
by: Yun, Vincent-Daniel, et al.
Published: (2025)
by: Yun, Vincent-Daniel, et al.
Published: (2025)
Mitigating Staleness in Asynchronous Pipeline Parallelism via Basis Rotation
by: Jung, Hyunji, et al.
Published: (2026)
by: Jung, Hyunji, et al.
Published: (2026)
Critical Influence of Overparameterization on Sharpness-aware Minimization
by: Shin, Sungbin, et al.
Published: (2023)
by: Shin, Sungbin, et al.
Published: (2023)
Sparse Autoencoders Do Not Find Canonical Units of Analysis
by: Leask, Patrick, et al.
Published: (2025)
by: Leask, Patrick, et al.
Published: (2025)
A Function-Centric Perspective on Flat and Sharp Minima
by: Mason-Williams, Israel, et al.
Published: (2025)
by: Mason-Williams, Israel, et al.
Published: (2025)
Generalized Gaussian Temporal Difference Error for Uncertainty-aware Reinforcement Learning
by: Kim, Seyeon, et al.
Published: (2024)
by: Kim, Seyeon, et al.
Published: (2024)
DIVER-1: Scaling Intracranial EEG Foundation Models for Transferable Representations
by: Han, Danny Dongyeop, et al.
Published: (2025)
by: Han, Danny Dongyeop, et al.
Published: (2025)
Steering Sparse Autoencoder Latents to Control Dynamic Head Pruning in Vision Transformers (Student Abstract)
by: Lee, Yousung, et al.
Published: (2026)
by: Lee, Yousung, et al.
Published: (2026)
Realizing Unaligned Block-wise Pruning for DNN Acceleration on Mobile Devices
by: Lee, Hayun, et al.
Published: (2024)
by: Lee, Hayun, et al.
Published: (2024)
Catalyst: a Novel Regularizer for Structured Pruning with Auxiliary Extension of Parameter Space
by: Jung, Jaeheun, et al.
Published: (2025)
by: Jung, Jaeheun, et al.
Published: (2025)
DIVER-0 : A Fully Channel Equivariant EEG Foundation Model
by: Han, Danny Dongyeop, et al.
Published: (2025)
by: Han, Danny Dongyeop, et al.
Published: (2025)
Sparse Model Soups: A Recipe for Improved Pruning via Model Averaging
by: Zimmer, Max, et al.
Published: (2023)
by: Zimmer, Max, et al.
Published: (2023)
A Unified Theory of Sparse Dictionary Learning in Mechanistic Interpretability: Piecewise Biconvexity and Spurious Minima
by: Tang, Yiming, et al.
Published: (2025)
by: Tang, Yiming, et al.
Published: (2025)
Data-Driven Lipschitz Continuity: A Cost-Effective Approach to Improve Adversarial Robustness
by: Chen, Erh-Chung, et al.
Published: (2024)
by: Chen, Erh-Chung, et al.
Published: (2024)
REAM: Merging Improves Pruning of Experts in LLMs
by: Jha, Saurav, et al.
Published: (2026)
by: Jha, Saurav, et al.
Published: (2026)
The Amazing Agent Race: Strong Tool Users, Weak Navigators
by: Kim, Zae Myung, et al.
Published: (2026)
by: Kim, Zae Myung, et al.
Published: (2026)
Identifying Sparsely Active Circuits Through Local Loss Landscape Decomposition
by: Chrisman, Brianna, et al.
Published: (2025)
by: Chrisman, Brianna, et al.
Published: (2025)
DreamSmooth: Improving Model-based Reinforcement Learning via Reward Smoothing
by: Lee, Vint, et al.
Published: (2023)
by: Lee, Vint, et al.
Published: (2023)
Find A Winning Sign: Sign Is All We Need to Win the Lottery
by: Oh, Junghun, et al.
Published: (2025)
by: Oh, Junghun, et al.
Published: (2025)
Multi-View Node Pruning for Accurate Graph Representation
by: Kim, Hanjin, et al.
Published: (2025)
by: Kim, Hanjin, et al.
Published: (2025)
Offline Reinforcement Learning with Universal Horizon Models
by: Chung, Hojun, et al.
Published: (2026)
by: Chung, Hojun, et al.
Published: (2026)
Dynamic Relational Priming Improves Transformer in Multivariate Time Series
by: Lee, Hunjae, et al.
Published: (2025)
by: Lee, Hunjae, et al.
Published: (2025)
Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization Challenges
by: Lee, Nayoung, et al.
Published: (2025)
by: Lee, Nayoung, et al.
Published: (2025)
Think Clearly: Improving Reasoning via Redundant Token Pruning
by: Choi, Daewon, et al.
Published: (2025)
by: Choi, Daewon, et al.
Published: (2025)
Locality-Aware Redundancy Pruning for LLM Depth Compression
by: Yun, Vincent-Daniel, et al.
Published: (2026)
by: Yun, Vincent-Daniel, et al.
Published: (2026)
IgPose: A Generative Data-Augmented Pipeline for Robust Immunoglobulin-Antigen Binding Prediction
by: Bui, Tien-Cuong, et al.
Published: (2026)
by: Bui, Tien-Cuong, et al.
Published: (2026)
Rethinking Pruning Large Language Models: Benefits and Pitfalls of Reconstruction Error Minimization
by: Shin, Sungbin, et al.
Published: (2024)
by: Shin, Sungbin, et al.
Published: (2024)
SparseSwaps: Tractable LLM Pruning Mask Refinement at Scale
by: Zimmer, Max, et al.
Published: (2025)
by: Zimmer, Max, et al.
Published: (2025)
The Right to be Forgotten in Pruning: Unveil Machine Unlearning on Sparse Models
by: Xiao, Yang, et al.
Published: (2025)
by: Xiao, Yang, et al.
Published: (2025)
AdaRank: Adaptive Rank Pruning for Enhanced Model Merging
by: Lee, Chanhyuk, et al.
Published: (2025)
by: Lee, Chanhyuk, et al.
Published: (2025)
Similar Items
-
SASSHA: Sharpness-aware Adaptive Second-order Optimization with Stable Hessian Approximation
by: Shin, Dahun, et al.
Published: (2025) -
The Unseen Frontier: Pushing the Limits of LLM Sparsity with Surrogate-Free ADMM
by: Lee, Kwanhee, et al.
Published: (2025) -
Zeroth-Order Optimization Finds Flat Minima
by: Zhang, Liang, et al.
Published: (2025) -
Are Flat Minima an Illusion?
by: Bennett, Michael Timothy
Published: (2026) -
Sparse Weight Averaging with Multiple Particles for Iterative Magnitude Pruning
by: Choi, Moonseok, et al.
Published: (2023)