Minor First, Major Last: A Depth-Induced Implicit Bias of Sharpness-Aware Minimization
Fuente:
arXiv
Saved in:
| Main Authors: | Moon, Chaewon, Si, Dongkuk, Yun, Chulhee |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Cost of Robustness: Tighter Bounds on Parameter Complexity for Robust Memorization in ReLU Nets
by: Kim, Yujun, et al.
Published: (2025)
by: Kim, Yujun, et al.
Published: (2025)
Implicit Bias of Per-sample Adam on Separable Data: Departure from the Full-batch Regime
by: Baek, Beomhan, et al.
Published: (2025)
by: Baek, Beomhan, et al.
Published: (2025)
Implicit Bias and Loss of Plasticity in Matrix Completion: Depth Promotes Low-Rankness
by: Shin, Baekrok, et al.
Published: (2026)
by: Shin, Baekrok, et al.
Published: (2026)
Unpacking the Implicit Norm Dynamics of Sharpness-Aware Minimization in Tensorized Models
by: Cao, Tianxiao, et al.
Published: (2025)
by: Cao, Tianxiao, et al.
Published: (2025)
Provable Benefit of Cutout and CutMix for Feature Learning
by: Oh, Junsoo, et al.
Published: (2024)
by: Oh, Junsoo, et al.
Published: (2024)
Stability Analysis of Sharpness-Aware Minimization
by: Kim, Hoki, et al.
Published: (2023)
by: Kim, Hoki, et al.
Published: (2023)
Navigating Potholes with Geometry-Aware Sharpness Minimization
by: Dufort-Labbé, Simon, et al.
Published: (2026)
by: Dufort-Labbé, Simon, et al.
Published: (2026)
Membership Privacy Risks of Sharpness Aware Minimization
by: Kim, Young In, et al.
Published: (2023)
by: Kim, Young In, et al.
Published: (2023)
Towards Understanding The Calibration Benefits of Sharpness-Aware Minimization
by: Tan, Chengli, et al.
Published: (2025)
by: Tan, Chengli, et al.
Published: (2025)
Revisiting Sharpness-Aware Minimization: A More Faithful and Effective Implementation
by: Chen, Jianlong, et al.
Published: (2026)
by: Chen, Jianlong, et al.
Published: (2026)
SAMOSA: Sharpness Aware Minimization for Open Set Active learning
by: Kim, Young In, et al.
Published: (2025)
by: Kim, Young In, et al.
Published: (2025)
Arithmetic Transformers Can Length-Generalize in Both Operand Length and Count
by: Cho, Hanseul, et al.
Published: (2024)
by: Cho, Hanseul, et al.
Published: (2024)
DASH: Warm-Starting Neural Network Training in Stationary Settings without Loss of Plasticity
by: Shin, Baekrok, et al.
Published: (2024)
by: Shin, Baekrok, et al.
Published: (2024)
Lightweight Dataset Pruning without Full Training via Example Difficulty and Prediction Uncertainty
by: Cho, Yeseul, et al.
Published: (2025)
by: Cho, Yeseul, et al.
Published: (2025)
Scaling Laws of SignSGD in Linear Regression: When Does It Outperform SGD?
by: Kim, Jihwan, et al.
Published: (2026)
by: Kim, Jihwan, et al.
Published: (2026)
X-SAM: Boosting Sharpness-Aware Minimization with Dominant-Eigenvector Gradient Correction
by: Duan, Hongru, et al.
Published: (2026)
by: Duan, Hongru, et al.
Published: (2026)
Sharpness-Aware Minimization in Logit Space Efficiently Enhances Direct Preference Optimization
by: Luo, Haocheng, et al.
Published: (2026)
by: Luo, Haocheng, et al.
Published: (2026)
FairSAM: Fair Classification on Corrupted Data Through Sharpness-Aware Minimization
by: Dai, Yucong, et al.
Published: (2025)
by: Dai, Yucong, et al.
Published: (2025)
Towards Understanding the Role of Sharpness-Aware Minimization Algorithms for Out-of-Distribution Generalization
by: Schapiro, Samuel, et al.
Published: (2024)
by: Schapiro, Samuel, et al.
Published: (2024)
Attributing Data for Sharpness-Aware Minimization
by: Ren, Chenyang, et al.
Published: (2025)
by: Ren, Chenyang, et al.
Published: (2025)
Fast Graph Sharpness-Aware Minimization for Enhancing and Accelerating Few-Shot Node Classification
by: Luo, Yihong, et al.
Published: (2024)
by: Luo, Yihong, et al.
Published: (2024)
Sharpness-Aware Minimization with Z-Score Gradient Filtering
by: Yun, Vincent-Daniel
Published: (2025)
by: Yun, Vincent-Daniel
Published: (2025)
Parameter Expanded Stochastic Gradient Markov Chain Monte Carlo
by: Kim, Hyunsu, et al.
Published: (2025)
by: Kim, Hyunsu, et al.
Published: (2025)
Convergence and Implicit Bias of Gradient Descent on Continual Linear Classification
by: Jung, Hyunji, et al.
Published: (2025)
by: Jung, Hyunji, et al.
Published: (2025)
Gradient Compression May Hurt Generalization: A Remedy by Synthetic Data Guided Sharpness Aware Minimization
by: Gu, Yujie, et al.
Published: (2026)
by: Gu, Yujie, et al.
Published: (2026)
Bi-LoRA: Efficient Sharpness-Aware Minimization for Fine-Tuning Large-Scale Models
by: Liu, Yuhang, et al.
Published: (2025)
by: Liu, Yuhang, et al.
Published: (2025)
Enhancing Robustness of Offline Reinforcement Learning Under Data Corruption via Sharpness-Aware Minimization
by: Xu, Le, et al.
Published: (2025)
by: Xu, Le, et al.
Published: (2025)
Understanding Sharpness Dynamics in NN Training with a Minimalist Example: The Effects of Dataset Difficulty, Depth, Stochasticity, and More
by: Yoo, Geonhui, et al.
Published: (2025)
by: Yoo, Geonhui, et al.
Published: (2025)
Through the River: Understanding the Benefit of Schedule-Free Methods for Language Model Training
by: Song, Minhak, et al.
Published: (2025)
by: Song, Minhak, et al.
Published: (2025)
On the Duality Between Sharpness-Aware Minimization and Adversarial Training
by: Zhang, Yihao, et al.
Published: (2024)
by: Zhang, Yihao, et al.
Published: (2024)
Stabilizing Sharpness-aware Minimization Through A Simple Renormalization Strategy
by: Tan, Chengli, et al.
Published: (2024)
by: Tan, Chengli, et al.
Published: (2024)
Sharpness-Aware Teleportation on Riemannian Manifolds
by: Truong, Tuan, et al.
Published: (2023)
by: Truong, Tuan, et al.
Published: (2023)
Sharpness-Aware Black-Box Optimization
by: Ye, Feiyang, et al.
Published: (2024)
by: Ye, Feiyang, et al.
Published: (2024)
Locality-Aware Redundancy Pruning for LLM Depth Compression
by: Yun, Vincent-Daniel, et al.
Published: (2026)
by: Yun, Vincent-Daniel, et al.
Published: (2026)
Flexible Bayesian Last Layer Models Using Implicit Priors and Diffusion Posterior Sampling
by: Xu, Jian, et al.
Published: (2024)
by: Xu, Jian, et al.
Published: (2024)
Disentangling Granularity: An Implicit Inductive Bias in Factorized VAEs
by: Chen, Zihao, et al.
Published: (2025)
by: Chen, Zihao, et al.
Published: (2025)
Correcting Performance Estimation Bias in Imbalanced Classification with Minority Subconcepts
by: Maxson, Taylor, et al.
Published: (2026)
by: Maxson, Taylor, et al.
Published: (2026)
Implicit Bias of the JKO Scheme
by: Halmos, Peter, et al.
Published: (2025)
by: Halmos, Peter, et al.
Published: (2025)
NCSAM Noise-Compensated Sharpness-Aware Minimization for Noisy Label Learning
by: Xu, Jiayu, et al.
Published: (2026)
by: Xu, Jiayu, et al.
Published: (2026)
Tractable Sharpness-Aware Learning of Probabilistic Circuits
by: Suresh, Hrithik, et al.
Published: (2025)
by: Suresh, Hrithik, et al.
Published: (2025)
Similar Items
-
The Cost of Robustness: Tighter Bounds on Parameter Complexity for Robust Memorization in ReLU Nets
by: Kim, Yujun, et al.
Published: (2025) -
Implicit Bias of Per-sample Adam on Separable Data: Departure from the Full-batch Regime
by: Baek, Beomhan, et al.
Published: (2025) -
Implicit Bias and Loss of Plasticity in Matrix Completion: Depth Promotes Low-Rankness
by: Shin, Baekrok, et al.
Published: (2026) -
Unpacking the Implicit Norm Dynamics of Sharpness-Aware Minimization in Tensorized Models
by: Cao, Tianxiao, et al.
Published: (2025) -
Provable Benefit of Cutout and CutMix for Feature Learning
by: Oh, Junsoo, et al.
Published: (2024)