Sharpness-Aware Minimization Enhances Feature Quality via Balanced Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Springer, Jacob Mitchell, Nagarajan, Vaishnavh, Raghunathan, Aditi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sharpness-Aware Pretraining Mitigates Catastrophic Forgetting
by: Watts, Ishaan, et al.
Published: (2026)
by: Watts, Ishaan, et al.
Published: (2026)
Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction
by: Nagarajan, Vaishnavh, et al.
Published: (2025)
by: Nagarajan, Vaishnavh, et al.
Published: (2025)
Understanding Catastrophic Forgetting in Language Models via Implicit Inference
by: Kotha, Suhas, et al.
Published: (2023)
by: Kotha, Suhas, et al.
Published: (2023)
The pitfalls of next-token prediction
by: Bachmann, Gregor, et al.
Published: (2024)
by: Bachmann, Gregor, et al.
Published: (2024)
Mitigating Bias in RAG: Controlling the Embedder
by: Kim, Taeyoun, et al.
Published: (2025)
by: Kim, Taeyoun, et al.
Published: (2025)
Repetition Improves Language Model Embeddings
by: Springer, Jacob Mitchell, et al.
Published: (2024)
by: Springer, Jacob Mitchell, et al.
Published: (2024)
Early Data Exposure Improves Robustness to Subsequent Fine-Tuning
by: Feng, Lawrence, et al.
Published: (2026)
by: Feng, Lawrence, et al.
Published: (2026)
Understanding the Failure Modes of Out-of-Distribution Generalization
by: Nagarajan, Vaishnavh, et al.
Published: (2020)
by: Nagarajan, Vaishnavh, et al.
Published: (2020)
Enhancing Sharpness-Aware Minimization by Learning Perturbation Radius
by: Wang, Xuehao, et al.
Published: (2024)
by: Wang, Xuehao, et al.
Published: (2024)
Deep sequence models tend to memorize geometrically; it is unclear why
by: Noroozizadeh, Shahriar, et al.
Published: (2025)
by: Noroozizadeh, Shahriar, et al.
Published: (2025)
Sharpness-Aware Minimization for Evolutionary Feature Construction in Regression
by: Zhang, Hengzhe, et al.
Published: (2024)
by: Zhang, Hengzhe, et al.
Published: (2024)
Agnostic Sharpness-Aware Minimization
by: Nguyen, Van-Anh, et al.
Published: (2024)
by: Nguyen, Van-Anh, et al.
Published: (2024)
Friendly Sharpness-Aware Minimization
by: Li, Tao, et al.
Published: (2024)
by: Li, Tao, et al.
Published: (2024)
Tilted Sharpness-Aware Minimization
by: Li, Tian, et al.
Published: (2024)
by: Li, Tian, et al.
Published: (2024)
Enhancing Robustness of Offline Reinforcement Learning Under Data Corruption via Sharpness-Aware Minimization
by: Xu, Le, et al.
Published: (2025)
by: Xu, Le, et al.
Published: (2025)
Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMs
by: Zhong, Ziqian, et al.
Published: (2025)
by: Zhong, Ziqian, et al.
Published: (2025)
On student-teacher deviations in distillation: does it pay to disobey?
by: Nagarajan, Vaishnavh, et al.
Published: (2023)
by: Nagarajan, Vaishnavh, et al.
Published: (2023)
Sharpness-Aware Minimization for Generalized Embedding Learning in Federated Recommendation
by: Yu, Fengyuan, et al.
Published: (2026)
by: Yu, Fengyuan, et al.
Published: (2026)
Asynchronous Sharpness-Aware Minimization For Fast and Accurate Deep Learning
by: Jo, Junhyuk, et al.
Published: (2025)
by: Jo, Junhyuk, et al.
Published: (2025)
Sharpness-Aware Minimization Can Hallucinate Minimizers
by: Park, Chanwoong, et al.
Published: (2025)
by: Park, Chanwoong, et al.
Published: (2025)
Sharpness-Aware Minimization Revisited: Weighted Sharpness as a Regularization Term
by: Yue, Yun, et al.
Published: (2023)
by: Yue, Yun, et al.
Published: (2023)
VASSO: Variance Suppression for Sharpness-Aware Minimization
by: Li, Bingcong, et al.
Published: (2025)
by: Li, Bingcong, et al.
Published: (2025)
GCSAM: Gradient Centralized Sharpness Aware Minimization
by: Hassan, Mohamed, et al.
Published: (2025)
by: Hassan, Mohamed, et al.
Published: (2025)
Why is SAM Robust to Label Noise?
by: Baek, Christina, et al.
Published: (2024)
by: Baek, Christina, et al.
Published: (2024)
Stability Analysis of Sharpness-Aware Minimization
by: Kim, Hoki, et al.
Published: (2023)
by: Kim, Hoki, et al.
Published: (2023)
A Universal Class of Sharpness-Aware Minimization Algorithms
by: Tahmasebi, Behrooz, et al.
Published: (2024)
by: Tahmasebi, Behrooz, et al.
Published: (2024)
LightSAM: Parameter-Agnostic Sharpness-Aware Minimization
by: Cheng, Yifei, et al.
Published: (2025)
by: Cheng, Yifei, et al.
Published: (2025)
Sharpness-Aware Minimization in Genetic Programming
by: Bakurov, Illya, et al.
Published: (2024)
by: Bakurov, Illya, et al.
Published: (2024)
Sharpness-Aware Minimization and the Edge of Stability
by: Long, Philip M., et al.
Published: (2023)
by: Long, Philip M., et al.
Published: (2023)
Preconditioned Sharpness-Aware Minimization: Unifying Analysis and a Novel Learning Algorithm
by: Zhang, Yilang, et al.
Published: (2025)
by: Zhang, Yilang, et al.
Published: (2025)
DGSAM: Domain Generalization via Individual Sharpness-Aware Minimization
by: Song, Youngjun, et al.
Published: (2025)
by: Song, Youngjun, et al.
Published: (2025)
Sharpness-Aware Minimization in Logit Space Efficiently Enhances Direct Preference Optimization
by: Luo, Haocheng, et al.
Published: (2026)
by: Luo, Haocheng, et al.
Published: (2026)
Membership Privacy Risks of Sharpness Aware Minimization
by: Kim, Young In, et al.
Published: (2023)
by: Kim, Young In, et al.
Published: (2023)
Navigating Potholes with Geometry-Aware Sharpness Minimization
by: Dufort-Labbé, Simon, et al.
Published: (2026)
by: Dufort-Labbé, Simon, et al.
Published: (2026)
Self-Trained Verification for Training- and Test-Time Self-Improvement
by: Wu, Chen Henry, et al.
Published: (2026)
by: Wu, Chen Henry, et al.
Published: (2026)
Understanding Finetuning for Factual Knowledge Extraction
by: Ghosal, Gaurav, et al.
Published: (2024)
by: Ghosal, Gaurav, et al.
Published: (2024)
ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test Cases
by: Zhong, Ziqian, et al.
Published: (2025)
by: Zhong, Ziqian, et al.
Published: (2025)
Efficient Sharpness-Aware Minimization for Molecular Graph Transformer Models
by: Wang, Yili, et al.
Published: (2024)
by: Wang, Yili, et al.
Published: (2024)
Implicit Regularization of Sharpness-Aware Minimization for Scale-Invariant Problems
by: Li, Bingcong, et al.
Published: (2024)
by: Li, Bingcong, et al.
Published: (2024)
1st-Order Magic: Analysis of Sharpness-Aware Minimization
by: Tiwary, Nalin, et al.
Published: (2024)
by: Tiwary, Nalin, et al.
Published: (2024)
Similar Items
-
Sharpness-Aware Pretraining Mitigates Catastrophic Forgetting
by: Watts, Ishaan, et al.
Published: (2026) -
Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction
by: Nagarajan, Vaishnavh, et al.
Published: (2025) -
Understanding Catastrophic Forgetting in Language Models via Implicit Inference
by: Kotha, Suhas, et al.
Published: (2023) -
The pitfalls of next-token prediction
by: Bachmann, Gregor, et al.
Published: (2024) -
Mitigating Bias in RAG: Controlling the Embedder
by: Kim, Taeyoun, et al.
Published: (2025)