SAMPa: Sharpness-aware Minimization Parallelized
Fuente:
arXiv
Saved in:
| Main Authors: | Xie, Wanyun, Pethick, Thomas, Cevher, Volkan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Stable Nonconvex-Nonconcave Training via Linear Interpolation
by: Pethick, Thomas, et al.
Published: (2023)
by: Pethick, Thomas, et al.
Published: (2023)
Optimistic Dual Averaging Unifies Modern Optimizers
by: Pethick, Thomas, et al.
Published: (2026)
by: Pethick, Thomas, et al.
Published: (2026)
Improving SAM Requires Rethinking its Optimization Formulation
by: Xie, Wanyun, et al.
Published: (2024)
by: Xie, Wanyun, et al.
Published: (2024)
Generalized Gradient Norm Clipping & Non-Euclidean $(L_0,L_1)$-Smoothness
by: Pethick, Thomas, et al.
Published: (2025)
by: Pethick, Thomas, et al.
Published: (2025)
Training Deep Learning Models with Norm-Constrained LMOs
by: Pethick, Thomas, et al.
Published: (2025)
by: Pethick, Thomas, et al.
Published: (2025)
MaD-Mix: Multi-Modal Data Mixtures via Latent Space Coupling for Vision-Language Model Training
by: Xie, Wanyun, et al.
Published: (2026)
by: Xie, Wanyun, et al.
Published: (2026)
Chameleon: A Flexible Data-mixing Framework for Language Model Pretraining and Finetuning
by: Xie, Wanyun, et al.
Published: (2025)
by: Xie, Wanyun, et al.
Published: (2025)
μP$^2$: Effective Sharpness Aware Minimization Requires Layerwise Perturbation Scaling
by: Haas, Moritz, et al.
Published: (2024)
by: Haas, Moritz, et al.
Published: (2024)
Training Neural Networks at Any Scale
by: Pethick, Thomas, et al.
Published: (2025)
by: Pethick, Thomas, et al.
Published: (2025)
Efficient Continual Finite-Sum Minimization
by: Mavrothalassitis, Ioannis, et al.
Published: (2024)
by: Mavrothalassitis, Ioannis, et al.
Published: (2024)
Best of Both Worlds: Regret Minimization versus Minimax Play
by: Müller, Adrian, et al.
Published: (2025)
by: Müller, Adrian, et al.
Published: (2025)
Learning with Norm Constrained, Over-parameterized, Two-layer Neural Networks
by: Liu, Fanghui, et al.
Published: (2024)
by: Liu, Fanghui, et al.
Published: (2024)
Imitation Learning in Discounted Linear MDPs without exploration assumptions
by: Viano, Luca, et al.
Published: (2024)
by: Viano, Luca, et al.
Published: (2024)
IL-SOAR : Imitation Learning with Soft Optimistic Actor cRitic
by: Viel, Stefano, et al.
Published: (2025)
by: Viel, Stefano, et al.
Published: (2025)
Efficient Large Language Model Inference with Neural Block Linearization
by: Erdogan, Mete, et al.
Published: (2025)
by: Erdogan, Mete, et al.
Published: (2025)
Critical Influence of Overparameterization on Sharpness-aware Minimization
by: Shin, Sungbin, et al.
Published: (2023)
by: Shin, Sungbin, et al.
Published: (2023)
MT-NAM: An Efficient and Adaptive Model for Epileptic Seizure Detection
by: Afzal, Arshia, et al.
Published: (2025)
by: Afzal, Arshia, et al.
Published: (2025)
High-Dimensional Kernel Methods under Covariate Shift: Data-Dependent Implicit Regularization
by: Chen, Yihang, et al.
Published: (2024)
by: Chen, Yihang, et al.
Published: (2024)
Split the Differences, Pool the Rest: Provably Efficient Multi-Objective Imitation
by: Sheebaelhamd, Ziyad, et al.
Published: (2026)
by: Sheebaelhamd, Ziyad, et al.
Published: (2026)
Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution
by: Barla, Adam, et al.
Published: (2026)
by: Barla, Adam, et al.
Published: (2026)
Easy Data Unlearning Bench
by: Rinberg, Roy, et al.
Published: (2026)
by: Rinberg, Roy, et al.
Published: (2026)
Adversarial Training for Defense Against Label Poisoning Attacks
by: Bal, Melis Ilayda, et al.
Published: (2025)
by: Bal, Melis Ilayda, et al.
Published: (2025)
ESLM: Risk-Averse Selective Language Modeling for Efficient Pretraining
by: Bal, Melis Ilayda, et al.
Published: (2025)
by: Bal, Melis Ilayda, et al.
Published: (2025)
Matching Multiple Experts: On the Exploitability of Multi-Agent Imitation Learning
by: Bergerault, Antoine, et al.
Published: (2026)
by: Bergerault, Antoine, et al.
Published: (2026)
Truly No-Regret Learning in Constrained MDPs
by: Müller, Adrian, et al.
Published: (2024)
by: Müller, Adrian, et al.
Published: (2024)
Learning Equilibria from Data: Provably Efficient Multi-Agent Imitation Learning
by: Freihaut, Till, et al.
Published: (2025)
by: Freihaut, Till, et al.
Published: (2025)
Accelerating Spectral Clustering under Fairness Constraints
by: Tonin, Francesco, et al.
Published: (2025)
by: Tonin, Francesco, et al.
Published: (2025)
Stabilizing Sharpness-aware Minimization Through A Simple Renormalization Strategy
by: Tan, Chengli, et al.
Published: (2024)
by: Tan, Chengli, et al.
Published: (2024)
Ascent Fails to Forget
by: Mavrothalassitis, Ioannis, et al.
Published: (2025)
by: Mavrothalassitis, Ioannis, et al.
Published: (2025)
REST: Efficient and Accelerated EEG Seizure Analysis through Residual State Updates
by: Afzal, Arshia, et al.
Published: (2024)
by: Afzal, Arshia, et al.
Published: (2024)
Certified Robustness Under Bounded Levenshtein Distance
by: Rocamora, Elias Abad, et al.
Published: (2025)
by: Rocamora, Elias Abad, et al.
Published: (2025)
Generalization of Scaled Deep ResNets in the Mean-Field Regime
by: Chen, Yihang, et al.
Published: (2024)
by: Chen, Yihang, et al.
Published: (2024)
Hadamard product in deep learning: Introduction, Advances and Challenges
by: Chrysos, Grigorios G, et al.
Published: (2025)
by: Chrysos, Grigorios G, et al.
Published: (2025)
On the Generalization of Stochastic Gradient Descent with Momentum
by: Ramezani-Kebrya, Ali, et al.
Published: (2018)
by: Ramezani-Kebrya, Ali, et al.
Published: (2018)
Revisiting SMoE Language Models by Evaluating Inefficiencies with Task Specific Expert Pruning
by: Sarkar, Soumajyoti, et al.
Published: (2024)
by: Sarkar, Soumajyoti, et al.
Published: (2024)
Leveraging the Context through Multi-Round Interactions for Jailbreaking Attacks
by: Cheng, Yixin, et al.
Published: (2024)
by: Cheng, Yixin, et al.
Published: (2024)
The Last Mile to Supervised Performance: Semi-Supervised Domain Adaptation for Semantic Segmentation
by: Morales-Brotons, Daniel, et al.
Published: (2024)
by: Morales-Brotons, Daniel, et al.
Published: (2024)
Multilinear Operator Networks
by: Cheng, Yixin, et al.
Published: (2024)
by: Cheng, Yixin, et al.
Published: (2024)
Faster Inference of Flow-Based Generative Models via Improved Data-Noise Coupling
by: Davtyan, Aram, et al.
Published: (2026)
by: Davtyan, Aram, et al.
Published: (2026)
Agnostic Sharpness-Aware Minimization
by: Nguyen, Van-Anh, et al.
Published: (2024)
by: Nguyen, Van-Anh, et al.
Published: (2024)
Similar Items
-
Stable Nonconvex-Nonconcave Training via Linear Interpolation
by: Pethick, Thomas, et al.
Published: (2023) -
Optimistic Dual Averaging Unifies Modern Optimizers
by: Pethick, Thomas, et al.
Published: (2026) -
Improving SAM Requires Rethinking its Optimization Formulation
by: Xie, Wanyun, et al.
Published: (2024) -
Generalized Gradient Norm Clipping & Non-Euclidean $(L_0,L_1)$-Smoothness
by: Pethick, Thomas, et al.
Published: (2025) -
Training Deep Learning Models with Norm-Constrained LMOs
by: Pethick, Thomas, et al.
Published: (2025)