Generalized Gradient Norm Clipping & Non-Euclidean $(L_0,L_1)$-Smoothness
Fuente:
arXiv
Saved in:
| Main Authors: | Pethick, Thomas, Xie, Wanyun, Erdogan, Mete, Antonakopoulos, Kimon, Silveti-Falls, Antonio, Cevher, Volkan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Training Deep Learning Models with Norm-Constrained LMOs
by: Pethick, Thomas, et al.
Published: (2025)
by: Pethick, Thomas, et al.
Published: (2025)
Training Neural Networks at Any Scale
by: Pethick, Thomas, et al.
Published: (2025)
by: Pethick, Thomas, et al.
Published: (2025)
Improving SAM Requires Rethinking its Optimization Formulation
by: Xie, Wanyun, et al.
Published: (2024)
by: Xie, Wanyun, et al.
Published: (2024)
Constrained Stochastic Spectral Preconditioning Converges for Nonconvex Objectives
by: Oikonomidis, Konstantinos, et al.
Published: (2026)
by: Oikonomidis, Konstantinos, et al.
Published: (2026)
SAMPa: Sharpness-aware Minimization Parallelized
by: Xie, Wanyun, et al.
Published: (2024)
by: Xie, Wanyun, et al.
Published: (2024)
Stable Nonconvex-Nonconcave Training via Linear Interpolation
by: Pethick, Thomas, et al.
Published: (2023)
by: Pethick, Thomas, et al.
Published: (2023)
Optimistic Dual Averaging Unifies Modern Optimizers
by: Pethick, Thomas, et al.
Published: (2026)
by: Pethick, Thomas, et al.
Published: (2026)
On the Generalization of Stochastic Gradient Descent with Momentum
by: Ramezani-Kebrya, Ali, et al.
Published: (2018)
by: Ramezani-Kebrya, Ali, et al.
Published: (2018)
On the Role of Batch Size in Stochastic Conditional Gradient Methods
by: Islamov, Rustem, et al.
Published: (2026)
by: Islamov, Rustem, et al.
Published: (2026)
Adaptive Conditional Gradient Descent
by: Khademi, Abbas, et al.
Published: (2025)
by: Khademi, Abbas, et al.
Published: (2025)
Efficient Large Language Model Inference with Neural Block Linearization
by: Erdogan, Mete, et al.
Published: (2025)
by: Erdogan, Mete, et al.
Published: (2025)
MaD-Mix: Multi-Modal Data Mixtures via Latent Space Coupling for Vision-Language Model Training
by: Xie, Wanyun, et al.
Published: (2026)
by: Xie, Wanyun, et al.
Published: (2026)
Chameleon: A Flexible Data-mixing Framework for Language Model Pretraining and Finetuning
by: Xie, Wanyun, et al.
Published: (2025)
by: Xie, Wanyun, et al.
Published: (2025)
Multi-Step Alignment as Markov Games: An Optimistic Online Gradient Descent Approach with Convergence Guarantees
by: Wu, Yongtao, et al.
Published: (2025)
by: Wu, Yongtao, et al.
Published: (2025)
Universal Gradient Methods for Stochastic Convex Optimization
by: Rodomanov, Anton, et al.
Published: (2024)
by: Rodomanov, Anton, et al.
Published: (2024)
Learning with Norm Constrained, Over-parameterized, Two-layer Neural Networks
by: Liu, Fanghui, et al.
Published: (2024)
by: Liu, Fanghui, et al.
Published: (2024)
Layer-wise Quantization for Quantized Optimistic Dual Averaging
by: Nguyen, Anh Duc, et al.
Published: (2025)
by: Nguyen, Anh Duc, et al.
Published: (2025)
Boosted Stochastic Frank-Wolfe for Constrained Nonconvex Optimization
by: Nandhan, Navil, et al.
Published: (2026)
by: Nandhan, Navil, et al.
Published: (2026)
Methods for Convex $(L_0,L_1)$-Smooth Optimization: Clipping, Acceleration, and Adaptivity
by: Gorbunov, Eduard, et al.
Published: (2024)
by: Gorbunov, Eduard, et al.
Published: (2024)
Tangent Space Fine-Tuning for Directional Preference Alignment in Large Language Models
by: Erdogan, Mete
Published: (2026)
by: Erdogan, Mete
Published: (2026)
Convergence of Clipped-SGD for Convex $(L_0,L_1)$-Smooth Optimization with Heavy-Tailed Noise
by: Chezhegov, Savelii, et al.
Published: (2025)
by: Chezhegov, Savelii, et al.
Published: (2025)
Near-Optimal Convergence of Accelerated Gradient Methods under Generalized and $(L_0, L_1)$-Smoothness
by: Tyurin, Alexander
Published: (2025)
by: Tyurin, Alexander
Published: (2025)
Score Broadcast and Decorrelation: A General Framework for Broadcast-Based Credit Assignment
by: Uzun, Mustafa, et al.
Published: (2026)
by: Uzun, Mustafa, et al.
Published: (2026)
Error Broadcast and Decorrelation as a Potential Artificial and Natural Learning Mechanism
by: Erdogan, Mete, et al.
Published: (2025)
by: Erdogan, Mete, et al.
Published: (2025)
IL-SOAR : Imitation Learning with Soft Optimistic Actor cRitic
by: Viel, Stefano, et al.
Published: (2025)
by: Viel, Stefano, et al.
Published: (2025)
Imitation Learning in Discounted Linear MDPs without exploration assumptions
by: Viano, Luca, et al.
Published: (2024)
by: Viano, Luca, et al.
Published: (2024)
To Clip or not to Clip: the Dynamics of SGD with Gradient Clipping in High-Dimensions
by: Marshall, Noah, et al.
Published: (2024)
by: Marshall, Noah, et al.
Published: (2024)
Gradient Clipping Beyond Vector Norms: A Spectral Approach for Matrix-Valued Parameters
by: Yukhimchuk, Alexander, et al.
Published: (2026)
by: Yukhimchuk, Alexander, et al.
Published: (2026)
Byzantine-Robust Optimization under $(L_0, L_1)$-Smoothness
by: Bolatov, Arman, et al.
Published: (2026)
by: Bolatov, Arman, et al.
Published: (2026)
SVD Based Least Squares for X-Ray Pneumonia Classification Using Deep Features
by: Erdogan, Mete, et al.
Published: (2025)
by: Erdogan, Mete, et al.
Published: (2025)
MT-NAM: An Efficient and Adaptive Model for Epileptic Seizure Detection
by: Afzal, Arshia, et al.
Published: (2025)
by: Afzal, Arshia, et al.
Published: (2025)
Split the Differences, Pool the Rest: Provably Efficient Multi-Objective Imitation
by: Sheebaelhamd, Ziyad, et al.
Published: (2026)
by: Sheebaelhamd, Ziyad, et al.
Published: (2026)
Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution
by: Barla, Adam, et al.
Published: (2026)
by: Barla, Adam, et al.
Published: (2026)
High-Dimensional Kernel Methods under Covariate Shift: Data-Dependent Implicit Regularization
by: Chen, Yihang, et al.
Published: (2024)
by: Chen, Yihang, et al.
Published: (2024)
Easy Data Unlearning Bench
by: Rinberg, Roy, et al.
Published: (2026)
by: Rinberg, Roy, et al.
Published: (2026)
Adversarial Training for Defense Against Label Poisoning Attacks
by: Bal, Melis Ilayda, et al.
Published: (2025)
by: Bal, Melis Ilayda, et al.
Published: (2025)
ESLM: Risk-Averse Selective Language Modeling for Efficient Pretraining
by: Bal, Melis Ilayda, et al.
Published: (2025)
by: Bal, Melis Ilayda, et al.
Published: (2025)
On the Interaction of Batch Noise, Adaptivity, and Compression, under $(L_0,L_1)$-Smoothness: An SDE Approach
by: Compagnoni, Enea Monzio, et al.
Published: (2025)
by: Compagnoni, Enea Monzio, et al.
Published: (2025)
Error Feedback under $(L_0,L_1)$-Smoothness: Normalization and Momentum
by: Khirirat, Sarit, et al.
Published: (2024)
by: Khirirat, Sarit, et al.
Published: (2024)
Generalization of Scaled Deep ResNets in the Mean-Field Regime
by: Chen, Yihang, et al.
Published: (2024)
by: Chen, Yihang, et al.
Published: (2024)
Similar Items
-
Training Deep Learning Models with Norm-Constrained LMOs
by: Pethick, Thomas, et al.
Published: (2025) -
Training Neural Networks at Any Scale
by: Pethick, Thomas, et al.
Published: (2025) -
Improving SAM Requires Rethinking its Optimization Formulation
by: Xie, Wanyun, et al.
Published: (2024) -
Constrained Stochastic Spectral Preconditioning Converges for Nonconvex Objectives
by: Oikonomidis, Konstantinos, et al.
Published: (2026) -
SAMPa: Sharpness-aware Minimization Parallelized
by: Xie, Wanyun, et al.
Published: (2024)