Training Deep Learning Models with Norm-Constrained LMOs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pethick, Thomas, Xie, Wanyun, Antonakopoulos, Kimon, Zhu, Zhenyu, Silveti-Falls, Antonio, Cevher, Volkan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Constrained Stochastic Spectral Preconditioning Converges for Nonconvex Objectives
von: Oikonomidis, Konstantinos, et al.
Veröffentlicht: (2026)
von: Oikonomidis, Konstantinos, et al.
Veröffentlicht: (2026)
Generalized Gradient Norm Clipping & Non-Euclidean $(L_0,L_1)$-Smoothness
von: Pethick, Thomas, et al.
Veröffentlicht: (2025)
von: Pethick, Thomas, et al.
Veröffentlicht: (2025)
Stable Nonconvex-Nonconcave Training via Linear Interpolation
von: Pethick, Thomas, et al.
Veröffentlicht: (2023)
von: Pethick, Thomas, et al.
Veröffentlicht: (2023)
Training Neural Networks at Any Scale
von: Pethick, Thomas, et al.
Veröffentlicht: (2025)
von: Pethick, Thomas, et al.
Veröffentlicht: (2025)
Improving SAM Requires Rethinking its Optimization Formulation
von: Xie, Wanyun, et al.
Veröffentlicht: (2024)
von: Xie, Wanyun, et al.
Veröffentlicht: (2024)
On the Role of Batch Size in Stochastic Conditional Gradient Methods
von: Islamov, Rustem, et al.
Veröffentlicht: (2026)
von: Islamov, Rustem, et al.
Veröffentlicht: (2026)
Adaptive Conditional Gradient Descent
von: Khademi, Abbas, et al.
Veröffentlicht: (2025)
von: Khademi, Abbas, et al.
Veröffentlicht: (2025)
Layer-wise Quantization for Quantized Optimistic Dual Averaging
von: Nguyen, Anh Duc, et al.
Veröffentlicht: (2025)
von: Nguyen, Anh Duc, et al.
Veröffentlicht: (2025)
SAMPa: Sharpness-aware Minimization Parallelized
von: Xie, Wanyun, et al.
Veröffentlicht: (2024)
von: Xie, Wanyun, et al.
Veröffentlicht: (2024)
Universal Gradient Methods for Stochastic Convex Optimization
von: Rodomanov, Anton, et al.
Veröffentlicht: (2024)
von: Rodomanov, Anton, et al.
Veröffentlicht: (2024)
Boosted Stochastic Frank-Wolfe for Constrained Nonconvex Optimization
von: Nandhan, Navil, et al.
Veröffentlicht: (2026)
von: Nandhan, Navil, et al.
Veröffentlicht: (2026)
Optimistic Dual Averaging Unifies Modern Optimizers
von: Pethick, Thomas, et al.
Veröffentlicht: (2026)
von: Pethick, Thomas, et al.
Veröffentlicht: (2026)
Advancing the lower bounds: An accelerated, stochastic, second-order method with optimal adaptation to inexactness
von: Agafonov, Artem, et al.
Veröffentlicht: (2023)
von: Agafonov, Artem, et al.
Veröffentlicht: (2023)
Adversarial Training Should Be Cast as a Non-Zero-Sum Game
von: Robey, Alexander, et al.
Veröffentlicht: (2023)
von: Robey, Alexander, et al.
Veröffentlicht: (2023)
Efficient Continual Finite-Sum Minimization
von: Mavrothalassitis, Ioannis, et al.
Veröffentlicht: (2024)
von: Mavrothalassitis, Ioannis, et al.
Veröffentlicht: (2024)
Implicit Bias of AdamW: $\ell_\infty$ Norm Constrained Optimization
von: Xie, Shuo, et al.
Veröffentlicht: (2024)
von: Xie, Shuo, et al.
Veröffentlicht: (2024)
Randomized algorithms and PAC bounds for inverse reinforcement learning in continuous spaces
von: Kamoutsi, Angeliki, et al.
Veröffentlicht: (2024)
von: Kamoutsi, Angeliki, et al.
Veröffentlicht: (2024)
Krylov Cubic Regularized Newton: A Subspace Second-Order Method with Dimension-Free Convergence Rate
von: Jiang, Ruichen, et al.
Veröffentlicht: (2024)
von: Jiang, Ruichen, et al.
Veröffentlicht: (2024)
MaD-Mix: Multi-Modal Data Mixtures via Latent Space Coupling for Vision-Language Model Training
von: Xie, Wanyun, et al.
Veröffentlicht: (2026)
von: Xie, Wanyun, et al.
Veröffentlicht: (2026)
Learning to Remove Cuts in Integer Linear Programming
von: Puigdemont, Pol, et al.
Veröffentlicht: (2024)
von: Puigdemont, Pol, et al.
Veröffentlicht: (2024)
MADA: Meta-Adaptive Optimizers through hyper-gradient Descent
von: Ozkara, Kaan, et al.
Veröffentlicht: (2024)
von: Ozkara, Kaan, et al.
Veröffentlicht: (2024)
Where Does Warm-Up Come From? Adaptive Scheduling for Norm-Constrained Optimizers
von: Riabinin, Artem, et al.
Veröffentlicht: (2026)
von: Riabinin, Artem, et al.
Veröffentlicht: (2026)
Learning Constrained Optimization with Deep Augmented Lagrangian Methods
von: Kotary, James, et al.
Veröffentlicht: (2024)
von: Kotary, James, et al.
Veröffentlicht: (2024)
Chameleon: A Flexible Data-mixing Framework for Language Model Pretraining and Finetuning
von: Xie, Wanyun, et al.
Veröffentlicht: (2025)
von: Xie, Wanyun, et al.
Veröffentlicht: (2025)
Towards The Implicit Bias on Multiclass Separable Data Under Norm Constraints
von: Xie, Shengping, et al.
Veröffentlicht: (2026)
von: Xie, Shengping, et al.
Veröffentlicht: (2026)
On the Generalization of Stochastic Gradient Descent with Momentum
von: Ramezani-Kebrya, Ali, et al.
Veröffentlicht: (2018)
von: Ramezani-Kebrya, Ali, et al.
Veröffentlicht: (2018)
Universal Architectures for the Learning of Polyhedral Norms and Convex Regularizers
von: Unser, Michael, et al.
Veröffentlicht: (2025)
von: Unser, Michael, et al.
Veröffentlicht: (2025)
Clipping Improves Adam-Norm and AdaGrad-Norm when the Noise Is Heavy-Tailed
von: Chezhegov, Savelii, et al.
Veröffentlicht: (2024)
von: Chezhegov, Savelii, et al.
Veröffentlicht: (2024)
The ADMM-PINNs Algorithmic Framework for Nonsmooth PDE-Constrained Optimization: A Deep Learning Approach
von: Song, Yongcun, et al.
Veröffentlicht: (2023)
von: Song, Yongcun, et al.
Veröffentlicht: (2023)
Multi-Step Alignment as Markov Games: An Optimistic Online Gradient Descent Approach with Convergence Guarantees
von: Wu, Yongtao, et al.
Veröffentlicht: (2025)
von: Wu, Yongtao, et al.
Veröffentlicht: (2025)
Exact and Heuristic Algorithms for Constrained Biclustering
von: Sudoso, Antonio M.
Veröffentlicht: (2025)
von: Sudoso, Antonio M.
Veröffentlicht: (2025)
Benchmarking Stochastic Approximation Algorithms for Fairness-Constrained Training of Deep Neural Networks
von: Kliachkin, Andrii, et al.
Veröffentlicht: (2025)
von: Kliachkin, Andrii, et al.
Veröffentlicht: (2025)
Old Optimizer, New Norm: An Anthology
von: Bernstein, Jeremy, et al.
Veröffentlicht: (2024)
von: Bernstein, Jeremy, et al.
Veröffentlicht: (2024)
Complexity of Classical Acceleration for $\ell_1$-Regularized PageRank
von: Fountoulakis, Kimon, et al.
Veröffentlicht: (2026)
von: Fountoulakis, Kimon, et al.
Veröffentlicht: (2026)
Learning to Stop: Deep Learning for Mean Field Optimal Stopping
von: Magnino, Lorenzo, et al.
Veröffentlicht: (2024)
von: Magnino, Lorenzo, et al.
Veröffentlicht: (2024)
Resilient Constrained Reinforcement Learning
von: Ding, Dongsheng, et al.
Veröffentlicht: (2023)
von: Ding, Dongsheng, et al.
Veröffentlicht: (2023)
On Learning for Ambiguous Chance Constrained Problems
von: Madhusudanarao, A Ch, et al.
Veröffentlicht: (2023)
von: Madhusudanarao, A Ch, et al.
Veröffentlicht: (2023)
Learning with Norm Constrained, Over-parameterized, Two-layer Neural Networks
von: Liu, Fanghui, et al.
Veröffentlicht: (2024)
von: Liu, Fanghui, et al.
Veröffentlicht: (2024)
Muon Optimizes Under Spectral Norm Constraints
von: Chen, Lizhang, et al.
Veröffentlicht: (2025)
von: Chen, Lizhang, et al.
Veröffentlicht: (2025)
Self-supervised Equality Embedded Deep Lagrange Dual for Approximate Constrained Optimization
von: Kim, Minsoo, et al.
Veröffentlicht: (2023)
von: Kim, Minsoo, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Constrained Stochastic Spectral Preconditioning Converges for Nonconvex Objectives
von: Oikonomidis, Konstantinos, et al.
Veröffentlicht: (2026) -
Generalized Gradient Norm Clipping & Non-Euclidean $(L_0,L_1)$-Smoothness
von: Pethick, Thomas, et al.
Veröffentlicht: (2025) -
Stable Nonconvex-Nonconcave Training via Linear Interpolation
von: Pethick, Thomas, et al.
Veröffentlicht: (2023) -
Training Neural Networks at Any Scale
von: Pethick, Thomas, et al.
Veröffentlicht: (2025) -
Improving SAM Requires Rethinking its Optimization Formulation
von: Xie, Wanyun, et al.
Veröffentlicht: (2024)