On the Generalization of Stochastic Gradient Descent with Momentum
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ramezani-Kebrya, Ali, Antonakopoulos, Kimon, Cevher, Volkan, Khisti, Ashish, Liang, Ben |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2018
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Layer-wise Quantization for Quantized Optimistic Dual Averaging
von: Nguyen, Anh Duc, et al.
Veröffentlicht: (2025)
von: Nguyen, Anh Duc, et al.
Veröffentlicht: (2025)
Multi-Step Alignment as Markov Games: An Optimistic Online Gradient Descent Approach with Convergence Guarantees
von: Wu, Yongtao, et al.
Veröffentlicht: (2025)
von: Wu, Yongtao, et al.
Veröffentlicht: (2025)
Generalized Gradient Norm Clipping & Non-Euclidean $(L_0,L_1)$-Smoothness
von: Pethick, Thomas, et al.
Veröffentlicht: (2025)
von: Pethick, Thomas, et al.
Veröffentlicht: (2025)
Improving SAM Requires Rethinking its Optimization Formulation
von: Xie, Wanyun, et al.
Veröffentlicht: (2024)
von: Xie, Wanyun, et al.
Veröffentlicht: (2024)
Addressing Label Shift in Distributed Learning via Entropy Regularization
von: Wu, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Wu, Zhiyuan, et al.
Veröffentlicht: (2025)
Constrained Stochastic Spectral Preconditioning Converges for Nonconvex Objectives
von: Oikonomidis, Konstantinos, et al.
Veröffentlicht: (2026)
von: Oikonomidis, Konstantinos, et al.
Veröffentlicht: (2026)
Universal Gradient Methods for Stochastic Convex Optimization
von: Rodomanov, Anton, et al.
Veröffentlicht: (2024)
von: Rodomanov, Anton, et al.
Veröffentlicht: (2024)
Training Neural Networks at Any Scale
von: Pethick, Thomas, et al.
Veröffentlicht: (2025)
von: Pethick, Thomas, et al.
Veröffentlicht: (2025)
Training Deep Learning Models with Norm-Constrained LMOs
von: Pethick, Thomas, et al.
Veröffentlicht: (2025)
von: Pethick, Thomas, et al.
Veröffentlicht: (2025)
Principled Operator Learning in Ocean Dynamics: The Role of Temporal Structure
von: Jahanmard, Vahidreza, et al.
Veröffentlicht: (2025)
von: Jahanmard, Vahidreza, et al.
Veröffentlicht: (2025)
Stochastic Gradient Descent with Momentum is Algorithmically Stable
von: Lei, Yunwen, et al.
Veröffentlicht: (2026)
von: Lei, Yunwen, et al.
Veröffentlicht: (2026)
Black-Box Detection of LLM-Generated Text Using Generalized Jensen-Shannon Divergence
von: Chen, Shuangyi, et al.
Veröffentlicht: (2025)
von: Chen, Shuangyi, et al.
Veröffentlicht: (2025)
Momentum Does Not Reduce Stochastic Noise in Stochastic Gradient Descent
von: Sato, Naoki, et al.
Veröffentlicht: (2024)
von: Sato, Naoki, et al.
Veröffentlicht: (2024)
Stochastic Normalized Gradient Descent with Momentum for Large-Batch Training
von: Zhao, Shen-Yi, et al.
Veröffentlicht: (2020)
von: Zhao, Shen-Yi, et al.
Veröffentlicht: (2020)
Aligning Attention with Human Rationales for Self-Explaining Hate Speech Detection
von: Eilertsen, Brage, et al.
Veröffentlicht: (2025)
von: Eilertsen, Brage, et al.
Veröffentlicht: (2025)
Increasing Batch Size Improves Convergence of Stochastic Gradient Descent with Momentum
von: Kamo, Keisuke, et al.
Veröffentlicht: (2025)
von: Kamo, Keisuke, et al.
Veröffentlicht: (2025)
On the Role of Batch Size in Stochastic Conditional Gradient Methods
von: Islamov, Rustem, et al.
Veröffentlicht: (2026)
von: Islamov, Rustem, et al.
Veröffentlicht: (2026)
Robust Gradient Descent via Heavy-Ball Momentum with Predictive Extrapolation
von: Ali, Sarwan
Veröffentlicht: (2025)
von: Ali, Sarwan
Veröffentlicht: (2025)
Towards Self-Explainable Document Visual Question Answering with Chain-of-Explanation Predictions
von: Indrehus, Kjetil, et al.
Veröffentlicht: (2026)
von: Indrehus, Kjetil, et al.
Veröffentlicht: (2026)
Channel Simulation and Distributed Compression with Ensemble Rejection Sampling
von: Phan, Buu, et al.
Veröffentlicht: (2025)
von: Phan, Buu, et al.
Veröffentlicht: (2025)
Continuous-Time Analysis of Heavy Ball Momentum in Min-Max Games
von: Feng, Yi, et al.
Veröffentlicht: (2025)
von: Feng, Yi, et al.
Veröffentlicht: (2025)
First and Second Order Approximations to Stochastic Gradient Descent Methods with Momentum Terms
von: Lu, Eric
Veröffentlicht: (2025)
von: Lu, Eric
Veröffentlicht: (2025)
Policy Mirror Descent with Lookahead
von: Protopapas, Kimon, et al.
Veröffentlicht: (2024)
von: Protopapas, Kimon, et al.
Veröffentlicht: (2024)
Algorithmic Stability of Stochastic Gradient Descent with Momentum under Heavy-Tailed Noise
von: Dang, Thanh, et al.
Veröffentlicht: (2025)
von: Dang, Thanh, et al.
Veröffentlicht: (2025)
List-Level Distribution Coupling with Applications to Speculative Decoding and Lossy Compression
von: Rowan, Joseph, et al.
Veröffentlicht: (2025)
von: Rowan, Joseph, et al.
Veröffentlicht: (2025)
Random Cycle Coding: Lossless Compression of Cluster Assignments via Bits-Back Coding
von: Severo, Daniel, et al.
Veröffentlicht: (2024)
von: Severo, Daniel, et al.
Veröffentlicht: (2024)
Learning with Norm Constrained, Over-parameterized, Two-layer Neural Networks
von: Liu, Fanghui, et al.
Veröffentlicht: (2024)
von: Liu, Fanghui, et al.
Veröffentlicht: (2024)
SAMPa: Sharpness-aware Minimization Parallelized
von: Xie, Wanyun, et al.
Veröffentlicht: (2024)
von: Xie, Wanyun, et al.
Veröffentlicht: (2024)
MaD-Mix: Multi-Modal Data Mixtures via Latent Space Coupling for Vision-Language Model Training
von: Xie, Wanyun, et al.
Veröffentlicht: (2026)
von: Xie, Wanyun, et al.
Veröffentlicht: (2026)
IL-SOAR : Imitation Learning with Soft Optimistic Actor cRitic
von: Viel, Stefano, et al.
Veröffentlicht: (2025)
von: Viel, Stefano, et al.
Veröffentlicht: (2025)
Imitation Learning in Discounted Linear MDPs without exploration assumptions
von: Viano, Luca, et al.
Veröffentlicht: (2024)
von: Viano, Luca, et al.
Veröffentlicht: (2024)
MADA: Meta-Adaptive Optimizers through hyper-gradient Descent
von: Ozkara, Kaan, et al.
Veröffentlicht: (2024)
von: Ozkara, Kaan, et al.
Veröffentlicht: (2024)
Cross-Tokenizer Likelihood Scoring Algorithms for Language Model Distillation
von: Phan, Buu, et al.
Veröffentlicht: (2025)
von: Phan, Buu, et al.
Veröffentlicht: (2025)
Stochastic Adaptive Gradient Descent Without Descent
von: Aujol, Jean-François, et al.
Veröffentlicht: (2025)
von: Aujol, Jean-François, et al.
Veröffentlicht: (2025)
Generalization Bounds of Stochastic Gradient Descent in Homogeneous Neural Networks
von: Ma, Wenquan, et al.
Veröffentlicht: (2026)
von: Ma, Wenquan, et al.
Veröffentlicht: (2026)
Chameleon: A Flexible Data-mixing Framework for Language Model Pretraining and Finetuning
von: Xie, Wanyun, et al.
Veröffentlicht: (2025)
von: Xie, Wanyun, et al.
Veröffentlicht: (2025)
Efficient Large Language Model Inference with Neural Block Linearization
von: Erdogan, Mete, et al.
Veröffentlicht: (2025)
von: Erdogan, Mete, et al.
Veröffentlicht: (2025)
Stable Nonconvex-Nonconcave Training via Linear Interpolation
von: Pethick, Thomas, et al.
Veröffentlicht: (2023)
von: Pethick, Thomas, et al.
Veröffentlicht: (2023)
Improving Energy Natural Gradient Descent through Woodbury, Momentum, and Randomization
von: Guzmán-Cordero, Andrés, et al.
Veröffentlicht: (2025)
von: Guzmán-Cordero, Andrés, et al.
Veröffentlicht: (2025)
Beyond Accuracy: Decomposing the Reasoning Efficiency of LLMs
von: Kaiser, Daniel, et al.
Veröffentlicht: (2026)
von: Kaiser, Daniel, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Layer-wise Quantization for Quantized Optimistic Dual Averaging
von: Nguyen, Anh Duc, et al.
Veröffentlicht: (2025) -
Multi-Step Alignment as Markov Games: An Optimistic Online Gradient Descent Approach with Convergence Guarantees
von: Wu, Yongtao, et al.
Veröffentlicht: (2025) -
Generalized Gradient Norm Clipping & Non-Euclidean $(L_0,L_1)$-Smoothness
von: Pethick, Thomas, et al.
Veröffentlicht: (2025) -
Improving SAM Requires Rethinking its Optimization Formulation
von: Xie, Wanyun, et al.
Veröffentlicht: (2024) -
Addressing Label Shift in Distributed Learning via Entropy Regularization
von: Wu, Zhiyuan, et al.
Veröffentlicht: (2025)