Generalization and Optimization of SGD with Lookahead
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Kangcheng, Lei, Yunwen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bootstrap SGD: Algorithmic Stability and Robustness
by: Christmann, Andreas, et al.
Published: (2024)
by: Christmann, Andreas, et al.
Published: (2024)
Minibatch and Local SGD: Algorithmic Stability and Linear Speedup in Generalization
by: Lei, Yunwen, et al.
Published: (2023)
by: Lei, Yunwen, et al.
Published: (2023)
Towards Initialization-dependent and Non-vacuous Generalization Bounds for Overparameterized Shallow Neural Networks
by: Lei, Yunwen, et al.
Published: (2026)
by: Lei, Yunwen, et al.
Published: (2026)
Generalization Bounds for Rank-sparse Neural Networks
by: Ledent, Antoine, et al.
Published: (2025)
by: Ledent, Antoine, et al.
Published: (2025)
Stability-based Generalization Analysis of Randomized Coordinate Descent for Pairwise Learning
by: Wu, Liang, et al.
Published: (2025)
by: Wu, Liang, et al.
Published: (2025)
Learning Theory of the SVRG: Generalization and Convergence Analysis
by: Lei, Yunwen, et al.
Published: (2026)
by: Lei, Yunwen, et al.
Published: (2026)
Lookahead Path Likelihood Optimization for Diffusion LLMs
by: Liu, Xuejie, et al.
Published: (2026)
by: Liu, Xuejie, et al.
Published: (2026)
Optimal Rates for Generalization of Gradient Descent for Deep ReLU Classification
by: Li, Yuanfan, et al.
Published: (2025)
by: Li, Yuanfan, et al.
Published: (2025)
Randomized Pairwise Learning with Adaptive Sampling: A PAC-Bayes Analysis
by: Zhou, Sijia, et al.
Published: (2025)
by: Zhou, Sijia, et al.
Published: (2025)
SGD with Partial Hessian for Deep Neural Networks Optimization
by: Sun, Ying, et al.
Published: (2024)
by: Sun, Ying, et al.
Published: (2024)
Generalization Analysis for Deep Contrastive Representation Learning
by: Hieu, Nong Minh, et al.
Published: (2024)
by: Hieu, Nong Minh, et al.
Published: (2024)
Next-Depth Lookahead Tree
by: Lee, Jaeho, et al.
Published: (2025)
by: Lee, Jaeho, et al.
Published: (2025)
Reinforcement Learning with Lookahead Information
by: Merlis, Nadav
Published: (2024)
by: Merlis, Nadav
Published: (2024)
EMA-Nesterov: Stabilizing Nesterov's Lookahead for Accelerated Deep Learning Optimization
by: Yau, Chung-Yiu, et al.
Published: (2026)
by: Yau, Chung-Yiu, et al.
Published: (2026)
Lookahead Counterfactual Fairness
by: Zuo, Zhiqun, et al.
Published: (2024)
by: Zuo, Zhiqun, et al.
Published: (2024)
Stability and Generalization for Decentralized Markov SGD
by: Wang, Jiahuan, et al.
Published: (2026)
by: Wang, Jiahuan, et al.
Published: (2026)
Stochastic Gradient Descent with Momentum is Algorithmically Stable
by: Lei, Yunwen, et al.
Published: (2026)
by: Lei, Yunwen, et al.
Published: (2026)
Generalization Guarantees of Gradient Descent for Multi-Layer Neural Networks
by: Wang, Puyu, et al.
Published: (2023)
by: Wang, Puyu, et al.
Published: (2023)
The Value of Reward Lookahead in Reinforcement Learning
by: Merlis, Nadav, et al.
Published: (2024)
by: Merlis, Nadav, et al.
Published: (2024)
Leveraging Flatness to Improve Information-Theoretic Generalization Bounds for SGD
by: Peng, Ze, et al.
Published: (2026)
by: Peng, Ze, et al.
Published: (2026)
Unveiling High-Probability Generalization in Decentralized SGD
by: Wang, Jiahuan, et al.
Published: (2026)
by: Wang, Jiahuan, et al.
Published: (2026)
On Discriminative Probabilistic Modeling for Self-Supervised Representation Learning
by: Wang, Bokun, et al.
Published: (2024)
by: Wang, Bokun, et al.
Published: (2024)
Causal Attention with Lookahead Keys
by: Song, Zhuoqing, et al.
Published: (2025)
by: Song, Zhuoqing, et al.
Published: (2025)
Policy Mirror Descent with Lookahead
by: Protopapas, Kimon, et al.
Published: (2024)
by: Protopapas, Kimon, et al.
Published: (2024)
Structured and Fast Optimization: The Kronecker SGD Algorithm
by: Song, Zhao, et al.
Published: (2023)
by: Song, Zhao, et al.
Published: (2023)
Topology-aware Generalization of Decentralized SGD
by: Zhu, Tongtian, et al.
Published: (2022)
by: Zhu, Tongtian, et al.
Published: (2022)
SLowcal-SGD: Slow Query Points Improve Local-SGD for Stochastic Convex Optimization
by: Dahan, Tehila, et al.
Published: (2023)
by: Dahan, Tehila, et al.
Published: (2023)
The Optimization Landscape of SGD Across the Feature Learning Strength
by: Atanasov, Alexander, et al.
Published: (2024)
by: Atanasov, Alexander, et al.
Published: (2024)
Scaling Speculative Decoding with Lookahead Reasoning
by: Fu, Yichao, et al.
Published: (2025)
by: Fu, Yichao, et al.
Published: (2025)
Lookahead identification in adversarial bandits: accuracy and memory bounds
by: Brukhim, Nataly, et al.
Published: (2026)
by: Brukhim, Nataly, et al.
Published: (2026)
Lookahead Drifting Model
by: Zhang, Guoqiang, et al.
Published: (2026)
by: Zhang, Guoqiang, et al.
Published: (2026)
Curvature-Informed SGD via General Purpose Lie-Group Preconditioners
by: Pooladzandi, Omead, et al.
Published: (2024)
by: Pooladzandi, Omead, et al.
Published: (2024)
EARL-BO: Reinforcement Learning for Multi-Step Lookahead, High-Dimensional Bayesian Optimization
by: Cheon, Mujin, et al.
Published: (2024)
by: Cheon, Mujin, et al.
Published: (2024)
Sign-SGD via Parameter-Free Optimization
by: Medyakov, Daniil, et al.
Published: (2025)
by: Medyakov, Daniil, et al.
Published: (2025)
Beyond Cross-Validation: Adaptive Parameter Selection for Kernel-Based Gradient Descents
by: Liu, Xiaotong, et al.
Published: (2026)
by: Liu, Xiaotong, et al.
Published: (2026)
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees
by: Xie, Shengping, et al.
Published: (2025)
by: Xie, Shengping, et al.
Published: (2025)
Understanding Lookahead Dynamics Through Laplace Transform
by: Sanyal, Aniket, et al.
Published: (2025)
by: Sanyal, Aniket, et al.
Published: (2025)
Thinking into the Future: Latent Lookahead Training for Transformers
by: Noci, Lorenzo, et al.
Published: (2026)
by: Noci, Lorenzo, et al.
Published: (2026)
Improved Stability and Generalization Guarantees of the Decentralized SGD Algorithm
by: Bars, Batiste Le, et al.
Published: (2023)
by: Bars, Batiste Le, et al.
Published: (2023)
Physics-aware deep learning framework for the limited aperture inverse obstacle scattering problem
by: Yin, Yunwen, et al.
Published: (2024)
by: Yin, Yunwen, et al.
Published: (2024)
Similar Items
-
Bootstrap SGD: Algorithmic Stability and Robustness
by: Christmann, Andreas, et al.
Published: (2024) -
Minibatch and Local SGD: Algorithmic Stability and Linear Speedup in Generalization
by: Lei, Yunwen, et al.
Published: (2023) -
Towards Initialization-dependent and Non-vacuous Generalization Bounds for Overparameterized Shallow Neural Networks
by: Lei, Yunwen, et al.
Published: (2026) -
Generalization Bounds for Rank-sparse Neural Networks
by: Ledent, Antoine, et al.
Published: (2025) -
Stability-based Generalization Analysis of Randomized Coordinate Descent for Pairwise Learning
by: Wu, Liang, et al.
Published: (2025)