Similar Items
In-context Learning and Gradient Descent Revisited
by: Deutch, Gilad, et al.
Published: (2023)
by: Deutch, Gilad, et al.
Published: (2023)
Learning Provably Improves the Convergence of Gradient Descent
by: Song, Qingyu, et al.
Published: (2025)
by: Song, Qingyu, et al.
Published: (2025)
AWP: Activation-Aware Weight Pruning and Quantization with Projected Gradient Descent
by: Liu, Jing, et al.
Published: (2025)
by: Liu, Jing, et al.
Published: (2025)
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization
by: Abuduweili, Abulikemu, et al.
Published: (2024)
by: Abuduweili, Abulikemu, et al.
Published: (2024)
Weighted Averaged Stochastic Gradient Descent: Asymptotic Normality and Optimality
by: Wei, Ziyang, et al.
Published: (2023)
by: Wei, Ziyang, et al.
Published: (2023)
PMGDA: A Preference-based Multiple Gradient Descent Algorithm
by: Zhang, Xiaoyuan, et al.
Published: (2024)
by: Zhang, Xiaoyuan, et al.
Published: (2024)
Dual Space Preconditioning for Gradient Descent in the Overparameterized Regime
by: Ghane, Reza, et al.
Published: (2026)
by: Ghane, Reza, et al.
Published: (2026)
Randomness and Interpolation Improve Gradient Descent
by: Li, Jiawen, et al.
Published: (2025)
by: Li, Jiawen, et al.
Published: (2025)
An Improved Empirical Fisher Approximation for Natural Gradient Descent
by: Wu, Xiaodong, et al.
Published: (2024)
by: Wu, Xiaodong, et al.
Published: (2024)
Stochastic Gradient Descent Revisited
by: Louzi, Azar
Published: (2024)
by: Louzi, Azar
Published: (2024)
Distributed Gradient Descent for Functional Learning
by: Yu, Zhan, et al.
Published: (2023)
by: Yu, Zhan, et al.
Published: (2023)
On Penalty-based Bilevel Gradient Descent Method
by: Shen, Han, et al.
Published: (2023)
by: Shen, Han, et al.
Published: (2023)
Learning Tree-Based Models with Gradient Descent
by: Marton, Sascha
Published: (2026)
by: Marton, Sascha
Published: (2026)
Occam Gradient Descent
by: Kausik, B. N.
Published: (2024)
by: Kausik, B. N.
Published: (2024)
Improving Energy Natural Gradient Descent through Woodbury, Momentum, and Randomization
by: Guzmán-Cordero, Andrés, et al.
Published: (2025)
by: Guzmán-Cordero, Andrés, et al.
Published: (2025)
Increasing Batch Size Improves Convergence of Stochastic Gradient Descent with Momentum
by: Kamo, Keisuke, et al.
Published: (2025)
by: Kamo, Keisuke, et al.
Published: (2025)
Gradient Descent Algorithm Survey
by: Fucheng, Deng, et al.
Published: (2025)
by: Fucheng, Deng, et al.
Published: (2025)
Learning Associative Memories with Gradient Descent
by: Cabannes, Vivien, et al.
Published: (2024)
by: Cabannes, Vivien, et al.
Published: (2024)
Revisiting Random Weight Perturbation for Efficiently Improving Generalization
by: Li, Tao, et al.
Published: (2024)
by: Li, Tao, et al.
Published: (2024)
Inverse-Free Fast Natural Gradient Descent Method for Deep Learning
by: Ou, Xinwei, et al.
Published: (2024)
by: Ou, Xinwei, et al.
Published: (2024)
Gradient Descent with Provably Tuned Learning-rate Schedules
by: Sharma, Dravyansh
Published: (2025)
by: Sharma, Dravyansh
Published: (2025)
Learning Curves of Stochastic Gradient Descent in Kernel Regression
by: Zhang, Haihan, et al.
Published: (2025)
by: Zhang, Haihan, et al.
Published: (2025)
Personalized Federated Learning with Exact Stochastic Gradient Descent
by: Nikoloutsopoulos, Sotirios, et al.
Published: (2022)
by: Nikoloutsopoulos, Sotirios, et al.
Published: (2022)
Towards Learning Stochastic Population Models by Gradient Descent
by: Kreikemeyer, Justin N., et al.
Published: (2024)
by: Kreikemeyer, Justin N., et al.
Published: (2024)
Partially Lazy Gradient Descent for Smoothed Online Learning
by: Mhaisen, Naram, et al.
Published: (2026)
by: Mhaisen, Naram, et al.
Published: (2026)
Revisiting Stochastic Approximation and Stochastic Gradient Descent
by: Karandikar, Rajeeva Laxman, et al.
Published: (2025)
by: Karandikar, Rajeeva Laxman, et al.
Published: (2025)
Online Statistical Inference for Contextual Bandits via Stochastic Gradient Descent
by: Chang, Xiangyu, et al.
Published: (2022)
by: Chang, Xiangyu, et al.
Published: (2022)
Revisiting Gradient Pruning: A Dual Realization for Defending against Gradient Attacks
by: Xue, Lulu, et al.
Published: (2024)
by: Xue, Lulu, et al.
Published: (2024)
Low-Tubal-Rank Tensor Recovery via Factorized Gradient Descent
by: Liu, Zhiyu, et al.
Published: (2024)
by: Liu, Zhiyu, et al.
Published: (2024)
Hybrid Coordinate Descent for Efficient Neural Network Learning Using Line Search and Gradient Descent
by: Hsiao, Yen-Che, et al.
Published: (2024)
by: Hsiao, Yen-Che, et al.
Published: (2024)
Stacking as Accelerated Gradient Descent
by: Agarwal, Naman, et al.
Published: (2024)
by: Agarwal, Naman, et al.
Published: (2024)
Quadratic Gradient: A Unified Framework Bridging Gradient Descent and Newton-Type Methods by Synthesizing Hessians and Gradients
by: Chiang, John
Published: (2022)
by: Chiang, John
Published: (2022)
Adjacent Leader Decentralized Stochastic Gradient Descent
by: He, Haoze, et al.
Published: (2024)
by: He, Haoze, et al.
Published: (2024)
Elastic Multi-Gradient Descent for Parallel Continual Learning
by: Lyu, Fan, et al.
Published: (2024)
by: Lyu, Fan, et al.
Published: (2024)
Weighted Low-rank Approximation via Stochastic Gradient Descent on Manifolds
by: Xu, Conglong, et al.
Published: (2025)
by: Xu, Conglong, et al.
Published: (2025)
Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs
by: Cheng, Wenhua, et al.
Published: (2023)
by: Cheng, Wenhua, et al.
Published: (2023)
Stochastic Adaptive Gradient Descent Without Descent
by: Aujol, Jean-François, et al.
Published: (2025)
by: Aujol, Jean-François, et al.
Published: (2025)
A Theoretical Analysis of Noise Geometry in Stochastic Gradient Descent
by: Wang, Mingze, et al.
Published: (2023)
by: Wang, Mingze, et al.
Published: (2023)
Enhancing Fractional Gradient Descent with Learned Optimizers
by: Sobotka, Jan, et al.
Published: (2025)
by: Sobotka, Jan, et al.
Published: (2025)
Corner Gradient Descent
by: Yarotsky, Dmitry
Published: (2025)
by: Yarotsky, Dmitry
Published: (2025)
Similar Items
-
In-context Learning and Gradient Descent Revisited
by: Deutch, Gilad, et al.
Published: (2023) -
Learning Provably Improves the Convergence of Gradient Descent
by: Song, Qingyu, et al.
Published: (2025) -
AWP: Activation-Aware Weight Pruning and Quantization with Projected Gradient Descent
by: Liu, Jing, et al.
Published: (2025) -
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization
by: Abuduweili, Abulikemu, et al.
Published: (2024) -
Weighted Averaged Stochastic Gradient Descent: Asymptotic Normality and Optimality
by: Wei, Ziyang, et al.
Published: (2023)