Large Stepsize Gradient Descent for Logistic Loss: Non-Monotonicity of the Loss Improves Optimization Efficiency
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Jingfeng, Bartlett, Peter L., Telgarsky, Matus, Yu, Bin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Benefits of Early Stopping in Gradient Descent for Overparameterized Logistic Regression
by: Wu, Jingfeng, et al.
Published: (2025)
by: Wu, Jingfeng, et al.
Published: (2025)
Large Stepsizes Accelerate Gradient Descent for Regularized Logistic Regression
by: Wu, Jingfeng, et al.
Published: (2025)
by: Wu, Jingfeng, et al.
Published: (2025)
Minimax Optimal Convergence of Gradient Descent in Logistic Regression via Large and Adaptive Stepsizes
by: Zhang, Ruiqi, et al.
Published: (2025)
by: Zhang, Ruiqi, et al.
Published: (2025)
Large Stepsize Gradient Descent for Non-Homogeneous Two-Layer Networks: Margin Improvement and Fast Optimization
by: Cai, Yuhang, et al.
Published: (2024)
by: Cai, Yuhang, et al.
Published: (2024)
Tight Bounds for Logistic Regression with Large Stepsize Gradient Descent in Low Dimension
by: Crawshaw, Michael, et al.
Published: (2026)
by: Crawshaw, Michael, et al.
Published: (2026)
Implicit Bias of Gradient Descent for Non-Homogeneous Deep Networks
by: Cai, Yuhang, et al.
Published: (2025)
by: Cai, Yuhang, et al.
Published: (2025)
On Achieving Optimal Adversarial Test Error
by: Li, Justin D., et al.
Published: (2023)
by: Li, Justin D., et al.
Published: (2023)
Transformers, parallel computation, and logarithmic depth
by: Sanford, Clayton, et al.
Published: (2024)
by: Sanford, Clayton, et al.
Published: (2024)
One-layer transformers fail to solve the induction heads task
by: Sanford, Clayton, et al.
Published: (2024)
by: Sanford, Clayton, et al.
Published: (2024)
Improved Scaling Laws in Linear Regression via Data Reuse
by: Lin, Licong, et al.
Published: (2025)
by: Lin, Licong, et al.
Published: (2025)
Gradient Descent on Logistic Regression with Non-Separable Data and Large Step Sizes
by: Meng, Si Yi, et al.
Published: (2024)
by: Meng, Si Yi, et al.
Published: (2024)
Negative Stepsizes Make Gradient-Descent-Ascent Converge
by: Shugart, Henry, et al.
Published: (2025)
by: Shugart, Henry, et al.
Published: (2025)
Coupling-based Convergence Diagnostic and Stepsize Scheme for Stochastic Gradient Descent
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
Astral Space: Convex Analysis at Infinity
by: Dudík, Miroslav, et al.
Published: (2022)
by: Dudík, Miroslav, et al.
Published: (2022)
Conformal Risk Control for Non-Monotonic Losses
by: Angelopoulos, Anastasios N.
Published: (2026)
by: Angelopoulos, Anastasios N.
Published: (2026)
Multiclass Loss Geometry Matters for Generalization of Gradient Descent in Separable Classification
by: Schliserman, Matan, et al.
Published: (2025)
by: Schliserman, Matan, et al.
Published: (2025)
Lai Loss: A Novel Loss for Gradient Control
by: Lai, YuFei
Published: (2024)
by: Lai, YuFei
Published: (2024)
Spectrum Extraction and Clipping for Implicitly Linear Layers
by: Boroojeny, Ali Ebrahimpour, et al.
Published: (2024)
by: Boroojeny, Ali Ebrahimpour, et al.
Published: (2024)
Online Structured Prediction with Fenchel--Young Losses and Improved Surrogate Regret for Online Multiclass Classification with Logistic Loss
by: Sakaue, Shinsaku, et al.
Published: (2024)
by: Sakaue, Shinsaku, et al.
Published: (2024)
In-Context Learning of a Linear Transformer Block: Benefits of the MLP Component and One-Step GD Initialization
by: Zhang, Ruiqi, et al.
Published: (2024)
by: Zhang, Ruiqi, et al.
Published: (2024)
Basic Inequalities for First-Order Optimization with Applications to Statistical Risk Analysis
by: Paik, Seunghoon, et al.
Published: (2025)
by: Paik, Seunghoon, et al.
Published: (2025)
Benign Overfitting without Linearity: Neural Network Classifiers Trained by Gradient Descent for Noisy Linear Data
by: Frei, Spencer, et al.
Published: (2022)
by: Frei, Spencer, et al.
Published: (2022)
Any-stepsize Gradient Descent for Separable Data under Fenchel-Young Losses
by: Bao, Han, et al.
Published: (2025)
by: Bao, Han, et al.
Published: (2025)
Constant Stepsize Local GD for Logistic Regression: Acceleration by Instability
by: Crawshaw, Michael, et al.
Published: (2025)
by: Crawshaw, Michael, et al.
Published: (2025)
From Logistic Regression to the Perceptron Algorithm: Exploring Gradient Descent with Large Step Sizes
by: Tyurin, Alexander
Published: (2024)
by: Tyurin, Alexander
Published: (2024)
Risk Comparisons in Linear Regression: Implicit Regularization Dominates Explicit Regularization
by: Wu, Jingfeng, et al.
Published: (2025)
by: Wu, Jingfeng, et al.
Published: (2025)
Quantitative Convergence Analysis of Projected Stochastic Gradient Descent for Non-Convex Losses via the Goldstein Subdifferential
by: Zheng, Yuping, et al.
Published: (2025)
by: Zheng, Yuping, et al.
Published: (2025)
Classification with Deep Neural Networks and Logistic Loss
by: Zhang, Zihan, et al.
Published: (2023)
by: Zhang, Zihan, et al.
Published: (2023)
AYLA: Amplifying Gradient Sensitivity via Loss Transformation in Non-Convex Optimization
by: Keslaki, Ben
Published: (2025)
by: Keslaki, Ben
Published: (2025)
Gradient Descent as Loss Landscape Navigation: a Normative Framework for Deriving Learning Rules
by: Vastola, John J., et al.
Published: (2025)
by: Vastola, John J., et al.
Published: (2025)
Modified Loss of Momentum Gradient Descent: Fine-Grained Analysis
by: Cattaneo, Matias D., et al.
Published: (2025)
by: Cattaneo, Matias D., et al.
Published: (2025)
An Improved Empirical Fisher Approximation for Natural Gradient Descent
by: Wu, Xiaodong, et al.
Published: (2024)
by: Wu, Xiaodong, et al.
Published: (2024)
The Space Complexity of Approximating Logistic Loss
by: Dexter, Gregory, et al.
Published: (2024)
by: Dexter, Gregory, et al.
Published: (2024)
Conformal Risk Control under Non-Monotone Losses: Theory and Finite-Sample Guarantees
by: Aldirawi, Tareq, et al.
Published: (2026)
by: Aldirawi, Tareq, et al.
Published: (2026)
Gradient Descent on Logistic Regression: Do Large Step-Sizes Work with Data on the Sphere?
by: Meng, Si Yi, et al.
Published: (2025)
by: Meng, Si Yi, et al.
Published: (2025)
MARINA-P: Superior Performance in Non-smooth Federated Optimization with Adaptive Stepsizes
by: Sokolov, Igor, et al.
Published: (2024)
by: Sokolov, Igor, et al.
Published: (2024)
Exponential Convergence of (Stochastic) Gradient Descent for Separable Logistic Regression
by: Kale, Sacchit, et al.
Published: (2026)
by: Kale, Sacchit, et al.
Published: (2026)
An Adaptive Stochastic Gradient Method with Non-negative Gauss-Newton Stepsizes
by: Orvieto, Antonio, et al.
Published: (2024)
by: Orvieto, Antonio, et al.
Published: (2024)
Transformers Efficiently Perform In-Context Logistic Regression via Normalized Gradient Descent
by: Zhang, Chenyang, et al.
Published: (2026)
by: Zhang, Chenyang, et al.
Published: (2026)
Reinforcement Learning in POMDP's via Direct Gradient Ascent
by: Baxter, Jonathan, et al.
Published: (2025)
by: Baxter, Jonathan, et al.
Published: (2025)
Similar Items
-
Benefits of Early Stopping in Gradient Descent for Overparameterized Logistic Regression
by: Wu, Jingfeng, et al.
Published: (2025) -
Large Stepsizes Accelerate Gradient Descent for Regularized Logistic Regression
by: Wu, Jingfeng, et al.
Published: (2025) -
Minimax Optimal Convergence of Gradient Descent in Logistic Regression via Large and Adaptive Stepsizes
by: Zhang, Ruiqi, et al.
Published: (2025) -
Large Stepsize Gradient Descent for Non-Homogeneous Two-Layer Networks: Margin Improvement and Fast Optimization
by: Cai, Yuhang, et al.
Published: (2024) -
Tight Bounds for Logistic Regression with Large Stepsize Gradient Descent in Low Dimension
by: Crawshaw, Michael, et al.
Published: (2026)