Small steps no more: Global convergence of stochastic gradient bandits for arbitrary learning rates
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mei, Jincheng, Dai, Bo, Agarwal, Alekh, Vaswani, Sharan, Raj, Anant, Szepesvari, Csaba, Schuurmans, Dale |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Ordering-based Conditions for Global Convergence of Policy Gradient Methods
von: Mei, Jincheng, et al.
Veröffentlicht: (2025)
von: Mei, Jincheng, et al.
Veröffentlicht: (2025)
Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation
von: Lin, Max Qiushi, et al.
Veröffentlicht: (2025)
von: Lin, Max Qiushi, et al.
Veröffentlicht: (2025)
Stochastic Gradient Succeeds for Bandits
von: Mei, Jincheng, et al.
Veröffentlicht: (2024)
von: Mei, Jincheng, et al.
Veröffentlicht: (2024)
Rectifying Regression in Reinforcement Learning
von: Ayoub, Alex, et al.
Veröffentlicht: (2025)
von: Ayoub, Alex, et al.
Veröffentlicht: (2025)
From Inverse Optimization to Feasibility to ERM
von: Mishra, Saurabh, et al.
Veröffentlicht: (2024)
von: Mishra, Saurabh, et al.
Veröffentlicht: (2024)
Ensemble sampling for linear bandits: small ensembles suffice
von: Janz, David, et al.
Veröffentlicht: (2023)
von: Janz, David, et al.
Veröffentlicht: (2023)
Towards Principled, Practical Policy Gradient for Bandits and Tabular MDPs
von: Lu, Michael, et al.
Veröffentlicht: (2024)
von: Lu, Michael, et al.
Veröffentlicht: (2024)
Learning to Reason Efficiently with Discounted Reinforcement Learning
von: Ayoub, Alex, et al.
Veröffentlicht: (2025)
von: Ayoub, Alex, et al.
Veröffentlicht: (2025)
Optimistic Actor-Critic with Parametric Policies for Linear Markov Decision Processes
von: Lin, Max Qiushi, et al.
Veröffentlicht: (2026)
von: Lin, Max Qiushi, et al.
Veröffentlicht: (2026)
Beyond Expectations: Learning with Stochastic Dominance Made Practical
von: Cen, Shicong, et al.
Veröffentlicht: (2024)
von: Cen, Shicong, et al.
Veröffentlicht: (2024)
Balancing optimism and pessimism in offline-to-online learning
von: Sentenac, Flore, et al.
Veröffentlicht: (2025)
von: Sentenac, Flore, et al.
Veröffentlicht: (2025)
Spectral Ghost in Representation Learning: from Component Analysis to Self-Supervised Learning
von: Dai, Bo, et al.
Veröffentlicht: (2026)
von: Dai, Bo, et al.
Veröffentlicht: (2026)
Glocal Smoothness: Line search and adaptive step sizes can help in theory too!
von: Fox, Curtis, et al.
Veröffentlicht: (2025)
von: Fox, Curtis, et al.
Veröffentlicht: (2025)
Sharper Guarantees for Misspecified Kernelized Bandit Optimization
von: Maran, Davide, et al.
Veröffentlicht: (2026)
von: Maran, Davide, et al.
Veröffentlicht: (2026)
Faster WIND: Accelerating Iterative Best-of-$N$ Distillation for LLM Alignment
von: Yang, Tong, et al.
Veröffentlicht: (2024)
von: Yang, Tong, et al.
Veröffentlicht: (2024)
Armijo Line-search Can Make (Stochastic) Gradient Descent Provably Faster
von: Vaswani, Sharan, et al.
Veröffentlicht: (2025)
von: Vaswani, Sharan, et al.
Veröffentlicht: (2025)
Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF
von: Cen, Shicong, et al.
Veröffentlicht: (2024)
von: Cen, Shicong, et al.
Veröffentlicht: (2024)
Trajectory Data Suffices for Statistically Efficient Learning in Offline RL with Linear $q^π$-Realizability and Concentrability
von: Tkachuk, Volodymyr, et al.
Veröffentlicht: (2024)
von: Tkachuk, Volodymyr, et al.
Veröffentlicht: (2024)
Sharp analysis of linear ensemble sampling
von: Akhavan, Arya, et al.
Veröffentlicht: (2026)
von: Akhavan, Arya, et al.
Veröffentlicht: (2026)
Autoregressive Large Language Models are Computationally Universal
von: Schuurmans, Dale, et al.
Veröffentlicht: (2024)
von: Schuurmans, Dale, et al.
Veröffentlicht: (2024)
Almost sure convergence rates of stochastic gradient methods under gradient domination
von: Weissmann, Simon, et al.
Veröffentlicht: (2024)
von: Weissmann, Simon, et al.
Veröffentlicht: (2024)
Dissecting Discrete Soft Actor-Critic: Limitations and Principled Alternatives
von: Asad, Reza, et al.
Veröffentlicht: (2025)
von: Asad, Reza, et al.
Veröffentlicht: (2025)
(Accelerated) Noise-adaptive Stochastic Heavy-Ball Momentum
von: Dang, Anh, et al.
Veröffentlicht: (2024)
von: Dang, Anh, et al.
Veröffentlicht: (2024)
Convergence of Steepest Descent and Adam under Non-Uniform Smoothness
von: Vaswani, Sharan, et al.
Veröffentlicht: (2026)
von: Vaswani, Sharan, et al.
Veröffentlicht: (2026)
Non-Linear Reinforcement Learning in Large Action Spaces: Structural Conditions and Sample-efficiency of Posterior Sampling
von: Agarwal, Alekh, et al.
Veröffentlicht: (2022)
von: Agarwal, Alekh, et al.
Veröffentlicht: (2022)
A proximal-gradient inertial algorithm with Tikhonov regularization: strong convergence to the minimal norm solution
von: László, Szilárd Csaba
Veröffentlicht: (2024)
von: László, Szilárd Csaba
Veröffentlicht: (2024)
Confident Natural Policy Gradient for Local Planning in $q_π$-realizable Constrained MDPs
von: Tian, Tian, et al.
Veröffentlicht: (2024)
von: Tian, Tian, et al.
Veröffentlicht: (2024)
Exponential convergence rates for momentum stochastic gradient descent in the overparametrized setting
von: Gess, Benjamin, et al.
Veröffentlicht: (2023)
von: Gess, Benjamin, et al.
Veröffentlicht: (2023)
Sample Complexity Bounds for Linear Constrained MDPs with a Generative Model
von: Liu, Xingtu, et al.
Veröffentlicht: (2025)
von: Liu, Xingtu, et al.
Veröffentlicht: (2025)
Towards Noise-adaptive, Problem-adaptive (Accelerated) Stochastic Gradient Descent
von: Vaswani, Sharan, et al.
Veröffentlicht: (2021)
von: Vaswani, Sharan, et al.
Veröffentlicht: (2021)
Spectral Representation-based Reinforcement Learning
von: Gao, Chenxiao, et al.
Veröffentlicht: (2025)
von: Gao, Chenxiao, et al.
Veröffentlicht: (2025)
Provable Representation with Efficient Planning for Partial Observable Reinforcement Learning
von: Zhang, Hongming, et al.
Veröffentlicht: (2023)
von: Zhang, Hongming, et al.
Veröffentlicht: (2023)
Non-convergence of Adam and other adaptive stochastic gradient descent optimization methods for non-vanishing learning rates
von: Dereich, Steffen, et al.
Veröffentlicht: (2024)
von: Dereich, Steffen, et al.
Veröffentlicht: (2024)
EM++: A parameter learning framework for stochastic switching systems
von: Wang, Renzi, et al.
Veröffentlicht: (2024)
von: Wang, Renzi, et al.
Veröffentlicht: (2024)
Target Networks and Over-parameterization Stabilize Off-policy Bootstrapping with Function Approximation
von: Che, Fengdi, et al.
Veröffentlicht: (2024)
von: Che, Fengdi, et al.
Veröffentlicht: (2024)
Exploration via linearly perturbed loss minimisation
von: Janz, David, et al.
Veröffentlicht: (2023)
von: Janz, David, et al.
Veröffentlicht: (2023)
Beyond Statistical Learning: Exact Learning Is Essential for General Intelligence
von: György, András, et al.
Veröffentlicht: (2025)
von: György, András, et al.
Veröffentlicht: (2025)
To See the Unseen: on the Generalization Ability of Transformers in Symbolic Reasoning
von: Lazić, Nevena, et al.
Veröffentlicht: (2026)
von: Lazić, Nevena, et al.
Veröffentlicht: (2026)
Strong convergence and fast rates for systems with Tikhonov regularization
von: Csetnek, Ernö Robert, et al.
Veröffentlicht: (2024)
von: Csetnek, Ernö Robert, et al.
Veröffentlicht: (2024)
Towards Parameter-Free Temporal Difference Learning
von: Li, Yunxiang, et al.
Veröffentlicht: (2026)
von: Li, Yunxiang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Ordering-based Conditions for Global Convergence of Policy Gradient Methods
von: Mei, Jincheng, et al.
Veröffentlicht: (2025) -
Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation
von: Lin, Max Qiushi, et al.
Veröffentlicht: (2025) -
Stochastic Gradient Succeeds for Bandits
von: Mei, Jincheng, et al.
Veröffentlicht: (2024) -
Rectifying Regression in Reinforcement Learning
von: Ayoub, Alex, et al.
Veröffentlicht: (2025) -
From Inverse Optimization to Feasibility to ERM
von: Mishra, Saurabh, et al.
Veröffentlicht: (2024)