Stochastic Gradient Succeeds for Bandits
Fuente:
arXiv
Guardado en:
| Autores principales: | Mei, Jincheng, Zhong, Zixin, Dai, Bo, Agarwal, Alekh, Szepesvari, Csaba, Schuurmans, Dale |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Ordering-based Conditions for Global Convergence of Policy Gradient Methods
por: Mei, Jincheng, et al.
Publicado: (2025)
por: Mei, Jincheng, et al.
Publicado: (2025)
Small steps no more: Global convergence of stochastic gradient bandits for arbitrary learning rates
por: Mei, Jincheng, et al.
Publicado: (2025)
por: Mei, Jincheng, et al.
Publicado: (2025)
Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation
por: Lin, Max Qiushi, et al.
Publicado: (2025)
por: Lin, Max Qiushi, et al.
Publicado: (2025)
Rectifying Regression in Reinforcement Learning
por: Ayoub, Alex, et al.
Publicado: (2025)
por: Ayoub, Alex, et al.
Publicado: (2025)
Learning to Reason Efficiently with Discounted Reinforcement Learning
por: Ayoub, Alex, et al.
Publicado: (2025)
por: Ayoub, Alex, et al.
Publicado: (2025)
Beyond Expectations: Learning with Stochastic Dominance Made Practical
por: Cen, Shicong, et al.
Publicado: (2024)
por: Cen, Shicong, et al.
Publicado: (2024)
Sharper Guarantees for Misspecified Kernelized Bandit Optimization
por: Maran, Davide, et al.
Publicado: (2026)
por: Maran, Davide, et al.
Publicado: (2026)
Faster WIND: Accelerating Iterative Best-of-$N$ Distillation for LLM Alignment
por: Yang, Tong, et al.
Publicado: (2024)
por: Yang, Tong, et al.
Publicado: (2024)
Spectral Ghost in Representation Learning: from Component Analysis to Self-Supervised Learning
por: Dai, Bo, et al.
Publicado: (2026)
por: Dai, Bo, et al.
Publicado: (2026)
Confident Natural Policy Gradient for Local Planning in $q_π$-realizable Constrained MDPs
por: Tian, Tian, et al.
Publicado: (2024)
por: Tian, Tian, et al.
Publicado: (2024)
Almost Free: Self-concordance in Natural Exponential Families and an Application to Bandits
por: Liu, Shuai, et al.
Publicado: (2024)
por: Liu, Shuai, et al.
Publicado: (2024)
Learning What to Recommend: Minimax Optimal Simple Regret in Logistic Bandits
por: Liu, Shuai, et al.
Publicado: (2026)
por: Liu, Shuai, et al.
Publicado: (2026)
Catoni Contextual Bandits are Robust to Heavy-tailed Rewards
por: Ye, Chenlu, et al.
Publicado: (2025)
por: Ye, Chenlu, et al.
Publicado: (2025)
Stochastic Gradient Descent for Gaussian Processes Done Right
por: Lin, Jihao Andreas, et al.
Publicado: (2023)
por: Lin, Jihao Andreas, et al.
Publicado: (2023)
Trajectory Data Suffices for Statistically Efficient Learning in Offline RL with Linear $q^π$-Realizability and Concentrability
por: Tkachuk, Volodymyr, et al.
Publicado: (2024)
por: Tkachuk, Volodymyr, et al.
Publicado: (2024)
Sharp analysis of linear ensemble sampling
por: Akhavan, Arya, et al.
Publicado: (2026)
por: Akhavan, Arya, et al.
Publicado: (2026)
Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF
por: Cen, Shicong, et al.
Publicado: (2024)
por: Cen, Shicong, et al.
Publicado: (2024)
Balancing optimism and pessimism in offline-to-online learning
por: Sentenac, Flore, et al.
Publicado: (2025)
por: Sentenac, Flore, et al.
Publicado: (2025)
Ensemble sampling for linear bandits: small ensembles suffice
por: Janz, David, et al.
Publicado: (2023)
por: Janz, David, et al.
Publicado: (2023)
Optimistic Policy Optimization is Provably Efficient in Non-stationary MDPs
por: Zhong, Han, et al.
Publicado: (2021)
por: Zhong, Han, et al.
Publicado: (2021)
Non-Linear Reinforcement Learning in Large Action Spaces: Structural Conditions and Sample-efficiency of Posterior Sampling
por: Agarwal, Alekh, et al.
Publicado: (2022)
por: Agarwal, Alekh, et al.
Publicado: (2022)
Exploration via linearly perturbed loss minimisation
por: Janz, David, et al.
Publicado: (2023)
por: Janz, David, et al.
Publicado: (2023)
Delightful Gradients Accelerate Corner Escape
por: Mei, Jincheng, et al.
Publicado: (2026)
por: Mei, Jincheng, et al.
Publicado: (2026)
Spectral Representation-based Reinforcement Learning
por: Gao, Chenxiao, et al.
Publicado: (2025)
por: Gao, Chenxiao, et al.
Publicado: (2025)
Provable Representation with Efficient Planning for Partial Observable Reinforcement Learning
por: Zhang, Hongming, et al.
Publicado: (2023)
por: Zhang, Hongming, et al.
Publicado: (2023)
Target Networks and Over-parameterization Stabilize Off-policy Bootstrapping with Function Approximation
por: Che, Fengdi, et al.
Publicado: (2024)
por: Che, Fengdi, et al.
Publicado: (2024)
Beyond Statistical Learning: Exact Learning Is Essential for General Intelligence
por: György, András, et al.
Publicado: (2025)
por: György, András, et al.
Publicado: (2025)
Eluder dimension: localise it!
por: Bakhtiari, Alireza, et al.
Publicado: (2026)
por: Bakhtiari, Alireza, et al.
Publicado: (2026)
Optimal Clustering with Bandit Feedback
por: Yang, Junwen, et al.
Publicado: (2022)
por: Yang, Junwen, et al.
Publicado: (2022)
Plastic Learning with Deep Fourier Features
por: Lewandowski, Alex, et al.
Publicado: (2024)
por: Lewandowski, Alex, et al.
Publicado: (2024)
Regret Minimization via Saddle Point Optimization
por: Kirschner, Johannes, et al.
Publicado: (2024)
por: Kirschner, Johannes, et al.
Publicado: (2024)
Offline Imitation Learning from Multiple Baselines with Applications to Compiler Optimization
por: Marinov, Teodor V., et al.
Publicado: (2024)
por: Marinov, Teodor V., et al.
Publicado: (2024)
Design Considerations in Offline Preference-based RL
por: Agarwal, Alekh, et al.
Publicado: (2025)
por: Agarwal, Alekh, et al.
Publicado: (2025)
To Believe or Not to Believe Your LLM
por: Yadkori, Yasin Abbasi, et al.
Publicado: (2024)
por: Yadkori, Yasin Abbasi, et al.
Publicado: (2024)
A Lyapunov Analysis of Softmax Policy Gradient for Stochastic Bandits
por: Lattimore, Tor
Publicado: (2026)
por: Lattimore, Tor
Publicado: (2026)
Semi-Bandit Learning for Monotone Stochastic Optimization
por: Agarwal, Arpit, et al.
Publicado: (2023)
por: Agarwal, Arpit, et al.
Publicado: (2023)
Toward Understanding In-context vs. In-weight Learning
por: Chan, Bryan, et al.
Publicado: (2024)
por: Chan, Bryan, et al.
Publicado: (2024)
Online Statistical Inference for Contextual Bandits via Stochastic Gradient Descent
por: Chang, Xiangyu, et al.
Publicado: (2022)
por: Chang, Xiangyu, et al.
Publicado: (2022)
Optimizing Pre-Training Data Mixtures with Mixtures of Data Expert Models
por: Belenki, Lior, et al.
Publicado: (2025)
por: Belenki, Lior, et al.
Publicado: (2025)
Mitigating Preference Hacking in Policy Optimization with Pessimism
por: Gupta, Dhawal, et al.
Publicado: (2025)
por: Gupta, Dhawal, et al.
Publicado: (2025)
Ejemplares similares
-
Ordering-based Conditions for Global Convergence of Policy Gradient Methods
por: Mei, Jincheng, et al.
Publicado: (2025) -
Small steps no more: Global convergence of stochastic gradient bandits for arbitrary learning rates
por: Mei, Jincheng, et al.
Publicado: (2025) -
Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation
por: Lin, Max Qiushi, et al.
Publicado: (2025) -
Rectifying Regression in Reinforcement Learning
por: Ayoub, Alex, et al.
Publicado: (2025) -
Learning to Reason Efficiently with Discounted Reinforcement Learning
por: Ayoub, Alex, et al.
Publicado: (2025)