Rectifying Regression in Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Ayoub, Alex, Szepesvári, David, Bakhtiari, Alireza, Szepesvári, Csaba, Schuurmans, Dale |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning to Reason Efficiently with Discounted Reinforcement Learning
by: Ayoub, Alex, et al.
Published: (2025)
by: Ayoub, Alex, et al.
Published: (2025)
Learning What to Recommend: Minimax Optimal Simple Regret in Logistic Bandits
by: Liu, Shuai, et al.
Published: (2026)
by: Liu, Shuai, et al.
Published: (2026)
Eluder dimension: localise it!
by: Bakhtiari, Alireza, et al.
Published: (2026)
by: Bakhtiari, Alireza, et al.
Published: (2026)
Exploration via linearly perturbed loss minimisation
by: Janz, David, et al.
Published: (2023)
by: Janz, David, et al.
Published: (2023)
Regret Minimization via Saddle Point Optimization
by: Kirschner, Johannes, et al.
Published: (2024)
by: Kirschner, Johannes, et al.
Published: (2024)
Ordering-based Conditions for Global Convergence of Policy Gradient Methods
by: Mei, Jincheng, et al.
Published: (2025)
by: Mei, Jincheng, et al.
Published: (2025)
Stochastic Gradient Succeeds for Bandits
by: Mei, Jincheng, et al.
Published: (2024)
by: Mei, Jincheng, et al.
Published: (2024)
Sharper Guarantees for Misspecified Kernelized Bandit Optimization
by: Maran, Davide, et al.
Published: (2026)
by: Maran, Davide, et al.
Published: (2026)
Almost Free: Self-concordance in Natural Exponential Families and an Application to Bandits
by: Liu, Shuai, et al.
Published: (2024)
by: Liu, Shuai, et al.
Published: (2024)
Sharp analysis of linear ensemble sampling
by: Akhavan, Arya, et al.
Published: (2026)
by: Akhavan, Arya, et al.
Published: (2026)
Trajectory Data Suffices for Statistically Efficient Learning in Offline RL with Linear $q^π$-Realizability and Concentrability
by: Tkachuk, Volodymyr, et al.
Published: (2024)
by: Tkachuk, Volodymyr, et al.
Published: (2024)
Small steps no more: Global convergence of stochastic gradient bandits for arbitrary learning rates
by: Mei, Jincheng, et al.
Published: (2025)
by: Mei, Jincheng, et al.
Published: (2025)
Ensemble sampling for linear bandits: small ensembles suffice
by: Janz, David, et al.
Published: (2023)
by: Janz, David, et al.
Published: (2023)
Switching the Loss Reduces the Cost in Batch (Offline) Reinforcement Learning
by: Ayoub, Alex, et al.
Published: (2024)
by: Ayoub, Alex, et al.
Published: (2024)
Balancing optimism and pessimism in offline-to-online learning
by: Sentenac, Flore, et al.
Published: (2025)
by: Sentenac, Flore, et al.
Published: (2025)
Confident Natural Policy Gradient for Local Planning in $q_π$-realizable Constrained MDPs
by: Tian, Tian, et al.
Published: (2024)
by: Tian, Tian, et al.
Published: (2024)
Beyond Statistical Learning: Exact Learning Is Essential for General Intelligence
by: György, András, et al.
Published: (2025)
by: György, András, et al.
Published: (2025)
Rethinking the Foundations for Continual Reinforcement Learning
by: Elelimy, Esraa, et al.
Published: (2025)
by: Elelimy, Esraa, et al.
Published: (2025)
LACONIC: Length-Aware Constrained Reinforcement Learning for LLM
by: Liu, Chang, et al.
Published: (2026)
by: Liu, Chang, et al.
Published: (2026)
Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation
by: Lin, Max Qiushi, et al.
Published: (2025)
by: Lin, Max Qiushi, et al.
Published: (2025)
Optimistic Policy Optimization is Provably Efficient in Non-stationary MDPs
by: Zhong, Han, et al.
Published: (2021)
by: Zhong, Han, et al.
Published: (2021)
To Believe or Not to Believe Your LLM
by: Yadkori, Yasin Abbasi, et al.
Published: (2024)
by: Yadkori, Yasin Abbasi, et al.
Published: (2024)
Optimistic Actor-Critic with Parametric Policies for Linear Markov Decision Processes
by: Lin, Max Qiushi, et al.
Published: (2026)
by: Lin, Max Qiushi, et al.
Published: (2026)
Frontier LLMs Still Struggle with Simple Reasoning Tasks
by: Malek, Alan, et al.
Published: (2025)
by: Malek, Alan, et al.
Published: (2025)
Plastic Learning with Deep Fourier Features
by: Lewandowski, Alex, et al.
Published: (2024)
by: Lewandowski, Alex, et al.
Published: (2024)
Reinforcement Teaching
by: Muslimani, Calarina, et al.
Published: (2022)
by: Muslimani, Calarina, et al.
Published: (2022)
Revisiting Subgradient Dominance in Robust MDPs: Counterexamples, Hardness, and Sufficient Conditions
by: Kitamura, Toshinori, et al.
Published: (2026)
by: Kitamura, Toshinori, et al.
Published: (2026)
Stochastic Gradient Descent for Gaussian Processes Done Right
by: Lin, Jihao Andreas, et al.
Published: (2023)
by: Lin, Jihao Andreas, et al.
Published: (2023)
Spectral Ghost in Representation Learning: from Component Analysis to Self-Supervised Learning
by: Dai, Bo, et al.
Published: (2026)
by: Dai, Bo, et al.
Published: (2026)
Spectral Representation-based Reinforcement Learning
by: Gao, Chenxiao, et al.
Published: (2025)
by: Gao, Chenxiao, et al.
Published: (2025)
LoopBench: Discovering Emergent Symmetry Breaking Strategies with LLM Swarms
by: Parsaee, Ali, et al.
Published: (2025)
by: Parsaee, Ali, et al.
Published: (2025)
Managing Temporal Resolution in Continuous Value Estimation: A Fundamental Trade-off
by: Zhang, Zichen, et al.
Published: (2022)
by: Zhang, Zichen, et al.
Published: (2022)
Provable Representation with Efficient Planning for Partial Observable Reinforcement Learning
by: Zhang, Hongming, et al.
Published: (2023)
by: Zhang, Hongming, et al.
Published: (2023)
Directions of Curvature as an Explanation for Loss of Plasticity
by: Lewandowski, Alex, et al.
Published: (2023)
by: Lewandowski, Alex, et al.
Published: (2023)
Toward Understanding In-context vs. In-weight Learning
by: Chan, Bryan, et al.
Published: (2024)
by: Chan, Bryan, et al.
Published: (2024)
To See the Unseen: on the Generalization Ability of Transformers in Symbolic Reasoning
by: Lazić, Nevena, et al.
Published: (2026)
by: Lazić, Nevena, et al.
Published: (2026)
Learning Continually by Spectral Regularization
by: Lewandowski, Alex, et al.
Published: (2024)
by: Lewandowski, Alex, et al.
Published: (2024)
Reinforcement Learning with Options and State Representation
by: Ghriss, Ayoub, et al.
Published: (2024)
by: Ghriss, Ayoub, et al.
Published: (2024)
A Survey of State Representation Learning for Deep Reinforcement Learning
by: Echchahed, Ayoub, et al.
Published: (2025)
by: Echchahed, Ayoub, et al.
Published: (2025)
Truncated Rectified Flow Policy for Reinforcement Learning with One-Step Sampling
by: Zhou, Xubin, et al.
Published: (2026)
by: Zhou, Xubin, et al.
Published: (2026)
Similar Items
-
Learning to Reason Efficiently with Discounted Reinforcement Learning
by: Ayoub, Alex, et al.
Published: (2025) -
Learning What to Recommend: Minimax Optimal Simple Regret in Logistic Bandits
by: Liu, Shuai, et al.
Published: (2026) -
Eluder dimension: localise it!
by: Bakhtiari, Alireza, et al.
Published: (2026) -
Exploration via linearly perturbed loss minimisation
by: Janz, David, et al.
Published: (2023) -
Regret Minimization via Saddle Point Optimization
by: Kirschner, Johannes, et al.
Published: (2024)