Optimistic Training and Convergence of Q-Learning -- Extended Version
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mehta, Prashant, Meyn, Sean |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Stability and Sensitivity Analysis of Relative Temporal-Difference Learning: Extended Version
von: Sakha, Masoud S., et al.
Veröffentlicht: (2026)
von: Sakha, Masoud S., et al.
Veröffentlicht: (2026)
Continuous-time Risk-sensitive Reinforcement Learning via Quadratic Variation Penalty
von: Jia, Yanwei
Veröffentlicht: (2024)
von: Jia, Yanwei
Veröffentlicht: (2024)
On the Rate of Gaussian Approximation for Linear Regression Problems
von: Khusainov, Marat, et al.
Veröffentlicht: (2025)
von: Khusainov, Marat, et al.
Veröffentlicht: (2025)
Sample Complexity of Policy Gradient for Log-Growth Control
von: Pan, Qiuhua, et al.
Veröffentlicht: (2026)
von: Pan, Qiuhua, et al.
Veröffentlicht: (2026)
Beyond Bellman: High-Order Generator Regression for Continuous-Time Policy Evaluation
von: Zheng, Yaowei, et al.
Veröffentlicht: (2026)
von: Zheng, Yaowei, et al.
Veröffentlicht: (2026)
Finite-Time Analysis of Projected Two-Time-Scale Stochastic Approximation
von: Bai, Yitao, et al.
Veröffentlicht: (2026)
von: Bai, Yitao, et al.
Veröffentlicht: (2026)
Continuous Policy and Value Iteration for Stochastic Control Problems and Its Convergence
von: Feng, Qi, et al.
Veröffentlicht: (2025)
von: Feng, Qi, et al.
Veröffentlicht: (2025)
Exploratory Randomization for Discrete-Time Linear Exponential Quadratic Gaussian (LEQG) Problem
von: Lleo, Sebastien, et al.
Veröffentlicht: (2025)
von: Lleo, Sebastien, et al.
Veröffentlicht: (2025)
Gaussian Approximation and Multiplier Bootstrap for Stochastic Gradient Descent
von: Sheshukova, Marina, et al.
Veröffentlicht: (2025)
von: Sheshukova, Marina, et al.
Veröffentlicht: (2025)
Asynchronous Stochastic Approximation with Applications to Average-Reward Reinforcement Learning
von: Yu, Huizhen, et al.
Veröffentlicht: (2024)
von: Yu, Huizhen, et al.
Veröffentlicht: (2024)
Mean--Variance Portfolio Selection by Continuous-Time Reinforcement Learning: Algorithms, Regret Analysis, and Empirical Study
von: Huang, Yilie, et al.
Veröffentlicht: (2024)
von: Huang, Yilie, et al.
Veröffentlicht: (2024)
Policy Gradient for Continuous-Time Mean-Field Control
von: Bayraktar, Erhan, et al.
Veröffentlicht: (2026)
von: Bayraktar, Erhan, et al.
Veröffentlicht: (2026)
Average-reward reinforcement learning in semi-Markov decision processes via relative value iteration
von: Yu, Huizhen, et al.
Veröffentlicht: (2025)
von: Yu, Huizhen, et al.
Veröffentlicht: (2025)
Stochastic Control with Signatures
von: Bank, P., et al.
Veröffentlicht: (2024)
von: Bank, P., et al.
Veröffentlicht: (2024)
A Note on Stability in Asynchronous Stochastic Approximation without Communication Delays
von: Yu, Huizhen, et al.
Veröffentlicht: (2023)
von: Yu, Huizhen, et al.
Veröffentlicht: (2023)
Exploratory Randomization for Discrete-Time Risk-Sensitive Benchmarked Investment Management with Reinforcement Learning
von: Lleo, Sebastien, et al.
Veröffentlicht: (2026)
von: Lleo, Sebastien, et al.
Veröffentlicht: (2026)
Convergence of policy gradient methods for finite-horizon exploratory linear-quadratic control problems
von: Giegrich, Michael, et al.
Veröffentlicht: (2022)
von: Giegrich, Michael, et al.
Veröffentlicht: (2022)
Optimal control of SDEs with merely measurable drift: an HJB approach
von: Du, Kai, et al.
Veröffentlicht: (2025)
von: Du, Kai, et al.
Veröffentlicht: (2025)
Does DQN Learn?
von: Gopalan, Aditya, et al.
Veröffentlicht: (2022)
von: Gopalan, Aditya, et al.
Veröffentlicht: (2022)
Neural Actor-Critic Methods for Hamilton-Jacobi-Bellman PDEs: Asymptotic Analysis and Numerical Studies
von: Cohen, Samuel N., et al.
Veröffentlicht: (2025)
von: Cohen, Samuel N., et al.
Veröffentlicht: (2025)
Infinite time horizon stochastic recursive control problems with jumps: dynamic programming and stochastic verification theorems
von: Luo, Sheng, et al.
Veröffentlicht: (2024)
von: Luo, Sheng, et al.
Veröffentlicht: (2024)
Learning to reflect: A unifying approach for data-driven stochastic control strategies
von: Christensen, Sören, et al.
Veröffentlicht: (2021)
von: Christensen, Sören, et al.
Veröffentlicht: (2021)
Markovian Foundations for Quasi-Stochastic Approximation with Applications to Extremum Seeking Control
von: Lauand, Caio Kalil, et al.
Veröffentlicht: (2022)
von: Lauand, Caio Kalil, et al.
Veröffentlicht: (2022)
The Score Kalman Filter
von: Iwasaki, Kaito, et al.
Veröffentlicht: (2026)
von: Iwasaki, Kaito, et al.
Veröffentlicht: (2026)
Continuous time Stochastic optimal control under discrete time partial observations
von: Bayer, Christian, et al.
Veröffentlicht: (2024)
von: Bayer, Christian, et al.
Veröffentlicht: (2024)
Discrete-Time Approximations of Controlled Diffusions with Infinite Horizon Discounted and Average Cost
von: Pradhan, Somnath, et al.
Veröffentlicht: (2025)
von: Pradhan, Somnath, et al.
Veröffentlicht: (2025)
Reflected stochastic recursive control problems with jumps: dynamic programming and stochastic verification theorems
von: Liu, Lu, et al.
Veröffentlicht: (2025)
von: Liu, Lu, et al.
Veröffentlicht: (2025)
Optimal Feedback Control in Social Networks in a McKean-Vlasov-Friedkin-Johnsen System
von: Pramanik, Paramahansa
Veröffentlicht: (2025)
von: Pramanik, Paramahansa
Veröffentlicht: (2025)
Open-loop and closed-loop solvabilities for zero-sum stochastic linear quadratic differential games of Markovian regime switching system
von: Wu, Fan, et al.
Veröffentlicht: (2024)
von: Wu, Fan, et al.
Veröffentlicht: (2024)
Stochastic linear-quadratic differential game with Markovian jumps in an infinite horizon
von: Wu, Fan, et al.
Veröffentlicht: (2024)
von: Wu, Fan, et al.
Veröffentlicht: (2024)
Reinforcement Learning, Optimal Control, and Bayesian Filtering in Data Assimilation
von: Hammoud, Abed
Veröffentlicht: (2026)
von: Hammoud, Abed
Veröffentlicht: (2026)
A measure-valued HJB perspective on Bayesian optimal adaptive control
von: Cox, Alexander M. G., et al.
Veröffentlicht: (2025)
von: Cox, Alexander M. G., et al.
Veröffentlicht: (2025)
Controllability and Vector Potential
von: Shankar, Shiva
Veröffentlicht: (2019)
von: Shankar, Shiva
Veröffentlicht: (2019)
Control randomisation approach for policy gradient and application to reinforcement learning in optimal switching
von: Denkert, Robert, et al.
Veröffentlicht: (2024)
von: Denkert, Robert, et al.
Veröffentlicht: (2024)
An optimal level of Stubbornness to win a soccer match
von: Pramanik, Paramahansa
Veröffentlicht: (2025)
von: Pramanik, Paramahansa
Veröffentlicht: (2025)
Path integral control under McKean-Vlasov dynamics
von: Bennett, Timothy
Veröffentlicht: (2024)
von: Bennett, Timothy
Veröffentlicht: (2024)
Stability of long run functionals with respect to stationary Markov controls
von: Stettner, Lukasz
Veröffentlicht: (2024)
von: Stettner, Lukasz
Veröffentlicht: (2024)
Reinforcement Learning in Real Option Models
von: Dianetti, Jodi, et al.
Veröffentlicht: (2026)
von: Dianetti, Jodi, et al.
Veröffentlicht: (2026)
Flow Matching for Efficient and Scalable Data Assimilation
von: Transue, Taos, et al.
Veröffentlicht: (2025)
von: Transue, Taos, et al.
Veröffentlicht: (2025)
Thompson Sampling for Infinite-Horizon Discounted Decision Processes
von: Adelman, Daniel, et al.
Veröffentlicht: (2024)
von: Adelman, Daniel, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Stability and Sensitivity Analysis of Relative Temporal-Difference Learning: Extended Version
von: Sakha, Masoud S., et al.
Veröffentlicht: (2026) -
Continuous-time Risk-sensitive Reinforcement Learning via Quadratic Variation Penalty
von: Jia, Yanwei
Veröffentlicht: (2024) -
On the Rate of Gaussian Approximation for Linear Regression Problems
von: Khusainov, Marat, et al.
Veröffentlicht: (2025) -
Sample Complexity of Policy Gradient for Log-Growth Control
von: Pan, Qiuhua, et al.
Veröffentlicht: (2026) -
Beyond Bellman: High-Order Generator Regression for Continuous-Time Policy Evaluation
von: Zheng, Yaowei, et al.
Veröffentlicht: (2026)