Convergence of off-policy TD(0) with linear function approximation for reversible Markov chains
Fuente:
arXiv
Salvato in:
| Autori principali: | Overmars, Maik, Goseling, Jasper, Boucherie, Richard |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Convergence of TD(0) under Polynomial Mixing with Nonlinear Function Approximation
di: Sridhar, Anupama, et al.
Pubblicazione: (2025)
di: Sridhar, Anupama, et al.
Pubblicazione: (2025)
Rates of Convergence in the Central Limit Theorem for Markov Chains, with an Application to TD Learning
di: Srikant, R.
Pubblicazione: (2024)
di: Srikant, R.
Pubblicazione: (2024)
Parameter-Free Federated TD Learning with Markov Noise in Heterogeneous Environments
di: Naskar, Ankur, et al.
Pubblicazione: (2025)
di: Naskar, Ankur, et al.
Pubblicazione: (2025)
BINDy -- Bayesian identification of nonlinear dynamics with reversible-jump Markov-chain Monte-Carlo
di: Champneys, Max D., et al.
Pubblicazione: (2024)
di: Champneys, Max D., et al.
Pubblicazione: (2024)
High-probability sample complexities for policy evaluation with linear function approximation
di: Li, Gen, et al.
Pubblicazione: (2023)
di: Li, Gen, et al.
Pubblicazione: (2023)
Convergence and concentration properties of constant step-size SGD through Markov chains
di: Merad, Ibrahim, et al.
Pubblicazione: (2023)
di: Merad, Ibrahim, et al.
Pubblicazione: (2023)
Adversarial bandit optimization for approximately linear functions
di: Cheng, Zhuoyu, et al.
Pubblicazione: (2025)
di: Cheng, Zhuoyu, et al.
Pubblicazione: (2025)
Multi-agent imitation learning with function approximation: Linear Markov games and beyond
di: Viano, Luca, et al.
Pubblicazione: (2026)
di: Viano, Luca, et al.
Pubblicazione: (2026)
One-Shot Averaging for Distributed TD($λ$) Under Markov Sampling
di: Tian, Haoxing, et al.
Pubblicazione: (2024)
di: Tian, Haoxing, et al.
Pubblicazione: (2024)
A Concentration Bound for TD(0) with Function Approximation
di: Chandak, Siddharth, et al.
Pubblicazione: (2023)
di: Chandak, Siddharth, et al.
Pubblicazione: (2023)
Exploring Adaptive MCTS with TD Learning in miniXCOM
di: Saadat, Kimiya, et al.
Pubblicazione: (2022)
di: Saadat, Kimiya, et al.
Pubblicazione: (2022)
Dimension lower bounds for linear approaches to function approximation
di: Hsu, Daniel
Pubblicazione: (2025)
di: Hsu, Daniel
Pubblicazione: (2025)
Distances for Markov chains from sample streams
di: Calo, Sergio, et al.
Pubblicazione: (2025)
di: Calo, Sergio, et al.
Pubblicazione: (2025)
On Convergence of Average-Reward Q-Learning in Weakly Communicating Markov Decision Processes
di: Wan, Yi, et al.
Pubblicazione: (2024)
di: Wan, Yi, et al.
Pubblicazione: (2024)
Rosenthal-type inequalities for linear statistics of Markov chains
di: Durmus, Alain, et al.
Pubblicazione: (2023)
di: Durmus, Alain, et al.
Pubblicazione: (2023)
Sampling Complexity of TD and PPO in RKHS
di: Zou, Lu, et al.
Pubblicazione: (2025)
di: Zou, Lu, et al.
Pubblicazione: (2025)
Advanced posterior analyses of hidden Markov models: finite Markov chain imbedding and hybrid decoding
di: Bæk, Zenia Elise Damgaard, et al.
Pubblicazione: (2025)
di: Bæk, Zenia Elise Damgaard, et al.
Pubblicazione: (2025)
Targeted stochastic gradient Markov chain Monte Carlo for hidden Markov models with rare latent states
di: Ou, Rihui, et al.
Pubblicazione: (2018)
di: Ou, Rihui, et al.
Pubblicazione: (2018)
Deep Learning for Computing Convergence Rates of Markov Chains
di: Qu, Yanlin, et al.
Pubblicazione: (2024)
di: Qu, Yanlin, et al.
Pubblicazione: (2024)
Markov flow policy -- deep MC
di: Soffair, Nitsan, et al.
Pubblicazione: (2024)
di: Soffair, Nitsan, et al.
Pubblicazione: (2024)
Smoothed functional-based gradient algorithms for off-policy reinforcement learning: A non-asymptotic viewpoint
di: Vijayan, Nithia, et al.
Pubblicazione: (2021)
di: Vijayan, Nithia, et al.
Pubblicazione: (2021)
Bridging the Gap Between Average and Discounted TD Learning
di: Tian, Haoxing, et al.
Pubblicazione: (2026)
di: Tian, Haoxing, et al.
Pubblicazione: (2026)
Differential privacy guarantees of Markov chain Monte Carlo algorithms
di: Bertazzi, Andrea, et al.
Pubblicazione: (2025)
di: Bertazzi, Andrea, et al.
Pubblicazione: (2025)
Coverage Improvement and Fast Convergence of On-policy Preference Learning
di: Kim, Juno, et al.
Pubblicazione: (2026)
di: Kim, Juno, et al.
Pubblicazione: (2026)
Improved off-policy training of diffusion samplers
di: Sendera, Marcin, et al.
Pubblicazione: (2024)
di: Sendera, Marcin, et al.
Pubblicazione: (2024)
On Gaussian approximation for entropy-regularized Q-learning with function approximation
di: Rubtsov, Artemy, et al.
Pubblicazione: (2026)
di: Rubtsov, Artemy, et al.
Pubblicazione: (2026)
Closing the gap between SVRG and TD-SVRG with Gradient Splitting
di: Mustafin, Arsenii, et al.
Pubblicazione: (2022)
di: Mustafin, Arsenii, et al.
Pubblicazione: (2022)
Variational Markov chain mixtures with automatic component selection
di: Miles, Christopher E., et al.
Pubblicazione: (2024)
di: Miles, Christopher E., et al.
Pubblicazione: (2024)
A policy gradient approach for Finite Horizon Constrained Markov Decision Processes
di: Guin, Soumyajit, et al.
Pubblicazione: (2022)
di: Guin, Soumyajit, et al.
Pubblicazione: (2022)
TD-Interpreter: Enhancing the Understanding of Timing Diagrams with Visual-Language Learning
di: He, Jie, et al.
Pubblicazione: (2025)
di: He, Jie, et al.
Pubblicazione: (2025)
True Online TD-Replan(lambda) Achieving Planning through Replaying
di: Altahhan, Abdulrahman
Pubblicazione: (2025)
di: Altahhan, Abdulrahman
Pubblicazione: (2025)
TD-JEPA: Latent-predictive Representations for Zero-Shot Reinforcement Learning
di: Bagatella, Marco, et al.
Pubblicazione: (2025)
di: Bagatella, Marco, et al.
Pubblicazione: (2025)
Convergence of policy gradient methods for finite-horizon exploratory linear-quadratic control problems
di: Giegrich, Michael, et al.
Pubblicazione: (2022)
di: Giegrich, Michael, et al.
Pubblicazione: (2022)
On the Global Convergence of Policy Gradient in Average Reward Markov Decision Processes
di: Kumar, Navdeep, et al.
Pubblicazione: (2024)
di: Kumar, Navdeep, et al.
Pubblicazione: (2024)
Variational Autoencoder for Generating Broader-Spectrum prior Proposals in Markov chain Monte Carlo Methods
di: Borges, Marcio, et al.
Pubblicazione: (2025)
di: Borges, Marcio, et al.
Pubblicazione: (2025)
A primal-dual perspective for distributed TD-learning
di: Lim, Han-Dong, et al.
Pubblicazione: (2023)
di: Lim, Han-Dong, et al.
Pubblicazione: (2023)
What Does Flow Matching Bring To TD Learning?
di: Agrawalla, Bhavya, et al.
Pubblicazione: (2026)
di: Agrawalla, Bhavya, et al.
Pubblicazione: (2026)
Finding good policies in average-reward Markov Decision Processes without prior knowledge
di: Tuynman, Adrienne, et al.
Pubblicazione: (2024)
di: Tuynman, Adrienne, et al.
Pubblicazione: (2024)
Finite time analysis of temporal difference learning with linear function approximation: Tail averaging and regularisation
di: Patil, Gandharv, et al.
Pubblicazione: (2022)
di: Patil, Gandharv, et al.
Pubblicazione: (2022)
Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization
di: Lei, Xing, et al.
Pubblicazione: (2025)
di: Lei, Xing, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Convergence of TD(0) under Polynomial Mixing with Nonlinear Function Approximation
di: Sridhar, Anupama, et al.
Pubblicazione: (2025) -
Rates of Convergence in the Central Limit Theorem for Markov Chains, with an Application to TD Learning
di: Srikant, R.
Pubblicazione: (2024) -
Parameter-Free Federated TD Learning with Markov Noise in Heterogeneous Environments
di: Naskar, Ankur, et al.
Pubblicazione: (2025) -
BINDy -- Bayesian identification of nonlinear dynamics with reversible-jump Markov-chain Monte-Carlo
di: Champneys, Max D., et al.
Pubblicazione: (2024) -
High-probability sample complexities for policy evaluation with linear function approximation
di: Li, Gen, et al.
Pubblicazione: (2023)