Towards Optimal Offline Reinforcement Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Mengmeng, Kuhn, Daniel, Sutter, Tobias |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Improved High-Probability Bounds for the Temporal Difference Learning Algorithm via Exponential Stability
por: Samsonov, Sergey, et al.
Publicado: (2023)
por: Samsonov, Sergey, et al.
Publicado: (2023)
Optimal Learning via Moderate Deviations Theory
por: Ganguly, Arnab, et al.
Publicado: (2023)
por: Ganguly, Arnab, et al.
Publicado: (2023)
Soft-constrained Schrodinger Bridge: a Stochastic Control Approach
por: Garg, Jhanvi, et al.
Publicado: (2024)
por: Garg, Jhanvi, et al.
Publicado: (2024)
A weak convergence approach to large deviations for stochastic approximations
por: Hult, Henrik, et al.
Publicado: (2025)
por: Hult, Henrik, et al.
Publicado: (2025)
First is the worst, second is the best? A Markov chain analysis of the basketball game knockout
por: Flatz, Andrew, et al.
Publicado: (2025)
por: Flatz, Andrew, et al.
Publicado: (2025)
Optimal withdrawals in a general diffusion model with control rates subject to a state-dependent upper bound
por: Guérin, Hélène, et al.
Publicado: (2024)
por: Guérin, Hélène, et al.
Publicado: (2024)
Optimal State Equation for the Control of a Diffusion with Two Distinct Dynamics
por: Chen, Zengjing, et al.
Publicado: (2024)
por: Chen, Zengjing, et al.
Publicado: (2024)
Mini-Batch Covariance, Diffusion Limits, and Oracle Complexity in Stochastic Gradient Descent: A Sampling-Design Perspective
por: Zantedeschi, Daniel, et al.
Publicado: (2026)
por: Zantedeschi, Daniel, et al.
Publicado: (2026)
Mean-Field Sampling for Cooperative Multi-Agent Reinforcement Learning
por: Anand, Emile, et al.
Publicado: (2024)
por: Anand, Emile, et al.
Publicado: (2024)
Policy Gradient Algorithms for Robust MDPs with Non-Rectangular Uncertainty Sets
por: Li, Mengmeng, et al.
Publicado: (2023)
por: Li, Mengmeng, et al.
Publicado: (2023)
A Large Deviations Perspective on Policy Gradient Algorithms
por: Jongeneel, Wouter, et al.
Publicado: (2023)
por: Jongeneel, Wouter, et al.
Publicado: (2023)
Optimal control of stochastic networks of $M/M/\infty$ queues with linear costs
por: Carratelli, Giovanni Pugliese, et al.
Publicado: (2025)
por: Carratelli, Giovanni Pugliese, et al.
Publicado: (2025)
Analysing heavy-tail properties of Stochastic Gradient Descent by means of Stochastic Recurrence Equations
por: Damek, Ewa, et al.
Publicado: (2024)
por: Damek, Ewa, et al.
Publicado: (2024)
Optimal error estimates of the stochastic parabolic optimal control problem with integral state constraint
por: Wang, Qiming, et al.
Publicado: (2024)
por: Wang, Qiming, et al.
Publicado: (2024)
Empirical Evaluation of Policy-Based Reinforcement Learning for Dynamic Service Control in an M/M/1 Queue
por: Walton, Joseph, et al.
Publicado: (2026)
por: Walton, Joseph, et al.
Publicado: (2026)
Thompson Sampling for Infinite-Horizon Discounted Decision Processes
por: Adelman, Daniel, et al.
Publicado: (2024)
por: Adelman, Daniel, et al.
Publicado: (2024)
Optimal resource allocation for maintaining system solvency
por: Guo, Gaoyue, et al.
Publicado: (2026)
por: Guo, Gaoyue, et al.
Publicado: (2026)
A Two-Timescale Primal-Dual Framework for Reinforcement Learning via Online Dual Variable Guidance
por: Wolter, Axel Friedrich, et al.
Publicado: (2025)
por: Wolter, Axel Friedrich, et al.
Publicado: (2025)
Optimal Control Problems Governed by MFSDEs with multi-defaults
por: Gou, Zhun, et al.
Publicado: (2020)
por: Gou, Zhun, et al.
Publicado: (2020)
Nonlocal Stochastic Optimal Control for Diffusion Processes: Existence, Maximum Principle and Financial Applications
por: Anita, Stefana-Lucia, et al.
Publicado: (2025)
por: Anita, Stefana-Lucia, et al.
Publicado: (2025)
Stochastic Modified Flows for Riemannian Stochastic Gradient Descent
por: Gess, Benjamin, et al.
Publicado: (2024)
por: Gess, Benjamin, et al.
Publicado: (2024)
Near Optimality of Lipschitz and Smooth Policies in Controlled Diffusions
por: Pradhan, Somnath, et al.
Publicado: (2024)
por: Pradhan, Somnath, et al.
Publicado: (2024)
Impulse control maximising average cost per unit time: a non-uniformly ergodic case
por: Palczewski, Jan, et al.
Publicado: (2016)
por: Palczewski, Jan, et al.
Publicado: (2016)
Long-Term Average Impulse and Singular Control of a Growth Model with Two Revenue Sources
por: Helmes, K. L., et al.
Publicado: (2026)
por: Helmes, K. L., et al.
Publicado: (2026)
Data-Driven Estimation of Conditional Expectations, Application to Optimal Stopping and Reinforcement Learning
por: Moustakides, George V.
Publicado: (2024)
por: Moustakides, George V.
Publicado: (2024)
Convergence Rate for the Last Iterate of Stochastic Gradient Descent Schemes
por: Hudiani, Marcel
Publicado: (2025)
por: Hudiani, Marcel
Publicado: (2025)
Optimal control for production inventory system with various cost criterion
por: Golui, Subrata, et al.
Publicado: (2022)
por: Golui, Subrata, et al.
Publicado: (2022)
A Framework for Exploring Social Interactions in Multiagent Decision-Making for Two-Queue Systems
por: Gaspard, Mallory E., et al.
Publicado: (2026)
por: Gaspard, Mallory E., et al.
Publicado: (2026)
An irreversible investment problem with a learning-by-doing feature
por: Ekström, Erik, et al.
Publicado: (2024)
por: Ekström, Erik, et al.
Publicado: (2024)
An Optimal Periodic Dividend and Risk Control Problem for an Insurance Company
por: Kelbert, Mark, et al.
Publicado: (2023)
por: Kelbert, Mark, et al.
Publicado: (2023)
Controlled Interacting Branching Diffusion Processes: A Viscosity Approach
por: Ocello, Antonio
Publicado: (2026)
por: Ocello, Antonio
Publicado: (2026)
On Strategic Measures and Optimality Properties in Discrete-Time Stochastic Control with Universally Measurable Policies
por: Yu, Huizhen
Publicado: (2022)
por: Yu, Huizhen
Publicado: (2022)
Controlled Interacting Branching Diffusion Processes: Relaxed Formulation in the Mean-Field Regime
por: Ocello, Antonio
Publicado: (2023)
por: Ocello, Antonio
Publicado: (2023)
Mean-Field Langevin Diffusions with Density-dependent Temperature
por: Huang, Yu-Jui, et al.
Publicado: (2025)
por: Huang, Yu-Jui, et al.
Publicado: (2025)
An Actor-Critic Framework for Continuous-Time Jump-Diffusion Controls with Normalizing Flows
por: Guo, Liya, et al.
Publicado: (2026)
por: Guo, Liya, et al.
Publicado: (2026)
Heavy-traffic limit of stationary distributions of a state-dependent queue
por: Kobayashi, Masahiro, et al.
Publicado: (2026)
por: Kobayashi, Masahiro, et al.
Publicado: (2026)
Singular stochastic control problems motivated by the optimal sustainable exploitation of an ecosystem
por: Liang, Gechun, et al.
Publicado: (2020)
por: Liang, Gechun, et al.
Publicado: (2020)
Zero-one Laws for a Control Problem with Random Action Sets
por: Flesch, János, et al.
Publicado: (2024)
por: Flesch, János, et al.
Publicado: (2024)
De Finetti's Control for Refracted Skew Brownian Motion
por: Gao, Zhongqin, et al.
Publicado: (2024)
por: Gao, Zhongqin, et al.
Publicado: (2024)
A Wasserstein Geometric Framework for Hebbian Plasticity
por: Tan, Ulrich
Publicado: (2026)
por: Tan, Ulrich
Publicado: (2026)
Ejemplares similares
-
Improved High-Probability Bounds for the Temporal Difference Learning Algorithm via Exponential Stability
por: Samsonov, Sergey, et al.
Publicado: (2023) -
Optimal Learning via Moderate Deviations Theory
por: Ganguly, Arnab, et al.
Publicado: (2023) -
Soft-constrained Schrodinger Bridge: a Stochastic Control Approach
por: Garg, Jhanvi, et al.
Publicado: (2024) -
A weak convergence approach to large deviations for stochastic approximations
por: Hult, Henrik, et al.
Publicado: (2025) -
First is the worst, second is the best? A Markov chain analysis of the basketball game knockout
por: Flatz, Andrew, et al.
Publicado: (2025)