Thompson Sampling for Infinite-Horizon Discounted Decision Processes
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Adelman, Daniel, Keceli, Cagla, Olivares-Nadal, Alba V. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On Strategic Measures and Optimality Properties in Discrete-Time Stochastic Control with Universally Measurable Policies
von: Yu, Huizhen
Veröffentlicht: (2022)
von: Yu, Huizhen
Veröffentlicht: (2022)
Measurized Markov Decision Processes
von: Adelman, Daniel, et al.
Veröffentlicht: (2024)
von: Adelman, Daniel, et al.
Veröffentlicht: (2024)
Blackwell optimality and policy stability for long-run risk sensitive stochastic control
von: Bäuerle, Nicole, et al.
Veröffentlicht: (2024)
von: Bäuerle, Nicole, et al.
Veröffentlicht: (2024)
Policy stability and ultimate stationarity in discounted risk-sensitive stochastic control
von: Bäuerle, Nicole, et al.
Veröffentlicht: (2026)
von: Bäuerle, Nicole, et al.
Veröffentlicht: (2026)
Reinforcement Learning Methods for the Stochastic Optimal Control of an Industrial Power-to-Heat System
von: Pilling, Eric, et al.
Veröffentlicht: (2024)
von: Pilling, Eric, et al.
Veröffentlicht: (2024)
Deep neural networks can provably solve Bellman equations for Markov decision processes without the curse of dimensionality
von: Jentzen, Arnulf, et al.
Veröffentlicht: (2025)
von: Jentzen, Arnulf, et al.
Veröffentlicht: (2025)
An Optimal-Control Approach to Infinite-Horizon Restless Bandits: Achieving Asymptotic Optimality with Minimal Assumptions
von: YAN, Chen
Veröffentlicht: (2024)
von: YAN, Chen
Veröffentlicht: (2024)
Properties of Turnpike Functions for Discounted Finite MDPs
von: Feinberg, Eugene A., et al.
Veröffentlicht: (2025)
von: Feinberg, Eugene A., et al.
Veröffentlicht: (2025)
Empirical Evaluation of Policy-Based Reinforcement Learning for Dynamic Service Control in an M/M/1 Queue
von: Walton, Joseph, et al.
Veröffentlicht: (2026)
von: Walton, Joseph, et al.
Veröffentlicht: (2026)
Single-Item Continuous-Review Inventory Models with Random Supplies
von: Helmes, K. L., et al.
Veröffentlicht: (2024)
von: Helmes, K. L., et al.
Veröffentlicht: (2024)
A Fisher-Rao gradient flow for entropy-regularised Markov decision processes in Polish spaces
von: Kerimkulov, Bekzhan, et al.
Veröffentlicht: (2023)
von: Kerimkulov, Bekzhan, et al.
Veröffentlicht: (2023)
A Stochastic Non-Zero-Sum Game of Controlling the Debt-to-GDP Ratio
von: Dammann, Felix, et al.
Veröffentlicht: (2023)
von: Dammann, Felix, et al.
Veröffentlicht: (2023)
Robust Ergodic Control of Jump-Diffusion Systems under Drift and Intensity Uncertainty
von: Azze, Abel, et al.
Veröffentlicht: (2026)
von: Azze, Abel, et al.
Veröffentlicht: (2026)
Policy Iteration for Exploratory Hamilton--Jacobi--Bellman Equations
von: Tran, Hung Vinh, et al.
Veröffentlicht: (2024)
von: Tran, Hung Vinh, et al.
Veröffentlicht: (2024)
Cost-optimal Management of a Residential Heating System With a Geothermal Energy Storage Under Uncertainty
von: Takam, Paul Honore, et al.
Veröffentlicht: (2025)
von: Takam, Paul Honore, et al.
Veröffentlicht: (2025)
Efficient Learning for Entropy-Regularized Markov Decision Processes via Multilevel Monte Carlo
von: Meunier, Matthieu, et al.
Veröffentlicht: (2025)
von: Meunier, Matthieu, et al.
Veröffentlicht: (2025)
Convex Regularization and Convergence of Policy Gradient Flows under Safety Constraints
von: Malo, Pekka, et al.
Veröffentlicht: (2024)
von: Malo, Pekka, et al.
Veröffentlicht: (2024)
Discrete-Time Approximations of Controlled Diffusions with Infinite Horizon Discounted and Average Cost
von: Pradhan, Somnath, et al.
Veröffentlicht: (2025)
von: Pradhan, Somnath, et al.
Veröffentlicht: (2025)
Sample Average Approximation for Stochastic Programming with Equality Constraints
von: Lew, Thomas, et al.
Veröffentlicht: (2022)
von: Lew, Thomas, et al.
Veröffentlicht: (2022)
Stability and Sensitivity Analysis of Relative Temporal-Difference Learning: Extended Version
von: Sakha, Masoud S., et al.
Veröffentlicht: (2026)
von: Sakha, Masoud S., et al.
Veröffentlicht: (2026)
Optimistic Training and Convergence of Q-Learning -- Extended Version
von: Mehta, Prashant, et al.
Veröffentlicht: (2026)
von: Mehta, Prashant, et al.
Veröffentlicht: (2026)
Model-free policy gradient for discrete-time mean-field control
von: Meunier, Matthieu, et al.
Veröffentlicht: (2026)
von: Meunier, Matthieu, et al.
Veröffentlicht: (2026)
An irreversible investment problem with a learning-by-doing feature
von: Ekström, Erik, et al.
Veröffentlicht: (2024)
von: Ekström, Erik, et al.
Veröffentlicht: (2024)
De Finetti's Control for Refracted Skew Brownian Motion
von: Gao, Zhongqin, et al.
Veröffentlicht: (2024)
von: Gao, Zhongqin, et al.
Veröffentlicht: (2024)
Average Cost Optimality of Partially Observed MDPS: Contraction of Non-linear Filters, Optimal Solutions and Approximations
von: Demirci, Yunus Emre, et al.
Veröffentlicht: (2023)
von: Demirci, Yunus Emre, et al.
Veröffentlicht: (2023)
A Representation Optimization Dichotomy, Lie-Algebraic Policy Optimization
von: KC, Sooraj, et al.
Veröffentlicht: (2026)
von: KC, Sooraj, et al.
Veröffentlicht: (2026)
On the saddle point of a zero-sum stopper vs. singular-controller game
von: Bovo, Andrea, et al.
Veröffentlicht: (2024)
von: Bovo, Andrea, et al.
Veröffentlicht: (2024)
Stochastic Optimal Control Problems for the Cost-Optimal Management of a Standalone Microgrid
von: Takam, Paul Honore, et al.
Veröffentlicht: (2025)
von: Takam, Paul Honore, et al.
Veröffentlicht: (2025)
Stopper vs. singular-controller games with degenerate diffusions
von: Bovo, Andrea, et al.
Veröffentlicht: (2023)
von: Bovo, Andrea, et al.
Veröffentlicht: (2023)
Finite-time horizon, stopper vs. singular-controller games on the half-line
von: Bovo, Andrea, et al.
Veröffentlicht: (2024)
von: Bovo, Andrea, et al.
Veröffentlicht: (2024)
Zero-sum stopper vs. singular-controller games with constrained control directions
von: Bovo, Andrea, et al.
Veröffentlicht: (2023)
von: Bovo, Andrea, et al.
Veröffentlicht: (2023)
Analysis and Optimization of Probabilities of Beneficial Mutation and Crossover Recombination in a Hamming Space
von: Belavkin, Roman V.
Veröffentlicht: (2025)
von: Belavkin, Roman V.
Veröffentlicht: (2025)
Long run control of nonhomogeneous Markov processes
von: Stettner, Łukasz
Veröffentlicht: (2025)
von: Stettner, Łukasz
Veröffentlicht: (2025)
Stochastic Control with Signatures
von: Bank, P., et al.
Veröffentlicht: (2024)
von: Bank, P., et al.
Veröffentlicht: (2024)
The one-shot problem: Solution to an open question of finite-fuel singular control with discretionary stopping
von: Moriarty, John, et al.
Veröffentlicht: (2024)
von: Moriarty, John, et al.
Veröffentlicht: (2024)
Exploratory Randomization for Discrete-Time Linear Exponential Quadratic Gaussian (LEQG) Problem
von: Lleo, Sebastien, et al.
Veröffentlicht: (2025)
von: Lleo, Sebastien, et al.
Veröffentlicht: (2025)
Data-driven rules for multidimensional reflection problems
von: Christensen, Sören, et al.
Veröffentlicht: (2023)
von: Christensen, Sören, et al.
Veröffentlicht: (2023)
Drift Control with Discretionary Stopping for a Diffusion
von: Beneš, Václav E., et al.
Veröffentlicht: (2024)
von: Beneš, Václav E., et al.
Veröffentlicht: (2024)
Learning to reflect: A unifying approach for data-driven stochastic control strategies
von: Christensen, Sören, et al.
Veröffentlicht: (2021)
von: Christensen, Sören, et al.
Veröffentlicht: (2021)
Reinforcement Learning in Real Option Models
von: Dianetti, Jodi, et al.
Veröffentlicht: (2026)
von: Dianetti, Jodi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
On Strategic Measures and Optimality Properties in Discrete-Time Stochastic Control with Universally Measurable Policies
von: Yu, Huizhen
Veröffentlicht: (2022) -
Measurized Markov Decision Processes
von: Adelman, Daniel, et al.
Veröffentlicht: (2024) -
Blackwell optimality and policy stability for long-run risk sensitive stochastic control
von: Bäuerle, Nicole, et al.
Veröffentlicht: (2024) -
Policy stability and ultimate stationarity in discounted risk-sensitive stochastic control
von: Bäuerle, Nicole, et al.
Veröffentlicht: (2026) -
Reinforcement Learning Methods for the Stochastic Optimal Control of an Industrial Power-to-Heat System
von: Pilling, Eric, et al.
Veröffentlicht: (2024)