The regret lower bound for communicating Markov Decision Processes
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Boone, Victor, Maillard, Odalric-Ambrym |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Asymptotically optimal regret in communicating Markov decision processes
von: Boone, Victor
Veröffentlicht: (2025)
von: Boone, Victor
Veröffentlicht: (2025)
Leveraging priors on distribution functions for multi-arm bandits
von: Vashishtha, Sumit, et al.
Veröffentlicht: (2025)
von: Vashishtha, Sumit, et al.
Veröffentlicht: (2025)
How Hard is it to Confuse a World Model?
von: Radji, Waris, et al.
Veröffentlicht: (2025)
von: Radji, Waris, et al.
Veröffentlicht: (2025)
The Confusing Instance Principle for Online Linear Quadratic Control
von: Radji, Waris, et al.
Veröffentlicht: (2025)
von: Radji, Waris, et al.
Veröffentlicht: (2025)
How to Shrink Confidence Sets for Many Equivalent Discrete Distributions?
von: Maillard, Odalric-Ambrym, et al.
Veröffentlicht: (2024)
von: Maillard, Odalric-Ambrym, et al.
Veröffentlicht: (2024)
A Continual Offline Reinforcement Learning Benchmark for Navigation Tasks
von: Kobanda, Anthony, et al.
Veröffentlicht: (2025)
von: Kobanda, Anthony, et al.
Veröffentlicht: (2025)
Asymptotically Optimal Problem-Dependent Bandit Policies for Transfer Learning
von: Prevost, Adrien, et al.
Veröffentlicht: (2025)
von: Prevost, Adrien, et al.
Veröffentlicht: (2025)
Pliable rejection sampling
von: Erraqabi, Akram, et al.
Veröffentlicht: (2026)
von: Erraqabi, Akram, et al.
Veröffentlicht: (2026)
Hierarchical Subspaces of Policies for Continual Offline Reinforcement Learning
von: Kobanda, Anthony, et al.
Veröffentlicht: (2024)
von: Kobanda, Anthony, et al.
Veröffentlicht: (2024)
Offline Goal-Conditioned Reinforcement Learning with Projective Quasimetric Planning
von: Kobanda, Anthony, et al.
Veröffentlicht: (2025)
von: Kobanda, Anthony, et al.
Veröffentlicht: (2025)
Logarithmic Regret of Exploration in Average Reward Markov Decision Processes
von: Boone, Victor, et al.
Veröffentlicht: (2025)
von: Boone, Victor, et al.
Veröffentlicht: (2025)
Provably Efficient Exploration in Reward Machines with Low Regret
von: Bourel, Hippolyte, et al.
Veröffentlicht: (2024)
von: Bourel, Hippolyte, et al.
Veröffentlicht: (2024)
Interpolation pour l'augmentation de donnees : Application à la gestion des adventices de la canne a sucre a la Reunion
von: Ferber, Frederick Fabre, et al.
Veröffentlicht: (2025)
von: Ferber, Frederick Fabre, et al.
Veröffentlicht: (2025)
AdaStop: adaptive statistical testing for sound comparisons of Deep RL agents
von: Mathieu, Timothée, et al.
Veröffentlicht: (2023)
von: Mathieu, Timothée, et al.
Veröffentlicht: (2023)
Power Mean Estimation in Stochastic Monte-Carlo Tree_Search
von: Dam, Tuan, et al.
Veröffentlicht: (2024)
von: Dam, Tuan, et al.
Veröffentlicht: (2024)
Logarithmic regret bounds for continuous-time average-reward Markov decision processes
von: Gao, Xuefeng, et al.
Veröffentlicht: (2022)
von: Gao, Xuefeng, et al.
Veröffentlicht: (2022)
A second order regret bound for NormalHedge
von: Freund, Yoav, et al.
Veröffentlicht: (2026)
von: Freund, Yoav, et al.
Veröffentlicht: (2026)
Monitored Markov Decision Processes
von: Parisi, Simone, et al.
Veröffentlicht: (2024)
von: Parisi, Simone, et al.
Veröffentlicht: (2024)
Kriging and Gaussian Process Interpolation for Georeferenced Data Augmentation
von: Ferber, Frédérick Fabre, et al.
Veröffentlicht: (2025)
von: Ferber, Frédérick Fabre, et al.
Veröffentlicht: (2025)
On the price of exact truthfulness in incentive-compatible online learning with bandit feedback: A regret lower bound for WSU-UX
von: Mortazavi, Ali, et al.
Veröffentlicht: (2024)
von: Mortazavi, Ali, et al.
Veröffentlicht: (2024)
Convergence to Nash Equilibrium and No-regret Guarantee in (Markov) Potential Games
von: Dong, Jing, et al.
Veröffentlicht: (2024)
von: Dong, Jing, et al.
Veröffentlicht: (2024)
Generalized Linear Markov Decision Process
von: Zhang, Sinian, et al.
Veröffentlicht: (2025)
von: Zhang, Sinian, et al.
Veröffentlicht: (2025)
Federated Control in Markov Decision Processes
von: Jin, Hao, et al.
Veröffentlicht: (2024)
von: Jin, Hao, et al.
Veröffentlicht: (2024)
Learning in Markov Decision Processes with Exogenous Dynamics
von: Maran, Davide, et al.
Veröffentlicht: (2026)
von: Maran, Davide, et al.
Veröffentlicht: (2026)
Optimal Decision Tree Policies for Markov Decision Processes
von: Vos, Daniël, et al.
Veröffentlicht: (2023)
von: Vos, Daniël, et al.
Veröffentlicht: (2023)
Policy Testing in Markov Decision Processes
von: Ariu, Kaito, et al.
Veröffentlicht: (2025)
von: Ariu, Kaito, et al.
Veröffentlicht: (2025)
Towards Blackwell Optimality: Bellman Optimality Is All You Can Get
von: Boone, Victor, et al.
Veröffentlicht: (2025)
von: Boone, Victor, et al.
Veröffentlicht: (2025)
Markov Decision Processes under External Temporal Processes
von: Ayyagari, Ranga Shaarad, et al.
Veröffentlicht: (2023)
von: Ayyagari, Ranga Shaarad, et al.
Veröffentlicht: (2023)
An Orthogonal Learner for Individualized Outcomes in Markov Decision Processes
von: Javurek, Emil, et al.
Veröffentlicht: (2025)
von: Javurek, Emil, et al.
Veröffentlicht: (2025)
Initial Distribution Sensitivity of Constrained Markov Decision Processes
von: Tercan, Alperen, et al.
Veröffentlicht: (2025)
von: Tercan, Alperen, et al.
Veröffentlicht: (2025)
Improving Controller Generalization with Dimensionless Markov Decision Processes
von: Charvet, Valentin, et al.
Veröffentlicht: (2025)
von: Charvet, Valentin, et al.
Veröffentlicht: (2025)
Model-Based Exploration in Monitored Markov Decision Processes
von: Kazemipour, Alireza, et al.
Veröffentlicht: (2025)
von: Kazemipour, Alireza, et al.
Veröffentlicht: (2025)
Horizon-Free Regret for Linear Markov Decision Processes
von: Zhang, Zihan, et al.
Veröffentlicht: (2024)
von: Zhang, Zihan, et al.
Veröffentlicht: (2024)
Learning Utilities from Demonstrations in Markov Decision Processes
von: Lazzati, Filippo, et al.
Veröffentlicht: (2024)
von: Lazzati, Filippo, et al.
Veröffentlicht: (2024)
Achieving Constant Regret in Linear Markov Decision Processes
von: Zhang, Weitong, et al.
Veröffentlicht: (2024)
von: Zhang, Weitong, et al.
Veröffentlicht: (2024)
Automated scientific minimization of regret
von: Binz, Marcel, et al.
Veröffentlicht: (2025)
von: Binz, Marcel, et al.
Veröffentlicht: (2025)
Concentration of Cumulative Reward in Markov Decision Processes
von: Sayedana, Borna, et al.
Veröffentlicht: (2024)
von: Sayedana, Borna, et al.
Veröffentlicht: (2024)
Policy Gradient for Robust Markov Decision Processes
von: Wang, Qiuhao, et al.
Veröffentlicht: (2024)
von: Wang, Qiuhao, et al.
Veröffentlicht: (2024)
No-Regret Thompson Sampling for Finite-Horizon Markov Decision Processes with Gaussian Processes
von: Bayrooti, Jasmine, et al.
Veröffentlicht: (2025)
von: Bayrooti, Jasmine, et al.
Veröffentlicht: (2025)
A Markov Decision Process for Variable Selection in Branch & Bound
von: Strang, Paul, et al.
Veröffentlicht: (2025)
von: Strang, Paul, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Asymptotically optimal regret in communicating Markov decision processes
von: Boone, Victor
Veröffentlicht: (2025) -
Leveraging priors on distribution functions for multi-arm bandits
von: Vashishtha, Sumit, et al.
Veröffentlicht: (2025) -
How Hard is it to Confuse a World Model?
von: Radji, Waris, et al.
Veröffentlicht: (2025) -
The Confusing Instance Principle for Online Linear Quadratic Control
von: Radji, Waris, et al.
Veröffentlicht: (2025) -
How to Shrink Confidence Sets for Many Equivalent Discrete Distributions?
von: Maillard, Odalric-Ambrym, et al.
Veröffentlicht: (2024)