Provably Efficient Exploration in Reward Machines with Low Regret
Fuente:
arXiv
Saved in:
| Main Authors: | Bourel, Hippolyte, Jonsson, Anders, Maillard, Odalric-Ambrym, Ma, Chenxiao, Talebi, Mohammad Sadegh |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How to Shrink Confidence Sets for Many Equivalent Discrete Distributions?
by: Maillard, Odalric-Ambrym, et al.
Published: (2024)
by: Maillard, Odalric-Ambrym, et al.
Published: (2024)
Leveraging priors on distribution functions for multi-arm bandits
by: Vashishtha, Sumit, et al.
Published: (2025)
by: Vashishtha, Sumit, et al.
Published: (2025)
How Hard is it to Confuse a World Model?
by: Radji, Waris, et al.
Published: (2025)
by: Radji, Waris, et al.
Published: (2025)
The regret lower bound for communicating Markov Decision Processes
by: Boone, Victor, et al.
Published: (2025)
by: Boone, Victor, et al.
Published: (2025)
The Confusing Instance Principle for Online Linear Quadratic Control
by: Radji, Waris, et al.
Published: (2025)
by: Radji, Waris, et al.
Published: (2025)
Power Mean Estimation in Stochastic Monte-Carlo Tree_Search
by: Dam, Tuan, et al.
Published: (2024)
by: Dam, Tuan, et al.
Published: (2024)
Tractable Offline Learning of Regular Decision Processes
by: Deb, Ahana, et al.
Published: (2024)
by: Deb, Ahana, et al.
Published: (2024)
A Continual Offline Reinforcement Learning Benchmark for Navigation Tasks
by: Kobanda, Anthony, et al.
Published: (2025)
by: Kobanda, Anthony, et al.
Published: (2025)
Asymptotically Optimal Problem-Dependent Bandit Policies for Transfer Learning
by: Prevost, Adrien, et al.
Published: (2025)
by: Prevost, Adrien, et al.
Published: (2025)
Recursive Entropic Risk Optimization in Discounted MDPs: Sample Complexity Bounds with a Generative Model
by: Mortensen, Oliver, et al.
Published: (2025)
by: Mortensen, Oliver, et al.
Published: (2025)
Provably Efficient Exploration in Quantum Reinforcement Learning with Logarithmic Worst-Case Regret
by: Zhong, Han, et al.
Published: (2023)
by: Zhong, Han, et al.
Published: (2023)
Hierarchical Subspaces of Policies for Continual Offline Reinforcement Learning
by: Kobanda, Anthony, et al.
Published: (2024)
by: Kobanda, Anthony, et al.
Published: (2024)
Pliable rejection sampling
by: Erraqabi, Akram, et al.
Published: (2026)
by: Erraqabi, Akram, et al.
Published: (2026)
Hierarchical Average-Reward Linearly-solvable Markov Decision Processes
by: Infante, Guillermo, et al.
Published: (2024)
by: Infante, Guillermo, et al.
Published: (2024)
Regret-Free Reinforcement Learning for LTL Specifications
by: Majumdar, Rupak, et al.
Published: (2024)
by: Majumdar, Rupak, et al.
Published: (2024)
Offline Goal-Conditioned Reinforcement Learning with Projective Quasimetric Planning
by: Kobanda, Anthony, et al.
Published: (2025)
by: Kobanda, Anthony, et al.
Published: (2025)
Provably Efficient Exploration in Inverse Constrained Reinforcement Learning
by: Yue, Bo, et al.
Published: (2024)
by: Yue, Bo, et al.
Published: (2024)
The Perils of Optimizing Learned Reward Functions: Low Training Error Does Not Guarantee Low Regret
by: Fluri, Lukas, et al.
Published: (2024)
by: Fluri, Lukas, et al.
Published: (2024)
The Terminal Representation in Reinforcement Learning
by: Esterhuysen, Amir, et al.
Published: (2026)
by: Esterhuysen, Amir, et al.
Published: (2026)
Kriging and Gaussian Process Interpolation for Georeferenced Data Augmentation
by: Ferber, Frédérick Fabre, et al.
Published: (2025)
by: Ferber, Frédérick Fabre, et al.
Published: (2025)
Bridging Distributional and Risk-sensitive Reinforcement Learning with Provable Regret Bounds
by: Liang, Hao, et al.
Published: (2022)
by: Liang, Hao, et al.
Published: (2022)
Provably Efficient Reward Transfer in Reinforcement Learning with Discrete Markov Decision Processes
by: Vora, Kevin, et al.
Published: (2025)
by: Vora, Kevin, et al.
Published: (2025)
Exploration by Random Reward Perturbation
by: Ma, Haozhe, et al.
Published: (2025)
by: Ma, Haozhe, et al.
Published: (2025)
Near-Optimal Reinforcement Learning with Shuffle Differential Privacy
by: Bai, Shaojie, et al.
Published: (2024)
by: Bai, Shaojie, et al.
Published: (2024)
Globally Optimal Hierarchical Reinforcement Learning for Linearly-Solvable Markov Decision Processes
by: Infante, Guillermo, et al.
Published: (2021)
by: Infante, Guillermo, et al.
Published: (2021)
Robust Yet Efficient Conformal Prediction Sets
by: Zargarbashi, Soroush H., et al.
Published: (2024)
by: Zargarbashi, Soroush H., et al.
Published: (2024)
Advantage Shaping as Surrogate Reward Maximization: Unifying Pass@K Policy Gradients
by: Thrampoulidis, Christos, et al.
Published: (2025)
by: Thrampoulidis, Christos, et al.
Published: (2025)
Provably Adaptive Average Reward Reinforcement Learning for Metric Spaces
by: Kar, Avik, et al.
Published: (2024)
by: Kar, Avik, et al.
Published: (2024)
Efficient Reinforcement Learning in Probabilistic Reward Machines
by: Lin, Xiaofeng, et al.
Published: (2024)
by: Lin, Xiaofeng, et al.
Published: (2024)
Regret Bounds and Reinforcement Learning Exploration of EXP-based Algorithms
by: Xu, Mengfan, et al.
Published: (2020)
by: Xu, Mengfan, et al.
Published: (2020)
Efficient Exploration at Scale
by: Asghari, Seyed Mohammad, et al.
Published: (2026)
by: Asghari, Seyed Mohammad, et al.
Published: (2026)
Kernel-Based Function Approximation for Average Reward Reinforcement Learning: An Optimist No-Regret Algorithm
by: Vakili, Sattar, et al.
Published: (2024)
by: Vakili, Sattar, et al.
Published: (2024)
Scalable Constrained Multi-Agent Reinforcement Learning via State Augmentation and Consensus for Separable Dynamics
by: Amaya-Corredor, Santiago, et al.
Published: (2026)
by: Amaya-Corredor, Santiago, et al.
Published: (2026)
Provable Reward-Agnostic Preference-Based Reinforcement Learning
by: Zhan, Wenhao, et al.
Published: (2023)
by: Zhan, Wenhao, et al.
Published: (2023)
Interpolation pour l'augmentation de donnees : Application à la gestion des adventices de la canne a sucre a la Reunion
by: Ferber, Frederick Fabre, et al.
Published: (2025)
by: Ferber, Frederick Fabre, et al.
Published: (2025)
MIR: Efficient Exploration in Episodic Multi-Agent Reinforcement Learning via Mutual Intrinsic Reward
by: Chen, Kesheng, et al.
Published: (2025)
by: Chen, Kesheng, et al.
Published: (2025)
Regret Analysis of Policy Gradient Algorithm for Infinite Horizon Average Reward Markov Decision Processes
by: Bai, Qinbo, et al.
Published: (2023)
by: Bai, Qinbo, et al.
Published: (2023)
Confidence-Controlled Exploration: Efficient Sparse-Reward Policy Learning for Robot Navigation
by: Patel, Bhrij, et al.
Published: (2023)
by: Patel, Bhrij, et al.
Published: (2023)
Neural Reward Machines
by: Umili, Elena, et al.
Published: (2024)
by: Umili, Elena, et al.
Published: (2024)
Numeric Reward Machines
by: Levina, Kristina, et al.
Published: (2024)
by: Levina, Kristina, et al.
Published: (2024)
Similar Items
-
How to Shrink Confidence Sets for Many Equivalent Discrete Distributions?
by: Maillard, Odalric-Ambrym, et al.
Published: (2024) -
Leveraging priors on distribution functions for multi-arm bandits
by: Vashishtha, Sumit, et al.
Published: (2025) -
How Hard is it to Confuse a World Model?
by: Radji, Waris, et al.
Published: (2025) -
The regret lower bound for communicating Markov Decision Processes
by: Boone, Victor, et al.
Published: (2025) -
The Confusing Instance Principle for Online Linear Quadratic Control
by: Radji, Waris, et al.
Published: (2025)