Pure Exploration for a Good Policy in Reinforcement Learning with Bandit Feedback
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Zitian, Cheung, Wang Chi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Episodic Contextual Bandits with Knapsacks under Conversion Models
by: Cheung, Wang Chi, et al.
Published: (2025)
by: Cheung, Wang Chi, et al.
Published: (2025)
Learning with a Budget: Identifying the Best Arm with Resource Constraints
by: Li, Zitian, et al.
Published: (2026)
by: Li, Zitian, et al.
Published: (2026)
Closing the Gap on the Sample Complexity of 1-Identification
by: Li, Zitian, et al.
Published: (2026)
by: Li, Zitian, et al.
Published: (2026)
Best Arm Identification with Resource Constraints
by: Li, Zitian, et al.
Published: (2024)
by: Li, Zitian, et al.
Published: (2024)
Near Optimal Non-asymptotic Sample Complexity of 1-Identification
by: Li, Zitian, et al.
Published: (2025)
by: Li, Zitian, et al.
Published: (2025)
Pure Exploration in Asynchronous Federated Bandits
by: Wang, Zichen, et al.
Published: (2023)
by: Wang, Zichen, et al.
Published: (2023)
Pure Exploration in Bandits with Linear Constraints
by: Carlsson, Emil, et al.
Published: (2023)
by: Carlsson, Emil, et al.
Published: (2023)
The Batch Complexity of Bandit Pure Exploration
by: Tuynman, Adrienne, et al.
Published: (2025)
by: Tuynman, Adrienne, et al.
Published: (2025)
Near Optimal Pure Exploration in Logistic Bandits
by: Rivera, Eduardo Ochoa, et al.
Published: (2024)
by: Rivera, Eduardo Ochoa, et al.
Published: (2024)
Pure Exploration with Feedback Graphs
by: Russo, Alessio, et al.
Published: (2025)
by: Russo, Alessio, et al.
Published: (2025)
Reward Maximization for Pure Exploration: Minimax Optimal Good Arm Identification for Nonparametric Multi-Armed Bandits
by: Cho, Brian, et al.
Published: (2024)
by: Cho, Brian, et al.
Published: (2024)
Online Bandits with (Biased) Offline Data: Adaptive Learning under Distribution Mismatch
by: Cheung, Wang Chi, et al.
Published: (2024)
by: Cheung, Wang Chi, et al.
Published: (2024)
Pure Exploration under Mediators' Feedback
by: Poiani, Riccardo, et al.
Published: (2023)
by: Poiani, Riccardo, et al.
Published: (2023)
Multi-thresholding Good Arm Identification with Bandit Feedback
by: Jiang, Xuanke, et al.
Published: (2025)
by: Jiang, Xuanke, et al.
Published: (2025)
Provably Efficient Reinforcement Learning for Adversarial Restless Multi-Armed Bandits with Unknown Transitions and Bandit Feedback
by: Xiong, Guojun, et al.
Published: (2024)
by: Xiong, Guojun, et al.
Published: (2024)
A Fast Algorithm for the Real-Valued Combinatorial Pure Exploration of Multi-Armed Bandit
by: Nakamura, Shintaro, et al.
Published: (2023)
by: Nakamura, Shintaro, et al.
Published: (2023)
Pure Exploration Beyond Reward Feedback: The Role of Post-Action Context
by: Shahverdikondori, Mohammad, et al.
Published: (2025)
by: Shahverdikondori, Mohammad, et al.
Published: (2025)
Learning to Route LLMs from Bandit Feedback: One Policy, Many Trade-offs
by: Wei, Wang, et al.
Published: (2025)
by: Wei, Wang, et al.
Published: (2025)
In-Context Learning for Pure Exploration
by: Russo, Alessio, et al.
Published: (2025)
by: Russo, Alessio, et al.
Published: (2025)
Best-of-Both-Worlds Policy Optimization for CMDPs with Bandit Feedback
by: Stradi, Francesco Emanuele, et al.
Published: (2024)
by: Stradi, Francesco Emanuele, et al.
Published: (2024)
Constrained Feedback Learning for Non-Stationary Multi-Armed Bandits
by: Li, Shaoang, et al.
Published: (2025)
by: Li, Shaoang, et al.
Published: (2025)
Towards Off-Policy Reinforcement Learning for Ranking Policies with Human Feedback
by: Xiao, Teng, et al.
Published: (2024)
by: Xiao, Teng, et al.
Published: (2024)
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback
by: Zhang, Zhen-Yu, et al.
Published: (2026)
by: Zhang, Zhen-Yu, et al.
Published: (2026)
Learning Equilibria in Matching Games with Bandit Feedback
by: Athanasopoulos, Andreas, et al.
Published: (2025)
by: Athanasopoulos, Andreas, et al.
Published: (2025)
Off-Policy Safe Reinforcement Learning with Constrained Optimistic Exploration
by: Li, Guopeng, et al.
Published: (2026)
by: Li, Guopeng, et al.
Published: (2026)
Uniform Last-Iterate Guarantee for Bandits and Reinforcement Learning
by: Liu, Junyan, et al.
Published: (2024)
by: Liu, Junyan, et al.
Published: (2024)
Graph Feedback Bandits with Similar Arms
by: Qi, Han, et al.
Published: (2024)
by: Qi, Han, et al.
Published: (2024)
Pure Exploration with Infinite Answers
by: Poiani, Riccardo, et al.
Published: (2025)
by: Poiani, Riccardo, et al.
Published: (2025)
Preference-based Pure Exploration
by: Shukla, Apurv, et al.
Published: (2024)
by: Shukla, Apurv, et al.
Published: (2024)
The Best Arm Evades: Near-optimal Multi-pass Streaming Lower Bounds for Pure Exploration in Multi-armed Bandits
by: Assadi, Sepehr, et al.
Published: (2023)
by: Assadi, Sepehr, et al.
Published: (2023)
Performative Prediction with Bandit Feedback: Learning through Reparameterization
by: Chen, Yatong, et al.
Published: (2023)
by: Chen, Yatong, et al.
Published: (2023)
Towards Efficient Online Exploration for Reinforcement Learning with Human Feedback
by: Li, Gen, et al.
Published: (2025)
by: Li, Gen, et al.
Published: (2025)
In-Context Learning for Pure Exploration in Continuous Spaces
by: Russo, Alessio, et al.
Published: (2026)
by: Russo, Alessio, et al.
Published: (2026)
Encouraging Good Processes Without the Need for Good Answers: Reinforcement Learning for LLM Agent Planning
by: Li, Zhiwei, et al.
Published: (2025)
by: Li, Zhiwei, et al.
Published: (2025)
Pessimistic Risk-Aware Policy Learning in Contextual Bandits
by: Wan, Yilong, et al.
Published: (2026)
by: Wan, Yilong, et al.
Published: (2026)
Consistency Models as a Rich and Efficient Policy Class for Reinforcement Learning
by: Ding, Zihan, et al.
Published: (2023)
by: Ding, Zihan, et al.
Published: (2023)
LPPG-RL: Lexicographically Projected Policy Gradient Reinforcement Learning with Subproblem Exploration
by: Qiu, Ruiyu, et al.
Published: (2025)
by: Qiu, Ruiyu, et al.
Published: (2025)
Infrequent Exploration in Linear Bandits
by: Lee, Harin, et al.
Published: (2025)
by: Lee, Harin, et al.
Published: (2025)
Nearest Neighbour with Bandit Feedback
by: Pasteris, Stephen, et al.
Published: (2023)
by: Pasteris, Stephen, et al.
Published: (2023)
Adversarial Bandits with Multi-User Delayed Feedback: Theory and Application
by: Li, Yandi, et al.
Published: (2023)
by: Li, Yandi, et al.
Published: (2023)
Similar Items
-
Episodic Contextual Bandits with Knapsacks under Conversion Models
by: Cheung, Wang Chi, et al.
Published: (2025) -
Learning with a Budget: Identifying the Best Arm with Resource Constraints
by: Li, Zitian, et al.
Published: (2026) -
Closing the Gap on the Sample Complexity of 1-Identification
by: Li, Zitian, et al.
Published: (2026) -
Best Arm Identification with Resource Constraints
by: Li, Zitian, et al.
Published: (2024) -
Near Optimal Non-asymptotic Sample Complexity of 1-Identification
by: Li, Zitian, et al.
Published: (2025)