Beyond Optimism: Exploration With Partially Observable Rewards
Fuente:
arXiv
Guardado en:
| Autores principales: | Parisi, Simone, Kazemipour, Alireza, Bowling, Michael |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Model-Based Exploration in Monitored Markov Decision Processes
por: Kazemipour, Alireza, et al.
Publicado: (2025)
por: Kazemipour, Alireza, et al.
Publicado: (2025)
Monitored Markov Decision Processes
por: Parisi, Simone, et al.
Publicado: (2024)
por: Parisi, Simone, et al.
Publicado: (2024)
Pure Exploration Beyond Reward Feedback: The Role of Post-Action Context
por: Shahverdikondori, Mohammad, et al.
Publicado: (2025)
por: Shahverdikondori, Mohammad, et al.
Publicado: (2025)
LLMs for Text-Based Exploration and Navigation Under Partial Observability
por: Sandfuchs, Stephan, et al.
Publicado: (2026)
por: Sandfuchs, Stephan, et al.
Publicado: (2026)
Partially Observable Reinforcement Learning with Memory Traces
por: Eberhard, Onno, et al.
Publicado: (2025)
por: Eberhard, Onno, et al.
Publicado: (2025)
Linear Bandits with Partially Observable Features
por: Kim, Wonyoung, et al.
Publicado: (2025)
por: Kim, Wonyoung, et al.
Publicado: (2025)
Inferring Reward Machines and Transition Machines from Partially Observable Markov Decision Processes
por: Wu, Yuly, et al.
Publicado: (2025)
por: Wu, Yuly, et al.
Publicado: (2025)
Improved Regret Bound for Safe Reinforcement Learning via Tighter Cost Pessimism and Reward Optimism
por: Yu, Kihyun, et al.
Publicado: (2024)
por: Yu, Kihyun, et al.
Publicado: (2024)
Thompson Sampling in Partially Observable Contextual Bandits
por: Park, Hongju, et al.
Publicado: (2024)
por: Park, Hongju, et al.
Publicado: (2024)
Partially Observable Contextual Bandits with Linear Payoffs
por: Zeng, Sihan, et al.
Publicado: (2024)
por: Zeng, Sihan, et al.
Publicado: (2024)
Directional Optimism for Safe Linear Bandits
por: Hutchinson, Spencer, et al.
Publicado: (2023)
por: Hutchinson, Spencer, et al.
Publicado: (2023)
Temporal Knowledge-Graph Memory in a Partially Observable Environment
por: Kim, Taewoon, et al.
Publicado: (2024)
por: Kim, Taewoon, et al.
Publicado: (2024)
Provable Partially Observable Reinforcement Learning with Privileged Information
por: Cai, Yang, et al.
Publicado: (2024)
por: Cai, Yang, et al.
Publicado: (2024)
Learning Causal States Under Partial Observability and Perturbation
por: Li, Na, et al.
Publicado: (2025)
por: Li, Na, et al.
Publicado: (2025)
Improved Bounds for Reward-Agnostic and Reward-Free Exploration
por: Ridel, Oran, et al.
Publicado: (2026)
por: Ridel, Oran, et al.
Publicado: (2026)
Minimax Optimal Reinforcement Learning with Quasi-Optimism
por: Lee, Harin, et al.
Publicado: (2025)
por: Lee, Harin, et al.
Publicado: (2025)
Belief-State RWKV for Reinforcement Learning under Partial Observability
por: Xiao, Liu
Publicado: (2026)
por: Xiao, Liu
Publicado: (2026)
Minimax-Optimal Policy Regret in Partially Observable Markov Games
por: Arora, Raman
Publicado: (2026)
por: Arora, Raman
Publicado: (2026)
Why Linear Recurrent Memory Works in Partially Observable Reinforcement Learning
por: Zhao, Yike, et al.
Publicado: (2026)
por: Zhao, Yike, et al.
Publicado: (2026)
Near-Optimal Partially Observable Reinforcement Learning with Partial Online State Information
por: Shi, Ming, et al.
Publicado: (2023)
por: Shi, Ming, et al.
Publicado: (2023)
Optimism in the Face of Ambiguity Principle for Multi-Armed Bandits
por: Li, Mengmeng, et al.
Publicado: (2024)
por: Li, Mengmeng, et al.
Publicado: (2024)
DROP: Distributional and Regular Optimism and Pessimism for Reinforcement Learning
por: Kobayashi, Taisuke
Publicado: (2024)
por: Kobayashi, Taisuke
Publicado: (2024)
Guided Policy Optimization under Partial Observability
por: Li, Yueheng, et al.
Publicado: (2025)
por: Li, Yueheng, et al.
Publicado: (2025)
On the Suboptimality of GP-UCB under Polynomial Effective Optimism
por: Wang, Wenjia, et al.
Publicado: (2023)
por: Wang, Wenjia, et al.
Publicado: (2023)
A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning
por: Adkins, Jacob, et al.
Publicado: (2024)
por: Adkins, Jacob, et al.
Publicado: (2024)
Network Topology Inference from Smooth Signals Under Partial Observability
por: Peng, Chuansen, et al.
Publicado: (2024)
por: Peng, Chuansen, et al.
Publicado: (2024)
Zero-Shot Reinforcement Learning Under Partial Observability
por: Jeen, Scott, et al.
Publicado: (2025)
por: Jeen, Scott, et al.
Publicado: (2025)
Multi-View Causal Representation Learning with Partial Observability
por: Yao, Dingling, et al.
Publicado: (2023)
por: Yao, Dingling, et al.
Publicado: (2023)
Mitigating Partial Observability in Sequential Decision Processes via the Lambda Discrepancy
por: Allen, Cameron, et al.
Publicado: (2024)
por: Allen, Cameron, et al.
Publicado: (2024)
Short-Term-to-Long-Term Memory Transfer for Knowledge Graphs under Partial Observability
por: Kim, Taewoon, et al.
Publicado: (2026)
por: Kim, Taewoon, et al.
Publicado: (2026)
Robustness in the Face of Partial Identifiability in Reward Learning
por: Lazzati, Filippo, et al.
Publicado: (2025)
por: Lazzati, Filippo, et al.
Publicado: (2025)
Beyond Binary: Turning Partial Success into Dense Verifiable Rewards for Reinforcement Learning in Code Generation
por: Wang, Longwen, et al.
Publicado: (2026)
por: Wang, Longwen, et al.
Publicado: (2026)
Exploration by Random Reward Perturbation
por: Ma, Haozhe, et al.
Publicado: (2025)
por: Ma, Haozhe, et al.
Publicado: (2025)
Exploration Through Introspection: A Self-Aware Reward Model
por: Petrowski, Michael, et al.
Publicado: (2026)
por: Petrowski, Michael, et al.
Publicado: (2026)
On the Dynamic Regret of Following the Regularized Leader: Optimism with History Pruning
por: Mhaisen, Naram, et al.
Publicado: (2025)
por: Mhaisen, Naram, et al.
Publicado: (2025)
Learning Interpretable Policies in Hindsight-Observable POMDPs through Partially Supervised Reinforcement Learning
por: Lanier, Michael, et al.
Publicado: (2024)
por: Lanier, Michael, et al.
Publicado: (2024)
Provably Efficient Partially Observable Risk-Sensitive Reinforcement Learning with Hindsight Observation
por: Zhang, Tonghe, et al.
Publicado: (2024)
por: Zhang, Tonghe, et al.
Publicado: (2024)
Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes
por: Arora, Ashok, et al.
Publicado: (2025)
por: Arora, Ashok, et al.
Publicado: (2025)
To Distill or Decide? Understanding the Algorithmic Trade-off in Partially Observable Reinforcement Learning
por: Song, Yuda, et al.
Publicado: (2025)
por: Song, Yuda, et al.
Publicado: (2025)
Streaming Reinforcement Learning under Partial Observability with Real-Time Recurrent Learning
por: Farr, Noah, et al.
Publicado: (2026)
por: Farr, Noah, et al.
Publicado: (2026)
Ejemplares similares
-
Model-Based Exploration in Monitored Markov Decision Processes
por: Kazemipour, Alireza, et al.
Publicado: (2025) -
Monitored Markov Decision Processes
por: Parisi, Simone, et al.
Publicado: (2024) -
Pure Exploration Beyond Reward Feedback: The Role of Post-Action Context
por: Shahverdikondori, Mohammad, et al.
Publicado: (2025) -
LLMs for Text-Based Exploration and Navigation Under Partial Observability
por: Sandfuchs, Stephan, et al.
Publicado: (2026) -
Partially Observable Reinforcement Learning with Memory Traces
por: Eberhard, Onno, et al.
Publicado: (2025)