Zero-Shot Reinforcement Learning Under Partial Observability
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915350127312896 |
|---|---|
| author | Jeen, Scott Bewley, Tom Cullen, Jonathan M. |
| author_facet | Jeen, Scott Bewley, Tom Cullen, Jonathan M. |
| contents | Recent work has shown that, under certain assumptions, zero-shot reinforcement learning (RL) methods can generalise to any unseen task in an environment after reward-free pre-training. Access to Markov states is one such assumption, yet, in many real-world applications, the Markov state is only partially observable. Here, we explore how the performance of standard zero-shot RL methods degrades when subjected to partially observability, and show that, as in single-task RL, memory-based architectures are an effective remedy. We evaluate our memory-based zero-shot RL methods in domains where the states, rewards and a change in dynamics are partially observed, and show improved performance over memory-free baselines. Our code is open-sourced via: https://enjeeneer.io/projects/bfms-with-memory/. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_15446 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Zero-Shot Reinforcement Learning Under Partial Observability Jeen, Scott Bewley, Tom Cullen, Jonathan M. Machine Learning Artificial Intelligence Recent work has shown that, under certain assumptions, zero-shot reinforcement learning (RL) methods can generalise to any unseen task in an environment after reward-free pre-training. Access to Markov states is one such assumption, yet, in many real-world applications, the Markov state is only partially observable. Here, we explore how the performance of standard zero-shot RL methods degrades when subjected to partially observability, and show that, as in single-task RL, memory-based architectures are an effective remedy. We evaluate our memory-based zero-shot RL methods in domains where the states, rewards and a change in dynamics are partially observed, and show improved performance over memory-free baselines. Our code is open-sourced via: https://enjeeneer.io/projects/bfms-with-memory/. |
| title | Zero-Shot Reinforcement Learning Under Partial Observability |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2506.15446 |