Zero-Shot Reinforcement Learning Under Partial Observability

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jeen, Scott, Bewley, Tom, Cullen, Jonathan M.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915350127312896
author Jeen, Scott
Bewley, Tom
Cullen, Jonathan M.
author_facet Jeen, Scott
Bewley, Tom
Cullen, Jonathan M.
contents Recent work has shown that, under certain assumptions, zero-shot reinforcement learning (RL) methods can generalise to any unseen task in an environment after reward-free pre-training. Access to Markov states is one such assumption, yet, in many real-world applications, the Markov state is only partially observable. Here, we explore how the performance of standard zero-shot RL methods degrades when subjected to partially observability, and show that, as in single-task RL, memory-based architectures are an effective remedy. We evaluate our memory-based zero-shot RL methods in domains where the states, rewards and a change in dynamics are partially observed, and show improved performance over memory-free baselines. Our code is open-sourced via: https://enjeeneer.io/projects/bfms-with-memory/.
format Preprint
id arxiv_https___arxiv_org_abs_2506_15446
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Zero-Shot Reinforcement Learning Under Partial Observability
Jeen, Scott
Bewley, Tom
Cullen, Jonathan M.
Machine Learning
Artificial Intelligence
Recent work has shown that, under certain assumptions, zero-shot reinforcement learning (RL) methods can generalise to any unseen task in an environment after reward-free pre-training. Access to Markov states is one such assumption, yet, in many real-world applications, the Markov state is only partially observable. Here, we explore how the performance of standard zero-shot RL methods degrades when subjected to partially observability, and show that, as in single-task RL, memory-based architectures are an effective remedy. We evaluate our memory-based zero-shot RL methods in domains where the states, rewards and a change in dynamics are partially observed, and show improved performance over memory-free baselines. Our code is open-sourced via: https://enjeeneer.io/projects/bfms-with-memory/.
title Zero-Shot Reinforcement Learning Under Partial Observability
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2506.15446