Reward Machines for Deep RL in Noisy and Uncertain Environments

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Andrew C., Chen, Zizhao, Klassen, Toryn Q., Vaezipoor, Pashootan, Icarte, Rodrigo Toro, McIlraith, Sheila A.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915103467634688
author Li, Andrew C.
Chen, Zizhao
Klassen, Toryn Q.
Vaezipoor, Pashootan
Icarte, Rodrigo Toro
McIlraith, Sheila A.
author_facet Li, Andrew C.
Chen, Zizhao
Klassen, Toryn Q.
Vaezipoor, Pashootan
Icarte, Rodrigo Toro
McIlraith, Sheila A.
contents Reward Machines provide an automaton-inspired structure for specifying instructions, safety constraints, and other temporally extended reward-worthy behaviour. By exposing the underlying structure of a reward function, they enable the decomposition of an RL task, leading to impressive gains in sample efficiency. Although Reward Machines and similar formal specifications have a rich history of application towards sequential decision-making problems, they critically rely on a ground-truth interpretation of the domain-specific vocabulary that forms the building blocks of the reward function--such ground-truth interpretations are elusive in the real world due in part to partial observability and noisy sensing. In this work, we explore the use of Reward Machines for Deep RL in noisy and uncertain environments. We characterize this problem as a POMDP and propose a suite of RL algorithms that exploit task structure under uncertain interpretation of the domain-specific vocabulary. Through theory and experiments, we expose pitfalls in naive approaches to this problem while simultaneously demonstrating how task structure can be successfully leveraged under noisy interpretations of the vocabulary.
format Preprint
id arxiv_https___arxiv_org_abs_2406_00120
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Reward Machines for Deep RL in Noisy and Uncertain Environments
Li, Andrew C.
Chen, Zizhao
Klassen, Toryn Q.
Vaezipoor, Pashootan
Icarte, Rodrigo Toro
McIlraith, Sheila A.
Machine Learning
Artificial Intelligence
Formal Languages and Automata Theory
I.2.0; I.2.6; I.2.4; F.4.3
Reward Machines provide an automaton-inspired structure for specifying instructions, safety constraints, and other temporally extended reward-worthy behaviour. By exposing the underlying structure of a reward function, they enable the decomposition of an RL task, leading to impressive gains in sample efficiency. Although Reward Machines and similar formal specifications have a rich history of application towards sequential decision-making problems, they critically rely on a ground-truth interpretation of the domain-specific vocabulary that forms the building blocks of the reward function--such ground-truth interpretations are elusive in the real world due in part to partial observability and noisy sensing. In this work, we explore the use of Reward Machines for Deep RL in noisy and uncertain environments. We characterize this problem as a POMDP and propose a suite of RL algorithms that exploit task structure under uncertain interpretation of the domain-specific vocabulary. Through theory and experiments, we expose pitfalls in naive approaches to this problem while simultaneously demonstrating how task structure can be successfully leveraged under noisy interpretations of the vocabulary.
title Reward Machines for Deep RL in Noisy and Uncertain Environments
topic Machine Learning
Artificial Intelligence
Formal Languages and Automata Theory
I.2.0; I.2.6; I.2.4; F.4.3
url https://arxiv.org/abs/2406.00120