Expediting Reinforcement Learning by Incorporating Knowledge About Temporal Causality in the Environment

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Corazza, Jan, Aria, Hadi Partovi, Neider, Daniel, Xu, Zhe
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908599456890880
author Corazza, Jan
Aria, Hadi Partovi
Neider, Daniel
Xu, Zhe
author_facet Corazza, Jan
Aria, Hadi Partovi
Neider, Daniel
Xu, Zhe
contents Reinforcement learning (RL) algorithms struggle with learning optimal policies for tasks where reward feedback is sparse and depends on a complex sequence of events in the environment. Probabilistic reward machines (PRMs) are finite-state formalisms that can capture temporal dependencies in the reward signal, along with nondeterministic task outcomes. While special RL algorithms can exploit this finite-state structure to expedite learning, PRMs remain difficult to modify and design by hand. This hinders the already difficult tasks of utilizing high-level causal knowledge about the environment, and transferring the reward formalism into a new domain with a different causal structure. This paper proposes a novel method to incorporate causal information in the form of Temporal Logic-based Causal Diagrams into the reward formalism, thereby expediting policy learning and aiding the transfer of task specifications to new environments. Furthermore, we provide a theoretical result about convergence to optimal policy for our method, and demonstrate its strengths empirically.
format Preprint
id arxiv_https___arxiv_org_abs_2510_15456
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Expediting Reinforcement Learning by Incorporating Knowledge About Temporal Causality in the Environment
Corazza, Jan
Aria, Hadi Partovi
Neider, Daniel
Xu, Zhe
Machine Learning
Artificial Intelligence
Reinforcement learning (RL) algorithms struggle with learning optimal policies for tasks where reward feedback is sparse and depends on a complex sequence of events in the environment. Probabilistic reward machines (PRMs) are finite-state formalisms that can capture temporal dependencies in the reward signal, along with nondeterministic task outcomes. While special RL algorithms can exploit this finite-state structure to expedite learning, PRMs remain difficult to modify and design by hand. This hinders the already difficult tasks of utilizing high-level causal knowledge about the environment, and transferring the reward formalism into a new domain with a different causal structure. This paper proposes a novel method to incorporate causal information in the form of Temporal Logic-based Causal Diagrams into the reward formalism, thereby expediting policy learning and aiding the transfer of task specifications to new environments. Furthermore, we provide a theoretical result about convergence to optimal policy for our method, and demonstrate its strengths empirically.
title Expediting Reinforcement Learning by Incorporating Knowledge About Temporal Causality in the Environment
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2510.15456