Active Reward Machine Inference From Raw State Trajectories

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Shehab, Mohamad Louai, Aspeel, Antoine, Ozay, Necmiye
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913016669274112
author Shehab, Mohamad Louai
Aspeel, Antoine
Ozay, Necmiye
author_facet Shehab, Mohamad Louai
Aspeel, Antoine
Ozay, Necmiye
contents Reward machines are automaton-like structures that capture the memory required to accomplish a multi-stage task. When combined with reinforcement learning or optimal control methods, they can be used to synthesize robot policies to achieve such tasks. However, specifying a reward machine by hand, including a labeling function capturing high-level features that the decisions are based on, can be a daunting task. This paper deals with the problem of learning reward machines directly from raw state and policy information. As opposed to existing works, we assume no access to observations of rewards, labels, or machine nodes, and show what trajectory data is sufficient for learning the reward machine in this information-scarce regime. We then extend the result to an active learning setting where we incrementally query trajectory extensions to improve data (and indirectly computational) efficiency. Results are demonstrated with several grid world examples.
format Preprint
id arxiv_https___arxiv_org_abs_2604_07480
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Active Reward Machine Inference From Raw State Trajectories
Shehab, Mohamad Louai
Aspeel, Antoine
Ozay, Necmiye
Robotics
Artificial Intelligence
Formal Languages and Automata Theory
Reward machines are automaton-like structures that capture the memory required to accomplish a multi-stage task. When combined with reinforcement learning or optimal control methods, they can be used to synthesize robot policies to achieve such tasks. However, specifying a reward machine by hand, including a labeling function capturing high-level features that the decisions are based on, can be a daunting task. This paper deals with the problem of learning reward machines directly from raw state and policy information. As opposed to existing works, we assume no access to observations of rewards, labels, or machine nodes, and show what trajectory data is sufficient for learning the reward machine in this information-scarce regime. We then extend the result to an active learning setting where we incrementally query trajectory extensions to improve data (and indirectly computational) efficiency. Results are demonstrated with several grid world examples.
title Active Reward Machine Inference From Raw State Trajectories
topic Robotics
Artificial Intelligence
Formal Languages and Automata Theory
url https://arxiv.org/abs/2604.07480