ARM-FM: Automated Reward Machines via Foundation Models for Compositional Reinforcement Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Castanyer, Roger Creus, Mohamed, Faisal, Castro, Pablo Samuel, Neary, Cyrus, Berseth, Glen
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912953315360768
author Castanyer, Roger Creus
Mohamed, Faisal
Castro, Pablo Samuel
Neary, Cyrus
Berseth, Glen
author_facet Castanyer, Roger Creus
Mohamed, Faisal
Castro, Pablo Samuel
Neary, Cyrus
Berseth, Glen
contents Reinforcement learning (RL) algorithms are highly sensitive to reward function specification, which remains a central challenge limiting their broad applicability. We present ARM-FM: Automated Reward Machines via Foundation Models, a framework for automated, compositional reward design in RL that leverages the high-level reasoning capabilities of foundation models (FMs). Reward machines (RMs) -- an automata-based formalism for reward specification -- are used as the mechanism for RL objective specification, and are automatically constructed via the use of FMs. The structured formalism of RMs yields effective task decompositions, while the use of FMs enables objective specifications in natural language. Concretely, we (i) use FMs to automatically generate RMs from natural language specifications; (ii) associate language embeddings with each RM automata-state to enable generalization across tasks; and (iii) provide empirical evidence of ARM-FM's effectiveness in a diverse suite of challenging environments, including evidence of zero-shot generalization.
format Preprint
id arxiv_https___arxiv_org_abs_2510_14176
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ARM-FM: Automated Reward Machines via Foundation Models for Compositional Reinforcement Learning
Castanyer, Roger Creus
Mohamed, Faisal
Castro, Pablo Samuel
Neary, Cyrus
Berseth, Glen
Artificial Intelligence
Machine Learning
Reinforcement learning (RL) algorithms are highly sensitive to reward function specification, which remains a central challenge limiting their broad applicability. We present ARM-FM: Automated Reward Machines via Foundation Models, a framework for automated, compositional reward design in RL that leverages the high-level reasoning capabilities of foundation models (FMs). Reward machines (RMs) -- an automata-based formalism for reward specification -- are used as the mechanism for RL objective specification, and are automatically constructed via the use of FMs. The structured formalism of RMs yields effective task decompositions, while the use of FMs enables objective specifications in natural language. Concretely, we (i) use FMs to automatically generate RMs from natural language specifications; (ii) associate language embeddings with each RM automata-state to enable generalization across tasks; and (iii) provide empirical evidence of ARM-FM's effectiveness in a diverse suite of challenging environments, including evidence of zero-shot generalization.
title ARM-FM: Automated Reward Machines via Foundation Models for Compositional Reinforcement Learning
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2510.14176