Ground-Compose-Reinforce: Grounding Language in Agentic Behaviours using Limited Data

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Andrew C., Klassen, Toryn Q., Wang, Andrew, Alamdari, Parand A., McIlraith, Sheila A.
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908613799313408
author Li, Andrew C.
Klassen, Toryn Q.
Wang, Andrew
Alamdari, Parand A.
McIlraith, Sheila A.
author_facet Li, Andrew C.
Klassen, Toryn Q.
Wang, Andrew
Alamdari, Parand A.
McIlraith, Sheila A.
contents Grounding language in perception and action is a key challenge when building situated agents that can interact with humans, or other agents, via language. In the past, addressing this challenge has required manually designing the language grounding or curating massive datasets that associate language with the environment. We propose Ground-Compose-Reinforce, an end-to-end, neurosymbolic framework for training RL agents directly from high-level task specifications--without manually designed reward functions or other domain-specific oracles, and without massive datasets. These task specifications take the form of Reward Machines, automata-based representations that capture high-level task structure and are in some cases autoformalizable from natural language. Critically, we show that Reward Machines can be grounded using limited data by exploiting compositionality. Experiments in a custom Meta-World domain with only 350 labelled pretraining trajectories show that our framework faithfully elicits complex behaviours from high-level specifications--including behaviours that never appear in pretraining--while non-compositional approaches fail.
format Preprint
id arxiv_https___arxiv_org_abs_2507_10741
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Ground-Compose-Reinforce: Grounding Language in Agentic Behaviours using Limited Data
Li, Andrew C.
Klassen, Toryn Q.
Wang, Andrew
Alamdari, Parand A.
McIlraith, Sheila A.
Machine Learning
Artificial Intelligence
I.2.6; I.2.4
Grounding language in perception and action is a key challenge when building situated agents that can interact with humans, or other agents, via language. In the past, addressing this challenge has required manually designing the language grounding or curating massive datasets that associate language with the environment. We propose Ground-Compose-Reinforce, an end-to-end, neurosymbolic framework for training RL agents directly from high-level task specifications--without manually designed reward functions or other domain-specific oracles, and without massive datasets. These task specifications take the form of Reward Machines, automata-based representations that capture high-level task structure and are in some cases autoformalizable from natural language. Critically, we show that Reward Machines can be grounded using limited data by exploiting compositionality. Experiments in a custom Meta-World domain with only 350 labelled pretraining trajectories show that our framework faithfully elicits complex behaviours from high-level specifications--including behaviours that never appear in pretraining--while non-compositional approaches fail.
title Ground-Compose-Reinforce: Grounding Language in Agentic Behaviours using Limited Data
topic Machine Learning
Artificial Intelligence
I.2.6; I.2.4
url https://arxiv.org/abs/2507.10741