Compositional Instruction Following with Language Models and Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cohen, Vanya, Tasse, Geraud Nangue, Gopalan, Nakul, James, Steven, Gombolay, Matthew, Mooney, Ray, Rosman, Benjamin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913660889202688
author Cohen, Vanya
Tasse, Geraud Nangue
Gopalan, Nakul
James, Steven
Gombolay, Matthew
Mooney, Ray
Rosman, Benjamin
author_facet Cohen, Vanya
Tasse, Geraud Nangue
Gopalan, Nakul
James, Steven
Gombolay, Matthew
Mooney, Ray
Rosman, Benjamin
contents Combining reinforcement learning with language grounding is challenging as the agent needs to explore the environment while simultaneously learning multiple language-conditioned tasks. To address this, we introduce a novel method: the compositionally-enabled reinforcement learning language agent (CERLLA). Our method reduces the sample complexity of tasks specified with language by leveraging compositional policy representations and a semantic parser trained using reinforcement learning and in-context learning. We evaluate our approach in an environment requiring function approximation and demonstrate compositional generalization to novel tasks. Our method significantly outperforms the previous best non-compositional baseline in terms of sample complexity on 162 tasks designed to test compositional generalization. Our model attains a higher success rate and learns in fewer steps than the non-compositional baseline. It reaches a success rate equal to an oracle policy's upper-bound performance of 92%. With the same number of environment steps, the baseline only reaches a success rate of 80%.
format Preprint
id arxiv_https___arxiv_org_abs_2501_12539
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Compositional Instruction Following with Language Models and Reinforcement Learning
Cohen, Vanya
Tasse, Geraud Nangue
Gopalan, Nakul
James, Steven
Gombolay, Matthew
Mooney, Ray
Rosman, Benjamin
Machine Learning
Computation and Language
Combining reinforcement learning with language grounding is challenging as the agent needs to explore the environment while simultaneously learning multiple language-conditioned tasks. To address this, we introduce a novel method: the compositionally-enabled reinforcement learning language agent (CERLLA). Our method reduces the sample complexity of tasks specified with language by leveraging compositional policy representations and a semantic parser trained using reinforcement learning and in-context learning. We evaluate our approach in an environment requiring function approximation and demonstrate compositional generalization to novel tasks. Our method significantly outperforms the previous best non-compositional baseline in terms of sample complexity on 162 tasks designed to test compositional generalization. Our model attains a higher success rate and learns in fewer steps than the non-compositional baseline. It reaches a success rate equal to an oracle policy's upper-bound performance of 92%. With the same number of environment steps, the baseline only reaches a success rate of 80%.
title Compositional Instruction Following with Language Models and Reinforcement Learning
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2501.12539