Compositional Instruction Following with Language Models and Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913660889202688 |
|---|---|
| author | Cohen, Vanya Tasse, Geraud Nangue Gopalan, Nakul James, Steven Gombolay, Matthew Mooney, Ray Rosman, Benjamin |
| author_facet | Cohen, Vanya Tasse, Geraud Nangue Gopalan, Nakul James, Steven Gombolay, Matthew Mooney, Ray Rosman, Benjamin |
| contents | Combining reinforcement learning with language grounding is challenging as the agent needs to explore the environment while simultaneously learning multiple language-conditioned tasks. To address this, we introduce a novel method: the compositionally-enabled reinforcement learning language agent (CERLLA). Our method reduces the sample complexity of tasks specified with language by leveraging compositional policy representations and a semantic parser trained using reinforcement learning and in-context learning. We evaluate our approach in an environment requiring function approximation and demonstrate compositional generalization to novel tasks. Our method significantly outperforms the previous best non-compositional baseline in terms of sample complexity on 162 tasks designed to test compositional generalization. Our model attains a higher success rate and learns in fewer steps than the non-compositional baseline. It reaches a success rate equal to an oracle policy's upper-bound performance of 92%. With the same number of environment steps, the baseline only reaches a success rate of 80%. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2501_12539 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Compositional Instruction Following with Language Models and Reinforcement Learning Cohen, Vanya Tasse, Geraud Nangue Gopalan, Nakul James, Steven Gombolay, Matthew Mooney, Ray Rosman, Benjamin Machine Learning Computation and Language Combining reinforcement learning with language grounding is challenging as the agent needs to explore the environment while simultaneously learning multiple language-conditioned tasks. To address this, we introduce a novel method: the compositionally-enabled reinforcement learning language agent (CERLLA). Our method reduces the sample complexity of tasks specified with language by leveraging compositional policy representations and a semantic parser trained using reinforcement learning and in-context learning. We evaluate our approach in an environment requiring function approximation and demonstrate compositional generalization to novel tasks. Our method significantly outperforms the previous best non-compositional baseline in terms of sample complexity on 162 tasks designed to test compositional generalization. Our model attains a higher success rate and learns in fewer steps than the non-compositional baseline. It reaches a success rate equal to an oracle policy's upper-bound performance of 92%. With the same number of environment steps, the baseline only reaches a success rate of 80%. |
| title | Compositional Instruction Following with Language Models and Reinforcement Learning |
| topic | Machine Learning Computation and Language |
| url | https://arxiv.org/abs/2501.12539 |