Meta-RL Induces Exploration in Language Agents

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jiang, Yulun, Jiang, Liangze, Teney, Damien, Moor, Michael, Brbic, Maria
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917320515911680
author Jiang, Yulun
Jiang, Liangze
Teney, Damien
Moor, Michael
Brbic, Maria
author_facet Jiang, Yulun
Jiang, Liangze
Teney, Damien
Moor, Michael
Brbic, Maria
contents Reinforcement learning (RL) has enabled the training of large language model (LLM) agents to interact with the environment and to solve multi-turn long-horizon tasks. However, the RL-trained agents often struggle in tasks that require active exploration and fail to efficiently adapt from trial-and-error experiences. In this paper, we present LaMer, a general Meta-RL framework that enables LLM agents to actively explore and learn from the environment feedback at test time. LaMer consists of two key components: (i) a cross-episode training framework to encourage exploration and long-term rewards optimization; and (ii) in-context policy adaptation via reflection, allowing the agent to adapt their policy from task feedback signal without gradient update. Experiments across diverse environments show that LaMer significantly improves performance over RL baselines, with 11%, 14%, and 19% performance gains on Sokoban, MineSweeper and Webshop, respectively. Moreover, LaMer also demonstrates better generalization to more challenging or previously unseen tasks compared to the RL-trained agents. Overall, our results demonstrate that Meta-RL provides a principled approach to induce exploration in language agents, enabling more robust adaptation to novel environments through learned exploration strategies.
format Preprint
id arxiv_https___arxiv_org_abs_2512_16848
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Meta-RL Induces Exploration in Language Agents
Jiang, Yulun
Jiang, Liangze
Teney, Damien
Moor, Michael
Brbic, Maria
Machine Learning
Artificial Intelligence
Reinforcement learning (RL) has enabled the training of large language model (LLM) agents to interact with the environment and to solve multi-turn long-horizon tasks. However, the RL-trained agents often struggle in tasks that require active exploration and fail to efficiently adapt from trial-and-error experiences. In this paper, we present LaMer, a general Meta-RL framework that enables LLM agents to actively explore and learn from the environment feedback at test time. LaMer consists of two key components: (i) a cross-episode training framework to encourage exploration and long-term rewards optimization; and (ii) in-context policy adaptation via reflection, allowing the agent to adapt their policy from task feedback signal without gradient update. Experiments across diverse environments show that LaMer significantly improves performance over RL baselines, with 11%, 14%, and 19% performance gains on Sokoban, MineSweeper and Webshop, respectively. Moreover, LaMer also demonstrates better generalization to more challenging or previously unseen tasks compared to the RL-trained agents. Overall, our results demonstrate that Meta-RL provides a principled approach to induce exploration in language agents, enabling more robust adaptation to novel environments through learned exploration strategies.
title Meta-RL Induces Exploration in Language Agents
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2512.16848