Exploration and Exploitation Errors Are Measurable for Language Model Agents

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Park, Jaden, Kim, Jungtaek, Jeong, Jongwon, Nowak, Robert D., Lee, Kangwook, Lee, Yong Jae
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918446995865600
author Park, Jaden
Kim, Jungtaek
Jeong, Jongwon
Nowak, Robert D.
Lee, Kangwook
Lee, Yong Jae
author_facet Park, Jaden
Kim, Jungtaek
Jeong, Jongwon
Nowak, Robert D.
Lee, Kangwook
Lee, Yong Jae
contents Language Model (LM) agents are increasingly used in complex open-ended decision-making tasks, from AI coding to physical AI. A core requirement in these settings is the ability to both explore the problem space and exploit acquired knowledge effectively. However, systematically distinguishing and quantifying exploration and exploitation from observed actions without access to the agent's internal policy remains challenging. To address this, we design controllable environments inspired by practical embodied AI scenarios. Each environment consists of a partially observable 2D grid map and an unknown task Directed Acyclic Graph (DAG). The map generation can be programmatically adjusted to emphasize exploration or exploitation difficulty. To enable policy-agnostic evaluation, we design a metric to quantify exploration and exploitation errors from agent's actions. We evaluate a variety of frontier LM agents and find that even state-of-the-art models struggle on our task, with different models exhibiting distinct failure modes. We further observe that reasoning models solve the task more effectively and show both exploration and exploitation can be significantly improved through minimal harness engineering. We release our code \href{https://github.com/jjj-madison/measurable-explore-exploit}{here}.
format Preprint
id arxiv_https___arxiv_org_abs_2604_13151
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Exploration and Exploitation Errors Are Measurable for Language Model Agents
Park, Jaden
Kim, Jungtaek
Jeong, Jongwon
Nowak, Robert D.
Lee, Kangwook
Lee, Yong Jae
Artificial Intelligence
Language Model (LM) agents are increasingly used in complex open-ended decision-making tasks, from AI coding to physical AI. A core requirement in these settings is the ability to both explore the problem space and exploit acquired knowledge effectively. However, systematically distinguishing and quantifying exploration and exploitation from observed actions without access to the agent's internal policy remains challenging. To address this, we design controllable environments inspired by practical embodied AI scenarios. Each environment consists of a partially observable 2D grid map and an unknown task Directed Acyclic Graph (DAG). The map generation can be programmatically adjusted to emphasize exploration or exploitation difficulty. To enable policy-agnostic evaluation, we design a metric to quantify exploration and exploitation errors from agent's actions. We evaluate a variety of frontier LM agents and find that even state-of-the-art models struggle on our task, with different models exhibiting distinct failure modes. We further observe that reasoning models solve the task more effectively and show both exploration and exploitation can be significantly improved through minimal harness engineering. We release our code \href{https://github.com/jjj-madison/measurable-explore-exploit}{here}.
title Exploration and Exploitation Errors Are Measurable for Language Model Agents
topic Artificial Intelligence
url https://arxiv.org/abs/2604.13151