A Behavioural and Representational Evaluation of Goal-Directedness in Language Model Agents

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Arghal, Raghu, Chen, Fade, Dalton, Niall, Kortukov, Evgenii, McNamara, Calum, Nalmpantis, Angelos, Nirvaan, Moksh, Sarti, Gabriele, Giulianelli, Mario
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914617173737472
author Arghal, Raghu
Chen, Fade
Dalton, Niall
Kortukov, Evgenii
McNamara, Calum
Nalmpantis, Angelos
Nirvaan, Moksh
Sarti, Gabriele
Giulianelli, Mario
author_facet Arghal, Raghu
Chen, Fade
Dalton, Niall
Kortukov, Evgenii
McNamara, Calum
Nalmpantis, Angelos
Nirvaan, Moksh
Sarti, Gabriele
Giulianelli, Mario
contents Understanding an agent's goals helps explain and predict its behaviour, yet there is no established methodology for reliably attributing goals to agentic systems. We propose a framework for evaluating goal-directedness that integrates behavioural evaluation with interpretability-based analyses of models' internal representations. As a case study, we examine an LLM agent navigating a 2D grid world towards a goal state. Behaviourally, we evaluate the agent against optimal policies across varying grid sizes, obstacle densities, and goal structures, finding that performance scales with task difficulty while remaining robust to difficulty-preserving transformations and multi-goal structures. We then use probing methods to decode internal representations of the environment and multi-step action plans. We find that the LLM agent non-linearly encodes a coarse spatial map, preserving approximate task-relevant cues about its position and the goal location; that its actions are broadly consistent with these internal representations; and that reasoning reorganises them, shifting from spatial cues towards immediate action selection. Our findings support the view that introspective examination is required beyond behavioural evaluations to characterise how agents represent and pursue their objectives.
format Preprint
id arxiv_https___arxiv_org_abs_2602_08964
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A Behavioural and Representational Evaluation of Goal-Directedness in Language Model Agents
Arghal, Raghu
Chen, Fade
Dalton, Niall
Kortukov, Evgenii
McNamara, Calum
Nalmpantis, Angelos
Nirvaan, Moksh
Sarti, Gabriele
Giulianelli, Mario
Machine Learning
Artificial Intelligence
Computation and Language
Computers and Society
Understanding an agent's goals helps explain and predict its behaviour, yet there is no established methodology for reliably attributing goals to agentic systems. We propose a framework for evaluating goal-directedness that integrates behavioural evaluation with interpretability-based analyses of models' internal representations. As a case study, we examine an LLM agent navigating a 2D grid world towards a goal state. Behaviourally, we evaluate the agent against optimal policies across varying grid sizes, obstacle densities, and goal structures, finding that performance scales with task difficulty while remaining robust to difficulty-preserving transformations and multi-goal structures. We then use probing methods to decode internal representations of the environment and multi-step action plans. We find that the LLM agent non-linearly encodes a coarse spatial map, preserving approximate task-relevant cues about its position and the goal location; that its actions are broadly consistent with these internal representations; and that reasoning reorganises them, shifting from spatial cues towards immediate action selection. Our findings support the view that introspective examination is required beyond behavioural evaluations to characterise how agents represent and pursue their objectives.
title A Behavioural and Representational Evaluation of Goal-Directedness in Language Model Agents
topic Machine Learning
Artificial Intelligence
Computation and Language
Computers and Society
url https://arxiv.org/abs/2602.08964