Evaluating the Goal-Directedness of Large Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Everitt, Tom, Garbacea, Cristina, Bellot, Alexis, Richens, Jonathan, Papadatos, Henry, Campos, Siméon, Shah, Rohin
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913796369416192
author Everitt, Tom
Garbacea, Cristina
Bellot, Alexis
Richens, Jonathan
Papadatos, Henry
Campos, Siméon
Shah, Rohin
author_facet Everitt, Tom
Garbacea, Cristina
Bellot, Alexis
Richens, Jonathan
Papadatos, Henry
Campos, Siméon
Shah, Rohin
contents To what extent do LLMs use their capabilities towards their given goal? We take this as a measure of their goal-directedness. We evaluate goal-directedness on tasks that require information gathering, cognitive effort, and plan execution, where we use subtasks to infer each model's relevant capabilities. Our evaluations of LLMs from Google DeepMind, OpenAI, and Anthropic show that goal-directedness is relatively consistent across tasks, differs from task performance, and is only moderately sensitive to motivational prompts. Notably, most models are not fully goal-directed. We hope our goal-directedness evaluations will enable better monitoring of LLM progress, and enable more deliberate design choices of agentic properties in LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2504_11844
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating the Goal-Directedness of Large Language Models
Everitt, Tom
Garbacea, Cristina
Bellot, Alexis
Richens, Jonathan
Papadatos, Henry
Campos, Siméon
Shah, Rohin
Artificial Intelligence
Computation and Language
Machine Learning
To what extent do LLMs use their capabilities towards their given goal? We take this as a measure of their goal-directedness. We evaluate goal-directedness on tasks that require information gathering, cognitive effort, and plan execution, where we use subtasks to infer each model's relevant capabilities. Our evaluations of LLMs from Google DeepMind, OpenAI, and Anthropic show that goal-directedness is relatively consistent across tasks, differs from task performance, and is only moderately sensitive to motivational prompts. Notably, most models are not fully goal-directed. We hope our goal-directedness evaluations will enable better monitoring of LLM progress, and enable more deliberate design choices of agentic properties in LLMs.
title Evaluating the Goal-Directedness of Large Language Models
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2504.11844