Shared Lexical Task Representations Explain Behavioral Variability In LLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yang, Zhuonan, Li, Jacob Xiaochen, Velez, Francisco Piedrahita, Todd, Eric, Bau, David, Littman, Michael L., Bach, Stephen H., Pavlick, Ellie
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913058279915520
author Yang, Zhuonan
Li, Jacob Xiaochen
Velez, Francisco Piedrahita
Todd, Eric
Bau, David
Littman, Michael L.
Bach, Stephen H.
Pavlick, Ellie
author_facet Yang, Zhuonan
Li, Jacob Xiaochen
Velez, Francisco Piedrahita
Todd, Eric
Bau, David
Littman, Michael L.
Bach, Stephen H.
Pavlick, Ellie
contents One of the most common complaints about large language models (LLMs) is their prompt sensitivity -- that is, the fact that their ability to perform a task or provide a correct answer to a question can depend unpredictably on the way the question is posed. We investigate this variation by comparing two very different but commonly-used styles of prompting: instruction-based prompts, which describe the task in natural language, and example-based prompts, which provide in-context few-shot demonstration pairs to illustrate the task. We find that, despite large variation in performance as a function of the prompt, the model engages some common underlying mechanisms across different prompts of a task. Specifically, we identify task-specific attention heads whose outputs literally describe the task -- which we dub lexical task heads -- and show that these heads are shared across prompting styles and trigger subsequent answer production. We further find that behavioral variation between prompts can be explained by the degree to which these heads are activated, and that failures are at least sometimes due to competing task representations that dilute the signal of the target task. Our results together present an increasingly clear picture of how LLMs' internal representations can explain behavior that otherwise seems idiosyncratic to users and developers.
format Preprint
id arxiv_https___arxiv_org_abs_2604_22027
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Shared Lexical Task Representations Explain Behavioral Variability In LLMs
Yang, Zhuonan
Li, Jacob Xiaochen
Velez, Francisco Piedrahita
Todd, Eric
Bau, David
Littman, Michael L.
Bach, Stephen H.
Pavlick, Ellie
Computation and Language
Artificial Intelligence
Machine Learning
One of the most common complaints about large language models (LLMs) is their prompt sensitivity -- that is, the fact that their ability to perform a task or provide a correct answer to a question can depend unpredictably on the way the question is posed. We investigate this variation by comparing two very different but commonly-used styles of prompting: instruction-based prompts, which describe the task in natural language, and example-based prompts, which provide in-context few-shot demonstration pairs to illustrate the task. We find that, despite large variation in performance as a function of the prompt, the model engages some common underlying mechanisms across different prompts of a task. Specifically, we identify task-specific attention heads whose outputs literally describe the task -- which we dub lexical task heads -- and show that these heads are shared across prompting styles and trigger subsequent answer production. We further find that behavioral variation between prompts can be explained by the degree to which these heads are activated, and that failures are at least sometimes due to competing task representations that dilute the signal of the target task. Our results together present an increasingly clear picture of how LLMs' internal representations can explain behavior that otherwise seems idiosyncratic to users and developers.
title Shared Lexical Task Representations Explain Behavioral Variability In LLMs
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2604.22027