Artificial Impressions: Evaluating Large Language Model Behavior Through the Lens of Trait Impressions

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Deas, Nicholas, McKeown, Kathleen
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917001536995328
author Deas, Nicholas
McKeown, Kathleen
author_facet Deas, Nicholas
McKeown, Kathleen
contents We introduce and study artificial impressions--patterns in LLMs' internal representations of prompts that resemble human impressions and stereotypes based on language. We fit linear probes on generated prompts to predict impressions according to the two-dimensional Stereotype Content Model (SCM). Using these probes, we study the relationship between impressions and downstream model behavior as well as prompt features that may inform such impressions. We find that LLMs inconsistently report impressions when prompted, but also that impressions are more consistently linearly decodable from their hidden representations. Additionally, we show that artificial impressions of prompts are predictive of the quality and use of hedging in model responses. We also investigate how particular content, stylistic, and dialectal features in prompts impact LLM impressions.
format Preprint
id arxiv_https___arxiv_org_abs_2510_08915
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Artificial Impressions: Evaluating Large Language Model Behavior Through the Lens of Trait Impressions
Deas, Nicholas
McKeown, Kathleen
Computation and Language
We introduce and study artificial impressions--patterns in LLMs' internal representations of prompts that resemble human impressions and stereotypes based on language. We fit linear probes on generated prompts to predict impressions according to the two-dimensional Stereotype Content Model (SCM). Using these probes, we study the relationship between impressions and downstream model behavior as well as prompt features that may inform such impressions. We find that LLMs inconsistently report impressions when prompted, but also that impressions are more consistently linearly decodable from their hidden representations. Additionally, we show that artificial impressions of prompts are predictive of the quality and use of hedging in model responses. We also investigate how particular content, stylistic, and dialectal features in prompts impact LLM impressions.
title Artificial Impressions: Evaluating Large Language Model Behavior Through the Lens of Trait Impressions
topic Computation and Language
url https://arxiv.org/abs/2510.08915