Human Psychometric Questionnaires Mischaracterize LLM Behavior

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Song, Woojung, Choi, Dongmin, Park, Yoonah, Han, Jongwook, Lee, Eun-Ju, Jo, Yohan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914616988139520
author Song, Woojung
Choi, Dongmin
Park, Yoonah
Han, Jongwook
Lee, Eun-Ju
Jo, Yohan
author_facet Song, Woojung
Choi, Dongmin
Park, Yoonah
Han, Jongwook
Lee, Eun-Ju
Jo, Yohan
contents We examine whether human psychometric questionnaires can serve as reliable tools for characterizing and predicting LLM behavior in everyday user interactions. We analyze eight open-source LLMs by comparing their value and personality profiles derived from two different methods: Likert self-reports on established questionnaires (PVQ-40/21 and BFI-44/10) and generation probabilities over value-laden responses to everyday user queries. The two profiles diverge substantially. Within-construct item consistency, often cited as evidence of stable LLM dispositions, disappears in generation probabilities. We attribute this gap to the fact that explicit lexical cues in established questionnaire items allow models to recognize the target construct and respond in alignment-consistent, socially desirable ways, whereas realistic user queries provide no such cues. In addition, demographic persona prompts shift models' responses to human questionnaires in ways consistent with real human patterns, but no such shifts appear in the generation probabilities of responses to realistic user queries, showing their limited ability to simulate the behaviors of target demographics in real-world user interactions. Overall, our study shows that human psychometric questionnaires are insufficient tools for predicting LLM behavior and suggests generation-based profiling as a more accurate measure.
format Preprint
id arxiv_https___arxiv_org_abs_2509_10078
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Human Psychometric Questionnaires Mischaracterize LLM Behavior
Song, Woojung
Choi, Dongmin
Park, Yoonah
Han, Jongwook
Lee, Eun-Ju
Jo, Yohan
Computation and Language
Artificial Intelligence
We examine whether human psychometric questionnaires can serve as reliable tools for characterizing and predicting LLM behavior in everyday user interactions. We analyze eight open-source LLMs by comparing their value and personality profiles derived from two different methods: Likert self-reports on established questionnaires (PVQ-40/21 and BFI-44/10) and generation probabilities over value-laden responses to everyday user queries. The two profiles diverge substantially. Within-construct item consistency, often cited as evidence of stable LLM dispositions, disappears in generation probabilities. We attribute this gap to the fact that explicit lexical cues in established questionnaire items allow models to recognize the target construct and respond in alignment-consistent, socially desirable ways, whereas realistic user queries provide no such cues. In addition, demographic persona prompts shift models' responses to human questionnaires in ways consistent with real human patterns, but no such shifts appear in the generation probabilities of responses to realistic user queries, showing their limited ability to simulate the behaviors of target demographics in real-world user interactions. Overall, our study shows that human psychometric questionnaires are insufficient tools for predicting LLM behavior and suggests generation-based profiling as a more accurate measure.
title Human Psychometric Questionnaires Mischaracterize LLM Behavior
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2509.10078