Revisiting the Reliability of Psychological Scales on Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Jen-tse, Jiao, Wenxiang, Lam, Man Ho, Li, Eric John, Wang, Wenxuan, Lyu, Michael R.
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914963721814016
author Huang, Jen-tse
Jiao, Wenxiang
Lam, Man Ho
Li, Eric John
Wang, Wenxuan
Lyu, Michael R.
author_facet Huang, Jen-tse
Jiao, Wenxiang
Lam, Man Ho
Li, Eric John
Wang, Wenxuan
Lyu, Michael R.
contents Recent research has focused on examining Large Language Models' (LLMs) characteristics from a psychological standpoint, acknowledging the necessity of understanding their behavioral characteristics. The administration of personality tests to LLMs has emerged as a noteworthy area in this context. However, the suitability of employing psychological scales, initially devised for humans, on LLMs is a matter of ongoing debate. Our study aims to determine the reliability of applying personality assessments to LLMs, explicitly investigating whether LLMs demonstrate consistent personality traits. Analysis of 2,500 settings per model, including GPT-3.5, GPT-4, Gemini-Pro, and LLaMA-3.1, reveals that various LLMs show consistency in responses to the Big Five Inventory, indicating a satisfactory level of reliability. Furthermore, our research explores the potential of GPT-3.5 to emulate diverse personalities and represent various groups-a capability increasingly sought after in social sciences for substituting human participants with LLMs to reduce costs. Our findings reveal that LLMs have the potential to represent different personalities with specific prompt instructions.
format Preprint
id arxiv_https___arxiv_org_abs_2305_19926
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Revisiting the Reliability of Psychological Scales on Large Language Models
Huang, Jen-tse
Jiao, Wenxiang
Lam, Man Ho
Li, Eric John
Wang, Wenxuan
Lyu, Michael R.
Computation and Language
Recent research has focused on examining Large Language Models' (LLMs) characteristics from a psychological standpoint, acknowledging the necessity of understanding their behavioral characteristics. The administration of personality tests to LLMs has emerged as a noteworthy area in this context. However, the suitability of employing psychological scales, initially devised for humans, on LLMs is a matter of ongoing debate. Our study aims to determine the reliability of applying personality assessments to LLMs, explicitly investigating whether LLMs demonstrate consistent personality traits. Analysis of 2,500 settings per model, including GPT-3.5, GPT-4, Gemini-Pro, and LLaMA-3.1, reveals that various LLMs show consistency in responses to the Big Five Inventory, indicating a satisfactory level of reliability. Furthermore, our research explores the potential of GPT-3.5 to emulate diverse personalities and represent various groups-a capability increasingly sought after in social sciences for substituting human participants with LLMs to reduce costs. Our findings reveal that LLMs have the potential to represent different personalities with specific prompt instructions.
title Revisiting the Reliability of Psychological Scales on Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2305.19926