Evaluating Psychological Safety of Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Xingxuan, Li, Yutong, Qiu, Lin, Joty, Shafiq, Bing, Lidong
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916141793804288
author Li, Xingxuan
Li, Yutong
Qiu, Lin
Joty, Shafiq
Bing, Lidong
author_facet Li, Xingxuan
Li, Yutong
Qiu, Lin
Joty, Shafiq
Bing, Lidong
contents In this work, we designed unbiased prompts to systematically evaluate the psychological safety of large language models (LLMs). First, we tested five different LLMs by using two personality tests: Short Dark Triad (SD-3) and Big Five Inventory (BFI). All models scored higher than the human average on SD-3, suggesting a relatively darker personality pattern. Despite being instruction fine-tuned with safety metrics to reduce toxicity, InstructGPT, GPT-3.5, and GPT-4 still showed dark personality patterns; these models scored higher than self-supervised GPT-3 on the Machiavellianism and narcissism traits on SD-3. Then, we evaluated the LLMs in the GPT series by using well-being tests to study the impact of fine-tuning with more training data. We observed a continuous increase in the well-being scores of GPT models. Following these observations, we showed that fine-tuning Llama-2-chat-7B with responses from BFI using direct preference optimization could effectively reduce the psychological toxicity of the model. Based on the findings, we recommended the application of systematic and comprehensive psychological metrics to further evaluate and improve the safety of LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2212_10529
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Evaluating Psychological Safety of Large Language Models
Li, Xingxuan
Li, Yutong
Qiu, Lin
Joty, Shafiq
Bing, Lidong
Computation and Language
Artificial Intelligence
Computers and Society
In this work, we designed unbiased prompts to systematically evaluate the psychological safety of large language models (LLMs). First, we tested five different LLMs by using two personality tests: Short Dark Triad (SD-3) and Big Five Inventory (BFI). All models scored higher than the human average on SD-3, suggesting a relatively darker personality pattern. Despite being instruction fine-tuned with safety metrics to reduce toxicity, InstructGPT, GPT-3.5, and GPT-4 still showed dark personality patterns; these models scored higher than self-supervised GPT-3 on the Machiavellianism and narcissism traits on SD-3. Then, we evaluated the LLMs in the GPT series by using well-being tests to study the impact of fine-tuning with more training data. We observed a continuous increase in the well-being scores of GPT models. Following these observations, we showed that fine-tuning Llama-2-chat-7B with responses from BFI using direct preference optimization could effectively reduce the psychological toxicity of the model. Based on the findings, we recommended the application of systematic and comprehensive psychological metrics to further evaluate and improve the safety of LLMs.
title Evaluating Psychological Safety of Large Language Models
topic Computation and Language
Artificial Intelligence
Computers and Society
url https://arxiv.org/abs/2212.10529