Inducing anxiety in large language models can induce bias

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Coda-Forno, Julian, Witte, Kristin, Jagadish, Akshay K., Binz, Marcel, Akata, Zeynep, Schulz, Eric
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929541845352448
author Coda-Forno, Julian
Witte, Kristin
Jagadish, Akshay K.
Binz, Marcel
Akata, Zeynep
Schulz, Eric
author_facet Coda-Forno, Julian
Witte, Kristin
Jagadish, Akshay K.
Binz, Marcel
Akata, Zeynep
Schulz, Eric
contents Large language models (LLMs) are transforming research on machine learning while galvanizing public debates. Understanding not only when these models work well and succeed but also why they fail and misbehave is of great societal relevance. We propose to turn the lens of psychiatry, a framework used to describe and modify maladaptive behavior, to the outputs produced by these models. We focus on twelve established LLMs and subject them to a questionnaire commonly used in psychiatry. Our results show that six of the latest LLMs respond robustly to the anxiety questionnaire, producing comparable anxiety scores to humans. Moreover, the LLMs' responses can be predictably changed by using anxiety-inducing prompts. Anxiety-induction not only influences LLMs' scores on an anxiety questionnaire but also influences their behavior in a previously-established benchmark measuring biases such as racism and ageism. Importantly, greater anxiety-inducing text leads to stronger increases in biases, suggesting that how anxiously a prompt is communicated to large language models has a strong influence on their behavior in applied settings. These results demonstrate the usefulness of methods taken from psychiatry for studying the capable algorithms to which we increasingly delegate authority and autonomy.
format Preprint
id arxiv_https___arxiv_org_abs_2304_11111
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Inducing anxiety in large language models can induce bias
Coda-Forno, Julian
Witte, Kristin
Jagadish, Akshay K.
Binz, Marcel
Akata, Zeynep
Schulz, Eric
Computation and Language
Artificial Intelligence
Machine Learning
Large language models (LLMs) are transforming research on machine learning while galvanizing public debates. Understanding not only when these models work well and succeed but also why they fail and misbehave is of great societal relevance. We propose to turn the lens of psychiatry, a framework used to describe and modify maladaptive behavior, to the outputs produced by these models. We focus on twelve established LLMs and subject them to a questionnaire commonly used in psychiatry. Our results show that six of the latest LLMs respond robustly to the anxiety questionnaire, producing comparable anxiety scores to humans. Moreover, the LLMs' responses can be predictably changed by using anxiety-inducing prompts. Anxiety-induction not only influences LLMs' scores on an anxiety questionnaire but also influences their behavior in a previously-established benchmark measuring biases such as racism and ageism. Importantly, greater anxiety-inducing text leads to stronger increases in biases, suggesting that how anxiously a prompt is communicated to large language models has a strong influence on their behavior in applied settings. These results demonstrate the usefulness of methods taken from psychiatry for studying the capable algorithms to which we increasingly delegate authority and autonomy.
title Inducing anxiety in large language models can induce bias
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2304.11111