Revealing Fine-Grained Values and Opinions in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wright, Dustin, Arora, Arnav, Borenstein, Nadav, Yadav, Srishti, Belongie, Serge, Augenstein, Isabelle
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911128508956672
author Wright, Dustin
Arora, Arnav
Borenstein, Nadav
Yadav, Srishti
Belongie, Serge
Augenstein, Isabelle
author_facet Wright, Dustin
Arora, Arnav
Borenstein, Nadav
Yadav, Srishti
Belongie, Serge
Augenstein, Isabelle
contents Uncovering latent values and opinions embedded in large language models (LLMs) can help identify biases and mitigate potential harm. Recently, this has been approached by prompting LLMs with survey questions and quantifying the stances in the outputs towards morally and politically charged statements. However, the stances generated by LLMs can vary greatly depending on how they are prompted, and there are many ways to argue for or against a given position. In this work, we propose to address this by analysing a large and robust dataset of 156k LLM responses to the 62 propositions of the Political Compass Test (PCT) generated by 6 LLMs using 420 prompt variations. We perform coarse-grained analysis of their generated stances and fine-grained analysis of the plain text justifications for those stances. For fine-grained analysis, we propose to identify tropes in the responses: semantically similar phrases that are recurrent and consistent across different prompts, revealing natural patterns in the text that a given LLM is prone to produce. We find that demographic features added to prompts significantly affect outcomes on the PCT, reflecting bias, as well as disparities between the results of tests when eliciting closed-form vs. open domain responses. Additionally, patterns in the plain text rationales via tropes show that similar justifications are repeatedly generated across models and prompts even with disparate stances.
format Preprint
id arxiv_https___arxiv_org_abs_2406_19238
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Revealing Fine-Grained Values and Opinions in Large Language Models
Wright, Dustin
Arora, Arnav
Borenstein, Nadav
Yadav, Srishti
Belongie, Serge
Augenstein, Isabelle
Computation and Language
Computers and Society
Machine Learning
Uncovering latent values and opinions embedded in large language models (LLMs) can help identify biases and mitigate potential harm. Recently, this has been approached by prompting LLMs with survey questions and quantifying the stances in the outputs towards morally and politically charged statements. However, the stances generated by LLMs can vary greatly depending on how they are prompted, and there are many ways to argue for or against a given position. In this work, we propose to address this by analysing a large and robust dataset of 156k LLM responses to the 62 propositions of the Political Compass Test (PCT) generated by 6 LLMs using 420 prompt variations. We perform coarse-grained analysis of their generated stances and fine-grained analysis of the plain text justifications for those stances. For fine-grained analysis, we propose to identify tropes in the responses: semantically similar phrases that are recurrent and consistent across different prompts, revealing natural patterns in the text that a given LLM is prone to produce. We find that demographic features added to prompts significantly affect outcomes on the PCT, reflecting bias, as well as disparities between the results of tests when eliciting closed-form vs. open domain responses. Additionally, patterns in the plain text rationales via tropes show that similar justifications are repeatedly generated across models and prompts even with disparate stances.
title Revealing Fine-Grained Values and Opinions in Large Language Models
topic Computation and Language
Computers and Society
Machine Learning
url https://arxiv.org/abs/2406.19238