Quantifying Risk Propensities of Large Language Models: Ethical Focus and Bias Detection through Role-Play

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zeng, Yifan, Kairong, Liang, Dong, Fangzhou, Zheng, Peijia
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913825904656384
author Zeng, Yifan
Kairong, Liang
Dong, Fangzhou
Zheng, Peijia
author_facet Zeng, Yifan
Kairong, Liang
Dong, Fangzhou
Zheng, Peijia
contents As Large Language Models (LLMs) become more prevalent, concerns about their safety, ethics, and potential biases have risen. Systematically evaluating LLMs' risk decision-making tendencies and attitudes, particularly in the ethical domain, has become crucial. This study innovatively applies the Domain-Specific Risk-Taking (DOSPERT) scale from cognitive science to LLMs and proposes a novel Ethical Decision-Making Risk Attitude Scale (EDRAS) to assess LLMs' ethical risk attitudes in depth. We further propose a novel approach integrating risk scales and role-playing to quantitatively evaluate systematic biases in LLMs. Through systematic evaluation and analysis of multiple mainstream LLMs, we assessed the "risk personalities" of LLMs across multiple domains, with a particular focus on the ethical domain, and revealed and quantified LLMs' systematic biases towards different groups. This research helps understand LLMs' risk decision-making and ensure their safe and reliable application. Our approach provides a tool for identifying and mitigating biases, contributing to fairer and more trustworthy AI systems. The code and data are available.
format Preprint
id arxiv_https___arxiv_org_abs_2411_08884
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Quantifying Risk Propensities of Large Language Models: Ethical Focus and Bias Detection through Role-Play
Zeng, Yifan
Kairong, Liang
Dong, Fangzhou
Zheng, Peijia
Computers and Society
Artificial Intelligence
Computation and Language
As Large Language Models (LLMs) become more prevalent, concerns about their safety, ethics, and potential biases have risen. Systematically evaluating LLMs' risk decision-making tendencies and attitudes, particularly in the ethical domain, has become crucial. This study innovatively applies the Domain-Specific Risk-Taking (DOSPERT) scale from cognitive science to LLMs and proposes a novel Ethical Decision-Making Risk Attitude Scale (EDRAS) to assess LLMs' ethical risk attitudes in depth. We further propose a novel approach integrating risk scales and role-playing to quantitatively evaluate systematic biases in LLMs. Through systematic evaluation and analysis of multiple mainstream LLMs, we assessed the "risk personalities" of LLMs across multiple domains, with a particular focus on the ethical domain, and revealed and quantified LLMs' systematic biases towards different groups. This research helps understand LLMs' risk decision-making and ensure their safe and reliable application. Our approach provides a tool for identifying and mitigating biases, contributing to fairer and more trustworthy AI systems. The code and data are available.
title Quantifying Risk Propensities of Large Language Models: Ethical Focus and Bias Detection through Role-Play
topic Computers and Society
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2411.08884