Walking in Others' Shoes: How Perspective-Taking Guides Large Language Models in Reducing Toxicity and Bias

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Xu, Rongwu, Zhou, Zi'an, Zhang, Tianwei, Qi, Zehan, Yao, Su, Xu, Ke, Xu, Wei, Qiu, Han
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866929429688614912
author Xu, Rongwu
Zhou, Zi'an
Zhang, Tianwei
Qi, Zehan
Yao, Su
Xu, Ke
Xu, Wei
Qiu, Han
author_facet Xu, Rongwu
Zhou, Zi'an
Zhang, Tianwei
Qi, Zehan
Yao, Su
Xu, Ke
Xu, Wei
Qiu, Han
contents The common toxicity and societal bias in contents generated by large language models (LLMs) necessitate strategies to reduce harm. Present solutions often demand white-box access to the model or substantial training, which is impractical for cutting-edge commercial LLMs. Moreover, prevailing prompting methods depend on external tool feedback and fail to simultaneously lessen toxicity and bias. Motivated by social psychology principles, we propose a novel strategy named \textbf{perspective-taking prompting (\textsc{PeT})} that inspires LLMs to integrate diverse human perspectives and self-regulate their responses. This self-correction mechanism can significantly diminish toxicity (up to $89\%$) and bias (up to $73\%$) in LLMs' responses. Rigorous evaluations and ablation studies are conducted on two commercial LLMs (ChatGPT and GLM) and three open-source LLMs, revealing \textsc{PeT}'s superiority in producing less harmful responses, outperforming five strong baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2407_15366
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Walking in Others' Shoes: How Perspective-Taking Guides Large Language Models in Reducing Toxicity and Bias
Xu, Rongwu
Zhou, Zi'an
Zhang, Tianwei
Qi, Zehan
Yao, Su
Xu, Ke
Xu, Wei
Qiu, Han
Computation and Language
Artificial Intelligence
Computers and Society
The common toxicity and societal bias in contents generated by large language models (LLMs) necessitate strategies to reduce harm. Present solutions often demand white-box access to the model or substantial training, which is impractical for cutting-edge commercial LLMs. Moreover, prevailing prompting methods depend on external tool feedback and fail to simultaneously lessen toxicity and bias. Motivated by social psychology principles, we propose a novel strategy named \textbf{perspective-taking prompting (\textsc{PeT})} that inspires LLMs to integrate diverse human perspectives and self-regulate their responses. This self-correction mechanism can significantly diminish toxicity (up to $89\%$) and bias (up to $73\%$) in LLMs' responses. Rigorous evaluations and ablation studies are conducted on two commercial LLMs (ChatGPT and GLM) and three open-source LLMs, revealing \textsc{PeT}'s superiority in producing less harmful responses, outperforming five strong baselines.
title Walking in Others' Shoes: How Perspective-Taking Guides Large Language Models in Reducing Toxicity and Bias
topic Computation and Language
Artificial Intelligence
Computers and Society
url https://arxiv.org/abs/2407.15366