Adaptable Moral Stances of Large Language Models on Sexist Content: Implications for Society and Gender Discourse

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guo, Rongchen, Nejadgholi, Isar, Dawkins, Hillary, Fraser, Kathleen C., Kiritchenko, Svetlana
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913524863729664
author Guo, Rongchen
Nejadgholi, Isar
Dawkins, Hillary
Fraser, Kathleen C.
Kiritchenko, Svetlana
author_facet Guo, Rongchen
Nejadgholi, Isar
Dawkins, Hillary
Fraser, Kathleen C.
Kiritchenko, Svetlana
contents This work provides an explanatory view of how LLMs can apply moral reasoning to both criticize and defend sexist language. We assessed eight large language models, all of which demonstrated the capability to provide explanations grounded in varying moral perspectives for both critiquing and endorsing views that reflect sexist assumptions. With both human and automatic evaluation, we show that all eight models produce comprehensible and contextually relevant text, which is helpful in understanding diverse views on how sexism is perceived. Also, through analysis of moral foundations cited by LLMs in their arguments, we uncover the diverse ideological perspectives in models' outputs, with some models aligning more with progressive or conservative views on gender roles and sexism. Based on our observations, we caution against the potential misuse of LLMs to justify sexist language. We also highlight that LLMs can serve as tools for understanding the roots of sexist beliefs and designing well-informed interventions. Given this dual capacity, it is crucial to monitor LLMs and design safety mechanisms for their use in applications that involve sensitive societal topics, such as sexism.
format Preprint
id arxiv_https___arxiv_org_abs_2410_00175
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Adaptable Moral Stances of Large Language Models on Sexist Content: Implications for Society and Gender Discourse
Guo, Rongchen
Nejadgholi, Isar
Dawkins, Hillary
Fraser, Kathleen C.
Kiritchenko, Svetlana
Computation and Language
This work provides an explanatory view of how LLMs can apply moral reasoning to both criticize and defend sexist language. We assessed eight large language models, all of which demonstrated the capability to provide explanations grounded in varying moral perspectives for both critiquing and endorsing views that reflect sexist assumptions. With both human and automatic evaluation, we show that all eight models produce comprehensible and contextually relevant text, which is helpful in understanding diverse views on how sexism is perceived. Also, through analysis of moral foundations cited by LLMs in their arguments, we uncover the diverse ideological perspectives in models' outputs, with some models aligning more with progressive or conservative views on gender roles and sexism. Based on our observations, we caution against the potential misuse of LLMs to justify sexist language. We also highlight that LLMs can serve as tools for understanding the roots of sexist beliefs and designing well-informed interventions. Given this dual capacity, it is crucial to monitor LLMs and design safety mechanisms for their use in applications that involve sensitive societal topics, such as sexism.
title Adaptable Moral Stances of Large Language Models on Sexist Content: Implications for Society and Gender Discourse
topic Computation and Language
url https://arxiv.org/abs/2410.00175