On The Role of Reasoning in the Identification of Subtle Stereotypes in Natural Language

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tian, Jacob-Junqi, Dige, Omkar, Emerson, D. B., Khattak, Faiza Khan
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913521414963200
author Tian, Jacob-Junqi
Dige, Omkar
Emerson, D. B.
Khattak, Faiza Khan
author_facet Tian, Jacob-Junqi
Dige, Omkar
Emerson, D. B.
Khattak, Faiza Khan
contents Large language models (LLMs) are trained on vast, uncurated datasets that contain various forms of biases and language reinforcing harmful stereotypes that may be subsequently inherited by the models themselves. Therefore, it is essential to examine and address biases in language models, integrating fairness into their development to ensure that these models do not perpetuate social biases. In this work, we demonstrate the importance of reasoning in zero-shot stereotype identification across several open-source LLMs. Accurate identification of stereotypical language is a complex task requiring a nuanced understanding of social structures, biases, and existing unfair generalizations about particular groups. While improved accuracy is observed through model scaling, the use of reasoning, especially multi-step reasoning, is crucial to consistent performance. Additionally, through a qualitative analysis of select reasoning traces, we highlight how reasoning improves not just accuracy, but also the interpretability of model decisions. This work firmly establishes reasoning as a critical component in automatic stereotype detection and is a first step towards stronger stereotype mitigation pipelines for LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2308_00071
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle On The Role of Reasoning in the Identification of Subtle Stereotypes in Natural Language
Tian, Jacob-Junqi
Dige, Omkar
Emerson, D. B.
Khattak, Faiza Khan
Computation and Language
Artificial Intelligence
Computers and Society
Machine Learning
68T50
Large language models (LLMs) are trained on vast, uncurated datasets that contain various forms of biases and language reinforcing harmful stereotypes that may be subsequently inherited by the models themselves. Therefore, it is essential to examine and address biases in language models, integrating fairness into their development to ensure that these models do not perpetuate social biases. In this work, we demonstrate the importance of reasoning in zero-shot stereotype identification across several open-source LLMs. Accurate identification of stereotypical language is a complex task requiring a nuanced understanding of social structures, biases, and existing unfair generalizations about particular groups. While improved accuracy is observed through model scaling, the use of reasoning, especially multi-step reasoning, is crucial to consistent performance. Additionally, through a qualitative analysis of select reasoning traces, we highlight how reasoning improves not just accuracy, but also the interpretability of model decisions. This work firmly establishes reasoning as a critical component in automatic stereotype detection and is a first step towards stronger stereotype mitigation pipelines for LLMs.
title On The Role of Reasoning in the Identification of Subtle Stereotypes in Natural Language
topic Computation and Language
Artificial Intelligence
Computers and Society
Machine Learning
68T50
url https://arxiv.org/abs/2308.00071