CEB: Compositional Evaluation Benchmark for Fairness in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Song, Wang, Peng, Zhou, Tong, Dong, Yushun, Tan, Zhen, Li, Jundong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916625188388864
author Wang, Song
Wang, Peng
Zhou, Tong
Dong, Yushun
Tan, Zhen
Li, Jundong
author_facet Wang, Song
Wang, Peng
Zhou, Tong
Dong, Yushun
Tan, Zhen
Li, Jundong
contents As Large Language Models (LLMs) are increasingly deployed to handle various natural language processing (NLP) tasks, concerns regarding the potential negative societal impacts of LLM-generated content have also arisen. To evaluate the biases exhibited by LLMs, researchers have recently proposed a variety of datasets. However, existing bias evaluation efforts often focus on only a particular type of bias and employ inconsistent evaluation metrics, leading to difficulties in comparison across different datasets and LLMs. To address these limitations, we collect a variety of datasets designed for the bias evaluation of LLMs, and further propose CEB, a Compositional Evaluation Benchmark that covers different types of bias across different social groups and tasks. The curation of CEB is based on our newly proposed compositional taxonomy, which characterizes each dataset from three dimensions: bias types, social groups, and tasks. By combining the three dimensions, we develop a comprehensive evaluation strategy for the bias in LLMs. Our experiments demonstrate that the levels of bias vary across these dimensions, thereby providing guidance for the development of specific bias mitigation methods.
format Preprint
id arxiv_https___arxiv_org_abs_2407_02408
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CEB: Compositional Evaluation Benchmark for Fairness in Large Language Models
Wang, Song
Wang, Peng
Zhou, Tong
Dong, Yushun
Tan, Zhen
Li, Jundong
Computation and Language
Machine Learning
As Large Language Models (LLMs) are increasingly deployed to handle various natural language processing (NLP) tasks, concerns regarding the potential negative societal impacts of LLM-generated content have also arisen. To evaluate the biases exhibited by LLMs, researchers have recently proposed a variety of datasets. However, existing bias evaluation efforts often focus on only a particular type of bias and employ inconsistent evaluation metrics, leading to difficulties in comparison across different datasets and LLMs. To address these limitations, we collect a variety of datasets designed for the bias evaluation of LLMs, and further propose CEB, a Compositional Evaluation Benchmark that covers different types of bias across different social groups and tasks. The curation of CEB is based on our newly proposed compositional taxonomy, which characterizes each dataset from three dimensions: bias types, social groups, and tasks. By combining the three dimensions, we develop a comprehensive evaluation strategy for the bias in LLMs. Our experiments demonstrate that the levels of bias vary across these dimensions, thereby providing guidance for the development of specific bias mitigation methods.
title CEB: Compositional Evaluation Benchmark for Fairness in Large Language Models
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2407.02408