McBE: A Multi-task Chinese Bias Evaluation Benchmark for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lan, Tian, Su, Xiangdong, Liu, Xu, Wang, Ruirui, Chang, Ke, Li, Jiang, Gao, Guanglai
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916884678443008
author Lan, Tian
Su, Xiangdong
Liu, Xu
Wang, Ruirui
Chang, Ke
Li, Jiang
Gao, Guanglai
author_facet Lan, Tian
Su, Xiangdong
Liu, Xu
Wang, Ruirui
Chang, Ke
Li, Jiang
Gao, Guanglai
contents As large language models (LLMs) are increasingly applied to various NLP tasks, their inherent biases are gradually disclosed. Therefore, measuring biases in LLMs is crucial to mitigate its ethical risks. However, most existing bias evaluation datasets focus on English and North American culture, and their bias categories are not fully applicable to other cultures. The datasets grounded in the Chinese language and culture are scarce. More importantly, these datasets usually only support single evaluation tasks and cannot evaluate the bias from multiple aspects in LLMs. To address these issues, we present a Multi-task Chinese Bias Evaluation Benchmark (McBE) that includes 4,077 bias evaluation instances, covering 12 single bias categories, 82 subcategories and introducing 5 evaluation tasks, providing extensive category coverage, content diversity, and measuring comprehensiveness. Additionally, we evaluate several popular LLMs from different series and with parameter sizes. In general, all these LLMs demonstrated varying degrees of bias. We conduct an in-depth analysis of results, offering novel insights into bias in LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2507_02088
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle McBE: A Multi-task Chinese Bias Evaluation Benchmark for Large Language Models
Lan, Tian
Su, Xiangdong
Liu, Xu
Wang, Ruirui
Chang, Ke
Li, Jiang
Gao, Guanglai
Computation and Language
As large language models (LLMs) are increasingly applied to various NLP tasks, their inherent biases are gradually disclosed. Therefore, measuring biases in LLMs is crucial to mitigate its ethical risks. However, most existing bias evaluation datasets focus on English and North American culture, and their bias categories are not fully applicable to other cultures. The datasets grounded in the Chinese language and culture are scarce. More importantly, these datasets usually only support single evaluation tasks and cannot evaluate the bias from multiple aspects in LLMs. To address these issues, we present a Multi-task Chinese Bias Evaluation Benchmark (McBE) that includes 4,077 bias evaluation instances, covering 12 single bias categories, 82 subcategories and introducing 5 evaluation tasks, providing extensive category coverage, content diversity, and measuring comprehensiveness. Additionally, we evaluate several popular LLMs from different series and with parameter sizes. In general, all these LLMs demonstrated varying degrees of bias. We conduct an in-depth analysis of results, offering novel insights into bias in LLMs.
title McBE: A Multi-task Chinese Bias Evaluation Benchmark for Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2507.02088