Overview of the NLPCC 2025 Shared Task: Gender Bias Mitigation Challenge

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Yizhi, Zhang, Ge, Hong, Hanhua, Wang, Yiwen, Lin, Chenghua
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916794074136576
author Li, Yizhi
Zhang, Ge
Hong, Hanhua
Wang, Yiwen
Lin, Chenghua
author_facet Li, Yizhi
Zhang, Ge
Hong, Hanhua
Wang, Yiwen
Lin, Chenghua
contents As natural language processing for gender bias becomes a significant interdisciplinary topic, the prevalent data-driven techniques, such as pre-trained language models, suffer from biased corpus. This case becomes more obvious regarding those languages with less fairness-related computational linguistic resources, such as Chinese. To this end, we propose a Chinese cOrpus foR Gender bIas Probing and Mitigation (CORGI-PM), which contains 32.9k sentences with high-quality labels derived by following an annotation scheme specifically developed for gender bias in the Chinese context. It is worth noting that CORGI-PM contains 5.2k gender-biased sentences along with the corresponding bias-eliminated versions rewritten by human annotators. We pose three challenges as a shared task to automate the mitigation of textual gender bias, which requires the models to detect, classify, and mitigate textual gender bias. In the literature, we present the results and analysis for the teams participating this shared task in NLPCC 2025.
format Preprint
id arxiv_https___arxiv_org_abs_2506_12574
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Overview of the NLPCC 2025 Shared Task: Gender Bias Mitigation Challenge
Li, Yizhi
Zhang, Ge
Hong, Hanhua
Wang, Yiwen
Lin, Chenghua
Computation and Language
As natural language processing for gender bias becomes a significant interdisciplinary topic, the prevalent data-driven techniques, such as pre-trained language models, suffer from biased corpus. This case becomes more obvious regarding those languages with less fairness-related computational linguistic resources, such as Chinese. To this end, we propose a Chinese cOrpus foR Gender bIas Probing and Mitigation (CORGI-PM), which contains 32.9k sentences with high-quality labels derived by following an annotation scheme specifically developed for gender bias in the Chinese context. It is worth noting that CORGI-PM contains 5.2k gender-biased sentences along with the corresponding bias-eliminated versions rewritten by human annotators. We pose three challenges as a shared task to automate the mitigation of textual gender bias, which requires the models to detect, classify, and mitigate textual gender bias. In the literature, we present the results and analysis for the teams participating this shared task in NLPCC 2025.
title Overview of the NLPCC 2025 Shared Task: Gender Bias Mitigation Challenge
topic Computation and Language
url https://arxiv.org/abs/2506.12574