GG-BBQ: German Gender Bias Benchmark for Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Satheesh, Shalaka, Klug, Katrin, Beckh, Katharina, Allende-Cid, Héctor, Houben, Sebastian, Hassan, Teena
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918101689303040
author Satheesh, Shalaka
Klug, Katrin
Beckh, Katharina
Allende-Cid, Héctor
Houben, Sebastian
Hassan, Teena
author_facet Satheesh, Shalaka
Klug, Katrin
Beckh, Katharina
Allende-Cid, Héctor
Houben, Sebastian
Hassan, Teena
contents Within the context of Natural Language Processing (NLP), fairness evaluation is often associated with the assessment of bias and reduction of associated harm. In this regard, the evaluation is usually carried out by using a benchmark dataset, for a task such as Question Answering, created for the measurement of bias in the model's predictions along various dimensions, including gender identity. In our work, we evaluate gender bias in German Large Language Models (LLMs) using the Bias Benchmark for Question Answering by Parrish et al. (2022) as a reference. Specifically, the templates in the gender identity subset of this English dataset were machine translated into German. The errors in the machine translated templates were then manually reviewed and corrected with the help of a language expert. We find that manual revision of the translation is crucial when creating datasets for gender bias evaluation because of the limitations of machine translation from English to a language such as German with grammatical gender. Our final dataset is comprised of two subsets: Subset-I, which consists of group terms related to gender identity, and Subset-II, where group terms are replaced with proper names. We evaluate several LLMs used for German NLP on this newly created dataset and report the accuracy and bias scores. The results show that all models exhibit bias, both along and against existing social stereotypes.
format Preprint
id arxiv_https___arxiv_org_abs_2507_16410
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GG-BBQ: German Gender Bias Benchmark for Question Answering
Satheesh, Shalaka
Klug, Katrin
Beckh, Katharina
Allende-Cid, Héctor
Houben, Sebastian
Hassan, Teena
Computation and Language
Computers and Society
Machine Learning
Within the context of Natural Language Processing (NLP), fairness evaluation is often associated with the assessment of bias and reduction of associated harm. In this regard, the evaluation is usually carried out by using a benchmark dataset, for a task such as Question Answering, created for the measurement of bias in the model's predictions along various dimensions, including gender identity. In our work, we evaluate gender bias in German Large Language Models (LLMs) using the Bias Benchmark for Question Answering by Parrish et al. (2022) as a reference. Specifically, the templates in the gender identity subset of this English dataset were machine translated into German. The errors in the machine translated templates were then manually reviewed and corrected with the help of a language expert. We find that manual revision of the translation is crucial when creating datasets for gender bias evaluation because of the limitations of machine translation from English to a language such as German with grammatical gender. Our final dataset is comprised of two subsets: Subset-I, which consists of group terms related to gender identity, and Subset-II, where group terms are replaced with proper names. We evaluate several LLMs used for German NLP on this newly created dataset and report the accuracy and bias scores. The results show that all models exhibit bias, both along and against existing social stereotypes.
title GG-BBQ: German Gender Bias Benchmark for Question Answering
topic Computation and Language
Computers and Society
Machine Learning
url https://arxiv.org/abs/2507.16410