EuroGEST: Investigating gender stereotypes in multilingual language models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Rowe, Jacqueline, Klimaszewski, Mateusz, Guillou, Liane, Vallor, Shannon, Birch, Alexandra
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908847958917120
author Rowe, Jacqueline
Klimaszewski, Mateusz
Guillou, Liane
Vallor, Shannon
Birch, Alexandra
author_facet Rowe, Jacqueline
Klimaszewski, Mateusz
Guillou, Liane
Vallor, Shannon
Birch, Alexandra
contents Large language models increasingly support multiple languages, yet most benchmarks for gender bias remain English-centric. We introduce EuroGEST, a dataset designed to measure gender-stereotypical reasoning in LLMs across English and 29 European languages. EuroGEST builds on an existing expert-informed benchmark covering 16 gender stereotypes, expanded in this work using translation tools, quality estimation metrics, and morphological heuristics. Human evaluations confirm that our data generation method results in high accuracy of both translations and gender labels across languages. We use EuroGEST to evaluate 24 multilingual language models from six model families, demonstrating that the strongest stereotypes in all models across all languages are that women are 'beautiful', 'empathetic' and 'neat' and men are 'leaders', 'strong, tough' and 'professional'. We also show that larger models encode gendered stereotypes more strongly and that instruction finetuning does not consistently reduce gendered stereotypes. Our work highlights the need for more multilingual studies of fairness in LLMs and offers scalable methods and resources to audit gender bias across languages.
format Preprint
id arxiv_https___arxiv_org_abs_2506_03867
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EuroGEST: Investigating gender stereotypes in multilingual language models
Rowe, Jacqueline
Klimaszewski, Mateusz
Guillou, Liane
Vallor, Shannon
Birch, Alexandra
Computation and Language
Large language models increasingly support multiple languages, yet most benchmarks for gender bias remain English-centric. We introduce EuroGEST, a dataset designed to measure gender-stereotypical reasoning in LLMs across English and 29 European languages. EuroGEST builds on an existing expert-informed benchmark covering 16 gender stereotypes, expanded in this work using translation tools, quality estimation metrics, and morphological heuristics. Human evaluations confirm that our data generation method results in high accuracy of both translations and gender labels across languages. We use EuroGEST to evaluate 24 multilingual language models from six model families, demonstrating that the strongest stereotypes in all models across all languages are that women are 'beautiful', 'empathetic' and 'neat' and men are 'leaders', 'strong, tough' and 'professional'. We also show that larger models encode gendered stereotypes more strongly and that instruction finetuning does not consistently reduce gendered stereotypes. Our work highlights the need for more multilingual studies of fairness in LLMs and offers scalable methods and resources to audit gender bias across languages.
title EuroGEST: Investigating gender stereotypes in multilingual language models
topic Computation and Language
url https://arxiv.org/abs/2506.03867