Colombian Waitresses y Jueces canadienses: Gender and Country Biases in Occupation Recommendations from LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rodríguez, Elisa Forcada, Perez-de-Viñaspre, Olatz, Campos, Jon Ander, Klakow, Dietrich, Gautam, Vagrant
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909707848908800
author Rodríguez, Elisa Forcada
Perez-de-Viñaspre, Olatz
Campos, Jon Ander
Klakow, Dietrich
Gautam, Vagrant
author_facet Rodríguez, Elisa Forcada
Perez-de-Viñaspre, Olatz
Campos, Jon Ander
Klakow, Dietrich
Gautam, Vagrant
contents One of the goals of fairness research in NLP is to measure and mitigate stereotypical biases that are propagated by NLP systems. However, such work tends to focus on single axes of bias (most often gender) and the English language. Addressing these limitations, we contribute the first study of multilingual intersecting country and gender biases, with a focus on occupation recommendations generated by large language models. We construct a benchmark of prompts in English, Spanish and German, where we systematically vary country and gender, using 25 countries and four pronoun sets. Then, we evaluate a suite of 5 Llama-based models on this benchmark, finding that LLMs encode significant gender and country biases. Notably, we find that even when models show parity for gender or country individually, intersectional occupational biases based on both country and gender persist. We also show that the prompting language significantly affects bias, and instruction-tuned models consistently demonstrate the lowest and most stable levels of bias. Our findings highlight the need for fairness researchers to use intersectional and multilingual lenses in their work.
format Preprint
id arxiv_https___arxiv_org_abs_2505_02456
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Colombian Waitresses y Jueces canadienses: Gender and Country Biases in Occupation Recommendations from LLMs
Rodríguez, Elisa Forcada
Perez-de-Viñaspre, Olatz
Campos, Jon Ander
Klakow, Dietrich
Gautam, Vagrant
Computation and Language
One of the goals of fairness research in NLP is to measure and mitigate stereotypical biases that are propagated by NLP systems. However, such work tends to focus on single axes of bias (most often gender) and the English language. Addressing these limitations, we contribute the first study of multilingual intersecting country and gender biases, with a focus on occupation recommendations generated by large language models. We construct a benchmark of prompts in English, Spanish and German, where we systematically vary country and gender, using 25 countries and four pronoun sets. Then, we evaluate a suite of 5 Llama-based models on this benchmark, finding that LLMs encode significant gender and country biases. Notably, we find that even when models show parity for gender or country individually, intersectional occupational biases based on both country and gender persist. We also show that the prompting language significantly affects bias, and instruction-tuned models consistently demonstrate the lowest and most stable levels of bias. Our findings highlight the need for fairness researchers to use intersectional and multilingual lenses in their work.
title Colombian Waitresses y Jueces canadienses: Gender and Country Biases in Occupation Recommendations from LLMs
topic Computation and Language
url https://arxiv.org/abs/2505.02456