StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jeune, Pierre Le, Duchesne, Étienne, Xiao, Weixuan, Palminteri, Stefano, Houssin, Bazire, Malézieux, Benoît, Dora, Matteo
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909035051089920
author Jeune, Pierre Le
Duchesne, Étienne
Xiao, Weixuan
Palminteri, Stefano
Houssin, Bazire
Malézieux, Benoît
Dora, Matteo
author_facet Jeune, Pierre Le
Duchesne, Étienne
Xiao, Weixuan
Palminteri, Stefano
Houssin, Bazire
Malézieux, Benoît
Dora, Matteo
contents Multilingual studies of social bias in open-ended LLM generation remain limited: most existing benchmarks are English-centric, template-based, or restricted to recognizing pre-specified stereotypes. We introduce StereoTales, a multilingual dataset and evaluation pipeline for systematically studying the emergence of social bias in open-ended LLM generation. The dataset covers 10 languages and 79 socio-demographic attributes, and comprises over 650k stories generated by 23 recent LLMs, each annotated with the socio-demographic profile of the protagonist across 19 dimensions. From these, we apply statistical tests to identify more than 1{,}500 over-represented associations, which we then rate for harmfulness through both a panel of humans (N = 247) and the same LLMs. We report three main findings. \textbf{(i)} Every model we evaluate emits consequential harmful stereotypes in open-ended generation, regardless of size or capabilities, and these associations are largely shared across providers rather than isolated misbehaviors. \textbf{(ii)} Prompt language strongly shapes which stereotypes appear: rather than transferring as a shared set of biases, harmful associations adapt culturally to the prompt language and amplify bias against locally salient protected groups. \textbf{(iii)} Human and LLM harmfulness judgments are broadly aligned (Spearman $ρ=0.62$), with disagreements concentrating on specific attribute classes rather than specific providers. To support further analyses, we release the evaluation code and the dataset, including model generations, attribute annotations, and harmfulness ratings.
format Preprint
id arxiv_https___arxiv_org_abs_2605_10442
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs
Jeune, Pierre Le
Duchesne, Étienne
Xiao, Weixuan
Palminteri, Stefano
Houssin, Bazire
Malézieux, Benoît
Dora, Matteo
Computers and Society
Artificial Intelligence
Computation and Language
Multilingual studies of social bias in open-ended LLM generation remain limited: most existing benchmarks are English-centric, template-based, or restricted to recognizing pre-specified stereotypes. We introduce StereoTales, a multilingual dataset and evaluation pipeline for systematically studying the emergence of social bias in open-ended LLM generation. The dataset covers 10 languages and 79 socio-demographic attributes, and comprises over 650k stories generated by 23 recent LLMs, each annotated with the socio-demographic profile of the protagonist across 19 dimensions. From these, we apply statistical tests to identify more than 1{,}500 over-represented associations, which we then rate for harmfulness through both a panel of humans (N = 247) and the same LLMs. We report three main findings. \textbf{(i)} Every model we evaluate emits consequential harmful stereotypes in open-ended generation, regardless of size or capabilities, and these associations are largely shared across providers rather than isolated misbehaviors. \textbf{(ii)} Prompt language strongly shapes which stereotypes appear: rather than transferring as a shared set of biases, harmful associations adapt culturally to the prompt language and amplify bias against locally salient protected groups. \textbf{(iii)} Human and LLM harmfulness judgments are broadly aligned (Spearman $ρ=0.62$), with disagreements concentrating on specific attribute classes rather than specific providers. To support further analyses, we release the evaluation code and the dataset, including model generations, attribute annotations, and harmfulness ratings.
title StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs
topic Computers and Society
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2605.10442