HESEIA: A community-based dataset for evaluating social biases in large language models, co-designed in real school settings in Latin America

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ivetta, Guido, Gomez, Marcos J., Martinelli, Sofía, Palombini, Pietro, Echeveste, M. Emilia, Mazzeo, Nair Carolina, Busaniche, Beatriz, Benotti, Luciana
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910976332267520
author Ivetta, Guido
Gomez, Marcos J.
Martinelli, Sofía
Palombini, Pietro
Echeveste, M. Emilia
Mazzeo, Nair Carolina
Busaniche, Beatriz
Benotti, Luciana
author_facet Ivetta, Guido
Gomez, Marcos J.
Martinelli, Sofía
Palombini, Pietro
Echeveste, M. Emilia
Mazzeo, Nair Carolina
Busaniche, Beatriz
Benotti, Luciana
contents Most resources for evaluating social biases in Large Language Models are developed without co-design from the communities affected by these biases, and rarely involve participatory approaches. We introduce HESEIA, a dataset of 46,499 sentences created in a professional development course. The course involved 370 high-school teachers and 5,370 students from 189 Latin-American schools. Unlike existing benchmarks, HESEIA captures intersectional biases across multiple demographic axes and school subjects. It reflects local contexts through the lived experience and pedagogical expertise of educators. Teachers used minimal pairs to create sentences that express stereotypes relevant to their school subjects and communities. We show the dataset diversity in term of demographic axes represented and also in terms of the knowledge areas included. We demonstrate that the dataset contains more stereotypes unrecognized by current LLMs than previous datasets. HESEIA is available to support bias assessments grounded in educational communities.
format Preprint
id arxiv_https___arxiv_org_abs_2505_24712
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HESEIA: A community-based dataset for evaluating social biases in large language models, co-designed in real school settings in Latin America
Ivetta, Guido
Gomez, Marcos J.
Martinelli, Sofía
Palombini, Pietro
Echeveste, M. Emilia
Mazzeo, Nair Carolina
Busaniche, Beatriz
Benotti, Luciana
Computation and Language
Computers and Society
Most resources for evaluating social biases in Large Language Models are developed without co-design from the communities affected by these biases, and rarely involve participatory approaches. We introduce HESEIA, a dataset of 46,499 sentences created in a professional development course. The course involved 370 high-school teachers and 5,370 students from 189 Latin-American schools. Unlike existing benchmarks, HESEIA captures intersectional biases across multiple demographic axes and school subjects. It reflects local contexts through the lived experience and pedagogical expertise of educators. Teachers used minimal pairs to create sentences that express stereotypes relevant to their school subjects and communities. We show the dataset diversity in term of demographic axes represented and also in terms of the knowledge areas included. We demonstrate that the dataset contains more stereotypes unrecognized by current LLMs than previous datasets. HESEIA is available to support bias assessments grounded in educational communities.
title HESEIA: A community-based dataset for evaluating social biases in large language models, co-designed in real school settings in Latin America
topic Computation and Language
Computers and Society
url https://arxiv.org/abs/2505.24712