Person-Centric Annotations of LAION-400M: Auditing Bias and Its Transfer to Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Girrbach, Leander, Alaniz, Stephan, Smith, Genevieve, Darrell, Trevor, Akata, Zeynep
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908918354018304
author Girrbach, Leander
Alaniz, Stephan
Smith, Genevieve
Darrell, Trevor
Akata, Zeynep
author_facet Girrbach, Leander
Alaniz, Stephan
Smith, Genevieve
Darrell, Trevor
Akata, Zeynep
contents Vision-language models trained on large-scale multimodal datasets show strong demographic biases, but the role of training data in producing these biases remains unclear. A major barrier has been the lack of demographic annotations in web-scale datasets such as LAION-400M. We address this gap by creating person-centric annotations for the full dataset, including over 276 million bounding boxes, perceived gender and race/ethnicity labels, and automatically generated captions. These annotations are produced through validated automatic labeling pipelines combining object detection, multimodal captioning, and finetuned classifiers. Using them, we uncover demographic imbalances and harmful associations, such as the disproportionate linking of men and individuals perceived as Black or Middle Eastern with crime-related and negative content. We also show that a linear fit predicts 60-70% of gender bias in CLIP and Stable Diffusion from direct co-occurrences in the data. Our resources establish the first large-scale empirical link between dataset composition and downstream model bias. Code is available at https://github.com/ExplainableML/LAION-400M-Person-Centric-Annotations.
format Preprint
id arxiv_https___arxiv_org_abs_2510_03721
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Person-Centric Annotations of LAION-400M: Auditing Bias and Its Transfer to Models
Girrbach, Leander
Alaniz, Stephan
Smith, Genevieve
Darrell, Trevor
Akata, Zeynep
Computer Vision and Pattern Recognition
Computation and Language
Computers and Society
Machine Learning
Vision-language models trained on large-scale multimodal datasets show strong demographic biases, but the role of training data in producing these biases remains unclear. A major barrier has been the lack of demographic annotations in web-scale datasets such as LAION-400M. We address this gap by creating person-centric annotations for the full dataset, including over 276 million bounding boxes, perceived gender and race/ethnicity labels, and automatically generated captions. These annotations are produced through validated automatic labeling pipelines combining object detection, multimodal captioning, and finetuned classifiers. Using them, we uncover demographic imbalances and harmful associations, such as the disproportionate linking of men and individuals perceived as Black or Middle Eastern with crime-related and negative content. We also show that a linear fit predicts 60-70% of gender bias in CLIP and Stable Diffusion from direct co-occurrences in the data. Our resources establish the first large-scale empirical link between dataset composition and downstream model bias. Code is available at https://github.com/ExplainableML/LAION-400M-Person-Centric-Annotations.
title Person-Centric Annotations of LAION-400M: Auditing Bias and Its Transfer to Models
topic Computer Vision and Pattern Recognition
Computation and Language
Computers and Society
Machine Learning
url https://arxiv.org/abs/2510.03721