Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rastogi, Charvi, Teh, Tian Huey, Mishra, Pushkar, Patel, Roma, Wang, Ding, Díaz, Mark, Parrish, Alicia, Davani, Aida Mostafazadeh, Ashwood, Zoe, Paganini, Michela, Prabhakaran, Vinodkumar, Rieser, Verena, Aroyo, Lora
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909694662017024
author Rastogi, Charvi
Teh, Tian Huey
Mishra, Pushkar
Patel, Roma
Wang, Ding
Díaz, Mark
Parrish, Alicia
Davani, Aida Mostafazadeh
Ashwood, Zoe
Paganini, Michela
Prabhakaran, Vinodkumar
Rieser, Verena
Aroyo, Lora
author_facet Rastogi, Charvi
Teh, Tian Huey
Mishra, Pushkar
Patel, Roma
Wang, Ding
Díaz, Mark
Parrish, Alicia
Davani, Aida Mostafazadeh
Ashwood, Zoe
Paganini, Michela
Prabhakaran, Vinodkumar
Rieser, Verena
Aroyo, Lora
contents Current text-to-image (T2I) models often fail to account for diverse human experiences, leading to misaligned systems. We advocate for pluralistic alignment, where an AI understands and is steerable towards diverse, and often conflicting, human values. Our work provides three core contributions to achieve this in T2I models. First, we introduce a novel dataset for Diverse Intersectional Visual Evaluation (DIVE) -- the first multimodal dataset for pluralistic alignment. It enable deep alignment to diverse safety perspectives through a large pool of demographically intersectional human raters who provided extensive feedback across 1000 prompts, with high replication, capturing nuanced safety perceptions. Second, we empirically confirm demographics as a crucial proxy for diverse viewpoints in this domain, revealing significant, context-dependent differences in harm perception that diverge from conventional evaluations. Finally, we discuss implications for building aligned T2I models, including efficient data collection strategies, LLM judgment capabilities, and model steerability towards diverse perspectives. This research offers foundational tools for more equitable and aligned T2I systems. Content Warning: The paper includes sensitive content that may be harmful.
format Preprint
id arxiv_https___arxiv_org_abs_2507_13383
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image Models
Rastogi, Charvi
Teh, Tian Huey
Mishra, Pushkar
Patel, Roma
Wang, Ding
Díaz, Mark
Parrish, Alicia
Davani, Aida Mostafazadeh
Ashwood, Zoe
Paganini, Michela
Prabhakaran, Vinodkumar
Rieser, Verena
Aroyo, Lora
Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Current text-to-image (T2I) models often fail to account for diverse human experiences, leading to misaligned systems. We advocate for pluralistic alignment, where an AI understands and is steerable towards diverse, and often conflicting, human values. Our work provides three core contributions to achieve this in T2I models. First, we introduce a novel dataset for Diverse Intersectional Visual Evaluation (DIVE) -- the first multimodal dataset for pluralistic alignment. It enable deep alignment to diverse safety perspectives through a large pool of demographically intersectional human raters who provided extensive feedback across 1000 prompts, with high replication, capturing nuanced safety perceptions. Second, we empirically confirm demographics as a crucial proxy for diverse viewpoints in this domain, revealing significant, context-dependent differences in harm perception that diverge from conventional evaluations. Finally, we discuss implications for building aligned T2I models, including efficient data collection strategies, LLM judgment capabilities, and model steerability towards diverse perspectives. This research offers foundational tools for more equitable and aligned T2I systems. Content Warning: The paper includes sensitive content that may be harmful.
title Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image Models
topic Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.13383