From Preferences to Prejudice: The Role of Alignment Tuning in Shaping Social Bias in Video Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cai, Zefan, Qiu, Haoyi, Zhao, Haozhe, Wan, Ke, Li, Jiachen, Gu, Jiuxiang, Xiao, Wen, Peng, Nanyun, Hu, Junjie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908827179286528
author Cai, Zefan
Qiu, Haoyi
Zhao, Haozhe
Wan, Ke
Li, Jiachen
Gu, Jiuxiang
Xiao, Wen
Peng, Nanyun
Hu, Junjie
author_facet Cai, Zefan
Qiu, Haoyi
Zhao, Haozhe
Wan, Ke
Li, Jiachen
Gu, Jiuxiang
Xiao, Wen
Peng, Nanyun
Hu, Junjie
contents Recent advances in video diffusion models have significantly enhanced text-to-video generation, particularly through alignment tuning using reward models trained on human preferences. While these methods improve visual quality, they can unintentionally encode and amplify social biases. To systematically trace how such biases evolve throughout the alignment pipeline, we introduce VideoBiasEval, a comprehensive diagnostic framework for evaluating social representation in video generation. Grounded in established social bias taxonomies, VideoBiasEval employs an event-based prompting strategy to disentangle semantic content (actions and contexts) from actor attributes (gender and ethnicity). It further introduces multi-granular metrics to evaluate (1) overall ethnicity bias, (2) gender bias conditioned on ethnicity, (3) distributional shifts in social attributes across model variants, and (4) the temporal persistence of bias within videos. Using this framework, we conduct the first end-to-end analysis connecting biases in human preference datasets, their amplification in reward models, and their propagation through alignment-tuned video diffusion models. Our results reveal that alignment tuning not only strengthens representational biases but also makes them temporally stable, producing smoother yet more stereotyped portrayals. These findings highlight the need for bias-aware evaluation and mitigation throughout the alignment process to ensure fair and socially responsible video generation.
format Preprint
id arxiv_https___arxiv_org_abs_2510_17247
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From Preferences to Prejudice: The Role of Alignment Tuning in Shaping Social Bias in Video Diffusion Models
Cai, Zefan
Qiu, Haoyi
Zhao, Haozhe
Wan, Ke
Li, Jiachen
Gu, Jiuxiang
Xiao, Wen
Peng, Nanyun
Hu, Junjie
Computation and Language
Computer Vision and Pattern Recognition
Recent advances in video diffusion models have significantly enhanced text-to-video generation, particularly through alignment tuning using reward models trained on human preferences. While these methods improve visual quality, they can unintentionally encode and amplify social biases. To systematically trace how such biases evolve throughout the alignment pipeline, we introduce VideoBiasEval, a comprehensive diagnostic framework for evaluating social representation in video generation. Grounded in established social bias taxonomies, VideoBiasEval employs an event-based prompting strategy to disentangle semantic content (actions and contexts) from actor attributes (gender and ethnicity). It further introduces multi-granular metrics to evaluate (1) overall ethnicity bias, (2) gender bias conditioned on ethnicity, (3) distributional shifts in social attributes across model variants, and (4) the temporal persistence of bias within videos. Using this framework, we conduct the first end-to-end analysis connecting biases in human preference datasets, their amplification in reward models, and their propagation through alignment-tuned video diffusion models. Our results reveal that alignment tuning not only strengthens representational biases but also makes them temporally stable, producing smoother yet more stereotyped portrayals. These findings highlight the need for bias-aware evaluation and mitigation throughout the alignment process to ensure fair and socially responsible video generation.
title From Preferences to Prejudice: The Role of Alignment Tuning in Shaping Social Bias in Video Diffusion Models
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.17247