Fair Text-to-Image Diffusion via Fair Mapping

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Jia, Hu, Lijie, Zhang, Jingfeng, Zheng, Tianhang, Zhang, Hua, Wang, Di
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910354621071360
author Li, Jia
Hu, Lijie
Zhang, Jingfeng
Zheng, Tianhang
Zhang, Hua
Wang, Di
author_facet Li, Jia
Hu, Lijie
Zhang, Jingfeng
Zheng, Tianhang
Zhang, Hua
Wang, Di
contents In this paper, we address the limitations of existing text-to-image diffusion models in generating demographically fair results when given human-related descriptions. These models often struggle to disentangle the target language context from sociocultural biases, resulting in biased image generation. To overcome this challenge, we propose Fair Mapping, a flexible, model-agnostic, and lightweight approach that modifies a pre-trained text-to-image diffusion model by controlling the prompt to achieve fair image generation. One key advantage of our approach is its high efficiency. It only requires updating an additional linear network with few parameters at a low computational cost. By developing a linear network that maps conditioning embeddings into a debiased space, we enable the generation of relatively balanced demographic results based on the specified text condition. With comprehensive experiments on face image generation, we show that our method significantly improves image generation fairness with almost the same image quality compared to conventional diffusion models when prompted with descriptions related to humans. By effectively addressing the issue of implicit language bias, our method produces more fair and diverse image outputs.
format Preprint
id arxiv_https___arxiv_org_abs_2311_17695
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Fair Text-to-Image Diffusion via Fair Mapping
Li, Jia
Hu, Lijie
Zhang, Jingfeng
Zheng, Tianhang
Zhang, Hua
Wang, Di
Computer Vision and Pattern Recognition
Artificial Intelligence
Computers and Society
Machine Learning
In this paper, we address the limitations of existing text-to-image diffusion models in generating demographically fair results when given human-related descriptions. These models often struggle to disentangle the target language context from sociocultural biases, resulting in biased image generation. To overcome this challenge, we propose Fair Mapping, a flexible, model-agnostic, and lightweight approach that modifies a pre-trained text-to-image diffusion model by controlling the prompt to achieve fair image generation. One key advantage of our approach is its high efficiency. It only requires updating an additional linear network with few parameters at a low computational cost. By developing a linear network that maps conditioning embeddings into a debiased space, we enable the generation of relatively balanced demographic results based on the specified text condition. With comprehensive experiments on face image generation, we show that our method significantly improves image generation fairness with almost the same image quality compared to conventional diffusion models when prompted with descriptions related to humans. By effectively addressing the issue of implicit language bias, our method produces more fair and diverse image outputs.
title Fair Text-to-Image Diffusion via Fair Mapping
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computers and Society
Machine Learning
url https://arxiv.org/abs/2311.17695