SCoFT: Self-Contrastive Fine-Tuning for Equitable Image Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Zhixuan, Schaldenbrand, Peter, Okogwu, Beverley-Claire, Peng, Wenxuan, Yun, Youngsik, Hundt, Andrew, Kim, Jihie, Oh, Jean
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911758522777600
author Liu, Zhixuan
Schaldenbrand, Peter
Okogwu, Beverley-Claire
Peng, Wenxuan
Yun, Youngsik
Hundt, Andrew
Kim, Jihie
Oh, Jean
author_facet Liu, Zhixuan
Schaldenbrand, Peter
Okogwu, Beverley-Claire
Peng, Wenxuan
Yun, Youngsik
Hundt, Andrew
Kim, Jihie
Oh, Jean
contents Accurate representation in media is known to improve the well-being of the people who consume it. Generative image models trained on large web-crawled datasets such as LAION are known to produce images with harmful stereotypes and misrepresentations of cultures. We improve inclusive representation in generated images by (1) engaging with communities to collect a culturally representative dataset that we call the Cross-Cultural Understanding Benchmark (CCUB) and (2) proposing a novel Self-Contrastive Fine-Tuning (SCoFT) method that leverages the model's known biases to self-improve. SCoFT is designed to prevent overfitting on small datasets, encode only high-level information from the data, and shift the generated distribution away from misrepresentations encoded in a pretrained model. Our user study conducted on 51 participants from 5 different countries based on their self-selected national cultural affiliation shows that fine-tuning on CCUB consistently generates images with higher cultural relevance and fewer stereotypes when compared to the Stable Diffusion baseline, which is further improved with our SCoFT technique.
format Preprint
id arxiv_https___arxiv_org_abs_2401_08053
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SCoFT: Self-Contrastive Fine-Tuning for Equitable Image Generation
Liu, Zhixuan
Schaldenbrand, Peter
Okogwu, Beverley-Claire
Peng, Wenxuan
Yun, Youngsik
Hundt, Andrew
Kim, Jihie
Oh, Jean
Computer Vision and Pattern Recognition
Accurate representation in media is known to improve the well-being of the people who consume it. Generative image models trained on large web-crawled datasets such as LAION are known to produce images with harmful stereotypes and misrepresentations of cultures. We improve inclusive representation in generated images by (1) engaging with communities to collect a culturally representative dataset that we call the Cross-Cultural Understanding Benchmark (CCUB) and (2) proposing a novel Self-Contrastive Fine-Tuning (SCoFT) method that leverages the model's known biases to self-improve. SCoFT is designed to prevent overfitting on small datasets, encode only high-level information from the data, and shift the generated distribution away from misrepresentations encoded in a pretrained model. Our user study conducted on 51 participants from 5 different countries based on their self-selected national cultural affiliation shows that fine-tuning on CCUB consistently generates images with higher cultural relevance and fewer stereotypes when compared to the Stable Diffusion baseline, which is further improved with our SCoFT technique.
title SCoFT: Self-Contrastive Fine-Tuning for Equitable Image Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2401.08053