From Descriptive Richness to Bias: Unveiling the Dark Side of Generative Image Caption Enrichment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hirota, Yusuke, Hachiuma, Ryo, Yang, Chao-Han Huck, Nakashima, Yuta
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909228029968384
author Hirota, Yusuke
Hachiuma, Ryo
Yang, Chao-Han Huck
Nakashima, Yuta
author_facet Hirota, Yusuke
Hachiuma, Ryo
Yang, Chao-Han Huck
Nakashima, Yuta
contents Large language models (LLMs) have enhanced the capacity of vision-language models to caption visual text. This generative approach to image caption enrichment further makes textual captions more descriptive, improving alignment with the visual context. However, while many studies focus on benefits of generative caption enrichment (GCE), are there any negative side effects? We compare standard-format captions and recent GCE processes from the perspectives of "gender bias" and "hallucination", showing that enriched captions suffer from increased gender bias and hallucination. Furthermore, models trained on these enriched captions amplify gender bias by an average of 30.9% and increase hallucination by 59.5%. This study serves as a caution against the trend of making captions more descriptive.
format Preprint
id arxiv_https___arxiv_org_abs_2406_13912
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle From Descriptive Richness to Bias: Unveiling the Dark Side of Generative Image Caption Enrichment
Hirota, Yusuke
Hachiuma, Ryo
Yang, Chao-Han Huck
Nakashima, Yuta
Computer Vision and Pattern Recognition
Large language models (LLMs) have enhanced the capacity of vision-language models to caption visual text. This generative approach to image caption enrichment further makes textual captions more descriptive, improving alignment with the visual context. However, while many studies focus on benefits of generative caption enrichment (GCE), are there any negative side effects? We compare standard-format captions and recent GCE processes from the perspectives of "gender bias" and "hallucination", showing that enriched captions suffer from increased gender bias and hallucination. Furthermore, models trained on these enriched captions amplify gender bias by an average of 30.9% and increase hallucination by 59.5%. This study serves as a caution against the trend of making captions more descriptive.
title From Descriptive Richness to Bias: Unveiling the Dark Side of Generative Image Caption Enrichment
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2406.13912