THRONE: An Object-based Hallucination Benchmark for the Free-form Generations of Large Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kaul, Prannay, Li, Zhizhong, Yang, Hao, Dukler, Yonatan, Swaminathan, Ashwin, Taylor, C. J., Soatto, Stefano
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909562476429312
author Kaul, Prannay
Li, Zhizhong
Yang, Hao
Dukler, Yonatan
Swaminathan, Ashwin
Taylor, C. J.
Soatto, Stefano
author_facet Kaul, Prannay
Li, Zhizhong
Yang, Hao
Dukler, Yonatan
Swaminathan, Ashwin
Taylor, C. J.
Soatto, Stefano
contents Mitigating hallucinations in large vision-language models (LVLMs) remains an open problem. Recent benchmarks do not address hallucinations in open-ended free-form responses, which we term "Type I hallucinations". Instead, they focus on hallucinations responding to very specific question formats -- typically a multiple-choice response regarding a particular object or attribute -- which we term "Type II hallucinations". Additionally, such benchmarks often require external API calls to models which are subject to change. In practice, we observe that a reduction in Type II hallucinations does not lead to a reduction in Type I hallucinations but rather that the two forms of hallucinations are often anti-correlated. To address this, we propose THRONE, a novel object-based automatic framework for quantitatively evaluating Type I hallucinations in LVLM free-form outputs. We use public language models (LMs) to identify hallucinations in LVLM responses and compute informative metrics. By evaluating a large selection of recent LVLMs using public datasets, we show that an improvement in existing metrics do not lead to a reduction in Type I hallucinations, and that established benchmarks for measuring Type I hallucinations are incomplete. Finally, we provide a simple and effective data augmentation method to reduce Type I and Type II hallucinations as a strong baseline. Code is now available at https://github.com/amazon-science/THRONE .
format Preprint
id arxiv_https___arxiv_org_abs_2405_05256
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle THRONE: An Object-based Hallucination Benchmark for the Free-form Generations of Large Vision-Language Models
Kaul, Prannay
Li, Zhizhong
Yang, Hao
Dukler, Yonatan
Swaminathan, Ashwin
Taylor, C. J.
Soatto, Stefano
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Mitigating hallucinations in large vision-language models (LVLMs) remains an open problem. Recent benchmarks do not address hallucinations in open-ended free-form responses, which we term "Type I hallucinations". Instead, they focus on hallucinations responding to very specific question formats -- typically a multiple-choice response regarding a particular object or attribute -- which we term "Type II hallucinations". Additionally, such benchmarks often require external API calls to models which are subject to change. In practice, we observe that a reduction in Type II hallucinations does not lead to a reduction in Type I hallucinations but rather that the two forms of hallucinations are often anti-correlated. To address this, we propose THRONE, a novel object-based automatic framework for quantitatively evaluating Type I hallucinations in LVLM free-form outputs. We use public language models (LMs) to identify hallucinations in LVLM responses and compute informative metrics. By evaluating a large selection of recent LVLMs using public datasets, we show that an improvement in existing metrics do not lead to a reduction in Type I hallucinations, and that established benchmarks for measuring Type I hallucinations are incomplete. Finally, we provide a simple and effective data augmentation method to reduce Type I and Type II hallucinations as a strong baseline. Code is now available at https://github.com/amazon-science/THRONE .
title THRONE: An Object-based Hallucination Benchmark for the Free-form Generations of Large Vision-Language Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2405.05256