AutoHallusion: Automatic Generation of Hallucination Benchmarks for Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Xiyang, Guan, Tianrui, Li, Dianqi, Huang, Shuaiyi, Liu, Xiaoyu, Wang, Xijun, Xian, Ruiqi, Shrivastava, Abhinav, Huang, Furong, Boyd-Graber, Jordan Lee, Zhou, Tianyi, Manocha, Dinesh
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914967602593792
author Wu, Xiyang
Guan, Tianrui
Li, Dianqi
Huang, Shuaiyi
Liu, Xiaoyu
Wang, Xijun
Xian, Ruiqi
Shrivastava, Abhinav
Huang, Furong
Boyd-Graber, Jordan Lee
Zhou, Tianyi
Manocha, Dinesh
author_facet Wu, Xiyang
Guan, Tianrui
Li, Dianqi
Huang, Shuaiyi
Liu, Xiaoyu
Wang, Xijun
Xian, Ruiqi
Shrivastava, Abhinav
Huang, Furong
Boyd-Graber, Jordan Lee
Zhou, Tianyi
Manocha, Dinesh
contents Large vision-language models (LVLMs) are prone to hallucinations, where certain contextual cues in an image can trigger the language module to produce overconfident and incorrect reasoning about abnormal or hypothetical objects. While some benchmarks have been developed to investigate LVLM hallucinations, they often rely on hand-crafted corner cases whose failure patterns may not generalize well. Additionally, fine-tuning on these examples could undermine their validity. To address this, we aim to scale up the number of cases through an automated approach, reducing human bias in crafting such corner cases. This motivates the development of AutoHallusion, the first automated benchmark generation approach that employs several key strategies to create a diverse range of hallucination examples. Our generated visual-question pairs pose significant challenges to LVLMs, requiring them to overcome contextual biases and distractions to arrive at correct answers. AutoHallusion enables us to create new benchmarks at the minimum cost and thus overcomes the fragility of hand-crafted benchmarks. It also reveals common failure patterns and reasons, providing key insights to detect, avoid, or control hallucinations. Comprehensive evaluations of top-tier LVLMs, e.g., GPT-4V(ision), Gemini Pro Vision, Claude 3, and LLaVA-1.5, show a 97.7% and 98.7% success rate of hallucination induction on synthetic and real-world datasets of AutoHallusion, paving the way for a long battle against hallucinations. The codebase and data can be accessed at https://github.com/wuxiyang1996/AutoHallusion.
format Preprint
id arxiv_https___arxiv_org_abs_2406_10900
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AutoHallusion: Automatic Generation of Hallucination Benchmarks for Vision-Language Models
Wu, Xiyang
Guan, Tianrui
Li, Dianqi
Huang, Shuaiyi
Liu, Xiaoyu
Wang, Xijun
Xian, Ruiqi
Shrivastava, Abhinav
Huang, Furong
Boyd-Graber, Jordan Lee
Zhou, Tianyi
Manocha, Dinesh
Computer Vision and Pattern Recognition
Computation and Language
Large vision-language models (LVLMs) are prone to hallucinations, where certain contextual cues in an image can trigger the language module to produce overconfident and incorrect reasoning about abnormal or hypothetical objects. While some benchmarks have been developed to investigate LVLM hallucinations, they often rely on hand-crafted corner cases whose failure patterns may not generalize well. Additionally, fine-tuning on these examples could undermine their validity. To address this, we aim to scale up the number of cases through an automated approach, reducing human bias in crafting such corner cases. This motivates the development of AutoHallusion, the first automated benchmark generation approach that employs several key strategies to create a diverse range of hallucination examples. Our generated visual-question pairs pose significant challenges to LVLMs, requiring them to overcome contextual biases and distractions to arrive at correct answers. AutoHallusion enables us to create new benchmarks at the minimum cost and thus overcomes the fragility of hand-crafted benchmarks. It also reveals common failure patterns and reasons, providing key insights to detect, avoid, or control hallucinations. Comprehensive evaluations of top-tier LVLMs, e.g., GPT-4V(ision), Gemini Pro Vision, Claude 3, and LLaVA-1.5, show a 97.7% and 98.7% success rate of hallucination induction on synthetic and real-world datasets of AutoHallusion, paving the way for a long battle against hallucinations. The codebase and data can be accessed at https://github.com/wuxiyang1996/AutoHallusion.
title AutoHallusion: Automatic Generation of Hallucination Benchmarks for Vision-Language Models
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2406.10900