Hal-Eval: A Universal and Fine-grained Hallucination Evaluation Framework for Large Vision Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Chaoya, Jia, Hongrui, Ye, Wei, Dong, Mengfan, Xu, Haiyang, Yan, Ming, Zhang, Ji, Zhang, Shikun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929582633910272
author Jiang, Chaoya
Jia, Hongrui
Ye, Wei
Dong, Mengfan
Xu, Haiyang
Yan, Ming
Zhang, Ji
Zhang, Shikun
author_facet Jiang, Chaoya
Jia, Hongrui
Ye, Wei
Dong, Mengfan
Xu, Haiyang
Yan, Ming
Zhang, Ji
Zhang, Shikun
contents Large Vision Language Models exhibit remarkable capabilities but struggle with hallucinations inconsistencies between images and their descriptions. Previous hallucination evaluation studies on LVLMs have identified hallucinations in terms of objects, attributes, and relations but overlooked complex hallucinations that create an entire narrative around a fictional entity. In this paper, we introduce a refined taxonomy of hallucinations, featuring a new category: Event Hallucination. We then utilize advanced LLMs to generate and filter fine grained hallucinatory data consisting of various types of hallucinations, with a particular focus on event hallucinations, laying the groundwork for integrating discriminative and generative evaluation methods within our universal evaluation framework. The proposed benchmark distinctively assesses LVLMs ability to tackle a broad spectrum of hallucinations, making it a reliable and comprehensive tool for gauging LVLMs efficacy in handling hallucinations. We will release our code and data.
format Preprint
id arxiv_https___arxiv_org_abs_2402_15721
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Hal-Eval: A Universal and Fine-grained Hallucination Evaluation Framework for Large Vision Language Models
Jiang, Chaoya
Jia, Hongrui
Ye, Wei
Dong, Mengfan
Xu, Haiyang
Yan, Ming
Zhang, Ji
Zhang, Shikun
Artificial Intelligence
Computation and Language
Large Vision Language Models exhibit remarkable capabilities but struggle with hallucinations inconsistencies between images and their descriptions. Previous hallucination evaluation studies on LVLMs have identified hallucinations in terms of objects, attributes, and relations but overlooked complex hallucinations that create an entire narrative around a fictional entity. In this paper, we introduce a refined taxonomy of hallucinations, featuring a new category: Event Hallucination. We then utilize advanced LLMs to generate and filter fine grained hallucinatory data consisting of various types of hallucinations, with a particular focus on event hallucinations, laying the groundwork for integrating discriminative and generative evaluation methods within our universal evaluation framework. The proposed benchmark distinctively assesses LVLMs ability to tackle a broad spectrum of hallucinations, making it a reliable and comprehensive tool for gauging LVLMs efficacy in handling hallucinations. We will release our code and data.
title Hal-Eval: A Universal and Fine-grained Hallucination Evaluation Framework for Large Vision Language Models
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2402.15721