AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sung-Bin, Kim, Hyun-Bin, Oh, Lee, JungMok, Senocak, Arda, Chung, Joon Son, Oh, Tae-Hyun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912278134128640
author Sung-Bin, Kim
Hyun-Bin, Oh
Lee, JungMok
Senocak, Arda
Chung, Joon Son
Oh, Tae-Hyun
author_facet Sung-Bin, Kim
Hyun-Bin, Oh
Lee, JungMok
Senocak, Arda
Chung, Joon Son
Oh, Tae-Hyun
contents Following the success of Large Language Models (LLMs), expanding their boundaries to new modalities represents a significant paradigm shift in multimodal understanding. Human perception is inherently multimodal, relying not only on text but also on auditory and visual cues for a complete understanding of the world. In recognition of this fact, audio-visual LLMs have recently emerged. Despite promising developments, the lack of dedicated benchmarks poses challenges for understanding and evaluating models. In this work, we show that audio-visual LLMs struggle to discern subtle relationships between audio and visual signals, leading to hallucinations and highlighting the need for reliable benchmarks. To address this, we introduce AVHBench, the first comprehensive benchmark specifically designed to evaluate the perception and comprehension capabilities of audio-visual LLMs. Our benchmark includes tests for assessing hallucinations, as well as the cross-modal matching and reasoning abilities of these models. Our results reveal that most existing audio-visual LLMs struggle with hallucinations caused by cross-interactions between modalities, due to their limited capacity to perceive complex multimodal signals and their relationships. Additionally, we demonstrate that simple training with our AVHBench improves robustness of audio-visual LLMs against hallucinations. Dataset: https://github.com/kaist-ami/AVHBench
format Preprint
id arxiv_https___arxiv_org_abs_2410_18325
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models
Sung-Bin, Kim
Hyun-Bin, Oh
Lee, JungMok
Senocak, Arda
Chung, Joon Son
Oh, Tae-Hyun
Computer Vision and Pattern Recognition
Following the success of Large Language Models (LLMs), expanding their boundaries to new modalities represents a significant paradigm shift in multimodal understanding. Human perception is inherently multimodal, relying not only on text but also on auditory and visual cues for a complete understanding of the world. In recognition of this fact, audio-visual LLMs have recently emerged. Despite promising developments, the lack of dedicated benchmarks poses challenges for understanding and evaluating models. In this work, we show that audio-visual LLMs struggle to discern subtle relationships between audio and visual signals, leading to hallucinations and highlighting the need for reliable benchmarks. To address this, we introduce AVHBench, the first comprehensive benchmark specifically designed to evaluate the perception and comprehension capabilities of audio-visual LLMs. Our benchmark includes tests for assessing hallucinations, as well as the cross-modal matching and reasoning abilities of these models. Our results reveal that most existing audio-visual LLMs struggle with hallucinations caused by cross-interactions between modalities, due to their limited capacity to perceive complex multimodal signals and their relationships. Additionally, we demonstrate that simple training with our AVHBench improves robustness of audio-visual LLMs against hallucinations. Dataset: https://github.com/kaist-ami/AVHBench
title AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.18325