EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Seth, Ashish, Tyagi, Utkarsh, Selvakumar, Ramaneswaran, Anand, Nishit, Kumar, Sonal, Ghosh, Sreyan, Duraiswami, Ramani, Agarwal, Chirag, Manocha, Dinesh
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912551120404480
author Seth, Ashish
Tyagi, Utkarsh
Selvakumar, Ramaneswaran
Anand, Nishit
Kumar, Sonal
Ghosh, Sreyan
Duraiswami, Ramani
Agarwal, Chirag
Manocha, Dinesh
author_facet Seth, Ashish
Tyagi, Utkarsh
Selvakumar, Ramaneswaran
Anand, Nishit
Kumar, Sonal
Ghosh, Sreyan
Duraiswami, Ramani
Agarwal, Chirag
Manocha, Dinesh
contents Multimodal Large Language Models (MLLMs) have demonstrated remarkable performance in complex multimodal tasks. While MLLMs excel at visual perception and reasoning in third-person and egocentric videos, they are prone to hallucinations, generating coherent yet inaccurate responses. We present EgoIllusion, a first benchmark to evaluate MLLM hallucinations in egocentric videos. EgoIllusion comprises 1,400 videos paired with 8,000 human-annotated open and closed-ended questions designed to trigger hallucinations in both visual and auditory cues in egocentric videos. Evaluations across ten MLLMs reveal significant challenges, including powerful models like GPT-4o and Gemini, achieving only 59% accuracy. EgoIllusion lays the foundation in developing robust benchmarks to evaluate the effectiveness of MLLMs and spurs the development of better egocentric MLLMs with reduced hallucination rates. Our benchmark will be open-sourced for reproducibility.
format Preprint
id arxiv_https___arxiv_org_abs_2508_12687
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding
Seth, Ashish
Tyagi, Utkarsh
Selvakumar, Ramaneswaran
Anand, Nishit
Kumar, Sonal
Ghosh, Sreyan
Duraiswami, Ramani
Agarwal, Chirag
Manocha, Dinesh
Artificial Intelligence
Computer Vision and Pattern Recognition
Multimodal Large Language Models (MLLMs) have demonstrated remarkable performance in complex multimodal tasks. While MLLMs excel at visual perception and reasoning in third-person and egocentric videos, they are prone to hallucinations, generating coherent yet inaccurate responses. We present EgoIllusion, a first benchmark to evaluate MLLM hallucinations in egocentric videos. EgoIllusion comprises 1,400 videos paired with 8,000 human-annotated open and closed-ended questions designed to trigger hallucinations in both visual and auditory cues in egocentric videos. Evaluations across ten MLLMs reveal significant challenges, including powerful models like GPT-4o and Gemini, achieving only 59% accuracy. EgoIllusion lays the foundation in developing robust benchmarks to evaluate the effectiveness of MLLMs and spurs the development of better egocentric MLLMs with reduced hallucination rates. Our benchmark will be open-sourced for reproducibility.
title EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding
topic Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.12687