EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Villa, Andrés, Alcázar, Juan León, Alfarra, Motasem, Araujo, Vladimir, Soto, Alvaro, Ghanem, Bernard
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912178087395328
author Villa, Andrés
Alcázar, Juan León
Alfarra, Motasem
Araujo, Vladimir
Soto, Alvaro
Ghanem, Bernard
author_facet Villa, Andrés
Alcázar, Juan León
Alfarra, Motasem
Araujo, Vladimir
Soto, Alvaro
Ghanem, Bernard
contents Large language models and vision transformers have demonstrated impressive zero-shot capabilities, enabling significant transferability in downstream tasks. The fusion of these models has resulted in multi-modal architectures with enhanced instructional capabilities. Despite incorporating vast image and language pre-training, these multi-modal architectures often generate responses that deviate from the ground truth in the image data. These failure cases are known as hallucinations. Current methods for mitigating hallucinations generally focus on regularizing the language component, improving the fusion module, or ensembling multiple visual encoders to improve visual representation. In this paper, we address the hallucination issue by directly enhancing the capabilities of the visual component. Our approach, named EAGLE, is fully agnostic to the LLM or fusion module and works as a post-pretraining approach that improves the grounding and language alignment of the visual encoder. We show that a straightforward reformulation of the original contrastive pre-training task results in an improved visual encoder that can be incorporated into the instructional multi-modal architecture without additional instructional training. As a result, EAGLE achieves a significant reduction in hallucinations across multiple challenging benchmarks and tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2501_02699
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models
Villa, Andrés
Alcázar, Juan León
Alfarra, Motasem
Araujo, Vladimir
Soto, Alvaro
Ghanem, Bernard
Computer Vision and Pattern Recognition
Artificial Intelligence
Large language models and vision transformers have demonstrated impressive zero-shot capabilities, enabling significant transferability in downstream tasks. The fusion of these models has resulted in multi-modal architectures with enhanced instructional capabilities. Despite incorporating vast image and language pre-training, these multi-modal architectures often generate responses that deviate from the ground truth in the image data. These failure cases are known as hallucinations. Current methods for mitigating hallucinations generally focus on regularizing the language component, improving the fusion module, or ensembling multiple visual encoders to improve visual representation. In this paper, we address the hallucination issue by directly enhancing the capabilities of the visual component. Our approach, named EAGLE, is fully agnostic to the LLM or fusion module and works as a post-pretraining approach that improves the grounding and language alignment of the visual encoder. We show that a straightforward reformulation of the original contrastive pre-training task results in an improved visual encoder that can be incorporated into the instructional multi-modal architecture without additional instructional training. As a result, EAGLE achieves a significant reduction in hallucinations across multiple challenging benchmarks and tasks.
title EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2501.02699