EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models
Fuente:
arXiv
Saved in:
| Main Authors: | Villa, Andrés, Alcázar, Juan León, Alfarra, Motasem, Araujo, Vladimir, Soto, Alvaro, Ghanem, Bernard |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Test-Time Adaptation for Combating Missing Modalities in Egocentric Videos
by: Ramazanova, Merey, et al.
Published: (2024)
by: Ramazanova, Merey, et al.
Published: (2024)
Behind the Magic, MERLIM: Multi-modal Evaluation Benchmark for Large Image-Language Models
by: Villa, Andrés, et al.
Published: (2023)
by: Villa, Andrés, et al.
Published: (2023)
MoDA: Modulation Adapter for Fine-Grained Visual Grounding in Instructional MLLMs
by: Barrios, Wayner, et al.
Published: (2025)
by: Barrios, Wayner, et al.
Published: (2025)
Forget Less, Retain More: A Lightweight Regularizer for Rehearsal-Based Continual Learning
by: Alssum, Lama, et al.
Published: (2025)
by: Alssum, Lama, et al.
Published: (2025)
ADVMEM: Adversarial Memory Initialization for Realistic Test-Time Adaptation via Tracklet-Based Benchmarking
by: Alhuwaider, Shyma, et al.
Published: (2025)
by: Alhuwaider, Shyma, et al.
Published: (2025)
SimCS: Simulation for Domain Incremental Online Continual Segmentation
by: Alfarra, Motasem, et al.
Published: (2022)
by: Alfarra, Motasem, et al.
Published: (2022)
CURE: Curriculum-guided Multi-task Training for Reliable Anatomy Grounded Report Generation
by: Messina, Pablo, et al.
Published: (2026)
by: Messina, Pablo, et al.
Published: (2026)
Deep Learning at the Intersection: Certified Robustness as a Tool for 3D Vision
by: S, Gabriel Pérez, et al.
Published: (2024)
by: S, Gabriel Pérez, et al.
Published: (2024)
Online Distillation with Continual Learning for Cyclic Domain Shifts
by: Houyon, Joachim, et al.
Published: (2023)
by: Houyon, Joachim, et al.
Published: (2023)
FedMedICL: Towards Holistic Evaluation of Distribution Shifts in Federated Medical Imaging
by: Alhamoud, Kumail, et al.
Published: (2024)
by: Alhamoud, Kumail, et al.
Published: (2024)
Evaluation of Test-Time Adaptation Under Computational Time Constraints
by: Alfarra, Motasem, et al.
Published: (2023)
by: Alfarra, Motasem, et al.
Published: (2023)
Compressed-Language Models for Understanding Compressed File Formats: a JPEG Exploration
by: Pérez, Juan C., et al.
Published: (2024)
by: Pérez, Juan C., et al.
Published: (2024)
MoCA-Video: Motion-Aware Concept Alignment for Consistent Video Editing
by: Zhang, Tong, et al.
Published: (2025)
by: Zhang, Tong, et al.
Published: (2025)
EAGLE: Elevating Geometric Reasoning through LLM-empowered Visual Instruction Tuning
by: Li, Zhihao, et al.
Published: (2024)
by: Li, Zhihao, et al.
Published: (2024)
EAGLE: Towards Efficient Arbitrary Referring Visual Prompts Comprehension for Multimodal Large Language Models
by: Zhang, Jiacheng, et al.
Published: (2024)
by: Zhang, Jiacheng, et al.
Published: (2024)
Exploring Missing Modality in Multimodal Egocentric Datasets
by: Ramazanova, Merey, et al.
Published: (2024)
by: Ramazanova, Merey, et al.
Published: (2024)
Mind-the-Glitch: Visual Correspondence for Detecting Inconsistencies in Subject-Driven Generation
by: Eldesokey, Abdelrahman, et al.
Published: (2025)
by: Eldesokey, Abdelrahman, et al.
Published: (2025)
Seeing to Ground: Visual Attention for Hallucination-Resilient MDLLMs
by: Narnaware, Vishal, et al.
Published: (2026)
by: Narnaware, Vishal, et al.
Published: (2026)
SAGE: Sink-Aware Grounded Decoding for Multimodal Hallucination Mitigation
by: Shukla, Tripti, et al.
Published: (2026)
by: Shukla, Tripti, et al.
Published: (2026)
Hallucinate, Ground, Repeat: A Framework for Generalized Visual Relationship Detection
by: Vellamcheti, Shanmukha, et al.
Published: (2025)
by: Vellamcheti, Shanmukha, et al.
Published: (2025)
Towards Faster and More Compact Foundation Models for Molecular Property Prediction
by: Ghunaim, Yasir, et al.
Published: (2025)
by: Ghunaim, Yasir, et al.
Published: (2025)
EAGLE: Expert-Augmented Attention Guidance for Tuning-Free Industrial Anomaly Detection in Multimodal Large Language Models
by: Peng, Xiaomeng, et al.
Published: (2026)
by: Peng, Xiaomeng, et al.
Published: (2026)
Adversarial Robustness for Visual Grounding of Multimodal Large Language Models
by: Gao, Kuofeng, et al.
Published: (2024)
by: Gao, Kuofeng, et al.
Published: (2024)
Hybrid Structure-from-Motion and Camera Relocalization for Enhanced Egocentric Localization
by: Mai, Jinjie, et al.
Published: (2024)
by: Mai, Jinjie, et al.
Published: (2024)
Look Carefully: Adaptive Visual Reinforcements in Multimodal Large Language Models for Hallucination Mitigation
by: Zhu, Xingyu, et al.
Published: (2026)
by: Zhu, Xingyu, et al.
Published: (2026)
Video Self-Stitching Graph Network for Temporal Action Localization
by: Zhao, Chen, et al.
Published: (2020)
by: Zhao, Chen, et al.
Published: (2020)
BOLT: Boost Large Vision-Language Model Without Training for Long-form Video Understanding
by: Liu, Shuming, et al.
Published: (2025)
by: Liu, Shuming, et al.
Published: (2025)
Cure or Poison? Embedding Instructions Visually Alters Hallucination in Vision-Language Models
by: Wang, Zhaochen, et al.
Published: (2025)
by: Wang, Zhaochen, et al.
Published: (2025)
Instruction-Aligned Visual Attention for Mitigating Hallucinations in Large Vision-Language Models
by: Li, Bin, et al.
Published: (2025)
by: Li, Bin, et al.
Published: (2025)
GroundVTS: Visual Token Sampling in Multimodal Large Language Models for Video Temporal Grounding
by: Fan, Rong, et al.
Published: (2026)
by: Fan, Rong, et al.
Published: (2026)
Localizing Before Answering: A Hallucination Evaluation Benchmark for Grounded Medical Multimodal LLMs
by: Nguyen, Dung, et al.
Published: (2025)
by: Nguyen, Dung, et al.
Published: (2025)
Pre-Training Multimodal Hallucination Detectors with Corrupted Grounding Data
by: Whitehead, Spencer, et al.
Published: (2024)
by: Whitehead, Spencer, et al.
Published: (2024)
The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio
by: Leng, Sicong, et al.
Published: (2024)
by: Leng, Sicong, et al.
Published: (2024)
Extracting Visual Facts from Intermediate Layers for Mitigating Hallucinations in Multimodal Large Language Models
by: Zhou, Haoran, et al.
Published: (2025)
by: Zhou, Haoran, et al.
Published: (2025)
GenView: Enhancing View Quality with Pretrained Generative Model for Self-Supervised Learning
by: Li, Xiaojie, et al.
Published: (2024)
by: Li, Xiaojie, et al.
Published: (2024)
Multimodal Reference Visual Grounding
by: Lu, Yangxiao, et al.
Published: (2025)
by: Lu, Yangxiao, et al.
Published: (2025)
Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models
by: Jin, Hyundong, et al.
Published: (2025)
by: Jin, Hyundong, et al.
Published: (2025)
Visual Attention Drifts,but Anchors Hold:Mitigating Hallucination in Multimodal Large Language Models via Cross-Layer Visual Anchors
by: Yang, Chengxu, et al.
Published: (2026)
by: Yang, Chengxu, et al.
Published: (2026)
Visual Question Answering Instruction: Unlocking Multimodal Large Language Model To Domain-Specific Visual Multitasks
by: Lee, Jusung, et al.
Published: (2024)
by: Lee, Jusung, et al.
Published: (2024)
Global Context or Local Detail? Adaptive Visual Grounding for Hallucination Mitigation
by: Jiang, Yubo, et al.
Published: (2026)
by: Jiang, Yubo, et al.
Published: (2026)
Similar Items
-
Test-Time Adaptation for Combating Missing Modalities in Egocentric Videos
by: Ramazanova, Merey, et al.
Published: (2024) -
Behind the Magic, MERLIM: Multi-modal Evaluation Benchmark for Large Image-Language Models
by: Villa, Andrés, et al.
Published: (2023) -
MoDA: Modulation Adapter for Fine-Grained Visual Grounding in Instructional MLLMs
by: Barrios, Wayner, et al.
Published: (2025) -
Forget Less, Retain More: A Lightweight Regularizer for Rehearsal-Based Continual Learning
by: Alssum, Lama, et al.
Published: (2025) -
ADVMEM: Adversarial Memory Initialization for Realistic Test-Time Adaptation via Tracklet-Based Benchmarking
by: Alhuwaider, Shyma, et al.
Published: (2025)