What's in Common? Multimodal Models Hallucinate When Reasoning Across Scenes
Fuente:
arXiv
Guardado en:
| Autores principales: | Ross, Candace, Bordes, Florian, Williams, Adina, Kirichenko, Polina, Ibrahim, Mark |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Understanding the Detrimental Class-level Effects of Data Augmentation
por: Kirichenko, Polina, et al.
Publicado: (2023)
por: Kirichenko, Polina, et al.
Publicado: (2023)
Modeling Caption Diversity in Contrastive Vision-Language Pretraining
por: Lavoie, Samuel, et al.
Publicado: (2024)
por: Lavoie, Samuel, et al.
Publicado: (2024)
The Impact of Coreset Selection on Spurious Correlations and Group Robustness
por: Dharmasiri, Amaya, et al.
Publicado: (2025)
por: Dharmasiri, Amaya, et al.
Publicado: (2025)
Decomposed evaluations of geographic disparities in text-to-image models
por: Sureddy, Abhishek, et al.
Publicado: (2024)
por: Sureddy, Abhishek, et al.
Publicado: (2024)
MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM
por: Dong, Bowen, et al.
Publicado: (2025)
por: Dong, Bowen, et al.
Publicado: (2025)
SceneAlign: Aligning Multimodal Reasoning to Scene Graphs in Complex Visual Scenes
por: Wang, Chuhan, et al.
Publicado: (2026)
por: Wang, Chuhan, et al.
Publicado: (2026)
Measuring Déjà vu Memorization Efficiently
por: Kokhlikyan, Narine, et al.
Publicado: (2025)
por: Kokhlikyan, Narine, et al.
Publicado: (2025)
Understanding the Role of Hallucination in Reinforcement Post-Training of Multimodal Reasoning Models
por: Zhang, Gengwei, et al.
Publicado: (2026)
por: Zhang, Gengwei, et al.
Publicado: (2026)
Text-to-Scene with Large Reasoning Models
por: Berdoz, Frédéric, et al.
Publicado: (2025)
por: Berdoz, Frédéric, et al.
Publicado: (2025)
Are Face Embeddings Compatible Across Deep Neural Network Models?
por: Rubab, Fizza, et al.
Publicado: (2026)
por: Rubab, Fizza, et al.
Publicado: (2026)
IndoorCrowd: A Multi-Scene Dataset for Human Detection, Segmentation, and Tracking with an Automated Annotation Pipeline
por: Nae, Sebastian-Ion, et al.
Publicado: (2026)
por: Nae, Sebastian-Ion, et al.
Publicado: (2026)
What Makes "Good" Distractors for Object Hallucination Evaluation in Large Vision-Language Models?
por: Xie, Ming-Kun, et al.
Publicado: (2025)
por: Xie, Ming-Kun, et al.
Publicado: (2025)
IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments
por: Bordes, Florian, et al.
Publicado: (2025)
por: Bordes, Florian, et al.
Publicado: (2025)
Understanding Multimodal Hallucination with Parameter-Free Representation Alignment
por: Wang, Yueqian, et al.
Publicado: (2024)
por: Wang, Yueqian, et al.
Publicado: (2024)
Hallucination-Aware Multimodal Benchmark for Gastrointestinal Image Analysis with Large Vision-Language Models
por: Khanal, Bidur, et al.
Publicado: (2025)
por: Khanal, Bidur, et al.
Publicado: (2025)
Eval Factsheets: A Structured Framework for Documenting AI Evaluations
por: Bordes, Florian, et al.
Publicado: (2025)
por: Bordes, Florian, et al.
Publicado: (2025)
An Image Dataset of Common Skin Diseases of Bangladesh and Benchmarking Performance with Machine Learning Models
por: Hossain, Sazzad, et al.
Publicado: (2026)
por: Hossain, Sazzad, et al.
Publicado: (2026)
A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs
por: Krojer, Benno, et al.
Publicado: (2025)
por: Krojer, Benno, et al.
Publicado: (2025)
LookWhen? Fast Video Recognition by Learning When, Where, and What to Compute
por: Salamatian, Ali, et al.
Publicado: (2026)
por: Salamatian, Ali, et al.
Publicado: (2026)
ConVis: Contrastive Decoding with Hallucination Visualization for Mitigating Hallucinations in Multimodal Large Language Models
por: Park, Yeji, et al.
Publicado: (2024)
por: Park, Yeji, et al.
Publicado: (2024)
HaloQuest: A Visual Hallucination Dataset for Advancing Multimodal Reasoning
por: Wang, Zhecan, et al.
Publicado: (2024)
por: Wang, Zhecan, et al.
Publicado: (2024)
Multimodal Machine Translation with Visual Scene Graph Pruning
por: Lu, Chenyu, et al.
Publicado: (2025)
por: Lu, Chenyu, et al.
Publicado: (2025)
Causal Decoding for Hallucination-Resistant Multimodal Large Language Models
por: Tan, Shiwei, et al.
Publicado: (2026)
por: Tan, Shiwei, et al.
Publicado: (2026)
Feedback-guided Data Synthesis for Imbalanced Classification
por: Hemmat, Reyhane Askari, et al.
Publicado: (2023)
por: Hemmat, Reyhane Askari, et al.
Publicado: (2023)
Multimodal Unsupervised Domain Generalization by Retrieving Across the Modality Gap
por: Liao, Christopher, et al.
Publicado: (2024)
por: Liao, Christopher, et al.
Publicado: (2024)
Woodpecker: Hallucination Correction for Multimodal Large Language Models
por: Yin, Shukang, et al.
Publicado: (2023)
por: Yin, Shukang, et al.
Publicado: (2023)
Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model
por: Liu, Benlin, et al.
Publicado: (2024)
por: Liu, Benlin, et al.
Publicado: (2024)
Mitigating Diffusion Model Hallucinations with Dynamic Guidance
por: Triaridis, Kostas, et al.
Publicado: (2025)
por: Triaridis, Kostas, et al.
Publicado: (2025)
Can Vision-Language Models See Squares? Text-Recognition Mediates Spatial Reasoning Across Three Model Families
por: Levental, Yuval
Publicado: (2026)
por: Levental, Yuval
Publicado: (2026)
Reefknot: A Comprehensive Benchmark for Relation Hallucination Evaluation, Analysis and Mitigation in Multimodal Large Language Models
por: Zheng, Kening, et al.
Publicado: (2024)
por: Zheng, Kening, et al.
Publicado: (2024)
Automatically Generating Visual Hallucination Test Cases for Multimodal Large Language Models
por: Liu, Zhongye, et al.
Publicado: (2024)
por: Liu, Zhongye, et al.
Publicado: (2024)
Detecting and Preventing Hallucinations in Large Vision Language Models
por: Gunjal, Anisha, et al.
Publicado: (2023)
por: Gunjal, Anisha, et al.
Publicado: (2023)
Insights on back marking for the automated identification of animals
por: Brunner, David, et al.
Publicado: (2026)
por: Brunner, David, et al.
Publicado: (2026)
Detect in Any Scene: An Agentic Framework for Object Detection with Experience-Aware Reasoning
por: Zhang, Wenlun, et al.
Publicado: (2026)
por: Zhang, Wenlun, et al.
Publicado: (2026)
GHOST: Hallucination-Inducing Image Generation for Multimodal LLMs
por: Parast, Aryan Yazdan, et al.
Publicado: (2025)
por: Parast, Aryan Yazdan, et al.
Publicado: (2025)
Steering the Verifiability of Multimodal AI Hallucinations
por: Pang, Jianhong, et al.
Publicado: (2026)
por: Pang, Jianhong, et al.
Publicado: (2026)
Lumos : Empowering Multimodal LLMs with Scene Text Recognition
por: Shenoy, Ashish, et al.
Publicado: (2024)
por: Shenoy, Ashish, et al.
Publicado: (2024)
ABC: Achieving Better Control of Multimodal Embeddings using VLMs
por: Schneider, Benjamin, et al.
Publicado: (2025)
por: Schneider, Benjamin, et al.
Publicado: (2025)
Online Self-Calibration Against Hallucination in Vision-Language Models
por: Chen, Minghui, et al.
Publicado: (2026)
por: Chen, Minghui, et al.
Publicado: (2026)
SaFeR-VLM: Toward Safety-aware Fine-grained Reasoning in Multimodal Models
por: Yi, Huahui, et al.
Publicado: (2025)
por: Yi, Huahui, et al.
Publicado: (2025)
Ejemplares similares
-
Understanding the Detrimental Class-level Effects of Data Augmentation
por: Kirichenko, Polina, et al.
Publicado: (2023) -
Modeling Caption Diversity in Contrastive Vision-Language Pretraining
por: Lavoie, Samuel, et al.
Publicado: (2024) -
The Impact of Coreset Selection on Spurious Correlations and Group Robustness
por: Dharmasiri, Amaya, et al.
Publicado: (2025) -
Decomposed evaluations of geographic disparities in text-to-image models
por: Sureddy, Abhishek, et al.
Publicado: (2024) -
MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM
por: Dong, Bowen, et al.
Publicado: (2025)