FINER: MLLMs Hallucinate under Fine-grained Negative Queries

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xiao, Rui, Kim, Sanghwan, Xian, Yongqin, Akata, Zeynep, Alaniz, Stephan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912972689899520
author Xiao, Rui
Kim, Sanghwan
Xian, Yongqin
Akata, Zeynep
Alaniz, Stephan
author_facet Xiao, Rui
Kim, Sanghwan
Xian, Yongqin
Akata, Zeynep
Alaniz, Stephan
contents Multimodal large language models (MLLMs) struggle with hallucinations, particularly with fine-grained queries, a challenge underrepresented by existing benchmarks that focus on coarse image-related questions. We introduce FIne-grained NEgative queRies (FINER), alongside two benchmarks: FINER-CompreCap and FINER-DOCCI. Using FINER, we analyze hallucinations across four settings: multi-object, multi-attribute, multi-relation, and ``what'' questions. Our benchmarks reveal that MLLMs hallucinate when fine-grained mismatches co-occur with genuinely present elements in the image. To address this, we propose FINER-Tuning, leveraging Direct Preference Optimization (DPO) on FINER-inspired data. Finetuning four frontier MLLMs with FINER-Tuning yields up to 24.2\% gains (InternVL3.5-14B) on hallucinations from our benchmarks, while simultaneously improving performance on eight existing hallucination suites, and enhancing general multimodal capabilities across six benchmarks. Code, benchmark, and models are available at \href{https://explainableml.github.io/finer-project/}{https://explainableml.github.io/finer-project/}.
format Preprint
id arxiv_https___arxiv_org_abs_2603_17662
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle FINER: MLLMs Hallucinate under Fine-grained Negative Queries
Xiao, Rui
Kim, Sanghwan
Xian, Yongqin
Akata, Zeynep
Alaniz, Stephan
Computer Vision and Pattern Recognition
Artificial Intelligence
Multimodal large language models (MLLMs) struggle with hallucinations, particularly with fine-grained queries, a challenge underrepresented by existing benchmarks that focus on coarse image-related questions. We introduce FIne-grained NEgative queRies (FINER), alongside two benchmarks: FINER-CompreCap and FINER-DOCCI. Using FINER, we analyze hallucinations across four settings: multi-object, multi-attribute, multi-relation, and ``what'' questions. Our benchmarks reveal that MLLMs hallucinate when fine-grained mismatches co-occur with genuinely present elements in the image. To address this, we propose FINER-Tuning, leveraging Direct Preference Optimization (DPO) on FINER-inspired data. Finetuning four frontier MLLMs with FINER-Tuning yields up to 24.2\% gains (InternVL3.5-14B) on hallucinations from our benchmarks, while simultaneously improving performance on eight existing hallucination suites, and enhancing general multimodal capabilities across six benchmarks. Code, benchmark, and models are available at \href{https://explainableml.github.io/finer-project/}{https://explainableml.github.io/finer-project/}.
title FINER: MLLMs Hallucinate under Fine-grained Negative Queries
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2603.17662