Judge Before Answer: Can MLLM Discern the False Premise in Question?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Jidong, Fang, Lingyong, Zhao, Haodong, Duan, Sufeng, Liu, Gongshen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917007330377728
author Li, Jidong
Fang, Lingyong
Zhao, Haodong
Duan, Sufeng
Liu, Gongshen
author_facet Li, Jidong
Fang, Lingyong
Zhao, Haodong
Duan, Sufeng
Liu, Gongshen
contents Multimodal large language models (MLLMs) have witnessed astonishing advancements in recent years. Despite these successes, MLLMs remain vulnerable to flase premise problems. However, existing benchmarks targeting this issue are limited in scope: they often lack fine-grained categorization, exhibit insufficient coverage, and thus fail to provide a rigorous evaluation of the ability of models to recognize false premises. To bridge this gap, we introduce a fully automated pipeline for constructing a comprehensive benchmark of false premise questions. Our method systematically categorizes the premises into three main types and thirteen subtypes according to the abilities required to identify the premises, resulting in the JBA dataset.Results show current MLLMs still struggle with false premise recognition. Building upon this benchmark, we further propose a recognition enhancement framework tailored to strengthen the robustness of MLLMs to detect false premises. Extensive experiments demonstrate that models trained with our framework achieve significant improvements in false premise recognition.
format Preprint
id arxiv_https___arxiv_org_abs_2510_10965
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Judge Before Answer: Can MLLM Discern the False Premise in Question?
Li, Jidong
Fang, Lingyong
Zhao, Haodong
Duan, Sufeng
Liu, Gongshen
Computation and Language
Artificial Intelligence
Multimodal large language models (MLLMs) have witnessed astonishing advancements in recent years. Despite these successes, MLLMs remain vulnerable to flase premise problems. However, existing benchmarks targeting this issue are limited in scope: they often lack fine-grained categorization, exhibit insufficient coverage, and thus fail to provide a rigorous evaluation of the ability of models to recognize false premises. To bridge this gap, we introduce a fully automated pipeline for constructing a comprehensive benchmark of false premise questions. Our method systematically categorizes the premises into three main types and thirteen subtypes according to the abilities required to identify the premises, resulting in the JBA dataset.Results show current MLLMs still struggle with false premise recognition. Building upon this benchmark, we further propose a recognition enhancement framework tailored to strengthen the robustness of MLLMs to detect false premises. Extensive experiments demonstrate that models trained with our framework achieve significant improvements in false premise recognition.
title Judge Before Answer: Can MLLM Discern the False Premise in Question?
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.10965