Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Miyai, Atsuyuki, Yang, Jingkang, Zhang, Jingyang, Ming, Yifei, Yu, Qing, Irie, Go, Li, Yixuan, Li, Hai, Liu, Ziwei, Aizawa, Kiyoharu
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908397984546816
author Miyai, Atsuyuki
Yang, Jingkang
Zhang, Jingyang
Ming, Yifei
Yu, Qing
Irie, Go
Li, Yixuan
Li, Hai
Liu, Ziwei
Aizawa, Kiyoharu
author_facet Miyai, Atsuyuki
Yang, Jingkang
Zhang, Jingyang
Ming, Yifei
Yu, Qing
Irie, Go
Li, Yixuan
Li, Hai
Liu, Ziwei
Aizawa, Kiyoharu
contents This paper introduces a novel task to evaluate the robust understanding capability of Large Multimodal Models (LMMs), termed $\textbf{Unsolvable Problem Detection (UPD)}$. Multiple-choice question answering (MCQA) is widely used to assess the understanding capability of LMMs, but it does not guarantee that LMMs truly comprehend the answer. UPD assesses the LMM's ability to withhold answers when encountering unsolvable problems of MCQA, verifying whether the model truly understands the answer. UPD encompasses three problems: Absent Answer Detection (AAD), Incompatible Answer Set Detection (IASD), and Incompatible Visual Question Detection (IVQD), covering unsolvable cases like answer-lacking or incompatible choices and image-question mismatches. For the evaluation, we introduce the MM-UPD Bench, a benchmark for assessing performance across various ability dimensions. Our experiments reveal that even most LMMs, which demonstrate adequate performance on existing benchmarks, struggle significantly with MM-UPD, underscoring a novel aspect of trustworthiness that current benchmarks have overlooked. A detailed analysis shows that LMMs have different bottlenecks and chain-of-thought and self-reflection improved performance for LMMs with the bottleneck in their LLM capability. We hope our insights will enhance the broader understanding and development of more reliable LMMs. The code is available at https://github.com/AtsuMiyai/UPD.
format Preprint
id arxiv_https___arxiv_org_abs_2403_20331
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal Models
Miyai, Atsuyuki
Yang, Jingkang
Zhang, Jingyang
Ming, Yifei
Yu, Qing
Irie, Go
Li, Yixuan
Li, Hai
Liu, Ziwei
Aizawa, Kiyoharu
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
This paper introduces a novel task to evaluate the robust understanding capability of Large Multimodal Models (LMMs), termed $\textbf{Unsolvable Problem Detection (UPD)}$. Multiple-choice question answering (MCQA) is widely used to assess the understanding capability of LMMs, but it does not guarantee that LMMs truly comprehend the answer. UPD assesses the LMM's ability to withhold answers when encountering unsolvable problems of MCQA, verifying whether the model truly understands the answer. UPD encompasses three problems: Absent Answer Detection (AAD), Incompatible Answer Set Detection (IASD), and Incompatible Visual Question Detection (IVQD), covering unsolvable cases like answer-lacking or incompatible choices and image-question mismatches. For the evaluation, we introduce the MM-UPD Bench, a benchmark for assessing performance across various ability dimensions. Our experiments reveal that even most LMMs, which demonstrate adequate performance on existing benchmarks, struggle significantly with MM-UPD, underscoring a novel aspect of trustworthiness that current benchmarks have overlooked. A detailed analysis shows that LMMs have different bottlenecks and chain-of-thought and self-reflection improved performance for LMMs with the bottleneck in their LLM capability. We hope our insights will enhance the broader understanding and development of more reliable LMMs. The code is available at https://github.com/AtsuMiyai/UPD.
title Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2403.20331