Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866908397984546816 |
|---|---|
| author | Miyai, Atsuyuki Yang, Jingkang Zhang, Jingyang Ming, Yifei Yu, Qing Irie, Go Li, Yixuan Li, Hai Liu, Ziwei Aizawa, Kiyoharu |
| author_facet | Miyai, Atsuyuki Yang, Jingkang Zhang, Jingyang Ming, Yifei Yu, Qing Irie, Go Li, Yixuan Li, Hai Liu, Ziwei Aizawa, Kiyoharu |
| contents | This paper introduces a novel task to evaluate the robust understanding capability of Large Multimodal Models (LMMs), termed $\textbf{Unsolvable Problem Detection (UPD)}$. Multiple-choice question answering (MCQA) is widely used to assess the understanding capability of LMMs, but it does not guarantee that LMMs truly comprehend the answer. UPD assesses the LMM's ability to withhold answers when encountering unsolvable problems of MCQA, verifying whether the model truly understands the answer. UPD encompasses three problems: Absent Answer Detection (AAD), Incompatible Answer Set Detection (IASD), and Incompatible Visual Question Detection (IVQD), covering unsolvable cases like answer-lacking or incompatible choices and image-question mismatches. For the evaluation, we introduce the MM-UPD Bench, a benchmark for assessing performance across various ability dimensions. Our experiments reveal that even most LMMs, which demonstrate adequate performance on existing benchmarks, struggle significantly with MM-UPD, underscoring a novel aspect of trustworthiness that current benchmarks have overlooked. A detailed analysis shows that LMMs have different bottlenecks and chain-of-thought and self-reflection improved performance for LMMs with the bottleneck in their LLM capability. We hope our insights will enhance the broader understanding and development of more reliable LMMs. The code is available at https://github.com/AtsuMiyai/UPD. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2403_20331 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal Models Miyai, Atsuyuki Yang, Jingkang Zhang, Jingyang Ming, Yifei Yu, Qing Irie, Go Li, Yixuan Li, Hai Liu, Ziwei Aizawa, Kiyoharu Computer Vision and Pattern Recognition Artificial Intelligence Computation and Language Machine Learning This paper introduces a novel task to evaluate the robust understanding capability of Large Multimodal Models (LMMs), termed $\textbf{Unsolvable Problem Detection (UPD)}$. Multiple-choice question answering (MCQA) is widely used to assess the understanding capability of LMMs, but it does not guarantee that LMMs truly comprehend the answer. UPD assesses the LMM's ability to withhold answers when encountering unsolvable problems of MCQA, verifying whether the model truly understands the answer. UPD encompasses three problems: Absent Answer Detection (AAD), Incompatible Answer Set Detection (IASD), and Incompatible Visual Question Detection (IVQD), covering unsolvable cases like answer-lacking or incompatible choices and image-question mismatches. For the evaluation, we introduce the MM-UPD Bench, a benchmark for assessing performance across various ability dimensions. Our experiments reveal that even most LMMs, which demonstrate adequate performance on existing benchmarks, struggle significantly with MM-UPD, underscoring a novel aspect of trustworthiness that current benchmarks have overlooked. A detailed analysis shows that LMMs have different bottlenecks and chain-of-thought and self-reflection improved performance for LMMs with the bottleneck in their LLM capability. We hope our insights will enhance the broader understanding and development of more reliable LMMs. The code is available at https://github.com/AtsuMiyai/UPD. |
| title | Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal Models |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2403.20331 |