MoHoBench: Assessing Honesty of Multimodal Large Language Models via Unanswerable Visual Questions

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhu, Yanxu, Duan, Shitong, Zhang, Xiangxu, Sang, Jitao, Zhang, Peng, Lu, Tun, Zhou, Xiao, Yao, Jing, Yi, Xiaoyuan, Xie, Xing
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917197599735808
author Zhu, Yanxu
Duan, Shitong
Zhang, Xiangxu
Sang, Jitao
Zhang, Peng
Lu, Tun
Zhou, Xiao
Yao, Jing
Yi, Xiaoyuan
Xie, Xing
author_facet Zhu, Yanxu
Duan, Shitong
Zhang, Xiangxu
Sang, Jitao
Zhang, Peng
Lu, Tun
Zhou, Xiao
Yao, Jing
Yi, Xiaoyuan
Xie, Xing
contents Recently Multimodal Large Language Models (MLLMs) have achieved considerable advancements in vision-language tasks, yet produce potentially harmful or untrustworthy content. Despite substantial work investigating the trustworthiness of language models, MMLMs' capability to act honestly, especially when faced with visually unanswerable questions, remains largely underexplored. This work presents the first systematic assessment of honesty behaviors across various MLLMs. We ground honesty in models' response behaviors to unanswerable visual questions, define four representative types of such questions, and construct MoHoBench, a large-scale MMLM honest benchmark, consisting of 12k+ visual question samples, whose quality is guaranteed by multi-stage filtering and human verification. Using MoHoBench, we benchmarked the honesty of 28 popular MMLMs and conducted a comprehensive analysis. Our findings show that: (1) most models fail to appropriately refuse to answer when necessary, and (2) MMLMs' honesty is not solely a language modeling issue, but is deeply influenced by visual information, necessitating the development of dedicated methods for multimodal honesty alignment. Therefore, we implemented initial alignment methods using supervised and preference learning to improve honesty behavior, providing a foundation for future work on trustworthy MLLMs. Our data and code can be found at https://github.com/yanxuzhu/MoHoBench.
format Preprint
id arxiv_https___arxiv_org_abs_2507_21503
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MoHoBench: Assessing Honesty of Multimodal Large Language Models via Unanswerable Visual Questions
Zhu, Yanxu
Duan, Shitong
Zhang, Xiangxu
Sang, Jitao
Zhang, Peng
Lu, Tun
Zhou, Xiao
Yao, Jing
Yi, Xiaoyuan
Xie, Xing
Artificial Intelligence
Computer Vision and Pattern Recognition
Recently Multimodal Large Language Models (MLLMs) have achieved considerable advancements in vision-language tasks, yet produce potentially harmful or untrustworthy content. Despite substantial work investigating the trustworthiness of language models, MMLMs' capability to act honestly, especially when faced with visually unanswerable questions, remains largely underexplored. This work presents the first systematic assessment of honesty behaviors across various MLLMs. We ground honesty in models' response behaviors to unanswerable visual questions, define four representative types of such questions, and construct MoHoBench, a large-scale MMLM honest benchmark, consisting of 12k+ visual question samples, whose quality is guaranteed by multi-stage filtering and human verification. Using MoHoBench, we benchmarked the honesty of 28 popular MMLMs and conducted a comprehensive analysis. Our findings show that: (1) most models fail to appropriately refuse to answer when necessary, and (2) MMLMs' honesty is not solely a language modeling issue, but is deeply influenced by visual information, necessitating the development of dedicated methods for multimodal honesty alignment. Therefore, we implemented initial alignment methods using supervised and preference learning to improve honesty behavior, providing a foundation for future work on trustworthy MLLMs. Our data and code can be found at https://github.com/yanxuzhu/MoHoBench.
title MoHoBench: Assessing Honesty of Multimodal Large Language Models via Unanswerable Visual Questions
topic Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.21503