The Side Effects of Being Smart: Safety Risks in MLLMs' Multi-Image Reasoning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chen, Renmiao, Lu, Yida, Cui, Shiyao, Ouyang, Xuan, Huang, Victor Shea-Jay, Zhang, Shumin, Pan, Chengwei, Qiu, Han, Huang, Minlie
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912835283451904
author Chen, Renmiao
Lu, Yida
Cui, Shiyao
Ouyang, Xuan
Huang, Victor Shea-Jay
Zhang, Shumin
Pan, Chengwei
Qiu, Han
Huang, Minlie
author_facet Chen, Renmiao
Lu, Yida
Cui, Shiyao
Ouyang, Xuan
Huang, Victor Shea-Jay
Zhang, Shumin
Pan, Chengwei
Qiu, Han
Huang, Minlie
contents As Multimodal Large Language Models (MLLMs) acquire stronger reasoning capabilities to handle complex, multi-image instructions, this advancement may pose new safety risks. We study this problem by introducing MIR-SafetyBench, the first benchmark focused on multi-image reasoning safety, which consists of 2,676 instances across a taxonomy of 9 multi-image relations. Our extensive evaluations on 19 MLLMs reveal a troubling trend: models with more advanced multi-image reasoning can be more vulnerable on MIR-SafetyBench. Beyond attack success rates, we find that many responses labeled as safe are superficial, often driven by misunderstanding or evasive, non-committal replies. We further observe that unsafe generations exhibit lower attention entropy than safe ones on average. This internal signature suggests a possible risk that models may over-focus on task solving while neglecting safety constraints. Our code and data are available at https://github.com/thu-coai/MIR-SafetyBench.
format Preprint
id arxiv_https___arxiv_org_abs_2601_14127
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle The Side Effects of Being Smart: Safety Risks in MLLMs' Multi-Image Reasoning
Chen, Renmiao
Lu, Yida
Cui, Shiyao
Ouyang, Xuan
Huang, Victor Shea-Jay
Zhang, Shumin
Pan, Chengwei
Qiu, Han
Huang, Minlie
Computer Vision and Pattern Recognition
Computation and Language
As Multimodal Large Language Models (MLLMs) acquire stronger reasoning capabilities to handle complex, multi-image instructions, this advancement may pose new safety risks. We study this problem by introducing MIR-SafetyBench, the first benchmark focused on multi-image reasoning safety, which consists of 2,676 instances across a taxonomy of 9 multi-image relations. Our extensive evaluations on 19 MLLMs reveal a troubling trend: models with more advanced multi-image reasoning can be more vulnerable on MIR-SafetyBench. Beyond attack success rates, we find that many responses labeled as safe are superficial, often driven by misunderstanding or evasive, non-committal replies. We further observe that unsafe generations exhibit lower attention entropy than safe ones on average. This internal signature suggests a possible risk that models may over-focus on task solving while neglecting safety constraints. Our code and data are available at https://github.com/thu-coai/MIR-SafetyBench.
title The Side Effects of Being Smart: Safety Risks in MLLMs' Multi-Image Reasoning
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2601.14127