Jailbreaking Multimodal Large Language Models using Multi-Clip Video

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kang, Choongwon, Sun, Seungjong, Jun, Hyunmin, Kim, Jang Hyun
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918535099318272
author Kang, Choongwon
Sun, Seungjong
Jun, Hyunmin
Kim, Jang Hyun
author_facet Kang, Choongwon
Sun, Seungjong
Jun, Hyunmin
Kim, Jang Hyun
contents As multimodal large language models (MLLMs) have advanced to process video inputs, concerns have emerged about their potential for malicious misuse. Prior jailbreak studies have shown that safety alignment in MLLMs can be bypassed through visual inputs, yet it remains unclear which properties of video inputs induce this vulnerability. To address this gap, we introduce Multi-Clip Video (MCV) SafetyBench, a dataset of 2,920 videos designed to evaluate how the diversity of video inputs affects the vulnerability of MLLMs. Each video consists of multiple short clips depicting diverse contexts related to a harmful query. Experiments on eight representative video MLLMs show that attack success consistently increases with the number of clips. Our results further indicate that the video modality is (1) more vulnerable than the image modality, (2) more vulnerable to dynamic videos than to static videos, and (3) more vulnerable when videos contain more diverse contexts. Building on these findings, we propose a defense strategy that leverages the relative robustness of the image modality.
format Preprint
id arxiv_https___arxiv_org_abs_2606_02111
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Jailbreaking Multimodal Large Language Models using Multi-Clip Video
Kang, Choongwon
Sun, Seungjong
Jun, Hyunmin
Kim, Jang Hyun
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
As multimodal large language models (MLLMs) have advanced to process video inputs, concerns have emerged about their potential for malicious misuse. Prior jailbreak studies have shown that safety alignment in MLLMs can be bypassed through visual inputs, yet it remains unclear which properties of video inputs induce this vulnerability. To address this gap, we introduce Multi-Clip Video (MCV) SafetyBench, a dataset of 2,920 videos designed to evaluate how the diversity of video inputs affects the vulnerability of MLLMs. Each video consists of multiple short clips depicting diverse contexts related to a harmful query. Experiments on eight representative video MLLMs show that attack success consistently increases with the number of clips. Our results further indicate that the video modality is (1) more vulnerable than the image modality, (2) more vulnerable to dynamic videos than to static videos, and (3) more vulnerable when videos contain more diverse contexts. Building on these findings, we propose a defense strategy that leverages the relative robustness of the image modality.
title Jailbreaking Multimodal Large Language Models using Multi-Clip Video
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2606.02111