SPARK: Jailbreaking T2V Models by Synergistically Prompting Auditory and Recontextualized Knowledge

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ying, Zonghao, Chen, Moyang, Li, Nizhang, Wang, Zhiqiang, Zhang, Wenxin, Zou, Quanchen, Jing, Zonglei, Liu, Aishan, Liu, Xianglong
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915838923112448
author Ying, Zonghao
Chen, Moyang
Li, Nizhang
Wang, Zhiqiang
Zhang, Wenxin
Zou, Quanchen
Jing, Zonglei
Liu, Aishan
Liu, Xianglong
author_facet Ying, Zonghao
Chen, Moyang
Li, Nizhang
Wang, Zhiqiang
Zhang, Wenxin
Zou, Quanchen
Jing, Zonglei
Liu, Aishan
Liu, Xianglong
contents Jailbreak attacks can circumvent model safety guardrails and reveal critical blind spots. Prior attacks on text-to-video (T2V) models typically add adversarial perturbations to obviously unsafe prompts, which are often easy to detect and defend. In contrast, we show that benign-looking prompts containing rich, implicit cues can induce T2V models to generate semantically unsafe videos that both violate policy and preserve the original (blocked) intent. To realize this, we propose SPARK, a jailbreak framework that leverages T2V models cross-modal associative patterns via a modular prompt design. Specifically, our prompts combine three components: neutral scene anchors, which provide the surface-level scene description extracted from the blocked intent to maintain plausibility; latent auditory triggers, textual descriptions of innocuous-sounding audio events (e.g., creaking, muffled noises) that exploit learned audio-visual co-occurrence priors to bias the model toward particular unsafe visual concepts; and stylistic modulators, cinematic directives (e.g., camera framing, atmosphere) that amplify and stabilize the latent trigger's effect. We formalize attack generation as a constrained optimization over the above modular prompt space and solve it with a guided search procedure that balances stealth and effectiveness. Extensive experiments over 7 T2V models demonstrate the efficacy of our attack, achieving a +23% improvement in average attack success rate in commercial models.
format Preprint
id arxiv_https___arxiv_org_abs_2511_13127
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SPARK: Jailbreaking T2V Models by Synergistically Prompting Auditory and Recontextualized Knowledge
Ying, Zonghao
Chen, Moyang
Li, Nizhang
Wang, Zhiqiang
Zhang, Wenxin
Zou, Quanchen
Jing, Zonglei
Liu, Aishan
Liu, Xianglong
Computer Vision and Pattern Recognition
Cryptography and Security
Jailbreak attacks can circumvent model safety guardrails and reveal critical blind spots. Prior attacks on text-to-video (T2V) models typically add adversarial perturbations to obviously unsafe prompts, which are often easy to detect and defend. In contrast, we show that benign-looking prompts containing rich, implicit cues can induce T2V models to generate semantically unsafe videos that both violate policy and preserve the original (blocked) intent. To realize this, we propose SPARK, a jailbreak framework that leverages T2V models cross-modal associative patterns via a modular prompt design. Specifically, our prompts combine three components: neutral scene anchors, which provide the surface-level scene description extracted from the blocked intent to maintain plausibility; latent auditory triggers, textual descriptions of innocuous-sounding audio events (e.g., creaking, muffled noises) that exploit learned audio-visual co-occurrence priors to bias the model toward particular unsafe visual concepts; and stylistic modulators, cinematic directives (e.g., camera framing, atmosphere) that amplify and stabilize the latent trigger's effect. We formalize attack generation as a constrained optimization over the above modular prompt space and solve it with a guided search procedure that balances stealth and effectiveness. Extensive experiments over 7 T2V models demonstrate the efficacy of our attack, achieving a +23% improvement in average attack success rate in commercial models.
title SPARK: Jailbreaking T2V Models by Synergistically Prompting Auditory and Recontextualized Knowledge
topic Computer Vision and Pattern Recognition
Cryptography and Security
url https://arxiv.org/abs/2511.13127