TEAR: Temporal-aware Automated Red-teaming for Text-to-Video Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: He, Jiaming, Hou, Guanyu, Li, Hongwei, Huang, Zhicong, Chen, Kangjie, Yu, Yi, Jiang, Wenbo, Xu, Guowen, Zhang, Tianwei
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912960966819840
author He, Jiaming
Hou, Guanyu
Li, Hongwei
Huang, Zhicong
Chen, Kangjie
Yu, Yi
Jiang, Wenbo
Xu, Guowen
Zhang, Tianwei
author_facet He, Jiaming
Hou, Guanyu
Li, Hongwei
Huang, Zhicong
Chen, Kangjie
Yu, Yi
Jiang, Wenbo
Xu, Guowen
Zhang, Tianwei
contents Text-to-Video (T2V) models are capable of synthesizing high-quality, temporally coherent dynamic video content, but the diverse generation also inherently introduces critical safety challenges. Existing safety evaluation methods,which focus on static image and text generation, are insufficient to capture the complex temporal dynamics in video generation. To address this, we propose a TEmporal-aware Automated Red-teaming framework, named TEAR, an automated framework designed to uncover safety risks specifically linked to the dynamic temporal sequencing of T2V models. TEAR employs a temporal-aware test generator optimized via a two-stage approach: initial generator training and temporal-aware online preference learning, to craft textually innocuous prompts that exploit temporal dynamics to elicit policy-violating video output. And a refine model is adopted to improve the prompt stealthiness and adversarial effectiveness cyclically. Extensive experimental evaluation demonstrates the effectiveness of TEAR across open-source and commercial T2V systems with over 80% attack success rate, a significant boost from prior best result of 57%.
format Preprint
id arxiv_https___arxiv_org_abs_2511_21145
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TEAR: Temporal-aware Automated Red-teaming for Text-to-Video Models
He, Jiaming
Hou, Guanyu
Li, Hongwei
Huang, Zhicong
Chen, Kangjie
Yu, Yi
Jiang, Wenbo
Xu, Guowen
Zhang, Tianwei
Computer Vision and Pattern Recognition
Text-to-Video (T2V) models are capable of synthesizing high-quality, temporally coherent dynamic video content, but the diverse generation also inherently introduces critical safety challenges. Existing safety evaluation methods,which focus on static image and text generation, are insufficient to capture the complex temporal dynamics in video generation. To address this, we propose a TEmporal-aware Automated Red-teaming framework, named TEAR, an automated framework designed to uncover safety risks specifically linked to the dynamic temporal sequencing of T2V models. TEAR employs a temporal-aware test generator optimized via a two-stage approach: initial generator training and temporal-aware online preference learning, to craft textually innocuous prompts that exploit temporal dynamics to elicit policy-violating video output. And a refine model is adopted to improve the prompt stealthiness and adversarial effectiveness cyclically. Extensive experimental evaluation demonstrates the effectiveness of TEAR across open-source and commercial T2V systems with over 80% attack success rate, a significant boost from prior best result of 57%.
title TEAR: Temporal-aware Automated Red-teaming for Text-to-Video Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.21145