VideoEraser: Concept Erasure in Text-to-Video Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Naen, Zhang, Jinghuai, Li, Changjiang, Chen, Zhi, Zhou, Chunyi, Li, Qingming, Du, Tianyu, Ji, Shouling
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911124371275776
author Xu, Naen
Zhang, Jinghuai
Li, Changjiang
Chen, Zhi
Zhou, Chunyi
Li, Qingming
Du, Tianyu
Ji, Shouling
author_facet Xu, Naen
Zhang, Jinghuai
Li, Changjiang
Chen, Zhi
Zhou, Chunyi
Li, Qingming
Du, Tianyu
Ji, Shouling
contents The rapid growth of text-to-video (T2V) diffusion models has raised concerns about privacy, copyright, and safety due to their potential misuse in generating harmful or misleading content. These models are often trained on numerous datasets, including unauthorized personal identities, artistic creations, and harmful materials, which can lead to uncontrolled production and distribution of such content. To address this, we propose VideoEraser, a training-free framework that prevents T2V diffusion models from generating videos with undesirable concepts, even when explicitly prompted with those concepts. Designed as a plug-and-play module, VideoEraser can seamlessly integrate with representative T2V diffusion models via a two-stage process: Selective Prompt Embedding Adjustment (SPEA) and Adversarial-Resilient Noise Guidance (ARNG). We conduct extensive evaluations across four tasks, including object erasure, artistic style erasure, celebrity erasure, and explicit content erasure. Experimental results show that VideoEraser consistently outperforms prior methods regarding efficacy, integrity, fidelity, robustness, and generalizability. Notably, VideoEraser achieves state-of-the-art performance in suppressing undesirable content during T2V generation, reducing it by 46% on average across four tasks compared to baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2508_15314
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VideoEraser: Concept Erasure in Text-to-Video Diffusion Models
Xu, Naen
Zhang, Jinghuai
Li, Changjiang
Chen, Zhi
Zhou, Chunyi
Li, Qingming
Du, Tianyu
Ji, Shouling
Computer Vision and Pattern Recognition
Artificial Intelligence
Cryptography and Security
The rapid growth of text-to-video (T2V) diffusion models has raised concerns about privacy, copyright, and safety due to their potential misuse in generating harmful or misleading content. These models are often trained on numerous datasets, including unauthorized personal identities, artistic creations, and harmful materials, which can lead to uncontrolled production and distribution of such content. To address this, we propose VideoEraser, a training-free framework that prevents T2V diffusion models from generating videos with undesirable concepts, even when explicitly prompted with those concepts. Designed as a plug-and-play module, VideoEraser can seamlessly integrate with representative T2V diffusion models via a two-stage process: Selective Prompt Embedding Adjustment (SPEA) and Adversarial-Resilient Noise Guidance (ARNG). We conduct extensive evaluations across four tasks, including object erasure, artistic style erasure, celebrity erasure, and explicit content erasure. Experimental results show that VideoEraser consistently outperforms prior methods regarding efficacy, integrity, fidelity, robustness, and generalizability. Notably, VideoEraser achieves state-of-the-art performance in suppressing undesirable content during T2V generation, reducing it by 46% on average across four tasks compared to baselines.
title VideoEraser: Concept Erasure in Text-to-Video Diffusion Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2508.15314