Video Unlearning via Low-Rank Refusal Vector

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Facchiano, Simone, Saravalle, Stefano, Migliarini, Matteo, De Matteis, Edoardo, Sampieri, Alessio, Pilzer, Andrea, Rodolà, Emanuele, Spinelli, Indro, Franco, Luca, Galasso, Fabio
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911409903763456
author Facchiano, Simone
Saravalle, Stefano
Migliarini, Matteo
De Matteis, Edoardo
Sampieri, Alessio
Pilzer, Andrea
Rodolà, Emanuele
Spinelli, Indro
Franco, Luca
Galasso, Fabio
author_facet Facchiano, Simone
Saravalle, Stefano
Migliarini, Matteo
De Matteis, Edoardo
Sampieri, Alessio
Pilzer, Andrea
Rodolà, Emanuele
Spinelli, Indro
Franco, Luca
Galasso, Fabio
contents Video generative models achieve high-quality synthesis from natural-language prompts by leveraging large-scale web data. However, this training paradigm inherently exposes them to unsafe biases and harmful concepts, introducing the risk of generating undesirable or illicit content. To mitigate unsafe generations, existing machine unlearning approaches either rely on filtering, and can therefore be bypassed, or they update model weights, but with costly fine-tuning or training-free closed-form edits. We propose the first training-free weight update framework for concept removal in video diffusion models. From five paired safe/unsafe prompts, our method estimates a refusal vector and integrates it into the model weights as a closed-form update. A contrastive low-rank factorization further disentangles the target concept from unrelated semantics, it ensures a selective concept suppression and it does not harm generation quality. Our approach reduces unsafe generations on the Open-Sora and ZeroScopeT2V models across the T2VSafetyBench and SafeSora benchmarks, with average reductions of 36.3% and 58.2% respectively, while preserving prompt alignment and video quality. This establishes an efficient and scalable solution for safe video generation without retraining nor any inference overhead. Project page: https://www.pinlab.org/video-unlearning.
format Preprint
id arxiv_https___arxiv_org_abs_2506_07891
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Video Unlearning via Low-Rank Refusal Vector
Facchiano, Simone
Saravalle, Stefano
Migliarini, Matteo
De Matteis, Edoardo
Sampieri, Alessio
Pilzer, Andrea
Rodolà, Emanuele
Spinelli, Indro
Franco, Luca
Galasso, Fabio
Computer Vision and Pattern Recognition
Video generative models achieve high-quality synthesis from natural-language prompts by leveraging large-scale web data. However, this training paradigm inherently exposes them to unsafe biases and harmful concepts, introducing the risk of generating undesirable or illicit content. To mitigate unsafe generations, existing machine unlearning approaches either rely on filtering, and can therefore be bypassed, or they update model weights, but with costly fine-tuning or training-free closed-form edits. We propose the first training-free weight update framework for concept removal in video diffusion models. From five paired safe/unsafe prompts, our method estimates a refusal vector and integrates it into the model weights as a closed-form update. A contrastive low-rank factorization further disentangles the target concept from unrelated semantics, it ensures a selective concept suppression and it does not harm generation quality. Our approach reduces unsafe generations on the Open-Sora and ZeroScopeT2V models across the T2VSafetyBench and SafeSora benchmarks, with average reductions of 36.3% and 58.2% respectively, while preserving prompt alignment and video quality. This establishes an efficient and scalable solution for safe video generation without retraining nor any inference overhead. Project page: https://www.pinlab.org/video-unlearning.
title Video Unlearning via Low-Rank Refusal Vector
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.07891