Immune2V: Image Immunization Against Dual-Stream Image-to-Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Long, Zeqian, Kara, Ozgur, Xue, Haotian, Chen, Yongxin, Rehg, James M.
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910123023138816
author Long, Zeqian
Kara, Ozgur
Xue, Haotian
Chen, Yongxin
Rehg, James M.
author_facet Long, Zeqian
Kara, Ozgur
Xue, Haotian
Chen, Yongxin
Rehg, James M.
contents Image-to-video (I2V) generation has the potential for societal harm because it enables the unauthorized animation of static images to create realistic deepfakes. While existing defenses effectively protect against static image manipulation, extending these to I2V generation remains underexplored and non-trivial. In this paper, we systematically analyze why modern I2V models are highly robust against naive image-level adversarial attacks (i.e., immunization). We observe that the video encoding process rapidly dilutes the adversarial noise across future frames, and the continuous text-conditioned guidance actively overrides the intended disruptive effect of the immunization. Building on these findings, we propose the Immune2V framework which enforces temporally balanced latent divergence at the encoder level to prevent signal dilution, and aligns intermediate generative representations with a precomputed collapse-inducing trajectory to counteract the text-guidance override. Extensive experiments demonstrate that Immune2V produces substantially stronger and more persistent degradation than adapted image-level baselines under the same imperceptibility budget.
format Preprint
id arxiv_https___arxiv_org_abs_2604_10837
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Immune2V: Image Immunization Against Dual-Stream Image-to-Video Generation
Long, Zeqian
Kara, Ozgur
Xue, Haotian
Chen, Yongxin
Rehg, James M.
Computer Vision and Pattern Recognition
Image-to-video (I2V) generation has the potential for societal harm because it enables the unauthorized animation of static images to create realistic deepfakes. While existing defenses effectively protect against static image manipulation, extending these to I2V generation remains underexplored and non-trivial. In this paper, we systematically analyze why modern I2V models are highly robust against naive image-level adversarial attacks (i.e., immunization). We observe that the video encoding process rapidly dilutes the adversarial noise across future frames, and the continuous text-conditioned guidance actively overrides the intended disruptive effect of the immunization. Building on these findings, we propose the Immune2V framework which enforces temporally balanced latent divergence at the encoder level to prevent signal dilution, and aligns intermediate generative representations with a precomputed collapse-inducing trajectory to counteract the text-guidance override. Extensive experiments demonstrate that Immune2V produces substantially stronger and more persistent degradation than adapted image-level baselines under the same imperceptibility budget.
title Immune2V: Image Immunization Against Dual-Stream Image-to-Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.10837