Edit as You See: Image-guided Video Editing via Masked Motion Modeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Zhi-Lin, Liu, Yixuan, Qin, Chujun, Wang, Zhongdao, Zhou, Dong, Li, Dong, Barsoum, Emad
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917887118147584
author Huang, Zhi-Lin
Liu, Yixuan
Qin, Chujun
Wang, Zhongdao
Zhou, Dong
Li, Dong
Barsoum, Emad
author_facet Huang, Zhi-Lin
Liu, Yixuan
Qin, Chujun
Wang, Zhongdao
Zhou, Dong
Li, Dong
Barsoum, Emad
contents Recent advancements in diffusion models have significantly facilitated text-guided video editing. However, there is a relative scarcity of research on image-guided video editing, a method that empowers users to edit videos by merely indicating a target object in the initial frame and providing an RGB image as reference, without relying on the text prompts. In this paper, we propose a novel Image-guided Video Editing Diffusion model, termed IVEDiff for the image-guided video editing. IVEDiff is built on top of image editing models, and is equipped with learnable motion modules to maintain the temporal consistency of edited video. Inspired by self-supervised learning concepts, we introduce a masked motion modeling fine-tuning strategy that empowers the motion module's capabilities for capturing inter-frame motion dynamics, while preserving the capabilities for intra-frame semantic correlations modeling of the base image editing model. Moreover, an optical-flow-guided motion reference network is proposed to ensure the accurate propagation of information between edited video frames, alleviating the misleading effects of invalid information. We also construct a benchmark to facilitate further research. The comprehensive experiments demonstrate that our method is able to generate temporally smooth edited videos while robustly dealing with various editing objects with high quality.
format Preprint
id arxiv_https___arxiv_org_abs_2501_04325
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Edit as You See: Image-guided Video Editing via Masked Motion Modeling
Huang, Zhi-Lin
Liu, Yixuan
Qin, Chujun
Wang, Zhongdao
Zhou, Dong
Li, Dong
Barsoum, Emad
Computer Vision and Pattern Recognition
Recent advancements in diffusion models have significantly facilitated text-guided video editing. However, there is a relative scarcity of research on image-guided video editing, a method that empowers users to edit videos by merely indicating a target object in the initial frame and providing an RGB image as reference, without relying on the text prompts. In this paper, we propose a novel Image-guided Video Editing Diffusion model, termed IVEDiff for the image-guided video editing. IVEDiff is built on top of image editing models, and is equipped with learnable motion modules to maintain the temporal consistency of edited video. Inspired by self-supervised learning concepts, we introduce a masked motion modeling fine-tuning strategy that empowers the motion module's capabilities for capturing inter-frame motion dynamics, while preserving the capabilities for intra-frame semantic correlations modeling of the base image editing model. Moreover, an optical-flow-guided motion reference network is proposed to ensure the accurate propagation of information between edited video frames, alleviating the misleading effects of invalid information. We also construct a benchmark to facilitate further research. The comprehensive experiments demonstrate that our method is able to generate temporally smooth edited videos while robustly dealing with various editing objects with high quality.
title Edit as You See: Image-guided Video Editing via Masked Motion Modeling
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.04325