Generative Video Propagation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Shaoteng, Wang, Tianyu, Wang, Jui-Hsien, Liu, Qing, Zhang, Zhifei, Lee, Joon-Young, Li, Yijun, Yu, Bei, Lin, Zhe, Kim, Soo Ye, Jia, Jiaya
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912170536599552
author Liu, Shaoteng
Wang, Tianyu
Wang, Jui-Hsien
Liu, Qing
Zhang, Zhifei
Lee, Joon-Young
Li, Yijun
Yu, Bei
Lin, Zhe
Kim, Soo Ye
Jia, Jiaya
author_facet Liu, Shaoteng
Wang, Tianyu
Wang, Jui-Hsien
Liu, Qing
Zhang, Zhifei
Lee, Joon-Young
Li, Yijun
Yu, Bei
Lin, Zhe
Kim, Soo Ye
Jia, Jiaya
contents Large-scale video generation models have the inherent ability to realistically model natural scenes. In this paper, we demonstrate that through a careful design of a generative video propagation framework, various video tasks can be addressed in a unified way by leveraging the generative power of such models. Specifically, our framework, GenProp, encodes the original video with a selective content encoder and propagates the changes made to the first frame using an image-to-video generation model. We propose a data generation scheme to cover multiple video tasks based on instance-level video segmentation datasets. Our model is trained by incorporating a mask prediction decoder head and optimizing a region-aware loss to aid the encoder to preserve the original content while the generation model propagates the modified region. This novel design opens up new possibilities: In editing scenarios, GenProp allows substantial changes to an object's shape; for insertion, the inserted objects can exhibit independent motion; for removal, GenProp effectively removes effects like shadows and reflections from the whole video; for tracking, GenProp is capable of tracking objects and their associated effects together. Experiment results demonstrate the leading performance of our model in various video tasks, and we further provide in-depth analyses of the proposed framework.
format Preprint
id arxiv_https___arxiv_org_abs_2412_19761
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Generative Video Propagation
Liu, Shaoteng
Wang, Tianyu
Wang, Jui-Hsien
Liu, Qing
Zhang, Zhifei
Lee, Joon-Young
Li, Yijun
Yu, Bei
Lin, Zhe
Kim, Soo Ye
Jia, Jiaya
Computer Vision and Pattern Recognition
Large-scale video generation models have the inherent ability to realistically model natural scenes. In this paper, we demonstrate that through a careful design of a generative video propagation framework, various video tasks can be addressed in a unified way by leveraging the generative power of such models. Specifically, our framework, GenProp, encodes the original video with a selective content encoder and propagates the changes made to the first frame using an image-to-video generation model. We propose a data generation scheme to cover multiple video tasks based on instance-level video segmentation datasets. Our model is trained by incorporating a mask prediction decoder head and optimizing a region-aware loss to aid the encoder to preserve the original content while the generation model propagates the modified region. This novel design opens up new possibilities: In editing scenarios, GenProp allows substantial changes to an object's shape; for insertion, the inserted objects can exhibit independent motion; for removal, GenProp effectively removes effects like shadows and reflections from the whole video; for tracking, GenProp is capable of tracking objects and their associated effects together. Experiment results demonstrate the leading performance of our model in various video tasks, and we further provide in-depth analyses of the proposed framework.
title Generative Video Propagation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.19761