Generative Video Propagation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912170536599552 |
|---|---|
| author | Liu, Shaoteng Wang, Tianyu Wang, Jui-Hsien Liu, Qing Zhang, Zhifei Lee, Joon-Young Li, Yijun Yu, Bei Lin, Zhe Kim, Soo Ye Jia, Jiaya |
| author_facet | Liu, Shaoteng Wang, Tianyu Wang, Jui-Hsien Liu, Qing Zhang, Zhifei Lee, Joon-Young Li, Yijun Yu, Bei Lin, Zhe Kim, Soo Ye Jia, Jiaya |
| contents | Large-scale video generation models have the inherent ability to realistically model natural scenes. In this paper, we demonstrate that through a careful design of a generative video propagation framework, various video tasks can be addressed in a unified way by leveraging the generative power of such models. Specifically, our framework, GenProp, encodes the original video with a selective content encoder and propagates the changes made to the first frame using an image-to-video generation model. We propose a data generation scheme to cover multiple video tasks based on instance-level video segmentation datasets. Our model is trained by incorporating a mask prediction decoder head and optimizing a region-aware loss to aid the encoder to preserve the original content while the generation model propagates the modified region. This novel design opens up new possibilities: In editing scenarios, GenProp allows substantial changes to an object's shape; for insertion, the inserted objects can exhibit independent motion; for removal, GenProp effectively removes effects like shadows and reflections from the whole video; for tracking, GenProp is capable of tracking objects and their associated effects together. Experiment results demonstrate the leading performance of our model in various video tasks, and we further provide in-depth analyses of the proposed framework. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_19761 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Generative Video Propagation Liu, Shaoteng Wang, Tianyu Wang, Jui-Hsien Liu, Qing Zhang, Zhifei Lee, Joon-Young Li, Yijun Yu, Bei Lin, Zhe Kim, Soo Ye Jia, Jiaya Computer Vision and Pattern Recognition Large-scale video generation models have the inherent ability to realistically model natural scenes. In this paper, we demonstrate that through a careful design of a generative video propagation framework, various video tasks can be addressed in a unified way by leveraging the generative power of such models. Specifically, our framework, GenProp, encodes the original video with a selective content encoder and propagates the changes made to the first frame using an image-to-video generation model. We propose a data generation scheme to cover multiple video tasks based on instance-level video segmentation datasets. Our model is trained by incorporating a mask prediction decoder head and optimizing a region-aware loss to aid the encoder to preserve the original content while the generation model propagates the modified region. This novel design opens up new possibilities: In editing scenarios, GenProp allows substantial changes to an object's shape; for insertion, the inserted objects can exhibit independent motion; for removal, GenProp effectively removes effects like shadows and reflections from the whole video; for tracking, GenProp is capable of tracking objects and their associated effects together. Experiment results demonstrate the leading performance of our model in various video tasks, and we further provide in-depth analyses of the proposed framework. |
| title | Generative Video Propagation |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2412.19761 |