CoCoCo: Improving Text-Guided Video Inpainting for Better Consistency, Controllability and Compatibility

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zi, Bojia, Zhao, Shihao, Qi, Xianbiao, Wang, Jianan, Shi, Yukai, Chen, Qianyu, Liang, Bin, Wong, Kam-Fai, Zhang, Lei
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917617400283136
author Zi, Bojia
Zhao, Shihao
Qi, Xianbiao
Wang, Jianan
Shi, Yukai
Chen, Qianyu
Liang, Bin
Wong, Kam-Fai
Zhang, Lei
author_facet Zi, Bojia
Zhao, Shihao
Qi, Xianbiao
Wang, Jianan
Shi, Yukai
Chen, Qianyu
Liang, Bin
Wong, Kam-Fai
Zhang, Lei
contents Recent advancements in video generation have been remarkable, yet many existing methods struggle with issues of consistency and poor text-video alignment. Moreover, the field lacks effective techniques for text-guided video inpainting, a stark contrast to the well-explored domain of text-guided image inpainting. To this end, this paper proposes a novel text-guided video inpainting model that achieves better consistency, controllability and compatibility. Specifically, we introduce a simple but efficient motion capture module to preserve motion consistency, and design an instance-aware region selection instead of a random region selection to obtain better textual controllability, and utilize a novel strategy to inject some personalized models into our CoCoCo model and thus obtain better model compatibility. Extensive experiments show that our model can generate high-quality video clips. Meanwhile, our model shows better motion consistency, textual controllability and model compatibility. More details are shown in [cococozibojia.github.io](cococozibojia.github.io).
format Preprint
id arxiv_https___arxiv_org_abs_2403_12035
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CoCoCo: Improving Text-Guided Video Inpainting for Better Consistency, Controllability and Compatibility
Zi, Bojia
Zhao, Shihao
Qi, Xianbiao
Wang, Jianan
Shi, Yukai
Chen, Qianyu
Liang, Bin
Wong, Kam-Fai
Zhang, Lei
Computer Vision and Pattern Recognition
Recent advancements in video generation have been remarkable, yet many existing methods struggle with issues of consistency and poor text-video alignment. Moreover, the field lacks effective techniques for text-guided video inpainting, a stark contrast to the well-explored domain of text-guided image inpainting. To this end, this paper proposes a novel text-guided video inpainting model that achieves better consistency, controllability and compatibility. Specifically, we introduce a simple but efficient motion capture module to preserve motion consistency, and design an instance-aware region selection instead of a random region selection to obtain better textual controllability, and utilize a novel strategy to inject some personalized models into our CoCoCo model and thus obtain better model compatibility. Extensive experiments show that our model can generate high-quality video clips. Meanwhile, our model shows better motion consistency, textual controllability and model compatibility. More details are shown in [cococozibojia.github.io](cococozibojia.github.io).
title CoCoCo: Improving Text-Guided Video Inpainting for Better Consistency, Controllability and Compatibility
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.12035