L-C4: Language-Based Video Colorization for Creative and Consistent Color

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chang, Zheng, Weng, Shuchen, Ouyang, Huan, Li, Yu, Li, Si, Shi, Boxin
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910681212649472
author Chang, Zheng
Weng, Shuchen
Ouyang, Huan
Li, Yu
Li, Si
Shi, Boxin
author_facet Chang, Zheng
Weng, Shuchen
Ouyang, Huan
Li, Yu
Li, Si
Shi, Boxin
contents Automatic video colorization is inherently an ill-posed problem because each monochrome frame has multiple optional color candidates. Previous exemplar-based video colorization methods restrict the user's imagination due to the elaborate retrieval process. Alternatively, conditional image colorization methods combined with post-processing algorithms still struggle to maintain temporal consistency. To address these issues, we present Language-based video Colorization for Creative and Consistent Colors (L-C4) to guide the colorization process using user-provided language descriptions. Our model is built upon a pre-trained cross-modality generative model, leveraging its comprehensive language understanding and robust color representation abilities. We introduce the cross-modality pre-fusion module to generate instance-aware text embeddings, enabling the application of creative colors. Additionally, we propose temporally deformable attention to prevent flickering or color shifts, and cross-clip fusion to maintain long-term color consistency. Extensive experimental results demonstrate that L-C4 outperforms relevant methods, achieving semantically accurate colors, unrestricted creative correspondence, and temporally robust consistency.
format Preprint
id arxiv_https___arxiv_org_abs_2410_04972
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle L-C4: Language-Based Video Colorization for Creative and Consistent Color
Chang, Zheng
Weng, Shuchen
Ouyang, Huan
Li, Yu
Li, Si
Shi, Boxin
Computer Vision and Pattern Recognition
Automatic video colorization is inherently an ill-posed problem because each monochrome frame has multiple optional color candidates. Previous exemplar-based video colorization methods restrict the user's imagination due to the elaborate retrieval process. Alternatively, conditional image colorization methods combined with post-processing algorithms still struggle to maintain temporal consistency. To address these issues, we present Language-based video Colorization for Creative and Consistent Colors (L-C4) to guide the colorization process using user-provided language descriptions. Our model is built upon a pre-trained cross-modality generative model, leveraging its comprehensive language understanding and robust color representation abilities. We introduce the cross-modality pre-fusion module to generate instance-aware text embeddings, enabling the application of creative colors. Additionally, we propose temporally deformable attention to prevent flickering or color shifts, and cross-clip fusion to maintain long-term color consistency. Extensive experimental results demonstrate that L-C4 outperforms relevant methods, achieving semantically accurate colors, unrestricted creative correspondence, and temporally robust consistency.
title L-C4: Language-Based Video Colorization for Creative and Consistent Color
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.04972