L-C4: Language-Based Video Colorization for Creative and Consistent Color
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866910681212649472 |
|---|---|
| author | Chang, Zheng Weng, Shuchen Ouyang, Huan Li, Yu Li, Si Shi, Boxin |
| author_facet | Chang, Zheng Weng, Shuchen Ouyang, Huan Li, Yu Li, Si Shi, Boxin |
| contents | Automatic video colorization is inherently an ill-posed problem because each monochrome frame has multiple optional color candidates. Previous exemplar-based video colorization methods restrict the user's imagination due to the elaborate retrieval process. Alternatively, conditional image colorization methods combined with post-processing algorithms still struggle to maintain temporal consistency. To address these issues, we present Language-based video Colorization for Creative and Consistent Colors (L-C4) to guide the colorization process using user-provided language descriptions. Our model is built upon a pre-trained cross-modality generative model, leveraging its comprehensive language understanding and robust color representation abilities. We introduce the cross-modality pre-fusion module to generate instance-aware text embeddings, enabling the application of creative colors. Additionally, we propose temporally deformable attention to prevent flickering or color shifts, and cross-clip fusion to maintain long-term color consistency. Extensive experimental results demonstrate that L-C4 outperforms relevant methods, achieving semantically accurate colors, unrestricted creative correspondence, and temporally robust consistency. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_04972 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | L-C4: Language-Based Video Colorization for Creative and Consistent Color Chang, Zheng Weng, Shuchen Ouyang, Huan Li, Yu Li, Si Shi, Boxin Computer Vision and Pattern Recognition Automatic video colorization is inherently an ill-posed problem because each monochrome frame has multiple optional color candidates. Previous exemplar-based video colorization methods restrict the user's imagination due to the elaborate retrieval process. Alternatively, conditional image colorization methods combined with post-processing algorithms still struggle to maintain temporal consistency. To address these issues, we present Language-based video Colorization for Creative and Consistent Colors (L-C4) to guide the colorization process using user-provided language descriptions. Our model is built upon a pre-trained cross-modality generative model, leveraging its comprehensive language understanding and robust color representation abilities. We introduce the cross-modality pre-fusion module to generate instance-aware text embeddings, enabling the application of creative colors. Additionally, we propose temporally deformable attention to prevent flickering or color shifts, and cross-clip fusion to maintain long-term color consistency. Extensive experimental results demonstrate that L-C4 outperforms relevant methods, achieving semantically accurate colors, unrestricted creative correspondence, and temporally robust consistency. |
| title | L-C4: Language-Based Video Colorization for Creative and Consistent Color |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2410.04972 |