Uni-Animator: Towards Unified Visual Colorization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Xinyuan, Xu, Yao, Wang, Shaowen, Song, Pengjie, Deng, Bowen
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914364224700416
author Chen, Xinyuan
Xu, Yao
Wang, Shaowen
Song, Pengjie
Deng, Bowen
author_facet Chen, Xinyuan
Xu, Yao
Wang, Shaowen
Song, Pengjie
Deng, Bowen
contents We propose Uni-Animator, a novel Diffusion Transformer (DiT)-based framework for unified image and video sketch colorization. Existing sketch colorization methods struggle to unify image and video tasks, suffering from imprecise color transfer with single or multiple references, inadequate preservation of high-frequency physical details, and compromised temporal coherence with motion artifacts in large-motion scenes. To tackle imprecise color transfer, we introduce visual reference enhancement via instance patch embedding, enabling precise alignment and fusion of reference color information. To resolve insufficient physical detail preservation, we design physical detail reinforcement using physical features that effectively capture and retain high-frequency textures. To mitigate motion-induced temporal inconsistency, we propose sketch-based dynamic RoPE encoding that adaptively models motion-aware spatial-temporal dependencies. Extensive experimental results demonstrate that Uni-Animator achieves competitive performance on both image and video sketch colorization, matching that of task-specific methods while unlocking unified cross-domain capabilities with high detail fidelity and robust temporal consistency.
format Preprint
id arxiv_https___arxiv_org_abs_2602_23191
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Uni-Animator: Towards Unified Visual Colorization
Chen, Xinyuan
Xu, Yao
Wang, Shaowen
Song, Pengjie
Deng, Bowen
Computer Vision and Pattern Recognition
We propose Uni-Animator, a novel Diffusion Transformer (DiT)-based framework for unified image and video sketch colorization. Existing sketch colorization methods struggle to unify image and video tasks, suffering from imprecise color transfer with single or multiple references, inadequate preservation of high-frequency physical details, and compromised temporal coherence with motion artifacts in large-motion scenes. To tackle imprecise color transfer, we introduce visual reference enhancement via instance patch embedding, enabling precise alignment and fusion of reference color information. To resolve insufficient physical detail preservation, we design physical detail reinforcement using physical features that effectively capture and retain high-frequency textures. To mitigate motion-induced temporal inconsistency, we propose sketch-based dynamic RoPE encoding that adaptively models motion-aware spatial-temporal dependencies. Extensive experimental results demonstrate that Uni-Animator achieves competitive performance on both image and video sketch colorization, matching that of task-specific methods while unlocking unified cross-domain capabilities with high detail fidelity and robust temporal consistency.
title Uni-Animator: Towards Unified Visual Colorization
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.23191