UNIC: Unified In-Context Video Editing

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ye, Zixuan, He, Xuanhua, Liu, Quande, Wang, Qiulin, Wang, Xintao, Wan, Pengfei, Zhang, Di, Gai, Kun, Chen, Qifeng, Luo, Wenhan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912413170794496
author Ye, Zixuan
He, Xuanhua
Liu, Quande
Wang, Qiulin
Wang, Xintao
Wan, Pengfei
Zhang, Di
Gai, Kun
Chen, Qifeng
Luo, Wenhan
author_facet Ye, Zixuan
He, Xuanhua
Liu, Quande
Wang, Qiulin
Wang, Xintao
Wan, Pengfei
Zhang, Di
Gai, Kun
Chen, Qifeng
Luo, Wenhan
contents Recent advances in text-to-video generation have sparked interest in generative video editing tasks. Previous methods often rely on task-specific architectures (e.g., additional adapter modules) or dedicated customizations (e.g., DDIM inversion), which limit the integration of versatile editing conditions and the unification of various editing tasks. In this paper, we introduce UNified In-Context Video Editing (UNIC), a simple yet effective framework that unifies diverse video editing tasks within a single model in an in-context manner. To achieve this unification, we represent the inputs of various video editing tasks as three types of tokens: the source video tokens, the noisy video latent, and the multi-modal conditioning tokens that vary according to the specific editing task. Based on this formulation, our key insight is to integrate these three types into a single consecutive token sequence and jointly model them using the native attention operations of DiT, thereby eliminating the need for task-specific adapter designs. Nevertheless, direct task unification under this framework is challenging, leading to severe token collisions and task confusion due to the varying video lengths and diverse condition modalities across tasks. To address these, we introduce task-aware RoPE to facilitate consistent temporal positional encoding, and condition bias that enables the model to clearly differentiate different editing tasks. This allows our approach to adaptively perform different video editing tasks by referring the source video and varying condition tokens "in context", and support flexible task composition. To validate our method, we construct a unified video editing benchmark containing six representative video editing tasks. Results demonstrate that our unified approach achieves superior performance on each task and exhibits emergent task composition abilities.
format Preprint
id arxiv_https___arxiv_org_abs_2506_04216
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle UNIC: Unified In-Context Video Editing
Ye, Zixuan
He, Xuanhua
Liu, Quande
Wang, Qiulin
Wang, Xintao
Wan, Pengfei
Zhang, Di
Gai, Kun
Chen, Qifeng
Luo, Wenhan
Computer Vision and Pattern Recognition
Recent advances in text-to-video generation have sparked interest in generative video editing tasks. Previous methods often rely on task-specific architectures (e.g., additional adapter modules) or dedicated customizations (e.g., DDIM inversion), which limit the integration of versatile editing conditions and the unification of various editing tasks. In this paper, we introduce UNified In-Context Video Editing (UNIC), a simple yet effective framework that unifies diverse video editing tasks within a single model in an in-context manner. To achieve this unification, we represent the inputs of various video editing tasks as three types of tokens: the source video tokens, the noisy video latent, and the multi-modal conditioning tokens that vary according to the specific editing task. Based on this formulation, our key insight is to integrate these three types into a single consecutive token sequence and jointly model them using the native attention operations of DiT, thereby eliminating the need for task-specific adapter designs. Nevertheless, direct task unification under this framework is challenging, leading to severe token collisions and task confusion due to the varying video lengths and diverse condition modalities across tasks. To address these, we introduce task-aware RoPE to facilitate consistent temporal positional encoding, and condition bias that enables the model to clearly differentiate different editing tasks. This allows our approach to adaptively perform different video editing tasks by referring the source video and varying condition tokens "in context", and support flexible task composition. To validate our method, we construct a unified video editing benchmark containing six representative video editing tasks. Results demonstrate that our unified approach achieves superior performance on each task and exhibits emergent task composition abilities.
title UNIC: Unified In-Context Video Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.04216