FullDiT2: Efficient In-Context Conditioning for Video Diffusion Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Xuanhua, Liu, Quande, Ye, Zixuan, Ye, Weicai, Wang, Qiulin, Wang, Xintao, Chen, Qifeng, Wan, Pengfei, Zhang, Di, Gai, Kun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908394113204224
author He, Xuanhua
Liu, Quande
Ye, Zixuan
Ye, Weicai
Wang, Qiulin
Wang, Xintao
Chen, Qifeng
Wan, Pengfei
Zhang, Di
Gai, Kun
author_facet He, Xuanhua
Liu, Quande
Ye, Zixuan
Ye, Weicai
Wang, Qiulin
Wang, Xintao
Chen, Qifeng
Wan, Pengfei
Zhang, Di
Gai, Kun
contents Fine-grained and efficient controllability on video diffusion transformers has raised increasing desires for the applicability. Recently, In-context Conditioning emerged as a powerful paradigm for unified conditional video generation, which enables diverse controls by concatenating varying context conditioning signals with noisy video latents into a long unified token sequence and jointly processing them via full-attention, e.g., FullDiT. Despite their effectiveness, these methods face quadratic computation overhead as task complexity increases, hindering practical deployment. In this paper, we study the efficiency bottleneck neglected in original in-context conditioning video generation framework. We begin with systematic analysis to identify two key sources of the computation inefficiencies: the inherent redundancy within context condition tokens and the computational redundancy in context-latent interactions throughout the diffusion process. Based on these insights, we propose FullDiT2, an efficient in-context conditioning framework for general controllability in both video generation and editing tasks, which innovates from two key perspectives. Firstly, to address the token redundancy, FullDiT2 leverages a dynamic token selection mechanism to adaptively identify important context tokens, reducing the sequence length for unified full-attention. Additionally, a selective context caching mechanism is devised to minimize redundant interactions between condition tokens and video latents. Extensive experiments on six diverse conditional video editing and generation tasks demonstrate that FullDiT2 achieves significant computation reduction and 2-3 times speedup in averaged time cost per diffusion step, with minimal degradation or even higher performance in video generation quality. The project page is at \href{https://fulldit2.github.io/}{https://fulldit2.github.io/}.
format Preprint
id arxiv_https___arxiv_org_abs_2506_04213
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FullDiT2: Efficient In-Context Conditioning for Video Diffusion Transformers
He, Xuanhua
Liu, Quande
Ye, Zixuan
Ye, Weicai
Wang, Qiulin
Wang, Xintao
Chen, Qifeng
Wan, Pengfei
Zhang, Di
Gai, Kun
Computer Vision and Pattern Recognition
Fine-grained and efficient controllability on video diffusion transformers has raised increasing desires for the applicability. Recently, In-context Conditioning emerged as a powerful paradigm for unified conditional video generation, which enables diverse controls by concatenating varying context conditioning signals with noisy video latents into a long unified token sequence and jointly processing them via full-attention, e.g., FullDiT. Despite their effectiveness, these methods face quadratic computation overhead as task complexity increases, hindering practical deployment. In this paper, we study the efficiency bottleneck neglected in original in-context conditioning video generation framework. We begin with systematic analysis to identify two key sources of the computation inefficiencies: the inherent redundancy within context condition tokens and the computational redundancy in context-latent interactions throughout the diffusion process. Based on these insights, we propose FullDiT2, an efficient in-context conditioning framework for general controllability in both video generation and editing tasks, which innovates from two key perspectives. Firstly, to address the token redundancy, FullDiT2 leverages a dynamic token selection mechanism to adaptively identify important context tokens, reducing the sequence length for unified full-attention. Additionally, a selective context caching mechanism is devised to minimize redundant interactions between condition tokens and video latents. Extensive experiments on six diverse conditional video editing and generation tasks demonstrate that FullDiT2 achieves significant computation reduction and 2-3 times speedup in averaged time cost per diffusion step, with minimal degradation or even higher performance in video generation quality. The project page is at \href{https://fulldit2.github.io/}{https://fulldit2.github.io/}.
title FullDiT2: Efficient In-Context Conditioning for Video Diffusion Transformers
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.04213