In-Context Audio Control of Video Diffusion Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Wenze, Ye, Weicai, Cai, Minghong, Liu, Quande, Wang, Xintao, Yue, Xiangyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911329951940608
author Liu, Wenze
Ye, Weicai
Cai, Minghong
Liu, Quande
Wang, Xintao
Yue, Xiangyu
author_facet Liu, Wenze
Ye, Weicai
Cai, Minghong
Liu, Quande
Wang, Xintao
Yue, Xiangyu
contents Recent advancements in video generation have seen a shift towards unified, transformer-based foundation models that can handle multiple conditional inputs in-context. However, these models have primarily focused on modalities like text, images, and depth maps, while strictly time-synchronous signals like audio have been underexplored. This paper introduces In-Context Audio Control of video diffusion transformers (ICAC), a framework that investigates the integration of audio signals for speech-driven video generation within a unified full-attention architecture, akin to FullDiT. We systematically explore three distinct mechanisms for injecting audio conditions: standard cross-attention, 2D self-attention, and unified 3D self-attention. Our findings reveal that while 3D attention offers the highest potential for capturing spatio-temporal audio-visual correlations, it presents significant training challenges. To overcome this, we propose a Masked 3D Attention mechanism that constrains the attention pattern to enforce temporal alignment, enabling stable training and superior performance. Our experiments demonstrate that this approach achieves strong lip synchronization and video quality, conditioned on an audio stream and reference images.
format Preprint
id arxiv_https___arxiv_org_abs_2512_18772
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle In-Context Audio Control of Video Diffusion Transformers
Liu, Wenze
Ye, Weicai
Cai, Minghong
Liu, Quande
Wang, Xintao
Yue, Xiangyu
Computer Vision and Pattern Recognition
Recent advancements in video generation have seen a shift towards unified, transformer-based foundation models that can handle multiple conditional inputs in-context. However, these models have primarily focused on modalities like text, images, and depth maps, while strictly time-synchronous signals like audio have been underexplored. This paper introduces In-Context Audio Control of video diffusion transformers (ICAC), a framework that investigates the integration of audio signals for speech-driven video generation within a unified full-attention architecture, akin to FullDiT. We systematically explore three distinct mechanisms for injecting audio conditions: standard cross-attention, 2D self-attention, and unified 3D self-attention. Our findings reveal that while 3D attention offers the highest potential for capturing spatio-temporal audio-visual correlations, it presents significant training challenges. To overcome this, we propose a Masked 3D Attention mechanism that constrains the attention pattern to enforce temporal alignment, enabling stable training and superior performance. Our experiments demonstrate that this approach achieves strong lip synchronization and video quality, conditioned on an audio stream and reference images.
title In-Context Audio Control of Video Diffusion Transformers
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.18772