Improving Joint Audio-Video Generation with Cross-Modal Context Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Bingqi, Lang, Linlong, Zhang, Ming, He, Dailan, Ge, Xingtong, Zhang, Yi, Song, Guanglu, Liu, Yu
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914408780791808
author Ma, Bingqi
Lang, Linlong
Zhang, Ming
He, Dailan
Ge, Xingtong
Zhang, Yi
Song, Guanglu
Liu, Yu
author_facet Ma, Bingqi
Lang, Linlong
Zhang, Ming
He, Dailan
Ge, Xingtong
Zhang, Yi
Song, Guanglu
Liu, Yu
contents The dual-stream transformer architecture-based joint audio-video generation method has become the dominant paradigm in current research. By incorporating pre-trained video diffusion models and audio diffusion models, along with a cross-modal interaction attention module, high-quality, temporally synchronized audio-video content can be generated with minimal training data. In this paper, we first revisit the dual-stream transformer paradigm and further analyze its limitations, including model manifold variations caused by the gating mechanism controlling cross-modal interactions, biases in multi-modal background regions introduced by cross-modal attention, and the inconsistencies in multi-modal classifier-free guidance (CFG) during training and inference, as well as conflicts between multiple conditions. To alleviate these issues, we propose Cross-Modal Context Learning (CCL), equipped with several carefully designed modules. Temporally Aligned RoPE and Partitioning (TARP) effectively enhances the temporal alignment between audio latent and video latent representations. The Learnable Context Tokens (LCT) and Dynamic Context Routing (DCR) in the Cross-Modal Context Attention (CCA) module provide stable unconditional anchors for cross-modal information, while dynamically routing based on different training tasks, further enhancing the model's convergence speed and generation quality. During inference, Unconditional Context Guidance (UCG) leverages the unconditional support provided by LCT to facilitate different forms of CFG, improving train-inference consistency and further alleviating conflicts. Through comprehensive evaluations, CCL achieves state-of-the-art performance compared with recent academic methods while requiring substantially fewer resources.
format Preprint
id arxiv_https___arxiv_org_abs_2603_18600
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Improving Joint Audio-Video Generation with Cross-Modal Context Learning
Ma, Bingqi
Lang, Linlong
Zhang, Ming
He, Dailan
Ge, Xingtong
Zhang, Yi
Song, Guanglu
Liu, Yu
Computer Vision and Pattern Recognition
The dual-stream transformer architecture-based joint audio-video generation method has become the dominant paradigm in current research. By incorporating pre-trained video diffusion models and audio diffusion models, along with a cross-modal interaction attention module, high-quality, temporally synchronized audio-video content can be generated with minimal training data. In this paper, we first revisit the dual-stream transformer paradigm and further analyze its limitations, including model manifold variations caused by the gating mechanism controlling cross-modal interactions, biases in multi-modal background regions introduced by cross-modal attention, and the inconsistencies in multi-modal classifier-free guidance (CFG) during training and inference, as well as conflicts between multiple conditions. To alleviate these issues, we propose Cross-Modal Context Learning (CCL), equipped with several carefully designed modules. Temporally Aligned RoPE and Partitioning (TARP) effectively enhances the temporal alignment between audio latent and video latent representations. The Learnable Context Tokens (LCT) and Dynamic Context Routing (DCR) in the Cross-Modal Context Attention (CCA) module provide stable unconditional anchors for cross-modal information, while dynamically routing based on different training tasks, further enhancing the model's convergence speed and generation quality. During inference, Unconditional Context Guidance (UCG) leverages the unconditional support provided by LCT to facilitate different forms of CFG, improving train-inference consistency and further alleviating conflicts. Through comprehensive evaluations, CCL achieves state-of-the-art performance compared with recent academic methods while requiring substantially fewer resources.
title Improving Joint Audio-Video Generation with Cross-Modal Context Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.18600