JavisDiT++: Unified Modeling and Optimization for Joint Audio-Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Kai, Zheng, Yanhao, Wang, Kai, Wu, Shengqiong, Zhang, Rongjunchen, Luo, Jiebo, Hatzinakos, Dimitrios, Liu, Ziwei, Fei, Hao, Chua, Tat-Seng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915811280551936
author Liu, Kai
Zheng, Yanhao
Wang, Kai
Wu, Shengqiong
Zhang, Rongjunchen
Luo, Jiebo
Hatzinakos, Dimitrios
Liu, Ziwei
Fei, Hao
Chua, Tat-Seng
author_facet Liu, Kai
Zheng, Yanhao
Wang, Kai
Wu, Shengqiong
Zhang, Rongjunchen
Luo, Jiebo
Hatzinakos, Dimitrios
Liu, Ziwei
Fei, Hao
Chua, Tat-Seng
contents AIGC has rapidly expanded from text-to-image generation toward high-quality multimodal synthesis across video and audio. Within this context, joint audio-video generation (JAVG) has emerged as a fundamental task that produces synchronized and semantically aligned sound and vision from textual descriptions. However, compared with advanced commercial models such as Veo3, existing open-source methods still suffer from limitations in generation quality, temporal synchrony, and alignment with human preferences. To bridge the gap, this paper presents JavisDiT++, a concise yet powerful framework for unified modeling and optimization of JAVG. First, we introduce a modality-specific mixture-of-experts (MS-MoE) design that enables cross-modal interaction efficacy while enhancing single-modal generation quality. Then, we propose a temporal-aligned RoPE (TA-RoPE) strategy to achieve explicit, frame-level synchronization between audio and video tokens. Besides, we develop an audio-video direct preference optimization (AV-DPO) method to align model outputs with human preference across quality, consistency, and synchrony dimensions. Built upon Wan2.1-1.3B-T2V, our model achieves state-of-the-art performance merely with around 1M public training entries, significantly outperforming prior approaches in both qualitative and quantitative evaluations. Comprehensive ablation studies have been conducted to validate the effectiveness of our proposed modules. All the code, model, and dataset are released at https://JavisVerse.github.io/JavisDiT2-page.
format Preprint
id arxiv_https___arxiv_org_abs_2602_19163
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle JavisDiT++: Unified Modeling and Optimization for Joint Audio-Video Generation
Liu, Kai
Zheng, Yanhao
Wang, Kai
Wu, Shengqiong
Zhang, Rongjunchen
Luo, Jiebo
Hatzinakos, Dimitrios
Liu, Ziwei
Fei, Hao
Chua, Tat-Seng
Computer Vision and Pattern Recognition
Multimedia
Sound
AIGC has rapidly expanded from text-to-image generation toward high-quality multimodal synthesis across video and audio. Within this context, joint audio-video generation (JAVG) has emerged as a fundamental task that produces synchronized and semantically aligned sound and vision from textual descriptions. However, compared with advanced commercial models such as Veo3, existing open-source methods still suffer from limitations in generation quality, temporal synchrony, and alignment with human preferences. To bridge the gap, this paper presents JavisDiT++, a concise yet powerful framework for unified modeling and optimization of JAVG. First, we introduce a modality-specific mixture-of-experts (MS-MoE) design that enables cross-modal interaction efficacy while enhancing single-modal generation quality. Then, we propose a temporal-aligned RoPE (TA-RoPE) strategy to achieve explicit, frame-level synchronization between audio and video tokens. Besides, we develop an audio-video direct preference optimization (AV-DPO) method to align model outputs with human preference across quality, consistency, and synchrony dimensions. Built upon Wan2.1-1.3B-T2V, our model achieves state-of-the-art performance merely with around 1M public training entries, significantly outperforming prior approaches in both qualitative and quantitative evaluations. Comprehensive ablation studies have been conducted to validate the effectiveness of our proposed modules. All the code, model, and dataset are released at https://JavisVerse.github.io/JavisDiT2-page.
title JavisDiT++: Unified Modeling and Optimization for Joint Audio-Video Generation
topic Computer Vision and Pattern Recognition
Multimedia
Sound
url https://arxiv.org/abs/2602.19163