Apollo: Unified Multi-Task Audio-Video Joint Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Jun, Qiang, Chunyu, Guo, Yuxin, Wang, Yiran, Zeng, Xijuan, Deng, Feng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908761176670208
author Wang, Jun
Qiang, Chunyu
Guo, Yuxin
Wang, Yiran
Zeng, Xijuan
Deng, Feng
author_facet Wang, Jun
Qiang, Chunyu
Guo, Yuxin
Wang, Yiran
Zeng, Xijuan
Deng, Feng
contents Audio-video joint generation has progressed rapidly, yet substantial challenges still remain. Non-commercial approaches still suffer audio-visual asynchrony, poor lip-speech alignment, and unimodal degradation, which can be stemmed from weak audio-visual correspondence modeling, limited generalization, and scarce high-quality dense-caption data. To address these issues, we introduce Apollo and delve into three axes--model architecture, training strategy, and data curation. Architecturally, we adopt a single-tower design with unified DiT blocks and an Omni-Full Attention mechanism, achieving tight audio-visual alignment and strong scalability. Training-wise, we adopt a progressive multitask regime--random modality masking to joint optimization across tasks, and a multistage curriculum, yielding robust representations, strengthening A-V aligned world knowledge, and preventing unimodal collapse. For datasets, we present the first large-scale audio-video dataset with dense captions, and introduce a novel automated data-construction pipeline which annotates and filters millions of diverse, high-quality, strictly aligned audio-video-caption triplets. Building on this, Apollo scales to large datasets, delivering high-fidelity, semantically and temporally aligned, instruction-following generation in both joint and unimodal settings while generalizing robustly to out-of-distribution scenarios. Across tasks, it substantially outperforms prior methods by a large margin and achieves performance comparable to Veo 3, offering a unified, scalable path toward next-generation audio-video synthesis.
format Preprint
id arxiv_https___arxiv_org_abs_2601_04151
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Apollo: Unified Multi-Task Audio-Video Joint Generation
Wang, Jun
Qiang, Chunyu
Guo, Yuxin
Wang, Yiran
Zeng, Xijuan
Deng, Feng
Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
Sound
Audio-video joint generation has progressed rapidly, yet substantial challenges still remain. Non-commercial approaches still suffer audio-visual asynchrony, poor lip-speech alignment, and unimodal degradation, which can be stemmed from weak audio-visual correspondence modeling, limited generalization, and scarce high-quality dense-caption data. To address these issues, we introduce Apollo and delve into three axes--model architecture, training strategy, and data curation. Architecturally, we adopt a single-tower design with unified DiT blocks and an Omni-Full Attention mechanism, achieving tight audio-visual alignment and strong scalability. Training-wise, we adopt a progressive multitask regime--random modality masking to joint optimization across tasks, and a multistage curriculum, yielding robust representations, strengthening A-V aligned world knowledge, and preventing unimodal collapse. For datasets, we present the first large-scale audio-video dataset with dense captions, and introduce a novel automated data-construction pipeline which annotates and filters millions of diverse, high-quality, strictly aligned audio-video-caption triplets. Building on this, Apollo scales to large datasets, delivering high-fidelity, semantically and temporally aligned, instruction-following generation in both joint and unimodal settings while generalizing robustly to out-of-distribution scenarios. Across tasks, it substantially outperforms prior methods by a large margin and achieves performance comparable to Veo 3, offering a unified, scalable path toward next-generation audio-video synthesis.
title Apollo: Unified Multi-Task Audio-Video Joint Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
Sound
url https://arxiv.org/abs/2601.04151