Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Cheng, Shihao, Zhang, Jiaxu, Song, Quanyue, Liu, Shansong, Guo, Zhizhi, Zhang, Xiaolei, Zhang, Chi, Li, Xuelong, Tu, Zhigang
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909029525094400
author Cheng, Shihao
Zhang, Jiaxu
Song, Quanyue
Liu, Shansong
Guo, Zhizhi
Zhang, Xiaolei
Zhang, Chi
Li, Xuelong
Tu, Zhigang
author_facet Cheng, Shihao
Zhang, Jiaxu
Song, Quanyue
Liu, Shansong
Guo, Zhizhi
Zhang, Xiaolei
Zhang, Chi
Li, Xuelong
Tu, Zhigang
contents Motion, speech, and sound effects are fundamental elements of human-centric videos, yet their heterogeneous temporal characteristics make joint generation highly challenging. Existing audio-video generation models often fail to maintain consistent alignment across these modalities, leading to noticeable mismatches between motion, speech, and environmental sounds. We present Unison, a unified framework that explicitly promotes coherence across the motion, speech, and sound modalities. Within the audio stream, Unison employs a semantic-guided harmonization strategy that decouples the generation of speech and sound-effect components. Leveraging bidirectional audio cross-attention and semantic-conditioned gating for semantic-driven adaptive recomposition, this approach effectively mitigates speech dominance and enhances acoustic clarity. For audio-motion synchronization, we propose a bidirectional cross-modal forcing strategy where the cleaner modality guides the noisier one through decoupled denoising schedules, reinforced by a progressive stabilization strategy. Extensive experiments demonstrate that Unison achieves state-of-the-art performance in both audio perceptual quality and cross-modal synchronization, highlighting the importance of explicit multimodal harmonization in human-centric video generation.
format Preprint
id arxiv_https___arxiv_org_abs_2605_08729
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation
Cheng, Shihao
Zhang, Jiaxu
Song, Quanyue
Liu, Shansong
Guo, Zhizhi
Zhang, Xiaolei
Zhang, Chi
Li, Xuelong
Tu, Zhigang
Computer Vision and Pattern Recognition
Graphics
Multimedia
Sound
Motion, speech, and sound effects are fundamental elements of human-centric videos, yet their heterogeneous temporal characteristics make joint generation highly challenging. Existing audio-video generation models often fail to maintain consistent alignment across these modalities, leading to noticeable mismatches between motion, speech, and environmental sounds. We present Unison, a unified framework that explicitly promotes coherence across the motion, speech, and sound modalities. Within the audio stream, Unison employs a semantic-guided harmonization strategy that decouples the generation of speech and sound-effect components. Leveraging bidirectional audio cross-attention and semantic-conditioned gating for semantic-driven adaptive recomposition, this approach effectively mitigates speech dominance and enhances acoustic clarity. For audio-motion synchronization, we propose a bidirectional cross-modal forcing strategy where the cleaner modality guides the noisier one through decoupled denoising schedules, reinforced by a progressive stabilization strategy. Extensive experiments demonstrate that Unison achieves state-of-the-art performance in both audio perceptual quality and cross-modal synchronization, highlighting the importance of explicit multimodal harmonization in human-centric video generation.
title Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation
topic Computer Vision and Pattern Recognition
Graphics
Multimedia
Sound
url https://arxiv.org/abs/2605.08729