Cosh-DiT: Co-Speech Gesture Video Synthesis via Hybrid Audio-Visual Diffusion Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Yasheng, Xu, Zhiliang, Zhou, Hang, Guan, Jiazhi, Yang, Quanwei, Wang, Kaisiyuan, Liang, Borong, Li, Yingying, Feng, Haocheng, Wang, Jingdong, Liu, Ziwei, Hideki, Koike
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917954873982976
author Sun, Yasheng
Xu, Zhiliang
Zhou, Hang
Guan, Jiazhi
Yang, Quanwei
Wang, Kaisiyuan
Liang, Borong
Li, Yingying
Feng, Haocheng
Wang, Jingdong
Liu, Ziwei
Hideki, Koike
author_facet Sun, Yasheng
Xu, Zhiliang
Zhou, Hang
Guan, Jiazhi
Yang, Quanwei
Wang, Kaisiyuan
Liang, Borong
Li, Yingying
Feng, Haocheng
Wang, Jingdong
Liu, Ziwei
Hideki, Koike
contents Co-speech gesture video synthesis is a challenging task that requires both probabilistic modeling of human gestures and the synthesis of realistic images that align with the rhythmic nuances of speech. To address these challenges, we propose Cosh-DiT, a Co-speech gesture video system with hybrid Diffusion Transformers that perform audio-to-motion and motion-to-video synthesis using discrete and continuous diffusion modeling, respectively. First, we introduce an audio Diffusion Transformer (Cosh-DiT-A) to synthesize expressive gesture dynamics synchronized with speech rhythms. To capture upper body, facial, and hand movement priors, we employ vector-quantized variational autoencoders (VQ-VAEs) to jointly learn their dependencies within a discrete latent space. Then, for realistic video synthesis conditioned on the generated speech-driven motion, we design a visual Diffusion Transformer (Cosh-DiT-V) that effectively integrates spatial and temporal contexts. Extensive experiments demonstrate that our framework consistently generates lifelike videos with expressive facial expressions and natural, smooth gestures that align seamlessly with speech.
format Preprint
id arxiv_https___arxiv_org_abs_2503_09942
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Cosh-DiT: Co-Speech Gesture Video Synthesis via Hybrid Audio-Visual Diffusion Transformers
Sun, Yasheng
Xu, Zhiliang
Zhou, Hang
Guan, Jiazhi
Yang, Quanwei
Wang, Kaisiyuan
Liang, Borong
Li, Yingying
Feng, Haocheng
Wang, Jingdong
Liu, Ziwei
Hideki, Koike
Computer Vision and Pattern Recognition
Co-speech gesture video synthesis is a challenging task that requires both probabilistic modeling of human gestures and the synthesis of realistic images that align with the rhythmic nuances of speech. To address these challenges, we propose Cosh-DiT, a Co-speech gesture video system with hybrid Diffusion Transformers that perform audio-to-motion and motion-to-video synthesis using discrete and continuous diffusion modeling, respectively. First, we introduce an audio Diffusion Transformer (Cosh-DiT-A) to synthesize expressive gesture dynamics synchronized with speech rhythms. To capture upper body, facial, and hand movement priors, we employ vector-quantized variational autoencoders (VQ-VAEs) to jointly learn their dependencies within a discrete latent space. Then, for realistic video synthesis conditioned on the generated speech-driven motion, we design a visual Diffusion Transformer (Cosh-DiT-V) that effectively integrates spatial and temporal contexts. Extensive experiments demonstrate that our framework consistently generates lifelike videos with expressive facial expressions and natural, smooth gestures that align seamlessly with speech.
title Cosh-DiT: Co-Speech Gesture Video Synthesis via Hybrid Audio-Visual Diffusion Transformers
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.09942