GestureHYDRA: Semantic Co-speech Gesture Synthesis via Hybrid Modality Diffusion Transformer and Cascaded-Synchronized Retrieval-Augmented Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Quanwei, Huang, Luying, Wang, Kaisiyuan, Guan, Jiazhi, He, Shengyi, Li, Fengguo, Zhou, Hang, Yu, Lingyun, Li, Yingying, Feng, Haocheng, Xie, Hongtao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917176262262784
author Yang, Quanwei
Huang, Luying
Wang, Kaisiyuan
Guan, Jiazhi
He, Shengyi
Li, Fengguo
Zhou, Hang
Yu, Lingyun
Li, Yingying
Feng, Haocheng
Xie, Hongtao
author_facet Yang, Quanwei
Huang, Luying
Wang, Kaisiyuan
Guan, Jiazhi
He, Shengyi
Li, Fengguo
Zhou, Hang
Yu, Lingyun
Li, Yingying
Feng, Haocheng
Xie, Hongtao
contents While increasing attention has been paid to co-speech gesture synthesis, most previous works neglect to investigate hand gestures with explicit and essential semantics. In this paper, we study co-speech gesture generation with an emphasis on specific hand gesture activation, which can deliver more instructional information than common body movements. To achieve this, we first build a high-quality dataset of 3D human body movements including a set of semantically explicit hand gestures that are commonly used by live streamers. Then we present a hybrid-modality gesture generation system GestureHYDRA built upon a hybrid-modality diffusion transformer architecture with novelly designed motion-style injective transformer layers, which enables advanced gesture modeling ability and versatile gesture operations. To guarantee these specific hand gestures can be activated, we introduce a cascaded retrieval-augmented generation strategy built upon a semantic gesture repository annotated for each subject and an adaptive audio-gesture synchronization mechanism, which substantially improves semantic gesture activation and production efficiency. Quantitative and qualitative experiments demonstrate that our proposed approach achieves superior performance over all the counterparts. The project page can be found at https://mumuwei.github.io/GestureHYDRA/.
format Preprint
id arxiv_https___arxiv_org_abs_2507_22731
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GestureHYDRA: Semantic Co-speech Gesture Synthesis via Hybrid Modality Diffusion Transformer and Cascaded-Synchronized Retrieval-Augmented Generation
Yang, Quanwei
Huang, Luying
Wang, Kaisiyuan
Guan, Jiazhi
He, Shengyi
Li, Fengguo
Zhou, Hang
Yu, Lingyun
Li, Yingying
Feng, Haocheng
Xie, Hongtao
Multimedia
While increasing attention has been paid to co-speech gesture synthesis, most previous works neglect to investigate hand gestures with explicit and essential semantics. In this paper, we study co-speech gesture generation with an emphasis on specific hand gesture activation, which can deliver more instructional information than common body movements. To achieve this, we first build a high-quality dataset of 3D human body movements including a set of semantically explicit hand gestures that are commonly used by live streamers. Then we present a hybrid-modality gesture generation system GestureHYDRA built upon a hybrid-modality diffusion transformer architecture with novelly designed motion-style injective transformer layers, which enables advanced gesture modeling ability and versatile gesture operations. To guarantee these specific hand gestures can be activated, we introduce a cascaded retrieval-augmented generation strategy built upon a semantic gesture repository annotated for each subject and an adaptive audio-gesture synchronization mechanism, which substantially improves semantic gesture activation and production efficiency. Quantitative and qualitative experiments demonstrate that our proposed approach achieves superior performance over all the counterparts. The project page can be found at https://mumuwei.github.io/GestureHYDRA/.
title GestureHYDRA: Semantic Co-speech Gesture Synthesis via Hybrid Modality Diffusion Transformer and Cascaded-Synchronized Retrieval-Augmented Generation
topic Multimedia
url https://arxiv.org/abs/2507.22731