DIDiffGes: Decoupled Semi-Implicit Diffusion Models for Real-time Gesture Generation from Speech
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Cheng, Yongkang, Huang, Shaoli, Chen, Xuelin, Ning, Jifeng, Gong, Mingming |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Conditional GAN for Enhancing Diffusion Models in Efficient and Authentic Global Gesture Generation from Audios
par: Cheng, Yongkang, et autres
Publié: (2024)
par: Cheng, Yongkang, et autres
Publié: (2024)
Freetalker: Controllable Speech and Text-Driven Gesture Generation Based on Diffusion Models for Enhanced Speaker Naturalness
par: Yang, Sicheng, et autres
Publié: (2024)
par: Yang, Sicheng, et autres
Publié: (2024)
Beat-It: Beat-Synchronized Multi-Condition 3D Dance Generation
par: Huang, Zikai, et autres
Publié: (2024)
par: Huang, Zikai, et autres
Publié: (2024)
Inter-Diffusion Generation Model of Speakers and Listeners for Effective Communication
par: Huang, Jinhe, et autres
Publié: (2025)
par: Huang, Jinhe, et autres
Publié: (2025)
ExpGest: Expressive Speaker Generation Using Diffusion Model and Hybrid Audio-Text Guidance
par: Cheng, Yongkang, et autres
Publié: (2024)
par: Cheng, Yongkang, et autres
Publié: (2024)
NAT: Neural Acoustic Transfer for Interactive Scenes in Real Time
par: Jin, Xutong, et autres
Publié: (2025)
par: Jin, Xutong, et autres
Publié: (2025)
Semantic Gesticulator: Semantics-Aware Co-Speech Gesture Synthesis
par: Zhang, Zeyi, et autres
Publié: (2024)
par: Zhang, Zeyi, et autres
Publié: (2024)
DiffSHEG: A Diffusion-Based Approach for Real-Time Speech-driven Holistic 3D Expression and Gesture Generation
par: Chen, Junming, et autres
Publié: (2024)
par: Chen, Junming, et autres
Publié: (2024)
Combining Genre Classification and Harmonic-Percussive Features with Diffusion Models for Music-Video Generation
par: Pina, Leonardo, et autres
Publié: (2024)
par: Pina, Leonardo, et autres
Publié: (2024)
It Takes Two: Real-time Co-Speech Two-person's Interaction Generation via Reactive Auto-regressive Diffusion Model
par: Shi, Mingyi, et autres
Publié: (2024)
par: Shi, Mingyi, et autres
Publié: (2024)
ELGAR: Expressive Cello Performance Motion Generation for Audio Rendition
par: Qiu, Zhiping, et autres
Publié: (2025)
par: Qiu, Zhiping, et autres
Publié: (2025)
Listen and Move: Improving GANs Coherency in Agnostic Sound-to-Video Generation
par: Redondo, Rafael
Publié: (2024)
par: Redondo, Rafael
Publié: (2024)
DGFM: Full Body Dance Generation Driven by Music Foundation Models
par: Liu, Xinran, et autres
Publié: (2025)
par: Liu, Xinran, et autres
Publié: (2025)
SyncViolinist: Music-Oriented Violin Motion Generation Based on Bowing and Fingering
par: Nishizawa, Hiroki, et autres
Publié: (2024)
par: Nishizawa, Hiroki, et autres
Publié: (2024)
READ: Real-time and Efficient Asynchronous Diffusion for Audio-driven Talking Head Generation
par: Wang, Haotian, et autres
Publié: (2025)
par: Wang, Haotian, et autres
Publié: (2025)
Text-Driven Voice Conversion via Latent State-Space Modeling
par: Li, Wen, et autres
Publié: (2025)
par: Li, Wen, et autres
Publié: (2025)
Content and Style Aware Audio-Driven Facial Animation
par: Liu, Qingju, et autres
Publié: (2024)
par: Liu, Qingju, et autres
Publié: (2024)
EnchantDance: Unveiling the Potential of Music-Driven Dance Movement
par: Han, Bo, et autres
Publié: (2023)
par: Han, Bo, et autres
Publié: (2023)
MusicScore: A Dataset for Music Score Modeling and Generation
par: Lin, Yuheng, et autres
Publié: (2024)
par: Lin, Yuheng, et autres
Publié: (2024)
RAP: Real-time Audio-driven Portrait Animation with Video Diffusion Transformer
par: Du, Fangyu, et autres
Publié: (2025)
par: Du, Fangyu, et autres
Publié: (2025)
DiM-Gestor: Co-Speech Gesture Generation with Adaptive Layer Normalization Mamba-2
par: Zhang, Fan, et autres
Publié: (2024)
par: Zhang, Fan, et autres
Publié: (2024)
MATHDance: Mamba-Transformer Architecture with Uniform Tokenization for High-Quality 3D Dance Generation
par: Yang, Kaixing, et autres
Publié: (2025)
par: Yang, Kaixing, et autres
Publié: (2025)
LLM Gesticulator: Leveraging Large Language Models for Scalable and Controllable Co-Speech Gesture Synthesis
par: Pang, Haozhou, et autres
Publié: (2024)
par: Pang, Haozhou, et autres
Publié: (2024)
DanceAnyWay: Synthesizing Beat-Guided 3D Dances with Randomized Temporal Contrastive Learning
par: Bhattacharya, Aneesh, et autres
Publié: (2023)
par: Bhattacharya, Aneesh, et autres
Publié: (2023)
MotionRAG-Diff: A Retrieval-Augmented Diffusion Framework for Long-Term Music-to-Dance Generation
par: Huang, Mingyang, et autres
Publié: (2025)
par: Huang, Mingyang, et autres
Publié: (2025)
Lodge: A Coarse to Fine Diffusion Network for Long Dance Generation Guided by the Characteristic Dance Primitives
par: Li, Ronghui, et autres
Publié: (2024)
par: Li, Ronghui, et autres
Publié: (2024)
Meta-Learning Empowered Meta-Face: Personalized Speaking Style Adaptation for Audio-Driven 3D Talking Face Animation
par: Zhou, Xukun, et autres
Publié: (2024)
par: Zhou, Xukun, et autres
Publié: (2024)
FürElise: Capturing and Physically Synthesizing Hand Motions of Piano Performance
par: Wang, Ruocheng, et autres
Publié: (2024)
par: Wang, Ruocheng, et autres
Publié: (2024)
Gaunt coefficients for complex and real spherical harmonics with applications to spherical array processing and Ambisonics
par: Politis, Archontis
Publié: (2024)
par: Politis, Archontis
Publié: (2024)
LAV: Audio-Driven Dynamic Visual Generation with Neural Compression and StyleGAN2
par: Jung, Jongmin, et autres
Publié: (2025)
par: Jung, Jongmin, et autres
Publié: (2025)
Audio-Plane: Audio Factorization Plane Gaussian Splatting for Real-Time Talking Head Synthesis
par: Shen, Shuai, et autres
Publié: (2025)
par: Shen, Shuai, et autres
Publié: (2025)
Integrating Representational Gestures into Automatically Generated Embodied Explanations and its Effects on Understanding and Interaction Quality
par: Robrecht, Amelie Sophie, et autres
Publié: (2024)
par: Robrecht, Amelie Sophie, et autres
Publié: (2024)
TCDiff++: An End-to-end Trajectory-Controllable Diffusion Model for Harmonious Music-Driven Group Choreography
par: Dai, Yuqin, et autres
Publié: (2025)
par: Dai, Yuqin, et autres
Publié: (2025)
Speech-driven Personalized Gesture Synthetics: Harnessing Automatic Fuzzy Feature Inference
par: Zhang, Fan, et autres
Publié: (2024)
par: Zhang, Fan, et autres
Publié: (2024)
MIDGET: Music Conditioned 3D Dance Generation
par: Wang, Jinwu, et autres
Publié: (2024)
par: Wang, Jinwu, et autres
Publié: (2024)
Audio is all in one: speech-driven gesture synthetics using WavLM pre-trained model
par: Zhang, Fan, et autres
Publié: (2023)
par: Zhang, Fan, et autres
Publié: (2023)
PointTalk: Audio-Driven Dynamic Lip Point Cloud for 3D Gaussian-based Talking Head Synthesis
par: Xie, Yifan, et autres
Publié: (2024)
par: Xie, Yifan, et autres
Publié: (2024)
Open Your Ears and Take a Look: A State-of-the-Art Report on the Integration of Sonification and Visualization
par: Enge, Kajetan, et autres
Publié: (2024)
par: Enge, Kajetan, et autres
Publié: (2024)
Tiny is not small enough: High-quality, low-resource facial animation models through hybrid knowledge distillation
par: Han, Zhen, et autres
Publié: (2025)
par: Han, Zhen, et autres
Publié: (2025)
A Comprehensive Multi-scale Approach for Speech and Dynamics Synchrony in Talking Head Generation
par: Airale, Louis, et autres
Publié: (2023)
par: Airale, Louis, et autres
Publié: (2023)
Documents similaires
-
Conditional GAN for Enhancing Diffusion Models in Efficient and Authentic Global Gesture Generation from Audios
par: Cheng, Yongkang, et autres
Publié: (2024) -
Freetalker: Controllable Speech and Text-Driven Gesture Generation Based on Diffusion Models for Enhanced Speaker Naturalness
par: Yang, Sicheng, et autres
Publié: (2024) -
Beat-It: Beat-Synchronized Multi-Condition 3D Dance Generation
par: Huang, Zikai, et autres
Publié: (2024) -
Inter-Diffusion Generation Model of Speakers and Listeners for Effective Communication
par: Huang, Jinhe, et autres
Publié: (2025) -
ExpGest: Expressive Speaker Generation Using Diffusion Model and Hybrid Audio-Text Guidance
par: Cheng, Yongkang, et autres
Publié: (2024)