DiffTED: One-shot Audio-driven TED Talk Video Generation with Diffusion-based Co-speech Gestures
Fuente:
arXiv
Saved in:
| Main Authors: | Hogue, Steven, Zhang, Chenxu, Daruger, Hamza, Tian, Yapeng, Guo, Xiaohu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters
by: Hogue, Steven, et al.
Published: (2024)
by: Hogue, Steven, et al.
Published: (2024)
TED-VITON: Transformer-Empowered Diffusion Models for Virtual Try-On
by: Wan, Zhenchen, et al.
Published: (2024)
by: Wan, Zhenchen, et al.
Published: (2024)
GRAN-TED: Generating Robust, Aligned, and Nuanced Text Embedding for Diffusion Models
by: Li, Bozhou, et al.
Published: (2025)
by: Li, Bozhou, et al.
Published: (2025)
StyleTalker: One-shot Style-based Audio-driven Talking Head Video Generation
by: Min, Dongchan, et al.
Published: (2022)
by: Min, Dongchan, et al.
Published: (2022)
New trends in knowledge dissemination: TED Talks
by: Giuseppina Scotto di Carlo
Published: (2014)
by: Giuseppina Scotto di Carlo
Published: (2014)
TED-4DGS: Temporally Activated and Embedding-based Deformation for 4DGS Compression
by: Ho, Cheng-Yuan, et al.
Published: (2025)
by: Ho, Cheng-Yuan, et al.
Published: (2025)
Co$^{3}$Gesture: Towards Coherent Concurrent Co-speech 3D Gesture Generation with Interactive Diffusion
by: Qi, Xingqun, et al.
Published: (2025)
by: Qi, Xingqun, et al.
Published: (2025)
AV-DiT: Efficient Audio-Visual Diffusion Transformer for Joint Audio and Video Generation
by: Wang, Kai, et al.
Published: (2024)
by: Wang, Kai, et al.
Published: (2024)
CoCoGesture: Toward Coherent Co-speech 3D Gesture Generation in the Wild
by: Qi, Xingqun, et al.
Published: (2024)
by: Qi, Xingqun, et al.
Published: (2024)
Computational Analysis of Speech Clarity Predicts Audience Engagement in TED Talks
by: Segal, Roni, et al.
Published: (2026)
by: Segal, Roni, et al.
Published: (2026)
Robust Active Speaker Detection in Noisy Environments
by: Vasireddy, Siva Sai Nagender, et al.
Published: (2024)
by: Vasireddy, Siva Sai Nagender, et al.
Published: (2024)
Co-speech Gesture Video Generation via Motion-Based Graph Retrieval
by: Song, Yafei, et al.
Published: (2025)
by: Song, Yafei, et al.
Published: (2025)
NeRFFaceSpeech: One-shot Audio-driven 3D Talking Head Synthesis via Generative Prior
by: Kim, Gihoon, et al.
Published: (2024)
by: Kim, Gihoon, et al.
Published: (2024)
TED: Accelerate Model Training by Internal Generalization
by: Xiao, Jinying, et al.
Published: (2024)
by: Xiao, Jinying, et al.
Published: (2024)
Text-to-Audio Generation Synchronized with Videos
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
OmniTalker: One-shot Real-time Text-Driven Talking Audio-Video Generation With Multimodal Style Mimicking
by: Wang, Zhongjian, et al.
Published: (2025)
by: Wang, Zhongjian, et al.
Published: (2025)
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation
by: Kushwaha, Saksham Singh, et al.
Published: (2024)
by: Kushwaha, Saksham Singh, et al.
Published: (2024)
Self-Supervised Learning of Deviation in Latent Representation for Co-speech Gesture Video Generation
by: Yang, Huan, et al.
Published: (2024)
by: Yang, Huan, et al.
Published: (2024)
InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing
by: Yang, Shaoshu, et al.
Published: (2025)
by: Yang, Shaoshu, et al.
Published: (2025)
Understanding Co-speech Gestures in-the-wild
by: Hegde, Sindhu B, et al.
Published: (2025)
by: Hegde, Sindhu B, et al.
Published: (2025)
ATL-Diff: Audio-Driven Talking Head Generation with Early Landmarks-Guide Noise Diffusion
by: Vo, Hoang-Son, et al.
Published: (2025)
by: Vo, Hoang-Son, et al.
Published: (2025)
Dimitra: Audio-driven Diffusion model for Expressive Talking Head Generation
by: Chopin, Baptiste, et al.
Published: (2025)
by: Chopin, Baptiste, et al.
Published: (2025)
Scaling Diffusion Mamba with Bidirectional SSMs for Efficient Image and Video Generation
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
HoloGest: Decoupled Diffusion and Motion Priors for Generating Holisticly Expressive Co-speech Gestures
by: Cheng, Yongkang, et al.
Published: (2025)
by: Cheng, Yongkang, et al.
Published: (2025)
EmotiveTalk: Expressive Talking Head Generation through Audio Information Decoupling and Emotional Video Diffusion
by: Wang, Haotian, et al.
Published: (2024)
by: Wang, Haotian, et al.
Published: (2024)
TANGO: Co-Speech Gesture Video Reenactment with Hierarchical Audio Motion Embedding and Diffusion Interpolation
by: Liu, Haiyang, et al.
Published: (2024)
by: Liu, Haiyang, et al.
Published: (2024)
Cosh-DiT: Co-Speech Gesture Video Synthesis via Hybrid Audio-Visual Diffusion Transformers
by: Sun, Yasheng, et al.
Published: (2025)
by: Sun, Yasheng, et al.
Published: (2025)
Weakly-Supervised Emotion Transition Learning for Diverse 3D Co-speech Gesture Generation
by: Qi, Xingqun, et al.
Published: (2023)
by: Qi, Xingqun, et al.
Published: (2023)
xTED: Cross-Domain Adaptation via Diffusion-Based Trajectory Editing
by: Niu, Haoyi, et al.
Published: (2024)
by: Niu, Haoyi, et al.
Published: (2024)
TED: Training-Free Experience Distillation for Multimodal Reasoning
by: Yuan, Shuozhi, et al.
Published: (2026)
by: Yuan, Shuozhi, et al.
Published: (2026)
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text
by: Pian, Weiguo, et al.
Published: (2026)
by: Pian, Weiguo, et al.
Published: (2026)
DiffMM: Efficient Method for Accurate Noisy and Sparse Trajectory Map Matching via One Step Diffusion
by: Han, Chenxu, et al.
Published: (2026)
by: Han, Chenxu, et al.
Published: (2026)
Audio-driven Gesture Generation via Deviation Feature in the Latent Space
by: Chen, Jiahui, et al.
Published: (2025)
by: Chen, Jiahui, et al.
Published: (2025)
Zero-shot Prompt-based Video Encoder for Surgical Gesture Recognition
by: Rao, Mingxing, et al.
Published: (2024)
by: Rao, Mingxing, et al.
Published: (2024)
From Waveforms to Pixels: A Survey on Audio-Visual Segmentation
by: Li, Jia, et al.
Published: (2025)
by: Li, Jia, et al.
Published: (2025)
TempoSyncDiff: Distilled Temporally-Consistent Diffusion for Low-Latency Audio-Driven Talking Head Generation
by: Mazumdar, Soumya, et al.
Published: (2026)
by: Mazumdar, Soumya, et al.
Published: (2026)
SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers
by: Fei, Zhengcong, et al.
Published: (2025)
by: Fei, Zhengcong, et al.
Published: (2025)
Audio-driven Talking Face Generation with Stabilized Synchronization Loss
by: Yaman, Dogucan, et al.
Published: (2023)
by: Yaman, Dogucan, et al.
Published: (2023)
Motion-example-controlled Co-speech Gesture Generation Leveraging Large Language Models
by: Chen, Bohong, et al.
Published: (2025)
by: Chen, Bohong, et al.
Published: (2025)
SignDiff: Diffusion Model for American Sign Language Production
by: Fang, Sen, et al.
Published: (2023)
by: Fang, Sen, et al.
Published: (2023)
Similar Items
-
Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters
by: Hogue, Steven, et al.
Published: (2024) -
TED-VITON: Transformer-Empowered Diffusion Models for Virtual Try-On
by: Wan, Zhenchen, et al.
Published: (2024) -
GRAN-TED: Generating Robust, Aligned, and Nuanced Text Embedding for Diffusion Models
by: Li, Bozhou, et al.
Published: (2025) -
StyleTalker: One-shot Style-based Audio-driven Talking Head Video Generation
by: Min, Dongchan, et al.
Published: (2022) -
New trends in knowledge dissemination: TED Talks
by: Giuseppina Scotto di Carlo
Published: (2014)