EasyGenNet: An Efficient Framework for Audio-Driven Gesture Video Generation Based on Diffusion Model
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Renda, Qi, Xiaohua, Ling, Qiang, Yu, Jun, Chen, Ziyi, Chang, Peng, Xiao, Mei HanJing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Data-free Knowledge Distillation with Diffusion Models
by: Qi, Xiaohua, et al.
Published: (2025)
by: Qi, Xiaohua, et al.
Published: (2025)
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation
by: Wang, Le, et al.
Published: (2025)
by: Wang, Le, et al.
Published: (2025)
EasyControl: Transfer ControlNet to Video Diffusion for Controllable Generation and Interpolation
by: Wang, Cong, et al.
Published: (2024)
by: Wang, Cong, et al.
Published: (2024)
EasyGen: Easing Multimodal Generation with BiDiffuser and LLMs
by: Zhao, Xiangyu, et al.
Published: (2023)
by: Zhao, Xiangyu, et al.
Published: (2023)
EmotionGesture: Audio-Driven Diverse Emotional Co-Speech 3D Gesture Generation
by: Qi, Xingqun, et al.
Published: (2023)
by: Qi, Xingqun, et al.
Published: (2023)
AudioScenic: Audio-Driven Video Scene Editing
by: Shen, Kaixin, et al.
Published: (2024)
by: Shen, Kaixin, et al.
Published: (2024)
Conditional GAN for Enhancing Diffusion Models in Efficient and Authentic Global Gesture Generation from Audios
by: Cheng, Yongkang, et al.
Published: (2024)
by: Cheng, Yongkang, et al.
Published: (2024)
JoyGen: Audio-Driven 3D Depth-Aware Talking-Face Video Editing
by: Wang, Qili, et al.
Published: (2025)
by: Wang, Qili, et al.
Published: (2025)
LibraGen: Playing a Balance Game in Subject-Driven Video Generation
by: Zhu, Jiahao, et al.
Published: (2026)
by: Zhu, Jiahao, et al.
Published: (2026)
AV-DiT: Efficient Audio-Visual Diffusion Transformer for Joint Audio and Video Generation
by: Wang, Kai, et al.
Published: (2024)
by: Wang, Kai, et al.
Published: (2024)
TANGO: Co-Speech Gesture Video Reenactment with Hierarchical Audio Motion Embedding and Diffusion Interpolation
by: Liu, Haiyang, et al.
Published: (2024)
by: Liu, Haiyang, et al.
Published: (2024)
How To Make Audio/Video as Easy To Use and Share as Text.
by: Spoerri, Anselm
Published: (2002)
by: Spoerri, Anselm
Published: (2002)
Cosh-DiT: Co-Speech Gesture Video Synthesis via Hybrid Audio-Visual Diffusion Transformers
by: Sun, Yasheng, et al.
Published: (2025)
by: Sun, Yasheng, et al.
Published: (2025)
ExGes: Expressive Human Motion Retrieval and Modulation for Audio-Driven Gesture Synthesis
by: Zhou, Xukun, et al.
Published: (2025)
by: Zhou, Xukun, et al.
Published: (2025)
Conversational Gesture Model (CGM): Extending Speaker‐Centric Audio‐Driven Motion Generation to Full Conversation Gestures
by: T. Koren, et al.
Published: (2026)
by: T. Koren, et al.
Published: (2026)
SyncDiff: Diffusion-based Talking Head Synthesis with Bottlenecked Temporal Visual Prior for Improved Synchronization
by: Fan, Xulin, et al.
Published: (2025)
by: Fan, Xulin, et al.
Published: (2025)
DiffTED: One-shot Audio-driven TED Talk Video Generation with Diffusion-based Co-speech Gestures
by: Hogue, Steven, et al.
Published: (2024)
by: Hogue, Steven, et al.
Published: (2024)
Wan-S2V: Audio-Driven Cinematic Video Generation
by: Gao, Xin, et al.
Published: (2025)
by: Gao, Xin, et al.
Published: (2025)
AudioGenX: Explainability on Text-to-Audio Generative Models
by: Kang, Hyunju, et al.
Published: (2025)
by: Kang, Hyunju, et al.
Published: (2025)
Using EasyNet in Libraries.
by: Still, Julie
Published: (1991)
by: Still, Julie
Published: (1991)
U-Nets as Belief Propagation: Efficient Classification, Denoising, and Diffusion in Generative Hierarchical Models
by: Mei, Song
Published: (2024)
by: Mei, Song
Published: (2024)
Input-Aware Sparse Attention for Real-Time Co-Speech Video Generation
by: Lu, Beijia, et al.
Published: (2025)
by: Lu, Beijia, et al.
Published: (2025)
SayAnything: Audio-Driven Lip Synchronization with Conditional Video Diffusion
by: Ma, Junxian, et al.
Published: (2025)
by: Ma, Junxian, et al.
Published: (2025)
Easy-Poly: An Easy Polyhedral Framework For 3D Multi-Object Tracking
by: Zhang, Peng, et al.
Published: (2025)
by: Zhang, Peng, et al.
Published: (2025)
DC-VideoGen: Efficient Video Generation with Deep Compression Video Autoencoder
by: Chen, Junyu, et al.
Published: (2025)
by: Chen, Junyu, et al.
Published: (2025)
Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity
by: Xi, Haocheng, et al.
Published: (2025)
by: Xi, Haocheng, et al.
Published: (2025)
StableAvatar: Infinite-Length Audio-Driven Avatar Video Generation
by: Tu, Shuyuan, et al.
Published: (2025)
by: Tu, Shuyuan, et al.
Published: (2025)
Diff-VS: Efficient Audio-Aware Diffusion U-Net for Vocals Separation
by: Yun-Ning, et al.
Published: (2026)
by: Yun-Ning, et al.
Published: (2026)
No compact split limit Ricci flow of type II from the blow-down
by: Zhao, Ziyi, et al.
Published: (2024)
by: Zhao, Ziyi, et al.
Published: (2024)
$4d$ steady gradient Ricci solitons with nonnegative curvature away from a compact set
by: Zhao, Ziyi, et al.
Published: (2023)
by: Zhao, Ziyi, et al.
Published: (2023)
Steady gradient Ricci solitons with nonnegative curvature operator away from a compact set
by: Zhao, Ziyi, et al.
Published: (2024)
by: Zhao, Ziyi, et al.
Published: (2024)
EMAGE: Towards Unified Holistic Co-Speech Gesture Generation via Expressive Masked Audio Gesture Modeling
by: Liu, Haiyang, et al.
Published: (2023)
by: Liu, Haiyang, et al.
Published: (2023)
Online Databases: The EasyNet Gateway.
by: Tenopir, Carol
Published: (1986)
by: Tenopir, Carol
Published: (1986)
ProAV-DiT: A Projected Latent Diffusion Transformer for Efficient Synchronized Audio-Video Generation
by: Sun, Jiahui, et al.
Published: (2025)
by: Sun, Jiahui, et al.
Published: (2025)
Temporal Residual Guided Diffusion Framework for Event-Driven Video Reconstruction
by: Zhu, Lin, et al.
Published: (2024)
by: Zhu, Lin, et al.
Published: (2024)
AudioRWKV: Efficient and Stable Bidirectional RWKV for Audio Pattern Recognition
by: Xiong, Jiayu, et al.
Published: (2025)
by: Xiong, Jiayu, et al.
Published: (2025)
AudCast: Audio-Driven Human Video Generation by Cascaded Diffusion Transformers
by: Guan, Jiazhi, et al.
Published: (2025)
by: Guan, Jiazhi, et al.
Published: (2025)
3MDiT: Unified Tri-Modal Diffusion Transformer for Text-Driven Synchronized Audio-Video Generation
by: Li, Yaoru, et al.
Published: (2025)
by: Li, Yaoru, et al.
Published: (2025)
StarGen: A Spatiotemporal Autoregression Framework with Video Diffusion Model for Scalable and Controllable Scene Generation
by: Zhai, Shangjin, et al.
Published: (2025)
by: Zhai, Shangjin, et al.
Published: (2025)
EasyVideoR1: Easier RL for Video Understanding
by: Qin, Chuanyu, et al.
Published: (2026)
by: Qin, Chuanyu, et al.
Published: (2026)
Similar Items
-
Data-free Knowledge Distillation with Diffusion Models
by: Qi, Xiaohua, et al.
Published: (2025) -
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation
by: Wang, Le, et al.
Published: (2025) -
EasyControl: Transfer ControlNet to Video Diffusion for Controllable Generation and Interpolation
by: Wang, Cong, et al.
Published: (2024) -
EasyGen: Easing Multimodal Generation with BiDiffuser and LLMs
by: Zhao, Xiangyu, et al.
Published: (2023) -
EmotionGesture: Audio-Driven Diverse Emotional Co-Speech 3D Gesture Generation
by: Qi, Xingqun, et al.
Published: (2023)