EasyGenNet: An Efficient Framework for Audio-Driven Gesture Video Generation Based on Diffusion Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Renda, Qi, Xiaohua, Ling, Qiang, Yu, Jun, Chen, Ziyi, Chang, Peng, Xiao, Mei HanJing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Data-free Knowledge Distillation with Diffusion Models
von: Qi, Xiaohua, et al.
Veröffentlicht: (2025)
von: Qi, Xiaohua, et al.
Veröffentlicht: (2025)
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation
von: Wang, Le, et al.
Veröffentlicht: (2025)
von: Wang, Le, et al.
Veröffentlicht: (2025)
EasyControl: Transfer ControlNet to Video Diffusion for Controllable Generation and Interpolation
von: Wang, Cong, et al.
Veröffentlicht: (2024)
von: Wang, Cong, et al.
Veröffentlicht: (2024)
EasyGen: Easing Multimodal Generation with BiDiffuser and LLMs
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2023)
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2023)
EmotionGesture: Audio-Driven Diverse Emotional Co-Speech 3D Gesture Generation
von: Qi, Xingqun, et al.
Veröffentlicht: (2023)
von: Qi, Xingqun, et al.
Veröffentlicht: (2023)
AudioScenic: Audio-Driven Video Scene Editing
von: Shen, Kaixin, et al.
Veröffentlicht: (2024)
von: Shen, Kaixin, et al.
Veröffentlicht: (2024)
Conditional GAN for Enhancing Diffusion Models in Efficient and Authentic Global Gesture Generation from Audios
von: Cheng, Yongkang, et al.
Veröffentlicht: (2024)
von: Cheng, Yongkang, et al.
Veröffentlicht: (2024)
JoyGen: Audio-Driven 3D Depth-Aware Talking-Face Video Editing
von: Wang, Qili, et al.
Veröffentlicht: (2025)
von: Wang, Qili, et al.
Veröffentlicht: (2025)
LibraGen: Playing a Balance Game in Subject-Driven Video Generation
von: Zhu, Jiahao, et al.
Veröffentlicht: (2026)
von: Zhu, Jiahao, et al.
Veröffentlicht: (2026)
AV-DiT: Efficient Audio-Visual Diffusion Transformer for Joint Audio and Video Generation
von: Wang, Kai, et al.
Veröffentlicht: (2024)
von: Wang, Kai, et al.
Veröffentlicht: (2024)
TANGO: Co-Speech Gesture Video Reenactment with Hierarchical Audio Motion Embedding and Diffusion Interpolation
von: Liu, Haiyang, et al.
Veröffentlicht: (2024)
von: Liu, Haiyang, et al.
Veröffentlicht: (2024)
How To Make Audio/Video as Easy To Use and Share as Text.
von: Spoerri, Anselm
Veröffentlicht: (2002)
von: Spoerri, Anselm
Veröffentlicht: (2002)
Cosh-DiT: Co-Speech Gesture Video Synthesis via Hybrid Audio-Visual Diffusion Transformers
von: Sun, Yasheng, et al.
Veröffentlicht: (2025)
von: Sun, Yasheng, et al.
Veröffentlicht: (2025)
ExGes: Expressive Human Motion Retrieval and Modulation for Audio-Driven Gesture Synthesis
von: Zhou, Xukun, et al.
Veröffentlicht: (2025)
von: Zhou, Xukun, et al.
Veröffentlicht: (2025)
Conversational Gesture Model (CGM): Extending Speaker‐Centric Audio‐Driven Motion Generation to Full Conversation Gestures
von: T. Koren, et al.
Veröffentlicht: (2026)
von: T. Koren, et al.
Veröffentlicht: (2026)
SyncDiff: Diffusion-based Talking Head Synthesis with Bottlenecked Temporal Visual Prior for Improved Synchronization
von: Fan, Xulin, et al.
Veröffentlicht: (2025)
von: Fan, Xulin, et al.
Veröffentlicht: (2025)
DiffTED: One-shot Audio-driven TED Talk Video Generation with Diffusion-based Co-speech Gestures
von: Hogue, Steven, et al.
Veröffentlicht: (2024)
von: Hogue, Steven, et al.
Veröffentlicht: (2024)
Wan-S2V: Audio-Driven Cinematic Video Generation
von: Gao, Xin, et al.
Veröffentlicht: (2025)
von: Gao, Xin, et al.
Veröffentlicht: (2025)
AudioGenX: Explainability on Text-to-Audio Generative Models
von: Kang, Hyunju, et al.
Veröffentlicht: (2025)
von: Kang, Hyunju, et al.
Veröffentlicht: (2025)
Using EasyNet in Libraries.
von: Still, Julie
Veröffentlicht: (1991)
von: Still, Julie
Veröffentlicht: (1991)
U-Nets as Belief Propagation: Efficient Classification, Denoising, and Diffusion in Generative Hierarchical Models
von: Mei, Song
Veröffentlicht: (2024)
von: Mei, Song
Veröffentlicht: (2024)
Input-Aware Sparse Attention for Real-Time Co-Speech Video Generation
von: Lu, Beijia, et al.
Veröffentlicht: (2025)
von: Lu, Beijia, et al.
Veröffentlicht: (2025)
SayAnything: Audio-Driven Lip Synchronization with Conditional Video Diffusion
von: Ma, Junxian, et al.
Veröffentlicht: (2025)
von: Ma, Junxian, et al.
Veröffentlicht: (2025)
Easy-Poly: An Easy Polyhedral Framework For 3D Multi-Object Tracking
von: Zhang, Peng, et al.
Veröffentlicht: (2025)
von: Zhang, Peng, et al.
Veröffentlicht: (2025)
DC-VideoGen: Efficient Video Generation with Deep Compression Video Autoencoder
von: Chen, Junyu, et al.
Veröffentlicht: (2025)
von: Chen, Junyu, et al.
Veröffentlicht: (2025)
Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity
von: Xi, Haocheng, et al.
Veröffentlicht: (2025)
von: Xi, Haocheng, et al.
Veröffentlicht: (2025)
StableAvatar: Infinite-Length Audio-Driven Avatar Video Generation
von: Tu, Shuyuan, et al.
Veröffentlicht: (2025)
von: Tu, Shuyuan, et al.
Veröffentlicht: (2025)
Diff-VS: Efficient Audio-Aware Diffusion U-Net for Vocals Separation
von: Yun-Ning, et al.
Veröffentlicht: (2026)
von: Yun-Ning, et al.
Veröffentlicht: (2026)
No compact split limit Ricci flow of type II from the blow-down
von: Zhao, Ziyi, et al.
Veröffentlicht: (2024)
von: Zhao, Ziyi, et al.
Veröffentlicht: (2024)
$4d$ steady gradient Ricci solitons with nonnegative curvature away from a compact set
von: Zhao, Ziyi, et al.
Veröffentlicht: (2023)
von: Zhao, Ziyi, et al.
Veröffentlicht: (2023)
Steady gradient Ricci solitons with nonnegative curvature operator away from a compact set
von: Zhao, Ziyi, et al.
Veröffentlicht: (2024)
von: Zhao, Ziyi, et al.
Veröffentlicht: (2024)
EMAGE: Towards Unified Holistic Co-Speech Gesture Generation via Expressive Masked Audio Gesture Modeling
von: Liu, Haiyang, et al.
Veröffentlicht: (2023)
von: Liu, Haiyang, et al.
Veröffentlicht: (2023)
Online Databases: The EasyNet Gateway.
von: Tenopir, Carol
Veröffentlicht: (1986)
von: Tenopir, Carol
Veröffentlicht: (1986)
ProAV-DiT: A Projected Latent Diffusion Transformer for Efficient Synchronized Audio-Video Generation
von: Sun, Jiahui, et al.
Veröffentlicht: (2025)
von: Sun, Jiahui, et al.
Veröffentlicht: (2025)
Temporal Residual Guided Diffusion Framework for Event-Driven Video Reconstruction
von: Zhu, Lin, et al.
Veröffentlicht: (2024)
von: Zhu, Lin, et al.
Veröffentlicht: (2024)
AudioRWKV: Efficient and Stable Bidirectional RWKV for Audio Pattern Recognition
von: Xiong, Jiayu, et al.
Veröffentlicht: (2025)
von: Xiong, Jiayu, et al.
Veröffentlicht: (2025)
AudCast: Audio-Driven Human Video Generation by Cascaded Diffusion Transformers
von: Guan, Jiazhi, et al.
Veröffentlicht: (2025)
von: Guan, Jiazhi, et al.
Veröffentlicht: (2025)
3MDiT: Unified Tri-Modal Diffusion Transformer for Text-Driven Synchronized Audio-Video Generation
von: Li, Yaoru, et al.
Veröffentlicht: (2025)
von: Li, Yaoru, et al.
Veröffentlicht: (2025)
StarGen: A Spatiotemporal Autoregression Framework with Video Diffusion Model for Scalable and Controllable Scene Generation
von: Zhai, Shangjin, et al.
Veröffentlicht: (2025)
von: Zhai, Shangjin, et al.
Veröffentlicht: (2025)
EasyVideoR1: Easier RL for Video Understanding
von: Qin, Chuanyu, et al.
Veröffentlicht: (2026)
von: Qin, Chuanyu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Data-free Knowledge Distillation with Diffusion Models
von: Qi, Xiaohua, et al.
Veröffentlicht: (2025) -
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation
von: Wang, Le, et al.
Veröffentlicht: (2025) -
EasyControl: Transfer ControlNet to Video Diffusion for Controllable Generation and Interpolation
von: Wang, Cong, et al.
Veröffentlicht: (2024) -
EasyGen: Easing Multimodal Generation with BiDiffuser and LLMs
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2023) -
EmotionGesture: Audio-Driven Diverse Emotional Co-Speech 3D Gesture Generation
von: Qi, Xingqun, et al.
Veröffentlicht: (2023)