Democratizing High-Fidelity Co-Speech Gesture Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Xu, Huang, Shaoli, Xie, Shenbo, Chen, Xuelin, Liu, Yifei, Ding, Changxing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Contextual Gesture: Co-Speech Gesture Video Generation through Context-aware Gesture Representation
by: Liu, Pinxin, et al.
Published: (2025)
by: Liu, Pinxin, et al.
Published: (2025)
Texture-Preserving Diffusion Models for High-Fidelity Virtual Try-On
by: Yang, Xu, et al.
Published: (2024)
by: Yang, Xu, et al.
Published: (2024)
EmotionGesture: Audio-Driven Diverse Emotional Co-Speech 3D Gesture Generation
by: Qi, Xingqun, et al.
Published: (2023)
by: Qi, Xingqun, et al.
Published: (2023)
Intentional Gesture: Deliver Your Intentions with Gestures for Speech
by: Liu, Pinxin, et al.
Published: (2025)
by: Liu, Pinxin, et al.
Published: (2025)
Taming Diffusion Probabilistic Models for Character Control
by: Chen, Rui, et al.
Published: (2024)
by: Chen, Rui, et al.
Published: (2024)
DIDiffGes: Decoupled Semi-Implicit Diffusion Models for Real-time Gesture Generation from Speech
by: Cheng, Yongkang, et al.
Published: (2025)
by: Cheng, Yongkang, et al.
Published: (2025)
Coordinating Multiple Conditions for Trajectory-Controlled Human Motion Generation
by: Cai, Deli, et al.
Published: (2026)
by: Cai, Deli, et al.
Published: (2026)
Guiding Human-Object Interactions with Rich Geometry and Relations
by: Xue, Mengqing, et al.
Published: (2025)
by: Xue, Mengqing, et al.
Published: (2025)
HoloGest: Decoupled Diffusion and Motion Priors for Generating Holisticly Expressive Co-speech Gestures
by: Cheng, Yongkang, et al.
Published: (2025)
by: Cheng, Yongkang, et al.
Published: (2025)
Towards Variable and Coordinated Holistic Co-Speech Motion Generation
by: Liu, Yifei, et al.
Published: (2024)
by: Liu, Yifei, et al.
Published: (2024)
Conditional GAN for Enhancing Diffusion Models in Efficient and Authentic Global Gesture Generation from Audios
by: Cheng, Yongkang, et al.
Published: (2024)
by: Cheng, Yongkang, et al.
Published: (2024)
MagDiff: Multi-Alignment Diffusion for High-Fidelity Video Generation and Editing
by: Zhao, Haoyu, et al.
Published: (2023)
by: Zhao, Haoyu, et al.
Published: (2023)
TeleBoost: A Systematic Alignment Framework for High-Fidelity, Controllable, and Robust Video Generation
by: Liang, Yuanzhi, et al.
Published: (2026)
by: Liang, Yuanzhi, et al.
Published: (2026)
CoLMDriver: LLM-based Negotiation Benefits Cooperative Autonomous Driving
by: Liu, Changxing, et al.
Published: (2025)
by: Liu, Changxing, et al.
Published: (2025)
Leveraging Speech for Gesture Detection in Multimodal Communication
by: Ghaleb, Esam, et al.
Published: (2024)
by: Ghaleb, Esam, et al.
Published: (2024)
Hunyuan3D 2.5: Towards High-Fidelity 3D Assets Generation with Ultimate Details
by: Lai, Zeqiang, et al.
Published: (2025)
by: Lai, Zeqiang, et al.
Published: (2025)
Effortless Active Labeling for Long-Term Test-Time Adaptation
by: Wang, Guowei, et al.
Published: (2025)
by: Wang, Guowei, et al.
Published: (2025)
From 2D Alignment to 3D Plausibility: Unifying Heterogeneous 2D Priors and Penetration-Free Diffusion for Occlusion-Robust Two-Hand Reconstruction
by: Han, Gaoge, et al.
Published: (2025)
by: Han, Gaoge, et al.
Published: (2025)
CoCoVideo: The High-Quality Commercial-Model-Based Contrastive Benchmark for AI-Generated Video Detection
by: Feng, Huidong, et al.
Published: (2026)
by: Feng, Huidong, et al.
Published: (2026)
This&That: Language-Gesture Controlled Video Generation for Robot Planning
by: Wang, Boyang, et al.
Published: (2024)
by: Wang, Boyang, et al.
Published: (2024)
Towards Unified Co-Speech Gesture Generation via Hierarchical Implicit Periodicity Learning
by: Guo, Xin, et al.
Published: (2025)
by: Guo, Xin, et al.
Published: (2025)
Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events
by: Liu, Xiaolin, et al.
Published: (2026)
by: Liu, Xiaolin, et al.
Published: (2026)
Phenotype-Guided Generative Model for High-Fidelity Cardiac MRI Synthesis: Advancing Pretraining and Clinical Applications
by: Li, Ziyu, et al.
Published: (2025)
by: Li, Ziyu, et al.
Published: (2025)
Flash-DMD: Towards High-Fidelity Few-Step Image Generation with Efficient Distillation and Joint Reinforcement Learning
by: Chen, Guanjie, et al.
Published: (2025)
by: Chen, Guanjie, et al.
Published: (2025)
RopeTP: Global Human Motion Recovery via Integrating Robust Pose Estimation with Diffusion Trajectory Prior
by: Liang, Mingjiang, et al.
Published: (2024)
by: Liang, Mingjiang, et al.
Published: (2024)
DIFFUMA: High-Fidelity Spatio-Temporal Video Prediction via Dual-Path Mamba and Diffusion Enhancement
by: Xie, Xinyu, et al.
Published: (2025)
by: Xie, Xinyu, et al.
Published: (2025)
Bilateral Collaboration with Large Vision-Language Models for Open Vocabulary Human-Object Interaction Detection
by: Hu, Yupeng, et al.
Published: (2025)
by: Hu, Yupeng, et al.
Published: (2025)
Video Quality Assessment for Online Processing: From Spatial to Temporal Sampling
by: Yan, Jiebin, et al.
Published: (2025)
by: Yan, Jiebin, et al.
Published: (2025)
Lightweight High-Fidelity Low-Bitrate Talking Face Compression for 3D Video Conference
by: Li, Jianglong, et al.
Published: (2026)
by: Li, Jianglong, et al.
Published: (2026)
DET-GS: Depth- and Edge-Aware Regularization for High-Fidelity 3D Gaussian Splatting
by: Huang, Zexu, et al.
Published: (2025)
by: Huang, Zexu, et al.
Published: (2025)
LiveGesture Streamable Co-Speech Gesture Generation Model
by: Saleem, Muhammad Usama, et al.
Published: (2026)
by: Saleem, Muhammad Usama, et al.
Published: (2026)
Evaluating Design Video Generation: Metrics for Compositional Fidelity
by: Deganutti, Adrienne, et al.
Published: (2026)
by: Deganutti, Adrienne, et al.
Published: (2026)
Conditional Video Generation for High-Efficiency Video Compression
by: Yi, Fangqiu, et al.
Published: (2025)
by: Yi, Fangqiu, et al.
Published: (2025)
VividFace: A Diffusion-Based Hybrid Framework for High-Fidelity Video Face Swapping
by: Shao, Hao, et al.
Published: (2024)
by: Shao, Hao, et al.
Published: (2024)
Generative AI-Driven High-Fidelity Human Motion Simulation
by: Iyer, Hari, et al.
Published: (2025)
by: Iyer, Hari, et al.
Published: (2025)
DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers
by: Wang, Lizhen, et al.
Published: (2025)
by: Wang, Lizhen, et al.
Published: (2025)
Factorized-Dreamer: Training A High-Quality Video Generator with Limited and Low-Quality Data
by: Yang, Tao, et al.
Published: (2024)
by: Yang, Tao, et al.
Published: (2024)
DreamText: High Fidelity Scene Text Synthesis
by: Wang, Yibin, et al.
Published: (2024)
by: Wang, Yibin, et al.
Published: (2024)
Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion Models
by: Huang, Zitong, et al.
Published: (2026)
by: Huang, Zitong, et al.
Published: (2026)
Ultra3D: Efficient and High-Fidelity 3D Generation with Part Attention
by: Chen, Yiwen, et al.
Published: (2025)
by: Chen, Yiwen, et al.
Published: (2025)
Similar Items
-
Contextual Gesture: Co-Speech Gesture Video Generation through Context-aware Gesture Representation
by: Liu, Pinxin, et al.
Published: (2025) -
Texture-Preserving Diffusion Models for High-Fidelity Virtual Try-On
by: Yang, Xu, et al.
Published: (2024) -
EmotionGesture: Audio-Driven Diverse Emotional Co-Speech 3D Gesture Generation
by: Qi, Xingqun, et al.
Published: (2023) -
Intentional Gesture: Deliver Your Intentions with Gestures for Speech
by: Liu, Pinxin, et al.
Published: (2025) -
Taming Diffusion Probabilistic Models for Character Control
by: Chen, Rui, et al.
Published: (2024)