Audio-driven Gesture Generation via Deviation Feature in the Latent Space
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Jiahui, Huan, Yang, Shi, Runhua, Ding, Chanfan, Mo, Xiaoqi, Xiong, Siyu, He, Yinong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Self-Supervised Learning of Deviation in Latent Representation for Co-speech Gesture Video Generation
by: Yang, Huan, et al.
Published: (2024)
by: Yang, Huan, et al.
Published: (2024)
YingVideo-MV: Music-Driven Multi-Stage Video Generation
by: Chen, Jiahui, et al.
Published: (2025)
by: Chen, Jiahui, et al.
Published: (2025)
Generalizing to Out-of-Sample Degradations via Model Reprogramming
by: Jiang, Runhua, et al.
Published: (2024)
by: Jiang, Runhua, et al.
Published: (2024)
EMAGE: Towards Unified Holistic Co-Speech Gesture Generation via Expressive Masked Audio Gesture Modeling
by: Liu, Haiyang, et al.
Published: (2023)
by: Liu, Haiyang, et al.
Published: (2023)
DiffTED: One-shot Audio-driven TED Talk Video Generation with Diffusion-based Co-speech Gestures
by: Hogue, Steven, et al.
Published: (2024)
by: Hogue, Steven, et al.
Published: (2024)
GestureLSM: Latent Shortcut based Co-Speech Gesture Generation with Spatial-Temporal Modeling
by: Liu, Pinxin, et al.
Published: (2025)
by: Liu, Pinxin, et al.
Published: (2025)
Autoregressive Image Generation Needs Only a Few Lines of Cached Tokens
by: Qin, Ziran, et al.
Published: (2025)
by: Qin, Ziran, et al.
Published: (2025)
Latent Transfer Attack: Adversarial Examples via Generative Latent Spaces
by: Shaar, Eitan, et al.
Published: (2026)
by: Shaar, Eitan, et al.
Published: (2026)
Hand2World: Autoregressive Egocentric Interaction Generation via Free-Space Hand Gestures
by: Wang, Yuxi, et al.
Published: (2026)
by: Wang, Yuxi, et al.
Published: (2026)
EmotionGesture: Audio-Driven Diverse Emotional Co-Speech 3D Gesture Generation
by: Qi, Xingqun, et al.
Published: (2023)
by: Qi, Xingqun, et al.
Published: (2023)
LiveGesture Streamable Co-Speech Gesture Generation Model
by: Saleem, Muhammad Usama, et al.
Published: (2026)
by: Saleem, Muhammad Usama, et al.
Published: (2026)
Cosh-DiT: Co-Speech Gesture Video Synthesis via Hybrid Audio-Visual Diffusion Transformers
by: Sun, Yasheng, et al.
Published: (2025)
by: Sun, Yasheng, et al.
Published: (2025)
CoordSpeaker: Exploiting Gesture Captioning for Coordinated Caption-Empowered Co-Speech Gesture Generation
by: Fang, Fengyi, et al.
Published: (2025)
by: Fang, Fengyi, et al.
Published: (2025)
EmoLat: Text-driven Image Sentiment Transfer via Emotion Latent Space
by: Zhang, Jing, et al.
Published: (2026)
by: Zhang, Jing, et al.
Published: (2026)
Robust Dreamer: Deviation-Aware Latent Gaussian Memory for Action-Controlled AR Video Generation
by: Chen, Hanlin, et al.
Published: (2026)
by: Chen, Hanlin, et al.
Published: (2026)
InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing
by: Yang, Shaoshu, et al.
Published: (2025)
by: Yang, Shaoshu, et al.
Published: (2025)
LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning
by: Dai, Yifan, et al.
Published: (2026)
by: Dai, Yifan, et al.
Published: (2026)
Narrating For You: Prompt-guided Audio-visual Narrating Face Generation Employing Multi-entangled Latent Space
by: Chandra, Aashish, et al.
Published: (2026)
by: Chandra, Aashish, et al.
Published: (2026)
EasyGenNet: An Efficient Framework for Audio-Driven Gesture Video Generation Based on Diffusion Model
by: Li, Renda, et al.
Published: (2025)
by: Li, Renda, et al.
Published: (2025)
Assessing Sample Quality via the Latent Space of Generative Models
by: Xu, Jingyi, et al.
Published: (2024)
by: Xu, Jingyi, et al.
Published: (2024)
TANGO: Co-Speech Gesture Video Reenactment with Hierarchical Audio Motion Embedding and Diffusion Interpolation
by: Liu, Haiyang, et al.
Published: (2024)
by: Liu, Haiyang, et al.
Published: (2024)
Gesture Classification in Artworks Using Contextual Image Features
by: Hussian, Azhar, et al.
Published: (2024)
by: Hussian, Azhar, et al.
Published: (2024)
Text-Driven Diverse Facial Texture Generation via Progressive Latent-Space Refinement
by: Wang, Chi, et al.
Published: (2024)
by: Wang, Chi, et al.
Published: (2024)
Prompt-to-Gesture: Measuring the Capabilities of Image-to-Video Deictic Gesture Generation
by: Ali, Hassan, et al.
Published: (2026)
by: Ali, Hassan, et al.
Published: (2026)
MPF-Net: Exposing High-Fidelity AI-Generated Video Forgeries via Hierarchical Manifold Deviation and Micro-Temporal Fluctuations
by: He, Xinan, et al.
Published: (2026)
by: He, Xinan, et al.
Published: (2026)
WavFlow: Audio Generation in Waveform Space
by: Zhou, Feiyan, et al.
Published: (2026)
by: Zhou, Feiyan, et al.
Published: (2026)
SwapTalk: Audio-Driven Talking Face Generation with One-Shot Customization in Latent Space
by: Zhang, Zeren, et al.
Published: (2024)
by: Zhang, Zeren, et al.
Published: (2024)
Democratizing High-Fidelity Co-Speech Gesture Video Generation
by: Yang, Xu, et al.
Published: (2025)
by: Yang, Xu, et al.
Published: (2025)
Detecting AI-Generated Forgeries via Iterative Manifold Deviation Amplification
by: Zhang, Jiangling, et al.
Published: (2026)
by: Zhang, Jiangling, et al.
Published: (2026)
InfoSculpt: Sculpting the Latent Space for Generalized Category Discovery
by: Liao, Wenwen, et al.
Published: (2026)
by: Liao, Wenwen, et al.
Published: (2026)
WeditGAN: Few-Shot Image Generation via Latent Space Relocation
by: Duan, Yuxuan, et al.
Published: (2023)
by: Duan, Yuxuan, et al.
Published: (2023)
Audio-driven Talking Face Generation with Stabilized Synchronization Loss
by: Yaman, Dogucan, et al.
Published: (2023)
by: Yaman, Dogucan, et al.
Published: (2023)
Continual Gesture Learning without Data via Synthetic Feature Sampling
by: Lu, Zhenyu, et al.
Published: (2024)
by: Lu, Zhenyu, et al.
Published: (2024)
Generative Human Motion Stylization in Latent Space
by: Guo, Chuan, et al.
Published: (2024)
by: Guo, Chuan, et al.
Published: (2024)
Enhancing Space-time Video Super-resolution via Spatial-temporal Feature Interaction
by: Yue, Zijie, et al.
Published: (2022)
by: Yue, Zijie, et al.
Published: (2022)
Detecting AI-Generated Images via Distributional Deviations from Real Images
by: Niu, Yakun, et al.
Published: (2026)
by: Niu, Yakun, et al.
Published: (2026)
ExGes: Expressive Human Motion Retrieval and Modulation for Audio-Driven Gesture Synthesis
by: Zhou, Xukun, et al.
Published: (2025)
by: Zhou, Xukun, et al.
Published: (2025)
Motion-example-controlled Co-speech Gesture Generation Leveraging Large Language Models
by: Chen, Bohong, et al.
Published: (2025)
by: Chen, Bohong, et al.
Published: (2025)
Towards Any-Quality Image Segmentation via Generative and Adaptive Latent Space Enhancement
by: Guo, Guangqian, et al.
Published: (2026)
by: Guo, Guangqian, et al.
Published: (2026)
LoomNet: Enhancing Multi-View Image Generation via Latent Space Weaving
by: Federico, Giulio, et al.
Published: (2025)
by: Federico, Giulio, et al.
Published: (2025)
Similar Items
-
Self-Supervised Learning of Deviation in Latent Representation for Co-speech Gesture Video Generation
by: Yang, Huan, et al.
Published: (2024) -
YingVideo-MV: Music-Driven Multi-Stage Video Generation
by: Chen, Jiahui, et al.
Published: (2025) -
Generalizing to Out-of-Sample Degradations via Model Reprogramming
by: Jiang, Runhua, et al.
Published: (2024) -
EMAGE: Towards Unified Holistic Co-Speech Gesture Generation via Expressive Masked Audio Gesture Modeling
by: Liu, Haiyang, et al.
Published: (2023) -
DiffTED: One-shot Audio-driven TED Talk Video Generation with Diffusion-based Co-speech Gestures
by: Hogue, Steven, et al.
Published: (2024)