SpeechAct: Towards Generating Whole-body Motion from Speech
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Jinsong, Zhu, Minjie, Zhang, Yuxiang, Liu, Yebin, Li, Kun |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PyMAF-X: Towards Well-aligned Full-body Model Regression from Monocular Images
by: Zhang, Hongwen, et al.
Published: (2022)
by: Zhang, Hongwen, et al.
Published: (2022)
MMGT: Motion Mask Guided Two-Stage Network for Co-Speech Gesture Video Generation
by: Wang, Siyuan, et al.
Published: (2025)
by: Wang, Siyuan, et al.
Published: (2025)
Towards Variable and Coordinated Holistic Co-Speech Motion Generation
by: Liu, Yifei, et al.
Published: (2024)
by: Liu, Yifei, et al.
Published: (2024)
Expressive Forecasting of 3D Whole-body Human Motions
by: Ding, Pengxiang, et al.
Published: (2023)
by: Ding, Pengxiang, et al.
Published: (2023)
Enabling Synergistic Full-Body Control in Prompt-Based Co-Speech Motion Generation
by: Chen, Bohong, et al.
Published: (2024)
by: Chen, Bohong, et al.
Published: (2024)
EchoMask: Speech-Queried Attention-based Mask Modeling for Holistic Co-Speech Motion Generation
by: Zhang, Xiangyue, et al.
Published: (2025)
by: Zhang, Xiangyue, et al.
Published: (2025)
Motion-X++: A Large-Scale Multimodal 3D Whole-body Human Motion Dataset
by: Zhang, Yuhong, et al.
Published: (2025)
by: Zhang, Yuhong, et al.
Published: (2025)
Motion-X: A Large-scale 3D Expressive Whole-body Human Motion Dataset
by: Lin, Jing, et al.
Published: (2023)
by: Lin, Jing, et al.
Published: (2023)
Mitigating Error Accumulation in Co-Speech Motion Generation via Global Rotation Diffusion and Multi-Level Constraints
by: Zhang, Xiangyue, et al.
Published: (2025)
by: Zhang, Xiangyue, et al.
Published: (2025)
FOF-X: Towards Real-time Detailed Human Reconstruction from a Single Image
by: Feng, Qiao, et al.
Published: (2024)
by: Feng, Qiao, et al.
Published: (2024)
HHMR: Holistic Hand Mesh Recovery by Enhancing the Multimodal Controllability of Graph Diffusion Models
by: Li, Mengcheng, et al.
Published: (2024)
by: Li, Mengcheng, et al.
Published: (2024)
ManiDext: Hand-Object Manipulation Synthesis via Continuous Correspondence Embeddings and Residual-Guided Diffusion
by: Zhang, Jiajun, et al.
Published: (2024)
by: Zhang, Jiajun, et al.
Published: (2024)
Holistic-Motion2D: Scalable Whole-body Human Motion Generation in 2D Space
by: Wang, Yuan, et al.
Published: (2024)
by: Wang, Yuan, et al.
Published: (2024)
SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios
by: Dang, Lingwei, et al.
Published: (2025)
by: Dang, Lingwei, et al.
Published: (2025)
T3M: Text Guided 3D Human Motion Synthesis from Speech
by: Peng, Wenshuo, et al.
Published: (2024)
by: Peng, Wenshuo, et al.
Published: (2024)
Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation
by: Cheng, Shihao, et al.
Published: (2026)
by: Cheng, Shihao, et al.
Published: (2026)
Preview WB-DH: Towards Whole Body Digital Human Bench for the Generation of Whole-body Talking Avatar Videos
by: Wang, Chaoyi, et al.
Published: (2025)
by: Wang, Chaoyi, et al.
Published: (2025)
A Plug-and-Play Physical Motion Restoration Approach for In-the-Wild High-Difficulty Motions
by: Zhang, Youliang, et al.
Published: (2024)
by: Zhang, Youliang, et al.
Published: (2024)
AnyAct: Towards Human Reenactment of Character Motion From Video
by: Chen, Liuhan, et al.
Published: (2026)
by: Chen, Liuhan, et al.
Published: (2026)
OmniMotion-X: Versatile Multimodal Whole-Body Motion Generation
by: Xu, Guowei, et al.
Published: (2025)
by: Xu, Guowei, et al.
Published: (2025)
SyncMV4D: Synchronized Multi-view Joint Diffusion of Appearance and Motion for Hand-Object Interaction Synthesis
by: Dang, Lingwei, et al.
Published: (2025)
by: Dang, Lingwei, et al.
Published: (2025)
MotionCraft: Crafting Whole-Body Motion with Plug-and-Play Multimodal Controls
by: Bian, Yuxuan, et al.
Published: (2024)
by: Bian, Yuxuan, et al.
Published: (2024)
4DEquine: Disentangling Motion and Appearance for 4D Equine Reconstruction from Monocular Video
by: Lyu, Jin, et al.
Published: (2026)
by: Lyu, Jin, et al.
Published: (2026)
TACO: Benchmarking Generalizable Bimanual Tool-ACtion-Object Understanding
by: Liu, Yun, et al.
Published: (2024)
by: Liu, Yun, et al.
Published: (2024)
M3Bench: Benchmarking Whole-body Motion Generation for Mobile Manipulation in 3D Scenes
by: Zhang, Zeyu, et al.
Published: (2024)
by: Zhang, Zeyu, et al.
Published: (2024)
UniAct: Unified Motion Generation and Action Streaming for Humanoid Robots
by: Jiang, Nan, et al.
Published: (2025)
by: Jiang, Nan, et al.
Published: (2025)
Layered 3D Human Generation via Semantic-Aware Diffusion Model
by: Wang, Yi, et al.
Published: (2023)
by: Wang, Yi, et al.
Published: (2023)
Co-Speech Gesture Video Generation via Motion-Decoupled Diffusion Model
by: He, Xu, et al.
Published: (2024)
by: He, Xu, et al.
Published: (2024)
Ins-HOI: Instance Aware Human-Object Interactions Recovery
by: Zhang, Jiajun, et al.
Published: (2023)
by: Zhang, Jiajun, et al.
Published: (2023)
SpeechForensics: Audio-Visual Speech Representation Learning for Face Forgery Detection
by: Liang, Yachao, et al.
Published: (2025)
by: Liang, Yachao, et al.
Published: (2025)
Lodge++: High-quality and Long Dance Generation with Vivid Choreography Patterns
by: Li, Ronghui, et al.
Published: (2024)
by: Li, Ronghui, et al.
Published: (2024)
OmniHands: Towards Robust 4D Hand Mesh Recovery via A Versatile Transformer
by: Lin, Dixuan, et al.
Published: (2024)
by: Lin, Dixuan, et al.
Published: (2024)
TANGO: Co-Speech Gesture Video Reenactment with Hierarchical Audio Motion Embedding and Diffusion Interpolation
by: Liu, Haiyang, et al.
Published: (2024)
by: Liu, Haiyang, et al.
Published: (2024)
KMTalk: Speech-Driven 3D Facial Animation with Key Motion Embedding
by: Xu, Zhihao, et al.
Published: (2024)
by: Xu, Zhihao, et al.
Published: (2024)
Learning Disentangled Speech- and Expression-Driven Blendshapes for 3D Talking Face Animation
by: Mao, Yuxiang, et al.
Published: (2025)
by: Mao, Yuxiang, et al.
Published: (2025)
EMAGE: Towards Unified Holistic Co-Speech Gesture Generation via Expressive Masked Audio Gesture Modeling
by: Liu, Haiyang, et al.
Published: (2023)
by: Liu, Haiyang, et al.
Published: (2023)
Human as Points: Explicit Point-based 3D Human Reconstruction from Single-view RGB Images
by: Tang, Yingzhi, et al.
Published: (2023)
by: Tang, Yingzhi, et al.
Published: (2023)
HOSIG: Full-Body Human-Object-Scene Interaction Generation with Hierarchical Scene Perception
by: Yao, Wei, et al.
Published: (2025)
by: Yao, Wei, et al.
Published: (2025)
X-NeMo: Expressive Neural Motion Reenactment via Disentangled Latent Attention
by: Zhao, Xiaochen, et al.
Published: (2025)
by: Zhao, Xiaochen, et al.
Published: (2025)
MimicParts: Part-aware Style Injection for Speech-Driven 3D Motion Generation
by: Liu, Lianlian, et al.
Published: (2025)
by: Liu, Lianlian, et al.
Published: (2025)
Similar Items
-
PyMAF-X: Towards Well-aligned Full-body Model Regression from Monocular Images
by: Zhang, Hongwen, et al.
Published: (2022) -
MMGT: Motion Mask Guided Two-Stage Network for Co-Speech Gesture Video Generation
by: Wang, Siyuan, et al.
Published: (2025) -
Towards Variable and Coordinated Holistic Co-Speech Motion Generation
by: Liu, Yifei, et al.
Published: (2024) -
Expressive Forecasting of 3D Whole-body Human Motions
by: Ding, Pengxiang, et al.
Published: (2023) -
Enabling Synergistic Full-Body Control in Prompt-Based Co-Speech Motion Generation
by: Chen, Bohong, et al.
Published: (2024)