The Quest for Generalizable Motion Generation: Data, Model, and Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Lin, Jing, Wang, Ruisi, Lu, Junzhe, Huang, Ziqi, Song, Guorui, Zeng, Ailing, Liu, Xian, Wei, Chen, Yin, Wanqi, Sun, Qingping, Cai, Zhongang, Yang, Lei, Liu, Ziwei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AiOS: All-in-One-Stage Expressive Human Pose and Shape Estimation
by: Sun, Qingping, et al.
Published: (2024)
by: Sun, Qingping, et al.
Published: (2024)
DPoser-X: Diffusion Model as Robust 3D Whole-body Human Pose Prior
by: Lu, Junzhe, et al.
Published: (2025)
by: Lu, Junzhe, et al.
Published: (2025)
SMPLest-X: Ultimate Scaling for Expressive Human Pose and Shape Estimation
by: Yin, Wanqi, et al.
Published: (2025)
by: Yin, Wanqi, et al.
Published: (2025)
WHAC: World-grounded Humans and Cameras
by: Yin, Wanqi, et al.
Published: (2024)
by: Yin, Wanqi, et al.
Published: (2024)
SMPLer-X: Scaling Up Expressive Human Pose and Shape Estimation
by: Cai, Zhongang, et al.
Published: (2023)
by: Cai, Zhongang, et al.
Published: (2023)
Demystifying Video Reasoning
by: Wang, Ruisi, et al.
Published: (2026)
by: Wang, Ruisi, et al.
Published: (2026)
Bridging Semantic and Kinematic Conditions with Diffusion-based Discrete Motion Tokenizer
by: Gu, Chenyang, et al.
Published: (2026)
by: Gu, Chenyang, et al.
Published: (2026)
Motion-X: A Large-scale 3D Expressive Whole-body Human Motion Dataset
by: Lin, Jing, et al.
Published: (2023)
by: Lin, Jing, et al.
Published: (2023)
Holistic Evaluation of Multimodal LLMs on Spatial Intelligence
by: Cai, Zhongang, et al.
Published: (2025)
by: Cai, Zhongang, et al.
Published: (2025)
ADHMR: Aligning Diffusion-based Human Mesh Recovery via Direct Preference Optimization
by: Shen, Wenhao, et al.
Published: (2025)
by: Shen, Wenhao, et al.
Published: (2025)
VLM-Guided Group Preference Alignment for Diffusion-based Human Mesh Recovery
by: Shen, Wenhao, et al.
Published: (2026)
by: Shen, Wenhao, et al.
Published: (2026)
MotionCraft: Crafting Whole-Body Motion with Plug-and-Play Multimodal Controls
by: Bian, Yuxuan, et al.
Published: (2024)
by: Bian, Yuxuan, et al.
Published: (2024)
DPoser: Diffusion Model as Robust 3D Human Pose Prior
by: Lu, Junzhe, et al.
Published: (2023)
by: Lu, Junzhe, et al.
Published: (2023)
Disco4D: Disentangled 4D Human Generation and Animation from a Single Image
by: Pang, Hui En, et al.
Published: (2024)
by: Pang, Hui En, et al.
Published: (2024)
Motion-X++: A Large-Scale Multimodal 3D Whole-body Human Motion Dataset
by: Zhang, Yuhong, et al.
Published: (2025)
by: Zhang, Yuhong, et al.
Published: (2025)
Large Motion Model for Unified Multi-Modal Motion Generation
by: Zhang, Mingyuan, et al.
Published: (2024)
by: Zhang, Mingyuan, et al.
Published: (2024)
UniMo: Unified Motion Generation and Understanding with Chain of Thought
by: Wang, Guocun, et al.
Published: (2026)
by: Wang, Guocun, et al.
Published: (2026)
GPAvatar: Generalizable and Precise Head Avatar from Image(s)
by: Chu, Xuangeng, et al.
Published: (2024)
by: Chu, Xuangeng, et al.
Published: (2024)
Deblur-Avatar: Animatable Avatars from Motion-Blurred Monocular Videos
by: Luo, Xianrui, et al.
Published: (2025)
by: Luo, Xianrui, et al.
Published: (2025)
Bridging the Gap between Human Motion and Action Semantics via Kinematic Phrases
by: Liu, Xinpeng, et al.
Published: (2023)
by: Liu, Xinpeng, et al.
Published: (2023)
MotionLLM: Understanding Human Behaviors from Human Motions and Videos
by: Chen, Ling-Hao, et al.
Published: (2024)
by: Chen, Ling-Hao, et al.
Published: (2024)
Scaling Spatial Intelligence with Multimodal Foundation Models
by: Cai, Zhongang, et al.
Published: (2025)
by: Cai, Zhongang, et al.
Published: (2025)
EchoMotion: Unified Human Video and Motion Generation via Dual-Modality Diffusion Transformer
by: Yang, Yuxiao, et al.
Published: (2025)
by: Yang, Yuxiao, et al.
Published: (2025)
Make Your Actor Talk: Generalizable and High-Fidelity Lip Sync with Motion and Appearance Disentanglement
by: Yu, Runyi, et al.
Published: (2024)
by: Yu, Runyi, et al.
Published: (2024)
SOLAMI: Social Vision-Language-Action Modeling for Immersive Interaction with 3D Autonomous Characters
by: Jiang, Jianping, et al.
Published: (2024)
by: Jiang, Jianping, et al.
Published: (2024)
Towards Fine-Grained Human Motion Video Captioning
by: Song, Guorui, et al.
Published: (2025)
by: Song, Guorui, et al.
Published: (2025)
Generalizable NGP-SR: Generalizable Neural Radiance Fields Super-Resolution via Neural Graph Primitives
by: Yuan, Wanqi, et al.
Published: (2026)
by: Yuan, Wanqi, et al.
Published: (2026)
Evaluation Agent: Efficient and Promptable Evaluation Framework for Visual Generative Models
by: Zhang, Fan, et al.
Published: (2024)
by: Zhang, Fan, et al.
Published: (2024)
Learning Generalizable Human Motion Generator with Reinforcement Learning
by: Mao, Yunyao, et al.
Published: (2024)
by: Mao, Yunyao, et al.
Published: (2024)
MARS: Enabling Autoregressive Models Multi-Token Generation
by: Jin, Ziqi, et al.
Published: (2026)
by: Jin, Ziqi, et al.
Published: (2026)
CoDAR: Continuous Diffusion Language Models are More Powerful Than You Think
by: Shen, Junzhe, et al.
Published: (2026)
by: Shen, Junzhe, et al.
Published: (2026)
Deep Learning Based Dynamic Environment Reconstruction for Vehicular ISAC Scenarios
by: Song, Junzhe, et al.
Published: (2025)
by: Song, Junzhe, et al.
Published: (2025)
DynamiCtrl: Rethinking the Basic Structure and the Role of Text for High-quality Human Image Animation
by: Zhao, Haoyu, et al.
Published: (2025)
by: Zhao, Haoyu, et al.
Published: (2025)
Playing for 3D Human Recovery
by: Cai, Zhongang, et al.
Published: (2021)
by: Cai, Zhongang, et al.
Published: (2021)
UniMotion: A Unified Motion Framework for Simulation, Prediction and Planning
by: Song, Nan, et al.
Published: (2026)
by: Song, Nan, et al.
Published: (2026)
Prompt Relay: Inference-Time Temporal Control for Multi-Event Video Generation
by: Chen, Gordon, et al.
Published: (2026)
by: Chen, Gordon, et al.
Published: (2026)
Recent Progress of the Application of Electropolymerization in Batteries and Supercapacitors: Specific Design of Functions in Electrodes
by: Shengxuan Lin, et al.
Published: (2024)
by: Shengxuan Lin, et al.
Published: (2024)
More Than Meets the Eye: A Semantics-Aware Traffic Augmentation Framework for Generalizable Website Fingerprinting
by: Xian, Youquan, et al.
Published: (2026)
by: Xian, Youquan, et al.
Published: (2026)
A Novel Site-Specific Inference Model for Urban Canyon Channels: From Measurements to Modeling
by: Song, Junzhe, et al.
Published: (2025)
by: Song, Junzhe, et al.
Published: (2025)
Deep Learning-Based Site-Specific Channel Modeling and Inference
by: Song, Junzhe, et al.
Published: (2026)
by: Song, Junzhe, et al.
Published: (2026)
Similar Items
-
AiOS: All-in-One-Stage Expressive Human Pose and Shape Estimation
by: Sun, Qingping, et al.
Published: (2024) -
DPoser-X: Diffusion Model as Robust 3D Whole-body Human Pose Prior
by: Lu, Junzhe, et al.
Published: (2025) -
SMPLest-X: Ultimate Scaling for Expressive Human Pose and Shape Estimation
by: Yin, Wanqi, et al.
Published: (2025) -
WHAC: World-grounded Humans and Cameras
by: Yin, Wanqi, et al.
Published: (2024) -
SMPLer-X: Scaling Up Expressive Human Pose and Shape Estimation
by: Cai, Zhongang, et al.
Published: (2023)