Rethinking Diffusion for Text-Driven Human Motion Generation: Redundant Representations, Evaluation, and Masked Autoregression
Fuente:
arXiv
Saved in:
| Main Authors: | Meng, Zichong, Xie, Yiming, Peng, Xiaogang, Han, Zeyu, Jiang, Huaizu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Absolute Coordinates Make Motion Generation Easy
by: Meng, Zichong, et al.
Published: (2025)
by: Meng, Zichong, et al.
Published: (2025)
HOI-Diff: Text-Driven Synthesis of 3D Human-Object Interactions using Diffusion Models
by: Peng, Xiaogang, et al.
Published: (2023)
by: Peng, Xiaogang, et al.
Published: (2023)
OmniControl: Control Any Joint at Any Time for Human Motion Generation
by: Xie, Yiming, et al.
Published: (2023)
by: Xie, Yiming, et al.
Published: (2023)
SMooDi: Stylized Motion Diffusion Model
by: Zhong, Lei, et al.
Published: (2024)
by: Zhong, Lei, et al.
Published: (2024)
SV4D 2.0: Enhancing Spatio-Temporal Consistency in Multi-View Video Diffusion for High-Quality 4D Generation
by: Yao, Chun-Han, et al.
Published: (2025)
by: Yao, Chun-Han, et al.
Published: (2025)
Text-driven Human Motion Generation with Motion Masked Diffusion Model
by: Chen, Xingyu
Published: (2024)
by: Chen, Xingyu
Published: (2024)
SV4D: Dynamic 3D Content Generation with Multi-Frame and Multi-View Consistency
by: Xie, Yiming, et al.
Published: (2024)
by: Xie, Yiming, et al.
Published: (2024)
Multi-granular body modeling with Redundancy-Free Spatiotemporal Fusion for Text-Driven Motion Generation
by: Zhan, Xingzu, et al.
Published: (2025)
by: Zhan, Xingzu, et al.
Published: (2025)
Zero-shot Referring Expression Comprehension via Structural Similarity Between Images and Captions
by: Han, Zeyu, et al.
Published: (2023)
by: Han, Zeyu, et al.
Published: (2023)
A Strong Baseline for Point Cloud Registration via Direct Superpoints Matching
by: Gupta, Aniket, et al.
Published: (2023)
by: Gupta, Aniket, et al.
Published: (2023)
OmniMotion: Multimodal Motion Generation with Continuous Masked Autoregression
by: Li, Zhe, et al.
Published: (2025)
by: Li, Zhe, et al.
Published: (2025)
Towards Robust and Controllable Text-to-Motion via Masked Autoregressive Diffusion
by: Zhang, Zongye, et al.
Published: (2025)
by: Zhang, Zongye, et al.
Published: (2025)
ReMoMask: Retrieval-Augmented Masked Motion Generation
by: Li, Zhengdao, et al.
Published: (2025)
by: Li, Zhengdao, et al.
Published: (2025)
Causal Motion Diffusion Models for Autoregressive Motion Generation
by: Yu, Qing, et al.
Published: (2026)
by: Yu, Qing, et al.
Published: (2026)
DartControl: A Diffusion-Based Autoregressive Motion Model for Real-Time Text-Driven Motion Control
by: Zhao, Kaifeng, et al.
Published: (2024)
by: Zhao, Kaifeng, et al.
Published: (2024)
MarDini: Masked Autoregressive Diffusion for Video Generation at Scale
by: Liu, Haozhe, et al.
Published: (2024)
by: Liu, Haozhe, et al.
Published: (2024)
Next-Scale Autoregressive Models for Text-to-Motion Generation
by: Zheng, Zhiwei, et al.
Published: (2026)
by: Zheng, Zhiwei, et al.
Published: (2026)
Diffusion Beats Autoregressive: An Evaluation of Compositional Generation in Text-to-Image Models
by: Marioriyad, Arash, et al.
Published: (2024)
by: Marioriyad, Arash, et al.
Published: (2024)
Unleashing the Potential of Large Language Models for Text-to-Image Generation through Autoregressive Representation Alignment
by: Xie, Xing, et al.
Published: (2025)
by: Xie, Xing, et al.
Published: (2025)
GUESS:GradUally Enriching SyntheSis for Text-Driven Human Motion Generation
by: Gao, Xuehao, et al.
Published: (2024)
by: Gao, Xuehao, et al.
Published: (2024)
StyleBlend: Enhancing Style-Specific Content Creation in Text-to-Image Diffusion Models
by: Chen, Zichong, et al.
Published: (2025)
by: Chen, Zichong, et al.
Published: (2025)
Rethinking UMM Visual Generation: Masked Modeling for Efficient Image-Only Pre-training
by: Sun, Peng, et al.
Published: (2026)
by: Sun, Peng, et al.
Published: (2026)
LASER: Layer-wise Scale Alignment for Training-Free Streaming 4D Reconstruction
by: Ding, Tianye, et al.
Published: (2025)
by: Ding, Tianye, et al.
Published: (2025)
LGTM: Local-to-Global Text-Driven Human Motion Diffusion Model
by: Sun, Haowen, et al.
Published: (2024)
by: Sun, Haowen, et al.
Published: (2024)
Token Painter: Training-Free Text-Guided Image Inpainting via Mask Autoregressive Models
by: Jiang, Longtao, et al.
Published: (2025)
by: Jiang, Longtao, et al.
Published: (2025)
MotionStreamer: Streaming Motion Generation via Diffusion-based Autoregressive Model in Causal Latent Space
by: Xiao, Lixing, et al.
Published: (2025)
by: Xiao, Lixing, et al.
Published: (2025)
Reconstruction-Anchored Diffusion Model for Text-to-Motion Generation
by: Liu, Yifei, et al.
Published: (2026)
by: Liu, Yifei, et al.
Published: (2026)
DCVNet: Dilated Cost Volume Networks for Fast Optical Flow
by: Jiang, Huaizu, et al.
Published: (2021)
by: Jiang, Huaizu, et al.
Published: (2021)
Local Action-Guided Motion Diffusion Model for Text-to-Motion Generation
by: Jin, Peng, et al.
Published: (2024)
by: Jin, Peng, et al.
Published: (2024)
Taming Teacher Forcing for Masked Autoregressive Video Generation
by: Zhou, Deyu, et al.
Published: (2025)
by: Zhou, Deyu, et al.
Published: (2025)
Motion Avatar: Generate Human and Animal Avatars with Arbitrary Motion
by: Zhang, Zeyu, et al.
Published: (2024)
by: Zhang, Zeyu, et al.
Published: (2024)
Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image Generation
by: Luo, Yifu, et al.
Published: (2025)
by: Luo, Yifu, et al.
Published: (2025)
MotionLab: Unified Human Motion Generation and Editing via the Motion-Condition-Motion Paradigm
by: Guo, Ziyan, et al.
Published: (2025)
by: Guo, Ziyan, et al.
Published: (2025)
LlamaSeg: Image Segmentation via Autoregressive Mask Generation
by: Deng, Jiru, et al.
Published: (2025)
by: Deng, Jiru, et al.
Published: (2025)
Fleximo: Towards Flexible Text-to-Human Motion Video Generation
by: Zhang, Yuhang, et al.
Published: (2024)
by: Zhang, Yuhang, et al.
Published: (2024)
Hawk: Leveraging Spatial Context for Faster Autoregressive Text-to-Image Generation
by: Chen, Zhi-Kai, et al.
Published: (2025)
by: Chen, Zhi-Kai, et al.
Published: (2025)
HINT: Hierarchical Interaction Modeling for Autoregressive Multi-Human Motion Generation
by: Liu, Mengge, et al.
Published: (2026)
by: Liu, Mengge, et al.
Published: (2026)
Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking
by: Zheng, Zirui, et al.
Published: (2025)
by: Zheng, Zirui, et al.
Published: (2025)
LaxMotion: Rethinking Supervision Granularity for 3D Human Motion Generation
by: Liu, Sheng, et al.
Published: (2025)
by: Liu, Sheng, et al.
Published: (2025)
Struct2D: A Perception-Guided Framework for Spatial Reasoning in MLLMs
by: Zhu, Fangrui, et al.
Published: (2025)
by: Zhu, Fangrui, et al.
Published: (2025)
Similar Items
-
Absolute Coordinates Make Motion Generation Easy
by: Meng, Zichong, et al.
Published: (2025) -
HOI-Diff: Text-Driven Synthesis of 3D Human-Object Interactions using Diffusion Models
by: Peng, Xiaogang, et al.
Published: (2023) -
OmniControl: Control Any Joint at Any Time for Human Motion Generation
by: Xie, Yiming, et al.
Published: (2023) -
SMooDi: Stylized Motion Diffusion Model
by: Zhong, Lei, et al.
Published: (2024) -
SV4D 2.0: Enhancing Spatio-Temporal Consistency in Multi-View Video Diffusion for High-Quality 4D Generation
by: Yao, Chun-Han, et al.
Published: (2025)