You Think, You ACT: The New Task of Arbitrary Text to Motion Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Runqi, Ma, Caoyuan, Li, Guopeng, Xu, Hanrui, Li, Yuke, Wang, Zheng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Think Before You Move: Latent Motion Reasoning for Text-to-Motion Generation
by: Qian, Yijie, et al.
Published: (2025)
by: Qian, Yijie, et al.
Published: (2025)
Behave Your Motion: Habit-preserved Cross-category Animal Motion Transfer
by: Zhang, Zhimin, et al.
Published: (2025)
by: Zhang, Zhimin, et al.
Published: (2025)
Move as You Say, Interact as You Can: Language-guided Human Motion Generation with Scene Affordance
by: Wang, Zan, et al.
Published: (2024)
by: Wang, Zan, et al.
Published: (2024)
Motion-R1: Enhancing Motion Generation with Decomposed Chain-of-Thought and RL Binding
by: Ouyang, Runqi, et al.
Published: (2025)
by: Ouyang, Runqi, et al.
Published: (2025)
Representation Entanglement for Generation: Training Diffusion Transformers Is Much Easier Than You Think
by: Wu, Ge, et al.
Published: (2025)
by: Wu, Ge, et al.
Published: (2025)
Edit Where You Mean: Region-Aware Adapter Injection for Mask-Free Local Image Editing
by: Cai, Honghao, et al.
Published: (2026)
by: Cai, Honghao, et al.
Published: (2026)
You Only Look at Once for Real-time and Generic Multi-Task
by: Wang, Jiayuan, et al.
Published: (2023)
by: Wang, Jiayuan, et al.
Published: (2023)
Get What You Want, Not What You Don't: Image Content Suppression for Text-to-Image Diffusion Models
by: Li, Senmao, et al.
Published: (2024)
by: Li, Senmao, et al.
Published: (2024)
Think Before You Act: A Two-Stage Framework for Mitigating Gender Bias Towards Vision-Language Tasks
by: Zhang, Yunqi, et al.
Published: (2024)
by: Zhang, Yunqi, et al.
Published: (2024)
Multimodal Representation Alignment for Image Generation: Text-Image Interleaved Control Is Easier Than You Think
by: Chen, Liang, et al.
Published: (2025)
by: Chen, Liang, et al.
Published: (2025)
Mind the Way You Select Negative Texts: Pursuing the Distance Consistency in OOD Detection with VLMs
by: Xu, Zhikang, et al.
Published: (2026)
by: Xu, Zhikang, et al.
Published: (2026)
EchoMimicV3: 1.3B Parameters are All You Need for Unified Multi-Modal and Multi-Task Human Animation
by: Meng, Rang, et al.
Published: (2025)
by: Meng, Rang, et al.
Published: (2025)
Think Before You Drive: World Model-Inspired Multimodal Grounding for Autonomous Vehicles
by: Liao, Haicheng, et al.
Published: (2025)
by: Liao, Haicheng, et al.
Published: (2025)
MotionShot: Adaptive Motion Transfer across Arbitrary Objects for Text-to-Video Generation
by: Liu, Yanchen, et al.
Published: (2025)
by: Liu, Yanchen, et al.
Published: (2025)
Improve Meta-learning for Few-Shot Text Classification with All You Can Acquire from the Tasks
by: Liu, Xinyue, et al.
Published: (2024)
by: Liu, Xinyue, et al.
Published: (2024)
Think Before You Diffuse: Infusing Physical Rules into Video Diffusion
by: Zhang, Ke, et al.
Published: (2025)
by: Zhang, Ke, et al.
Published: (2025)
ZeroFlow: Overcoming Catastrophic Forgetting is Easier than You Think
by: Feng, Tao, et al.
Published: (2025)
by: Feng, Tao, et al.
Published: (2025)
Now You See It, Now You Don't - Instant Concept Erasure for Safe Text-to-Image and Video Generation
by: Biswas, Shristi Das, et al.
Published: (2025)
by: Biswas, Shristi Das, et al.
Published: (2025)
HumanNeRF-SE: A Simple yet Effective Approach to Animate HumanNeRF with Diverse Poses
by: Ma, Caoyuan, et al.
Published: (2023)
by: Ma, Caoyuan, et al.
Published: (2023)
Fine-tuning MLLMs Without Forgetting Is Easier Than You Think
by: Li, He, et al.
Published: (2026)
by: Li, He, et al.
Published: (2026)
Subjective Camera 1.0: Bridging Human Cognition and Visual Reconstruction through Sequence-Aware Sketch-Guided Diffusion
by: Chen, Haoyang, et al.
Published: (2025)
by: Chen, Haoyang, et al.
Published: (2025)
Edit as You See: Image-guided Video Editing via Masked Motion Modeling
by: Huang, Zhi-Lin, et al.
Published: (2025)
by: Huang, Zhi-Lin, et al.
Published: (2025)
PRNet: Original Information Is All You Have
by: Zheng, PeiHuang, et al.
Published: (2025)
by: Zheng, PeiHuang, et al.
Published: (2025)
Motion Avatar: Generate Human and Animal Avatars with Arbitrary Motion
by: Zhang, Zeyu, et al.
Published: (2024)
by: Zhang, Zeyu, et al.
Published: (2024)
VideoScore2: Think before You Score in Generative Video Evaluation
by: He, Xuan, et al.
Published: (2025)
by: He, Xuan, et al.
Published: (2025)
MoTe: Learning Motion-Text Diffusion Model for Multiple Generation Tasks
by: Wu, Yiming, et al.
Published: (2024)
by: Wu, Yiming, et al.
Published: (2024)
AgenticOCR: Parsing Only What You Need for Efficient Retrieval-Augmented Generation
by: Wang, Zhengren, et al.
Published: (2026)
by: Wang, Zhengren, et al.
Published: (2026)
Think-as-You-See: Streaming Chain-of-Thought Reasoning for Large Vision-Language Models
by: Zhang, Jialiang, et al.
Published: (2026)
by: Zhang, Jialiang, et al.
Published: (2026)
Is What You Ask For What You Get? Investigating Concept Associations in Text-to-Image Models
by: Magid, Salma Abdel, et al.
Published: (2024)
by: Magid, Salma Abdel, et al.
Published: (2024)
What You Have is What You Track: Adaptive and Robust Multimodal Tracking
by: Tan, Yuedong, et al.
Published: (2025)
by: Tan, Yuedong, et al.
Published: (2025)
Is Discretization Fusion All You Need for Collaborative Perception?
by: Yang, Kang, et al.
Published: (2025)
by: Yang, Kang, et al.
Published: (2025)
You See it, You Got it: Learning 3D Creation on Pose-Free Videos at Scale
by: Ma, Baorui, et al.
Published: (2024)
by: Ma, Baorui, et al.
Published: (2024)
Aligning Text to Image in Diffusion Models is Easier Than You Think
by: Lee, Jaa-Yeon, et al.
Published: (2025)
by: Lee, Jaa-Yeon, et al.
Published: (2025)
Think While You Generate: Discrete Diffusion with Planned Denoising
by: Liu, Sulin, et al.
Published: (2024)
by: Liu, Sulin, et al.
Published: (2024)
DIMO: Diverse 3D Motion Generation for Arbitrary Objects
by: Mou, Linzhan, et al.
Published: (2025)
by: Mou, Linzhan, et al.
Published: (2025)
Get In Video: Add Anything You Want to the Video
by: Zhuang, Shaobin, et al.
Published: (2025)
by: Zhuang, Shaobin, et al.
Published: (2025)
V-GRPO: Online Reinforcement Learning for Denoising Generative Models Is Easier than You Think
by: Tang, Bingda, et al.
Published: (2026)
by: Tang, Bingda, et al.
Published: (2026)
YouDream: Generating Anatomically Controllable Consistent Text-to-3D Animals
by: Mishra, Sandeep, et al.
Published: (2024)
by: Mishra, Sandeep, et al.
Published: (2024)
Extrapolating and Decoupling Image-to-Video Generation Models: Motion Modeling is Easier Than You Think
by: Tian, Jie, et al.
Published: (2025)
by: Tian, Jie, et al.
Published: (2025)
IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation
by: Li, Yuan-Ming, et al.
Published: (2025)
by: Li, Yuan-Ming, et al.
Published: (2025)
Similar Items
-
Think Before You Move: Latent Motion Reasoning for Text-to-Motion Generation
by: Qian, Yijie, et al.
Published: (2025) -
Behave Your Motion: Habit-preserved Cross-category Animal Motion Transfer
by: Zhang, Zhimin, et al.
Published: (2025) -
Move as You Say, Interact as You Can: Language-guided Human Motion Generation with Scene Affordance
by: Wang, Zan, et al.
Published: (2024) -
Motion-R1: Enhancing Motion Generation with Decomposed Chain-of-Thought and RL Binding
by: Ouyang, Runqi, et al.
Published: (2025) -
Representation Entanglement for Generation: Training Diffusion Transformers Is Much Easier Than You Think
by: Wu, Ge, et al.
Published: (2025)