Multi-granular body modeling with Redundancy-Free Spatiotemporal Fusion for Text-Driven Motion Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhan, Xingzu, Xie, Chen, Chen, Honghang, Sun, Haoran, Mai, Xiaochun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
T2M Mamba: Motion Periodicity-Saliency Coupling Approach for Stable Text-Driven Motion Generation
by: Zhan, Xingzu, et al.
Published: (2026)
by: Zhan, Xingzu, et al.
Published: (2026)
Rethinking Diffusion for Text-Driven Human Motion Generation: Redundant Representations, Evaluation, and Masked Autoregression
by: Meng, Zichong, et al.
Published: (2024)
by: Meng, Zichong, et al.
Published: (2024)
MotionMERGE: A Multi-granular Framework for Human Motion Editing, Reasoning, Generation, and Explanation
by: Wu, Bizhu, et al.
Published: (2026)
by: Wu, Bizhu, et al.
Published: (2026)
TCMA: Text-Conditioned Multi-granularity Alignment for Drone Cross-Modal Text-Video Retrieval
by: Zhao, Zixu, et al.
Published: (2025)
by: Zhao, Zixu, et al.
Published: (2025)
GUESS:GradUally Enriching SyntheSis for Text-Driven Human Motion Generation
by: Gao, Xuehao, et al.
Published: (2024)
by: Gao, Xuehao, et al.
Published: (2024)
Infinite Motion: Extended Motion Generation via Long Text Instructions
by: Li, Mengtian, et al.
Published: (2024)
by: Li, Mengtian, et al.
Published: (2024)
Topology-Agnostic Animal Motion Generation from Text Prompt
by: Chen, Keyi, et al.
Published: (2025)
by: Chen, Keyi, et al.
Published: (2025)
Text-driven Human Motion Generation with Motion Masked Diffusion Model
by: Chen, Xingyu
Published: (2024)
by: Chen, Xingyu
Published: (2024)
RainFusion: Adaptive Video Generation Acceleration via Multi-Dimensional Visual Redundancy
by: Chen, Aiyue, et al.
Published: (2025)
by: Chen, Aiyue, et al.
Published: (2025)
CrowdMoGen: Zero-Shot Text-Driven Collective Motion Generation
by: Cao, Yukang, et al.
Published: (2024)
by: Cao, Yukang, et al.
Published: (2024)
Contrastive Learning-Driven Traffic Sign Perception: Multi-Modal Fusion of Text and Vision
by: Lu, Qiang, et al.
Published: (2025)
by: Lu, Qiang, et al.
Published: (2025)
MotionShot: Adaptive Motion Transfer across Arbitrary Objects for Text-to-Video Generation
by: Liu, Yanchen, et al.
Published: (2025)
by: Liu, Yanchen, et al.
Published: (2025)
FastVMT: Eliminating Redundancy in Video Motion Transfer
by: Ma, Yue, et al.
Published: (2026)
by: Ma, Yue, et al.
Published: (2026)
Setting the Stage: Text-Driven Scene-Consistent Image Generation
by: Xie, Cong, et al.
Published: (2025)
by: Xie, Cong, et al.
Published: (2025)
InterFusion: Text-Driven Generation of 3D Human-Object Interaction
by: Dai, Sisi, et al.
Published: (2024)
by: Dai, Sisi, et al.
Published: (2024)
GenM$^3$: Generative Pretrained Multi-path Motion Model for Text Conditional Human Motion Generation
by: Shi, Junyu, et al.
Published: (2025)
by: Shi, Junyu, et al.
Published: (2025)
SIGMA: Semantic-Difference Instruction-Grounding Mask Annotator for Text-Driven Image Manipulation Localization
by: Zhuang, Peiyu, et al.
Published: (2026)
by: Zhuang, Peiyu, et al.
Published: (2026)
FreeGraftor: Training-Free Cross-Image Feature Grafting for Subject-Driven Text-to-Image Generation
by: Yao, Zebin, et al.
Published: (2025)
by: Yao, Zebin, et al.
Published: (2025)
Free-T2M: Robust Text-to-Motion Generation for Humanoid Robots via Frequency-Domain
by: Chen, Wenshuo, et al.
Published: (2025)
by: Chen, Wenshuo, et al.
Published: (2025)
Think Before You Move: Latent Motion Reasoning for Text-to-Motion Generation
by: Qian, Yijie, et al.
Published: (2025)
by: Qian, Yijie, et al.
Published: (2025)
Towards Open Domain Text-Driven Synthesis of Multi-Person Motions
by: Shan, Mengyi, et al.
Published: (2024)
by: Shan, Mengyi, et al.
Published: (2024)
MotionBooth: Motion-Aware Customized Text-to-Video Generation
by: Wu, Jianzong, et al.
Published: (2024)
by: Wu, Jianzong, et al.
Published: (2024)
Unified Number-Free Text-to-Motion Generation Via Flow Matching
by: Huang, Guanhe, et al.
Published: (2026)
by: Huang, Guanhe, et al.
Published: (2026)
M2DAO-Talker: Harmonizing Multi-granular Motion Decoupling and Alternating Optimization for Talking-head Generation
by: Jiang, Kui, et al.
Published: (2025)
by: Jiang, Kui, et al.
Published: (2025)
RedMotion: Motion Prediction via Redundancy Reduction
by: Wagner, Royden, et al.
Published: (2023)
by: Wagner, Royden, et al.
Published: (2023)
SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
by: Ge, Yuying, et al.
Published: (2024)
by: Ge, Yuying, et al.
Published: (2024)
AlterMOMA: Fusion Redundancy Pruning for Camera-LiDAR Fusion Models with Alternative Modality Masking
by: Sun, Shiqi, et al.
Published: (2024)
by: Sun, Shiqi, et al.
Published: (2024)
SDDF: Specificity-Driven Dynamic Focusing for Open-Vocabulary Camouflaged Object Detection
by: Liang, Jiaming, et al.
Published: (2026)
by: Liang, Jiaming, et al.
Published: (2026)
UniCtrl: Improving the Spatiotemporal Consistency of Text-to-Video Diffusion Models via Training-Free Unified Attention Control
by: Xia, Tian, et al.
Published: (2024)
by: Xia, Tian, et al.
Published: (2024)
Dissecting RGB-D Learning for Improved Multi-modal Fusion
by: Chen, Hao, et al.
Published: (2023)
by: Chen, Hao, et al.
Published: (2023)
Spatiotemporal-Untrammelled Mixture of Experts for Multi-Person Motion Prediction
by: Yin, Zheng, et al.
Published: (2025)
by: Yin, Zheng, et al.
Published: (2025)
Local Action-Guided Motion Diffusion Model for Text-to-Motion Generation
by: Jin, Peng, et al.
Published: (2024)
by: Jin, Peng, et al.
Published: (2024)
MotionClone: Training-Free Motion Cloning for Controllable Video Generation
by: Ling, Pengyang, et al.
Published: (2024)
by: Ling, Pengyang, et al.
Published: (2024)
Reframing Music-Driven 2D Dance Pose Generation as Multi-Channel Image Generation
by: Zhang, Yan, et al.
Published: (2025)
by: Zhang, Yan, et al.
Published: (2025)
SATO: Stable Text-to-Motion Framework
by: Chen, Wenshuo, et al.
Published: (2024)
by: Chen, Wenshuo, et al.
Published: (2024)
SP-Det: Self-Prompted Dual-Text Fusion for Generalized Multi-Label Lesion Detection
by: Xu, Qing, et al.
Published: (2025)
by: Xu, Qing, et al.
Published: (2025)
SVGThinker: Instruction-Aligned and Reasoning-Driven Text-to-SVG Generation
by: Chen, Hanqi, et al.
Published: (2025)
by: Chen, Hanqi, et al.
Published: (2025)
Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation
by: Wang, Wenjing, et al.
Published: (2023)
by: Wang, Wenjing, et al.
Published: (2023)
FreeMotion: A Unified Framework for Number-free Text-to-Motion Synthesis
by: Fan, Ke, et al.
Published: (2024)
by: Fan, Ke, et al.
Published: (2024)
MotionRL: Align Text-to-Motion Generation to Human Preferences with Multi-Reward Reinforcement Learning
by: Liu, Xiaoyang, et al.
Published: (2024)
by: Liu, Xiaoyang, et al.
Published: (2024)
Similar Items
-
T2M Mamba: Motion Periodicity-Saliency Coupling Approach for Stable Text-Driven Motion Generation
by: Zhan, Xingzu, et al.
Published: (2026) -
Rethinking Diffusion for Text-Driven Human Motion Generation: Redundant Representations, Evaluation, and Masked Autoregression
by: Meng, Zichong, et al.
Published: (2024) -
MotionMERGE: A Multi-granular Framework for Human Motion Editing, Reasoning, Generation, and Explanation
by: Wu, Bizhu, et al.
Published: (2026) -
TCMA: Text-Conditioned Multi-granularity Alignment for Drone Cross-Modal Text-Video Retrieval
by: Zhao, Zixu, et al.
Published: (2025) -
GUESS:GradUally Enriching SyntheSis for Text-Driven Human Motion Generation
by: Gao, Xuehao, et al.
Published: (2024)