Generating Attribute-Aware Human Motions from Textual Prompt
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Xinghan, Xu, Kun, Li, Fei, Sheng, Cao, Yu, Jiazhong, Mu, Yadong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Text-controlled Motion Mamba: Text-Instructed Temporal Grounding of Human Motion
by: Wang, Xinghan, et al.
Published: (2024)
by: Wang, Xinghan, et al.
Published: (2024)
Bringing Textual Prompt to AI-Generated Image Quality Assessment
by: Qu, Bowen, et al.
Published: (2024)
by: Qu, Bowen, et al.
Published: (2024)
Human Motion Video Generation: A Survey
by: Xue, Haiwei, et al.
Published: (2025)
by: Xue, Haiwei, et al.
Published: (2025)
Biomechanics-Guided Residual Approach to Generalizable Human Motion Generation and Estimation
by: Kang, Zixi, et al.
Published: (2025)
by: Kang, Zixi, et al.
Published: (2025)
PP-Motion: Physical-Perceptual Fidelity Evaluation for Human Motion Generation
by: Zhao, Sihan, et al.
Published: (2025)
by: Zhao, Sihan, et al.
Published: (2025)
SOSControl: Enhancing Human Motion Generation through Saliency-Aware Symbolic Orientation and Timing Control
by: Au, Ho Yin, et al.
Published: (2025)
by: Au, Ho Yin, et al.
Published: (2025)
MTFusion: Reconstructing Any 3D Object from Single Image Using Multi-word Textual Inversion
by: Liu, Yu, et al.
Published: (2024)
by: Liu, Yu, et al.
Published: (2024)
Controllable Complex Human Motion Video Generation via Text-to-Skeleton Cascades
by: Taghipour, Ashkan, et al.
Published: (2026)
by: Taghipour, Ashkan, et al.
Published: (2026)
TOL: Textual Localization with OpenStreetMap
by: Liao, Youqi, et al.
Published: (2026)
by: Liao, Youqi, et al.
Published: (2026)
MoRAG -- Multi-Fusion Retrieval Augmented Generation for Human Motion
by: Kalakonda, Sai Shashank, et al.
Published: (2024)
by: Kalakonda, Sai Shashank, et al.
Published: (2024)
DBDH: A Dual-Branch Dual-Head Neural Network for Invisible Embedded Regions Localization
by: Zhao, Chengxin, et al.
Published: (2024)
by: Zhao, Chengxin, et al.
Published: (2024)
TA-V2A: Textually Assisted Video-to-Audio Generation
by: You, Yuhuan, et al.
Published: (2025)
by: You, Yuhuan, et al.
Published: (2025)
MotionDreamer: One-to-Many Motion Synthesis with Localized Generative Masked Transformer
by: Wang, Yilin, et al.
Published: (2025)
by: Wang, Yilin, et al.
Published: (2025)
SequencePAR: Understanding Pedestrian Attributes via A Sequence Generation Paradigm
by: Jin, Jiandong, et al.
Published: (2023)
by: Jin, Jiandong, et al.
Published: (2023)
MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance
by: Zhang, Yuang, et al.
Published: (2024)
by: Zhang, Yuang, et al.
Published: (2024)
Exploring Mutual Cross-Modal Attention for Context-Aware Human Affordance Generation
by: Roy, Prasun, et al.
Published: (2025)
by: Roy, Prasun, et al.
Published: (2025)
MotionPro: A Precise Motion Controller for Image-to-Video Generation
by: Zhang, Zhongwei, et al.
Published: (2025)
by: Zhang, Zhongwei, et al.
Published: (2025)
StableMoFusion: Towards Robust and Efficient Diffusion-based Motion Generation Framework
by: Huang, Yiheng, et al.
Published: (2024)
by: Huang, Yiheng, et al.
Published: (2024)
MSVBench: Towards Human-Level Evaluation of Multi-Shot Video Generation
by: Shi, Haoyuan, et al.
Published: (2026)
by: Shi, Haoyuan, et al.
Published: (2026)
Holistic Visual-Textual Sentiment Analysis with Prior Models
by: Chen, Junyu, et al.
Published: (2022)
by: Chen, Junyu, et al.
Published: (2022)
Radio Frequency Signal based Human Silhouette Segmentation: A Sequential Diffusion Approach
by: Wen, Penghui, et al.
Published: (2024)
by: Wen, Penghui, et al.
Published: (2024)
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues
by: Feng, X., et al.
Published: (2024)
by: Feng, X., et al.
Published: (2024)
Efficient and Generic Point Model for Lossless Point Cloud Attribute Compression
by: You, Kang, et al.
Published: (2024)
by: You, Kang, et al.
Published: (2024)
Bridging the Gap: Sketch-Aware Interpolation Network for High-Quality Animation Sketch Inbetweening
by: Shen, Jiaming, et al.
Published: (2023)
by: Shen, Jiaming, et al.
Published: (2023)
Dual Attribute-Spatial Relation Alignment for 3D Visual Grounding
by: Xu, Yue, et al.
Published: (2024)
by: Xu, Yue, et al.
Published: (2024)
GAIA: Zero-shot Talking Avatar Generation
by: He, Tianyu, et al.
Published: (2023)
by: He, Tianyu, et al.
Published: (2023)
ViMo: Generating Motions from Casual Videos
by: Qiu, Liangdong, et al.
Published: (2024)
by: Qiu, Liangdong, et al.
Published: (2024)
Hand1000: Generating Realistic Hands from Text with Only 1,000 Images
by: Zhang, Haozhuo, et al.
Published: (2024)
by: Zhang, Haozhuo, et al.
Published: (2024)
Prompt-Aware Adaptive Elastic Weight Consolidation for Continual Learning in Medical Vision-Language Models
by: Gao, Ziyuan, et al.
Published: (2025)
by: Gao, Ziyuan, et al.
Published: (2025)
ReactDiff: Latent Diffusion for Facial Reaction Generation
by: Li, Jiaming, et al.
Published: (2025)
by: Li, Jiaming, et al.
Published: (2025)
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering
by: Zhou, Sheng, et al.
Published: (2025)
by: Zhou, Sheng, et al.
Published: (2025)
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
by: Chen, Liyang, et al.
Published: (2025)
by: Chen, Liyang, et al.
Published: (2025)
Hierarchical Textual Knowledge for Enhanced Image Clustering
by: Zhong, Yijie, et al.
Published: (2026)
by: Zhong, Yijie, et al.
Published: (2026)
Fine-grained Textual Inversion Network for Zero-Shot Composed Image Retrieval
by: Lin, Haoqiang, et al.
Published: (2025)
by: Lin, Haoqiang, et al.
Published: (2025)
Deep Compositional Phase Diffusion for Long Motion Sequence Generation
by: Au, Ho Yin, et al.
Published: (2025)
by: Au, Ho Yin, et al.
Published: (2025)
Generative Frame Sampler for Long Video Understanding
by: Yao, Linli, et al.
Published: (2025)
by: Yao, Linli, et al.
Published: (2025)
DPC: Dual-Prompt Collaboration for Tuning Vision-Language Models
by: Li, Haoyang, et al.
Published: (2025)
by: Li, Haoyang, et al.
Published: (2025)
Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation
by: Cheng, Shihao, et al.
Published: (2026)
by: Cheng, Shihao, et al.
Published: (2026)
AudCast: Audio-Driven Human Video Generation by Cascaded Diffusion Transformers
by: Guan, Jiazhi, et al.
Published: (2025)
by: Guan, Jiazhi, et al.
Published: (2025)
REArtGS: Reconstructing and Generating Articulated Objects via 3D Gaussian Splatting with Geometric and Motion Constraints
by: Wu, Di, et al.
Published: (2025)
by: Wu, Di, et al.
Published: (2025)
Similar Items
-
Text-controlled Motion Mamba: Text-Instructed Temporal Grounding of Human Motion
by: Wang, Xinghan, et al.
Published: (2024) -
Bringing Textual Prompt to AI-Generated Image Quality Assessment
by: Qu, Bowen, et al.
Published: (2024) -
Human Motion Video Generation: A Survey
by: Xue, Haiwei, et al.
Published: (2025) -
Biomechanics-Guided Residual Approach to Generalizable Human Motion Generation and Estimation
by: Kang, Zixi, et al.
Published: (2025) -
PP-Motion: Physical-Perceptual Fidelity Evaluation for Human Motion Generation
by: Zhao, Sihan, et al.
Published: (2025)