PlanMoGPT: Flow-Enhanced Progressive Planning for Text to Motion Synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | Jin, Chuhao, Li, Haosen, Zhang, Bingzi, Liu, Che, Wang, Xiting, Song, Ruihua, Huang, Wenbing, Qin, Ying, Zhang, Fuzheng, Zhang, Di |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Latency Effects on Multi-Dimensional QoE in Networked VR Whiteboards
by: Song, Jiarun, et al.
Published: (2026)
by: Song, Jiarun, et al.
Published: (2026)
SentiAvatar: Towards Expressive and Interactive Digital Humans
by: Jin, Chuhao, et al.
Published: (2026)
by: Jin, Chuhao, et al.
Published: (2026)
TPIFM: A Task-Aware Model for Evaluating Perceptual Interaction Fluency in Remote AR Collaboration
by: Song, Jiarun, et al.
Published: (2026)
by: Song, Jiarun, et al.
Published: (2026)
VSpeechLM: A Visual Speech Language Model for Visual Text-to-Speech Task
by: Wang, Yuyue, et al.
Published: (2025)
by: Wang, Yuyue, et al.
Published: (2025)
CustomDancer: Customized Dance Recommendation by Text-Dance Retrieval
by: Qin, Yawen, et al.
Published: (2026)
by: Qin, Yawen, et al.
Published: (2026)
See or Guess: Counterfactually Regularized Image Captioning
by: Cao, Qian, et al.
Published: (2024)
by: Cao, Qian, et al.
Published: (2024)
Dynamic Multimodal Expression Generation for LLM-Driven Pedagogical Agents: From User Experience Perspective
by: Wan, Ninghao, et al.
Published: (2026)
by: Wan, Ninghao, et al.
Published: (2026)
EyEar: Learning Audio Synchronized Human Gaze Trajectory Based on Physics-Informed Dynamics
by: Liu, Xiaochuan, et al.
Published: (2025)
by: Liu, Xiaochuan, et al.
Published: (2025)
SIDQL: An Efficient Keyframe Extraction and Motion Reconstruction Framework in Motion Capture
by: Zhang, Xuling, et al.
Published: (2024)
by: Zhang, Xuling, et al.
Published: (2024)
Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation
by: Wang, Yongqi, et al.
Published: (2025)
by: Wang, Yongqi, et al.
Published: (2025)
ViMo: Generating Motions from Casual Videos
by: Qiu, Liangdong, et al.
Published: (2024)
by: Qiu, Liangdong, et al.
Published: (2024)
Loc4Plan: Locating Before Planning for Outdoor Vision and Language Navigation
by: Tian, Huilin, et al.
Published: (2024)
by: Tian, Huilin, et al.
Published: (2024)
AniME: Adaptive Multi-Agent Planning for Long Animation Generation
by: Zhang, Lisai, et al.
Published: (2025)
by: Zhang, Lisai, et al.
Published: (2025)
StableMoFusion: Towards Robust and Efficient Diffusion-based Motion Generation Framework
by: Huang, Yiheng, et al.
Published: (2024)
by: Huang, Yiheng, et al.
Published: (2024)
HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language Models
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
Video DataFlywheel: Resolving the Impossible Data Trinity in Video-Language Understanding
by: Wang, Xiao, et al.
Published: (2024)
by: Wang, Xiao, et al.
Published: (2024)
TeMTG: Text-Enhanced Multi-Hop Temporal Graph Modeling for Audio-Visual Video Parsing
by: Chen, Yaru, et al.
Published: (2025)
by: Chen, Yaru, et al.
Published: (2025)
Enhancing Neural Adaptive Wireless Video Streaming via Lower-Layer Information Exposure and Online Tuning
by: Zhao, Lingzhi, et al.
Published: (2025)
by: Zhao, Lingzhi, et al.
Published: (2025)
UniMuMo: Unified Text, Music and Motion Generation
by: Yang, Han, et al.
Published: (2024)
by: Yang, Han, et al.
Published: (2024)
Think-Then-React: Towards Unconstrained Human Action-to-Reaction Generation
by: Tan, Wenhui, et al.
Published: (2025)
by: Tan, Wenhui, et al.
Published: (2025)
Text-Only Data Synthesis for Vision Language Model Training
by: Yu, Xiaomin, et al.
Published: (2025)
by: Yu, Xiaomin, et al.
Published: (2025)
Hyperbolic Multimodal Generative Representation Learning for Generalized Zero-Shot Multimodal Information Extraction
by: Zhou, Baohang, et al.
Published: (2026)
by: Zhou, Baohang, et al.
Published: (2026)
Audio Matters Too! Enhancing Markerless Motion Capture with Audio Signals for String Performance Capture
by: Jin, Yitong, et al.
Published: (2024)
by: Jin, Yitong, et al.
Published: (2024)
Contrast then Memorize: Semantic Neighbor Retrieval-Enhanced Inductive Multimodal Knowledge Graph Completion
by: Zhao, Yu, et al.
Published: (2024)
by: Zhao, Yu, et al.
Published: (2024)
Harmony-Aware Music-driven Motion Synthesis with Perceptual Constraint on UGC Datasets
by: Wu, Xinyi, et al.
Published: (2025)
by: Wu, Xinyi, et al.
Published: (2025)
Prototypical Prompting for Text-to-image Person Re-identification
by: Yan, Shuanglin, et al.
Published: (2024)
by: Yan, Shuanglin, et al.
Published: (2024)
MMoFusion: Multi-modal Co-Speech Motion Generation with Diffusion Model
by: Wang, Sen, et al.
Published: (2024)
by: Wang, Sen, et al.
Published: (2024)
Text-controlled Motion Mamba: Text-Instructed Temporal Grounding of Human Motion
by: Wang, Xinghan, et al.
Published: (2024)
by: Wang, Xinghan, et al.
Published: (2024)
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting
by: Ilaslan, Muhammet Furkan, et al.
Published: (2024)
by: Ilaslan, Muhammet Furkan, et al.
Published: (2024)
RFNNS: Robust Fixed Neural Network Steganography with Universal Text-to-Image Models
by: Cheng, Yu, et al.
Published: (2025)
by: Cheng, Yu, et al.
Published: (2025)
COPA: Efficient Vision-Language Pre-training Through Collaborative Object- and Patch-Text Alignment
by: Jiang, Chaoya, et al.
Published: (2023)
by: Jiang, Chaoya, et al.
Published: (2023)
MoTAS: MoE-Guided Feature Selection from TTS-Augmented Speech for Enhanced Multimodal Alzheimer's Early Screening
by: Shao, Yongqi, et al.
Published: (2025)
by: Shao, Yongqi, et al.
Published: (2025)
Rendering-Oriented 3D Point Cloud Attribute Compression using Sparse Tensor-based Transformer
by: Huo, Xiao, et al.
Published: (2024)
by: Huo, Xiao, et al.
Published: (2024)
TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos
by: Kong, Fanheng, et al.
Published: (2025)
by: Kong, Fanheng, et al.
Published: (2025)
Design and Development of Laughter Recognition System Based on Multimodal Fusion and Deep Learning
by: Zhao, Fuzheng, et al.
Published: (2024)
by: Zhao, Fuzheng, et al.
Published: (2024)
TF-Mamba: Text-enhanced Fusion Mamba with Missing Modalities for Robust Multimodal Sentiment Analysis
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
REAL: Realism Evaluation of Text-to-Image Generation Models for Effective Data Augmentation
by: Li, Ran, et al.
Published: (2025)
by: Li, Ran, et al.
Published: (2025)
What's Wrong with the Bottom-up Methods in Arbitrary-shape Scene Text Detection
by: Xu, Chengpei, et al.
Published: (2021)
by: Xu, Chengpei, et al.
Published: (2021)
Towards Structure-aware Model for Multi-modal Knowledge Graph Completion
by: Li, Linyu, et al.
Published: (2025)
by: Li, Linyu, et al.
Published: (2025)
LoVA: Long-form Video-to-Audio Generation
by: Cheng, Xin, et al.
Published: (2024)
by: Cheng, Xin, et al.
Published: (2024)
Similar Items
-
Latency Effects on Multi-Dimensional QoE in Networked VR Whiteboards
by: Song, Jiarun, et al.
Published: (2026) -
SentiAvatar: Towards Expressive and Interactive Digital Humans
by: Jin, Chuhao, et al.
Published: (2026) -
TPIFM: A Task-Aware Model for Evaluating Perceptual Interaction Fluency in Remote AR Collaboration
by: Song, Jiarun, et al.
Published: (2026) -
VSpeechLM: A Visual Speech Language Model for Visual Text-to-Speech Task
by: Wang, Yuyue, et al.
Published: (2025) -
CustomDancer: Customized Dance Recommendation by Text-Dance Retrieval
by: Qin, Yawen, et al.
Published: (2026)