PlanMoGPT: Flow-Enhanced Progressive Planning for Text to Motion Synthesis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jin, Chuhao, Li, Haosen, Zhang, Bingzi, Liu, Che, Wang, Xiting, Song, Ruihua, Huang, Wenbing, Qin, Ying, Zhang, Fuzheng, Zhang, Di
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913906659688448
author Jin, Chuhao
Li, Haosen
Zhang, Bingzi
Liu, Che
Wang, Xiting
Song, Ruihua
Huang, Wenbing
Qin, Ying
Zhang, Fuzheng
Zhang, Di
author_facet Jin, Chuhao
Li, Haosen
Zhang, Bingzi
Liu, Che
Wang, Xiting
Song, Ruihua
Huang, Wenbing
Qin, Ying
Zhang, Fuzheng
Zhang, Di
contents Recent advances in large language models (LLMs) have enabled breakthroughs in many multimodal generation tasks, but a significant performance gap still exists in text-to-motion generation, where LLM-based methods lag far behind non-LLM methods. We identify the granularity of motion tokenization as a critical bottleneck: fine-grained tokenization induces local dependency issues, where LLMs overemphasize short-term coherence at the expense of global semantic alignment, while coarse-grained tokenization sacrifices motion details. To resolve this issue, we propose PlanMoGPT, an LLM-based framework integrating progressive planning and flow-enhanced fine-grained motion tokenization. First, our progressive planning mechanism leverages LLMs' autoregressive capabilities to hierarchically generate motion tokens by starting from sparse global plans and iteratively refining them into full sequences. Second, our flow-enhanced tokenizer doubles the downsampling resolution and expands the codebook size by eight times, minimizing detail loss during discretization, while a flow-enhanced decoder recovers motion nuances. Extensive experiments on text-to-motion benchmarks demonstrate that it achieves state-of-the-art performance, improving FID scores by 63.8% (from 0.380 to 0.141) on long-sequence generation while enhancing motion diversity by 49.9% compared to existing methods. The proposed framework successfully resolves the diversity-quality trade-off that plagues current non-LLM approaches, establishing new standards for text-to-motion generation.
format Preprint
id arxiv_https___arxiv_org_abs_2506_17912
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PlanMoGPT: Flow-Enhanced Progressive Planning for Text to Motion Synthesis
Jin, Chuhao
Li, Haosen
Zhang, Bingzi
Liu, Che
Wang, Xiting
Song, Ruihua
Huang, Wenbing
Qin, Ying
Zhang, Fuzheng
Zhang, Di
Computer Vision and Pattern Recognition
Multimedia
Recent advances in large language models (LLMs) have enabled breakthroughs in many multimodal generation tasks, but a significant performance gap still exists in text-to-motion generation, where LLM-based methods lag far behind non-LLM methods. We identify the granularity of motion tokenization as a critical bottleneck: fine-grained tokenization induces local dependency issues, where LLMs overemphasize short-term coherence at the expense of global semantic alignment, while coarse-grained tokenization sacrifices motion details. To resolve this issue, we propose PlanMoGPT, an LLM-based framework integrating progressive planning and flow-enhanced fine-grained motion tokenization. First, our progressive planning mechanism leverages LLMs' autoregressive capabilities to hierarchically generate motion tokens by starting from sparse global plans and iteratively refining them into full sequences. Second, our flow-enhanced tokenizer doubles the downsampling resolution and expands the codebook size by eight times, minimizing detail loss during discretization, while a flow-enhanced decoder recovers motion nuances. Extensive experiments on text-to-motion benchmarks demonstrate that it achieves state-of-the-art performance, improving FID scores by 63.8% (from 0.380 to 0.141) on long-sequence generation while enhancing motion diversity by 49.9% compared to existing methods. The proposed framework successfully resolves the diversity-quality trade-off that plagues current non-LLM approaches, establishing new standards for text-to-motion generation.
title PlanMoGPT: Flow-Enhanced Progressive Planning for Text to Motion Synthesis
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2506.17912