ActionPlan: Future-Aware Streaming Motion Synthesis via Frame-Level Action Planning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nazarenus, Eric, Li, Chuqiao, He, Yannan, Xie, Xianghui, Lenssen, Jan Eric, Pons-Moll, Gerard
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908885443411968
author Nazarenus, Eric
Li, Chuqiao
He, Yannan
Xie, Xianghui
Lenssen, Jan Eric
Pons-Moll, Gerard
author_facet Nazarenus, Eric
Li, Chuqiao
He, Yannan
Xie, Xianghui
Lenssen, Jan Eric
Pons-Moll, Gerard
contents We present ActionPlan, a unified motion diffusion framework that bridges real-time streaming with high-quality offline generation within a single model. The core idea is to introduce a per-frame action plan: the model predicts frame-level text latents that act as dense semantic anchors throughout denoising, and uses them to denoise the full motion sequence with combined semantic and motion cues. To support this structured workflow, we design latent-specific diffusion steps, allowing each motion latent to be denoised independently and sampled in flexible orders at inference. As a result, ActionPlan can run in a history-conditioned, future-aware mode for real-time streaming, while also supporting high-quality offline generation. The same mechanism further enables zero-shot motion editing and in-betweening without additional models. Experiments demonstrate that our real-time streaming is 5.25x faster while also achieving 18% motion quality improvement over the best previous method in terms of FID.
format Preprint
id arxiv_https___arxiv_org_abs_2603_13500
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ActionPlan: Future-Aware Streaming Motion Synthesis via Frame-Level Action Planning
Nazarenus, Eric
Li, Chuqiao
He, Yannan
Xie, Xianghui
Lenssen, Jan Eric
Pons-Moll, Gerard
Computer Vision and Pattern Recognition
We present ActionPlan, a unified motion diffusion framework that bridges real-time streaming with high-quality offline generation within a single model. The core idea is to introduce a per-frame action plan: the model predicts frame-level text latents that act as dense semantic anchors throughout denoising, and uses them to denoise the full motion sequence with combined semantic and motion cues. To support this structured workflow, we design latent-specific diffusion steps, allowing each motion latent to be denoised independently and sampled in flexible orders at inference. As a result, ActionPlan can run in a history-conditioned, future-aware mode for real-time streaming, while also supporting high-quality offline generation. The same mechanism further enables zero-shot motion editing and in-betweening without additional models. Experiments demonstrate that our real-time streaming is 5.25x faster while also achieving 18% motion quality improvement over the best previous method in terms of FID.
title ActionPlan: Future-Aware Streaming Motion Synthesis via Frame-Level Action Planning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.13500