Real-Time Motion-Controllable Autoregressive Video Diffusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Kesen, Shi, Jiaxin, Zhu, Beier, Zhou, Junbao, Shen, Xiaolong, Zhou, Yuan, Sun, Qianru, Zhang, Hanwang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911494772359168
author Zhao, Kesen
Shi, Jiaxin
Zhu, Beier
Zhou, Junbao
Shen, Xiaolong
Zhou, Yuan
Sun, Qianru
Zhang, Hanwang
author_facet Zhao, Kesen
Shi, Jiaxin
Zhu, Beier
Zhou, Junbao
Shen, Xiaolong
Zhou, Yuan
Sun, Qianru
Zhang, Hanwang
contents Real-time motion-controllable video generation remains challenging due to the inherent latency of bidirectional diffusion models and the lack of effective autoregressive (AR) approaches. Existing AR video diffusion models are limited to simple control signals or text-to-video generation, and often suffer from quality degradation and motion artifacts in few-step generation. To address these challenges, we propose AR-Drag, the first RL-enhanced few-step AR video diffusion model for real-time image-to-video generation with diverse motion control. We first fine-tune a base I2V model to support basic motion control, then further improve it via reinforcement learning with a trajectory-based reward model. Our design preserves the Markov property through a Self-Rollout mechanism and accelerates training by selectively introducing stochasticity in denoising steps. Extensive experiments demonstrate that AR-Drag achieves high visual fidelity and precise motion alignment, significantly reducing latency compared with state-of-the-art motion-controllable VDMs, while using only 1.3B parameters. Additional visualizations can be found on our project page: https://kesenzhao.github.io/AR-Drag.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2510_08131
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Real-Time Motion-Controllable Autoregressive Video Diffusion
Zhao, Kesen
Shi, Jiaxin
Zhu, Beier
Zhou, Junbao
Shen, Xiaolong
Zhou, Yuan
Sun, Qianru
Zhang, Hanwang
Computer Vision and Pattern Recognition
Real-time motion-controllable video generation remains challenging due to the inherent latency of bidirectional diffusion models and the lack of effective autoregressive (AR) approaches. Existing AR video diffusion models are limited to simple control signals or text-to-video generation, and often suffer from quality degradation and motion artifacts in few-step generation. To address these challenges, we propose AR-Drag, the first RL-enhanced few-step AR video diffusion model for real-time image-to-video generation with diverse motion control. We first fine-tune a base I2V model to support basic motion control, then further improve it via reinforcement learning with a trajectory-based reward model. Our design preserves the Markov property through a Self-Rollout mechanism and accelerates training by selectively introducing stochasticity in denoising steps. Extensive experiments demonstrate that AR-Drag achieves high visual fidelity and precise motion alignment, significantly reducing latency compared with state-of-the-art motion-controllable VDMs, while using only 1.3B parameters. Additional visualizations can be found on our project page: https://kesenzhao.github.io/AR-Drag.github.io/.
title Real-Time Motion-Controllable Autoregressive Video Diffusion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.08131