ATI: Any Trajectory Instruction for Controllable Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Angtian, Huang, Haibin, Fang, Jacob Zhiyuan, Yang, Yiding, Ma, Chongyang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909644913377280
author Wang, Angtian
Huang, Haibin
Fang, Jacob Zhiyuan
Yang, Yiding
Ma, Chongyang
author_facet Wang, Angtian
Huang, Haibin
Fang, Jacob Zhiyuan
Yang, Yiding
Ma, Chongyang
contents We propose a unified framework for motion control in video generation that seamlessly integrates camera movement, object-level translation, and fine-grained local motion using trajectory-based inputs. In contrast to prior methods that address these motion types through separate modules or task-specific designs, our approach offers a cohesive solution by projecting user-defined trajectories into the latent space of pre-trained image-to-video generation models via a lightweight motion injector. Users can specify keypoints and their motion paths to control localized deformations, entire object motion, virtual camera dynamics, or combinations of these. The injected trajectory signals guide the generative process to produce temporally consistent and semantically aligned motion sequences. Our framework demonstrates superior performance across multiple video motion control tasks, including stylized motion effects (e.g., motion brushes), dynamic viewpoint changes, and precise local motion manipulation. Experiments show that our method provides significantly better controllability and visual quality compared to prior approaches and commercial solutions, while remaining broadly compatible with various state-of-the-art video generation backbones. Project page: https://anytraj.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2505_22944
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ATI: Any Trajectory Instruction for Controllable Video Generation
Wang, Angtian
Huang, Haibin
Fang, Jacob Zhiyuan
Yang, Yiding
Ma, Chongyang
Computer Vision and Pattern Recognition
Artificial Intelligence
We propose a unified framework for motion control in video generation that seamlessly integrates camera movement, object-level translation, and fine-grained local motion using trajectory-based inputs. In contrast to prior methods that address these motion types through separate modules or task-specific designs, our approach offers a cohesive solution by projecting user-defined trajectories into the latent space of pre-trained image-to-video generation models via a lightweight motion injector. Users can specify keypoints and their motion paths to control localized deformations, entire object motion, virtual camera dynamics, or combinations of these. The injected trajectory signals guide the generative process to produce temporally consistent and semantically aligned motion sequences. Our framework demonstrates superior performance across multiple video motion control tasks, including stylized motion effects (e.g., motion brushes), dynamic viewpoint changes, and precise local motion manipulation. Experiments show that our method provides significantly better controllability and visual quality compared to prior approaches and commercial solutions, while remaining broadly compatible with various state-of-the-art video generation backbones. Project page: https://anytraj.github.io/.
title ATI: Any Trajectory Instruction for Controllable Video Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2505.22944