FlexTraj: Image-to-Video Generation with Flexible Point Trajectory Control

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Zhiyuan, Wang, Can, Chen, Dongdong, Liao, Jing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912639878168576
author Zhang, Zhiyuan
Wang, Can
Chen, Dongdong
Liao, Jing
author_facet Zhang, Zhiyuan
Wang, Can
Chen, Dongdong
Liao, Jing
contents We present FlexTraj, a framework for image-to-video generation with flexible point trajectory control. FlexTraj introduces a unified point-based motion representation that encodes each point with a segmentation ID, a temporally consistent trajectory ID, and an optional color channel for appearance cues, enabling both dense and sparse trajectory control. Instead of injecting trajectory conditions into the video generator through token concatenation or ControlNet, FlexTraj employs an efficient sequence-concatenation scheme that achieves faster convergence, stronger controllability, and more efficient inference, while maintaining robustness under unaligned conditions. To train such a unified point trajectory-controlled video generator, FlexTraj adopts an annealing training strategy that gradually reduces reliance on complete supervision and aligned condition. Experimental results demonstrate that FlexTraj enables multi-granularity, alignment-agnostic trajectory control for video generation, supporting various applications such as motion cloning, drag-based image-to-video, motion interpolation, camera redirection, flexible action control and mesh animations.
format Preprint
id arxiv_https___arxiv_org_abs_2510_08527
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FlexTraj: Image-to-Video Generation with Flexible Point Trajectory Control
Zhang, Zhiyuan
Wang, Can
Chen, Dongdong
Liao, Jing
Computer Vision and Pattern Recognition
We present FlexTraj, a framework for image-to-video generation with flexible point trajectory control. FlexTraj introduces a unified point-based motion representation that encodes each point with a segmentation ID, a temporally consistent trajectory ID, and an optional color channel for appearance cues, enabling both dense and sparse trajectory control. Instead of injecting trajectory conditions into the video generator through token concatenation or ControlNet, FlexTraj employs an efficient sequence-concatenation scheme that achieves faster convergence, stronger controllability, and more efficient inference, while maintaining robustness under unaligned conditions. To train such a unified point trajectory-controlled video generator, FlexTraj adopts an annealing training strategy that gradually reduces reliance on complete supervision and aligned condition. Experimental results demonstrate that FlexTraj enables multi-granularity, alignment-agnostic trajectory control for video generation, supporting various applications such as motion cloning, drag-based image-to-video, motion interpolation, camera redirection, flexible action control and mesh animations.
title FlexTraj: Image-to-Video Generation with Flexible Point Trajectory Control
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.08527