DiTraj: training-free trajectory control for video diffusion transformer

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lei, Cheng, Zhang, Jiayu, Ma, Yue, Wang, Xinyu, Chen, Long, Tang, Liang, Yan, Yiqiang, Su, Fei, Zhao, Zhicheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909814411493376
author Lei, Cheng
Zhang, Jiayu
Ma, Yue
Wang, Xinyu
Chen, Long
Tang, Liang
Yan, Yiqiang
Su, Fei
Zhao, Zhicheng
author_facet Lei, Cheng
Zhang, Jiayu
Ma, Yue
Wang, Xinyu
Chen, Long
Tang, Liang
Yan, Yiqiang
Su, Fei
Zhao, Zhicheng
contents Diffusion Transformers (DiT)-based video generation models with 3D full attention exhibit strong generative capabilities. Trajectory control represents a user-friendly task in the field of controllable video generation. However, existing methods either require substantial training resources or are specifically designed for U-Net, do not take advantage of the superior performance of DiT. To address these issues, we propose DiTraj, a simple but effective training-free framework for trajectory control in text-to-video generation, tailored for DiT. Specifically, first, to inject the object's trajectory, we propose foreground-background separation guidance: we use the Large Language Model (LLM) to convert user-provided prompts into foreground and background prompts, which respectively guide the generation of foreground and background regions in the video. Then, we analyze 3D full attention and explore the tight correlation between inter-token attention scores and position embedding. Based on this, we propose inter-frame Spatial-Temporal Decoupled 3D-RoPE (STD-RoPE). By modifying only foreground tokens' position embedding, STD-RoPE eliminates their cross-frame spatial discrepancies, strengthening cross-frame attention among them and thus enhancing trajectory control. Additionally, we achieve 3D-aware trajectory control by regulating the density of position embedding. Extensive experiments demonstrate that our method outperforms previous methods in both video quality and trajectory controllability.
format Preprint
id arxiv_https___arxiv_org_abs_2509_21839
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DiTraj: training-free trajectory control for video diffusion transformer
Lei, Cheng
Zhang, Jiayu
Ma, Yue
Wang, Xinyu
Chen, Long
Tang, Liang
Yan, Yiqiang
Su, Fei
Zhao, Zhicheng
Computer Vision and Pattern Recognition
Artificial Intelligence
Diffusion Transformers (DiT)-based video generation models with 3D full attention exhibit strong generative capabilities. Trajectory control represents a user-friendly task in the field of controllable video generation. However, existing methods either require substantial training resources or are specifically designed for U-Net, do not take advantage of the superior performance of DiT. To address these issues, we propose DiTraj, a simple but effective training-free framework for trajectory control in text-to-video generation, tailored for DiT. Specifically, first, to inject the object's trajectory, we propose foreground-background separation guidance: we use the Large Language Model (LLM) to convert user-provided prompts into foreground and background prompts, which respectively guide the generation of foreground and background regions in the video. Then, we analyze 3D full attention and explore the tight correlation between inter-token attention scores and position embedding. Based on this, we propose inter-frame Spatial-Temporal Decoupled 3D-RoPE (STD-RoPE). By modifying only foreground tokens' position embedding, STD-RoPE eliminates their cross-frame spatial discrepancies, strengthening cross-frame attention among them and thus enhancing trajectory control. Additionally, we achieve 3D-aware trajectory control by regulating the density of position embedding. Extensive experiments demonstrate that our method outperforms previous methods in both video quality and trajectory controllability.
title DiTraj: training-free trajectory control for video diffusion transformer
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2509.21839