TGT: Text-Grounded Trajectories for Locally Controlled Video Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Guofeng, Wang, Angtian, Fang, Jacob Zhiyuan, Jiang, Liming, Yang, Haotian, Liu, Bo, Yang, Yiding, Chen, Guang, Wen, Longyin, Yuille, Alan, Ma, Chongyang
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911216328245248
author Zhang, Guofeng
Wang, Angtian
Fang, Jacob Zhiyuan
Jiang, Liming
Yang, Haotian
Liu, Bo
Yang, Yiding
Chen, Guang
Wen, Longyin
Yuille, Alan
Ma, Chongyang
author_facet Zhang, Guofeng
Wang, Angtian
Fang, Jacob Zhiyuan
Jiang, Liming
Yang, Haotian
Liu, Bo
Yang, Yiding
Chen, Guang
Wen, Longyin
Yuille, Alan
Ma, Chongyang
contents Text-to-video generation has advanced rapidly in visual fidelity, whereas standard methods still have limited ability to control the subject composition of generated scenes. Prior work shows that adding localized text control signals, such as bounding boxes or segmentation masks, can help. However, these methods struggle in complex scenarios and degrade in multi-object settings, offering limited precision and lacking a clear correspondence between individual trajectories and visual entities as the number of controllable objects increases. We introduce Text-Grounded Trajectories (TGT), a framework that conditions video generation on trajectories paired with localized text descriptions. We propose Location-Aware Cross-Attention (LACA) to integrate these signals and adopt a dual-CFG scheme to separately modulate local and global text guidance. In addition, we develop a data processing pipeline that produces trajectories with localized descriptions of tracked entities, and we annotate two million high quality video clips to train TGT. Together, these components enable TGT to use point trajectories as intuitive motion handles, pairing each trajectory with text to control both appearance and motion. Extensive experiments show that TGT achieves higher visual quality, more accurate text alignment, and improved motion controllability compared with prior approaches. Website: https://textgroundedtraj.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2510_15104
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TGT: Text-Grounded Trajectories for Locally Controlled Video Generation
Zhang, Guofeng
Wang, Angtian
Fang, Jacob Zhiyuan
Jiang, Liming
Yang, Haotian
Liu, Bo
Yang, Yiding
Chen, Guang
Wen, Longyin
Yuille, Alan
Ma, Chongyang
Computer Vision and Pattern Recognition
Text-to-video generation has advanced rapidly in visual fidelity, whereas standard methods still have limited ability to control the subject composition of generated scenes. Prior work shows that adding localized text control signals, such as bounding boxes or segmentation masks, can help. However, these methods struggle in complex scenarios and degrade in multi-object settings, offering limited precision and lacking a clear correspondence between individual trajectories and visual entities as the number of controllable objects increases. We introduce Text-Grounded Trajectories (TGT), a framework that conditions video generation on trajectories paired with localized text descriptions. We propose Location-Aware Cross-Attention (LACA) to integrate these signals and adopt a dual-CFG scheme to separately modulate local and global text guidance. In addition, we develop a data processing pipeline that produces trajectories with localized descriptions of tracked entities, and we annotate two million high quality video clips to train TGT. Together, these components enable TGT to use point trajectories as intuitive motion handles, pairing each trajectory with text to control both appearance and motion. Extensive experiments show that TGT achieves higher visual quality, more accurate text alignment, and improved motion controllability compared with prior approaches. Website: https://textgroundedtraj.github.io.
title TGT: Text-Grounded Trajectories for Locally Controlled Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.15104