ATI: Any Trajectory Instruction for Controllable Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Angtian, Huang, Haibin, Fang, Jacob Zhiyuan, Yang, Yiding, Ma, Chongyang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MAGREF: Masked Guidance for Any-Reference Video Generation with Subject Disentanglement
by: Deng, Yufan, et al.
Published: (2025)
by: Deng, Yufan, et al.
Published: (2025)
CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance
by: Deng, Yufan, et al.
Published: (2025)
by: Deng, Yufan, et al.
Published: (2025)
TGT: Text-Grounded Trajectories for Locally Controlled Video Generation
by: Zhang, Guofeng, et al.
Published: (2025)
by: Zhang, Guofeng, et al.
Published: (2025)
VIVA: VLM-Guided Instruction-Based Video Editing with Reward Optimization
by: Cong, Xiaoyan, et al.
Published: (2025)
by: Cong, Xiaoyan, et al.
Published: (2025)
In-Video Instructions: Visual Signals as Generative Control
by: Fang, Gongfan, et al.
Published: (2025)
by: Fang, Gongfan, et al.
Published: (2025)
HECTOR: Hybrid Editable Compositional Object References for Video Generation
by: Zhang, Guofeng, et al.
Published: (2026)
by: Zhang, Guofeng, et al.
Published: (2026)
Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation
by: Wu, Shengqiong, et al.
Published: (2025)
by: Wu, Shengqiong, et al.
Published: (2025)
Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering
by: Wang, Xingrui, et al.
Published: (2024)
by: Wang, Xingrui, et al.
Published: (2024)
Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision
by: Zohar, Orr, et al.
Published: (2024)
by: Zohar, Orr, et al.
Published: (2024)
AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation
by: Gu, Yuchao, et al.
Published: (2026)
by: Gu, Yuchao, et al.
Published: (2026)
Depth Any Video with Scalable Synthetic Data
by: Yang, Honghui, et al.
Published: (2024)
by: Yang, Honghui, et al.
Published: (2024)
X2SAM: Any Segmentation in Images and Videos
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
Any-to-Bokeh: Arbitrary-Subject Video Refocusing with Video Diffusion Model
by: Yang, Yang, et al.
Published: (2025)
by: Yang, Yang, et al.
Published: (2025)
LumosX: Relate Any Identities with Their Attributes for Personalized Video Generation
by: Xing, Jiazheng, et al.
Published: (2026)
by: Xing, Jiazheng, et al.
Published: (2026)
TeleBoost: A Systematic Alignment Framework for High-Fidelity, Controllable, and Robust Video Generation
by: Liang, Yuanzhi, et al.
Published: (2026)
by: Liang, Yuanzhi, et al.
Published: (2026)
AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks
by: Ku, Max, et al.
Published: (2024)
by: Ku, Max, et al.
Published: (2024)
Coordinating Multiple Conditions for Trajectory-Controlled Human Motion Generation
by: Cai, Deli, et al.
Published: (2026)
by: Cai, Deli, et al.
Published: (2026)
Recognize Any Regions
by: Yang, Haosen, et al.
Published: (2023)
by: Yang, Haosen, et al.
Published: (2023)
Controllable Navigation Instruction Generation with Chain of Thought Prompting
by: Kong, Xianghao, et al.
Published: (2024)
by: Kong, Xianghao, et al.
Published: (2024)
VideoPainter: Any-length Video Inpainting and Editing with Plug-and-Play Context Control
by: Bian, Yuxuan, et al.
Published: (2025)
by: Bian, Yuxuan, et al.
Published: (2025)
AnyTrans: Translate AnyText in the Image with Large Scale Models
by: Qian, Zhipeng, et al.
Published: (2024)
by: Qian, Zhipeng, et al.
Published: (2024)
DeMansia: Mamba Never Forgets Any Tokens
by: Fang, Ricky
Published: (2024)
by: Fang, Ricky
Published: (2024)
Think over Trajectories: Leveraging Video Generation to Reconstruct GPS Trajectories from Cellular Signaling
by: Zhang, Ruixing, et al.
Published: (2026)
by: Zhang, Ruixing, et al.
Published: (2026)
VEGGIE: Instructional Editing and Reasoning Video Concepts with Grounded Generation
by: Yu, Shoubin, et al.
Published: (2025)
by: Yu, Shoubin, et al.
Published: (2025)
VIMI: Grounding Video Generation through Multi-modal Instruction
by: Fang, Yuwei, et al.
Published: (2024)
by: Fang, Yuwei, et al.
Published: (2024)
Detecting AI-Generated Video via Frame Consistency
by: Ma, Long, et al.
Published: (2024)
by: Ma, Long, et al.
Published: (2024)
Animate Any Character in Any World
by: Wang, Yitong, et al.
Published: (2025)
by: Wang, Yitong, et al.
Published: (2025)
Generative Unlearning for Any Identity
by: Seo, Juwon, et al.
Published: (2024)
by: Seo, Juwon, et al.
Published: (2024)
YTCommentQA: Video Question Answerability in Instructional Videos
by: Yang, Saelyne, et al.
Published: (2024)
by: Yang, Saelyne, et al.
Published: (2024)
FaceEditTalker: Controllable Talking Head Generation with Facial Attribute Editing
by: Feng, Guanwen, et al.
Published: (2025)
by: Feng, Guanwen, et al.
Published: (2025)
AnyMo: Scaling Any-Modality Conditional Motion Generation with Masked Modeling
by: Li, Yiheng, et al.
Published: (2026)
by: Li, Yiheng, et al.
Published: (2026)
GenXD: Generating Any 3D and 4D Scenes
by: Zhao, Yuyang, et al.
Published: (2024)
by: Zhao, Yuyang, et al.
Published: (2024)
Towards Predicting Any Human Trajectory In Context
by: Fujii, Ryo, et al.
Published: (2025)
by: Fujii, Ryo, et al.
Published: (2025)
X-SAM: From Segment Anything to Any Segmentation
by: Wang, Hao, et al.
Published: (2025)
by: Wang, Hao, et al.
Published: (2025)
VideoGUI: A Benchmark for GUI Automation from Instructional Videos
by: Lin, Kevin Qinghong, et al.
Published: (2024)
by: Lin, Kevin Qinghong, et al.
Published: (2024)
Your One-Stop Solution for AI-Generated Video Detection
by: Ma, Long, et al.
Published: (2026)
by: Ma, Long, et al.
Published: (2026)
Any-to-Any Vision-Language Model for Multimodal X-ray Imaging and Radiological Report Generation
by: Molino, Daniele, et al.
Published: (2025)
by: Molino, Daniele, et al.
Published: (2025)
FSVideo: Fast Speed Video Diffusion Model in a Highly-Compressed Latent Space
by: FSVideo Team, et al.
Published: (2026)
by: FSVideo Team, et al.
Published: (2026)
DepthPilot: From Controllability to Interpretability in Colonoscopy Video Generation
by: Fu, Junhu, et al.
Published: (2026)
by: Fu, Junhu, et al.
Published: (2026)
AnySR: Realizing Image Super-Resolution as Any-Scale, Any-Resource
by: Zhan, Wengyi, et al.
Published: (2024)
by: Zhan, Wengyi, et al.
Published: (2024)
Similar Items
-
MAGREF: Masked Guidance for Any-Reference Video Generation with Subject Disentanglement
by: Deng, Yufan, et al.
Published: (2025) -
CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance
by: Deng, Yufan, et al.
Published: (2025) -
TGT: Text-Grounded Trajectories for Locally Controlled Video Generation
by: Zhang, Guofeng, et al.
Published: (2025) -
VIVA: VLM-Guided Instruction-Based Video Editing with Reward Optimization
by: Cong, Xiaoyan, et al.
Published: (2025) -
In-Video Instructions: Visual Signals as Generative Control
by: Fang, Gongfan, et al.
Published: (2025)