TrackGo: A Flexible and Efficient Method for Controllable Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Haitao, Wang, Chuang, Nie, Rui, Liu, Jinlin, Yu, Dongdong, Yu, Qian, Wang, Changhu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Taming Camera-Controlled Video Generation with Verifiable Geometry Reward
by: Wang, Zhaoqing, et al.
Published: (2025)
by: Wang, Zhaoqing, et al.
Published: (2025)
MMGen: Unified Multi-modal Image Generation and Understanding in One Go
by: Wang, Jiepeng, et al.
Published: (2025)
by: Wang, Jiepeng, et al.
Published: (2025)
FlexTraj: Image-to-Video Generation with Flexible Point Trajectory Control
by: Zhang, Zhiyuan, et al.
Published: (2025)
by: Zhang, Zhiyuan, et al.
Published: (2025)
ViewCraft3D: High-Fidelity and View-Consistent 3D Vector Graphics Synthesis
by: Wang, Chuang, et al.
Published: (2025)
by: Wang, Chuang, et al.
Published: (2025)
OneWorld: Taming Scene Generation with 3D Unified Representation Autoencoder
by: Gao, Sensen, et al.
Published: (2026)
by: Gao, Sensen, et al.
Published: (2026)
LaVin-DiT: Large Vision Diffusion Transformer
by: Wang, Zhaoqing, et al.
Published: (2024)
by: Wang, Zhaoqing, et al.
Published: (2024)
SVGDreamer: Text Guided SVG Generation with Diffusion Model
by: Xing, Ximing, et al.
Published: (2023)
by: Xing, Ximing, et al.
Published: (2023)
SVGDreamer++: Advancing Editability and Diversity in Text-Guided SVG Generation
by: Xing, Ximing, et al.
Published: (2024)
by: Xing, Ximing, et al.
Published: (2024)
Deep Understanding of Soccer Match Videos
by: Xu, Shikun, et al.
Published: (2024)
by: Xu, Shikun, et al.
Published: (2024)
Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT
by: Liu, Dongyang, et al.
Published: (2025)
by: Liu, Dongyang, et al.
Published: (2025)
Inversion-by-Inversion: Exemplar-based Sketch-to-Photo Synthesis via Stochastic Differential Equations without Training
by: Xing, Ximing, et al.
Published: (2023)
by: Xing, Ximing, et al.
Published: (2023)
DiffSketcher: Text Guided Vector Sketch Synthesis through Latent Diffusion Models
by: Xing, Ximing, et al.
Published: (2023)
by: Xing, Ximing, et al.
Published: (2023)
VAnim: Rendering-Aware Sparse State Modeling for Structure-Preserving Vector Animation
by: Liang, Guotao, et al.
Published: (2026)
by: Liang, Guotao, et al.
Published: (2026)
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding
by: Zhang, Yunzhu, et al.
Published: (2025)
by: Zhang, Yunzhu, et al.
Published: (2025)
CoopTrack: Exploring End-to-End Learning for Efficient Cooperative Sequential Perception
by: Zhong, Jiaru, et al.
Published: (2025)
by: Zhong, Jiaru, et al.
Published: (2025)
OmniTracker: Unifying Object Tracking by Tracking-with-Detection
by: Wang, Junke, et al.
Published: (2023)
by: Wang, Junke, et al.
Published: (2023)
Disentangling Foreground and Background Motion for Enhanced Realism in Human Video Generation
by: Liu, Jinlin, et al.
Published: (2024)
by: Liu, Jinlin, et al.
Published: (2024)
EasyControl: Adding Efficient and Flexible Control for Diffusion Transformer
by: Zhang, Yuxuan, et al.
Published: (2025)
by: Zhang, Yuxuan, et al.
Published: (2025)
MyGo: Consistent and Controllable Multi-View Driving Video Generation with Camera Control
by: Yao, Yining, et al.
Published: (2024)
by: Yao, Yining, et al.
Published: (2024)
Dimension-Reduction Attack! Video Generative Models are Experts on Controllable Image Synthesis
by: Cao, Hengyuan, et al.
Published: (2025)
by: Cao, Hengyuan, et al.
Published: (2025)
FreqTrack: Frequency Learning based Vision Transformer for RGB-Event Object Tracking
by: You, Jinlin, et al.
Published: (2026)
by: You, Jinlin, et al.
Published: (2026)
PARE: Pruning and Adaptive Routing for Efficient Video Generation
by: Wang, Yutong, et al.
Published: (2026)
by: Wang, Yutong, et al.
Published: (2026)
Text-Animator: Controllable Visual Text Video Generation
by: Liu, Lin, et al.
Published: (2024)
by: Liu, Lin, et al.
Published: (2024)
TVG: A Training-free Transition Video Generation Method with Diffusion Models
by: Zhang, Rui, et al.
Published: (2024)
by: Zhang, Rui, et al.
Published: (2024)
CamCo: Camera-Controllable 3D-Consistent Image-to-Video Generation
by: Xu, Dejia, et al.
Published: (2024)
by: Xu, Dejia, et al.
Published: (2024)
OmniVid: A Generative Framework for Universal Video Understanding
by: Wang, Junke, et al.
Published: (2024)
by: Wang, Junke, et al.
Published: (2024)
GoTrack: Generic 6DoF Object Pose Refinement and Tracking
by: Nguyen, Van Nguyen, et al.
Published: (2025)
by: Nguyen, Van Nguyen, et al.
Published: (2025)
Make Your Training Flexible: Towards Deployment-Efficient Video Models
by: Wang, Chenting, et al.
Published: (2025)
by: Wang, Chenting, et al.
Published: (2025)
FlexEControl: Flexible and Efficient Multimodal Control for Text-to-Image Generation
by: He, Xuehai, et al.
Published: (2024)
by: He, Xuehai, et al.
Published: (2024)
VideoRFT: Incentivizing Video Reasoning Capability in MLLMs via Reinforced Fine-Tuning
by: Wang, Qi, et al.
Published: (2025)
by: Wang, Qi, et al.
Published: (2025)
I4VGen: Image as Free Stepping Stone for Text-to-Video Generation
by: Guo, Xiefan, et al.
Published: (2024)
by: Guo, Xiefan, et al.
Published: (2024)
ControlNeXt: Powerful and Efficient Control for Image and Video Generation
by: Peng, Bohao, et al.
Published: (2024)
by: Peng, Bohao, et al.
Published: (2024)
360DVD: Controllable Panorama Video Generation with 360-Degree Video Diffusion Model
by: Wang, Qian, et al.
Published: (2024)
by: Wang, Qian, et al.
Published: (2024)
LongDiff: Training-Free Long Video Generation in One Go
by: Li, Zhuoling, et al.
Published: (2025)
by: Li, Zhuoling, et al.
Published: (2025)
DyST-XL: Dynamic Layout Planning and Content Control for Compositional Text-to-Video Generation
by: He, Weijie, et al.
Published: (2025)
by: He, Weijie, et al.
Published: (2025)
Motion Control for Enhanced Complex Action Video Generation
by: Zhou, Qiang, et al.
Published: (2024)
by: Zhou, Qiang, et al.
Published: (2024)
Render-in-the-Loop: Vector Graphics Generation via Visual Self-Feedback
by: Liang, Guotao, et al.
Published: (2026)
by: Liang, Guotao, et al.
Published: (2026)
ShotDirector: Directorially Controllable Multi-Shot Video Generation with Cinematographic Transitions
by: Wu, Xiaoxue, et al.
Published: (2025)
by: Wu, Xiaoxue, et al.
Published: (2025)
GoMAvatar: Efficient Animatable Human Modeling from Monocular Video Using Gaussians-on-Mesh
by: Wen, Jing, et al.
Published: (2024)
by: Wen, Jing, et al.
Published: (2024)
FlexiFilm: Long Video Generation with Flexible Conditions
by: Ouyang, Yichen, et al.
Published: (2024)
by: Ouyang, Yichen, et al.
Published: (2024)
Similar Items
-
Taming Camera-Controlled Video Generation with Verifiable Geometry Reward
by: Wang, Zhaoqing, et al.
Published: (2025) -
MMGen: Unified Multi-modal Image Generation and Understanding in One Go
by: Wang, Jiepeng, et al.
Published: (2025) -
FlexTraj: Image-to-Video Generation with Flexible Point Trajectory Control
by: Zhang, Zhiyuan, et al.
Published: (2025) -
ViewCraft3D: High-Fidelity and View-Consistent 3D Vector Graphics Synthesis
by: Wang, Chuang, et al.
Published: (2025) -
OneWorld: Taming Scene Generation with 3D Unified Representation Autoencoder
by: Gao, Sensen, et al.
Published: (2026)