Boosting Camera Motion Control for Video Diffusion Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Cheong, Soon Yau, Ceylan, Duygu, Mustafa, Armin, Gilbert, Andrew, Huang, Chun-Hao Paul |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ViscoNet: Bridging and Harmonizing Visual and Textual Conditioning for ControlNet
by: Cheong, Soon Yau, et al.
Published: (2023)
by: Cheong, Soon Yau, et al.
Published: (2023)
Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation
by: Jeong, Hyeonho, et al.
Published: (2024)
by: Jeong, Hyeonho, et al.
Published: (2024)
JOG3R: Towards 3D-Consistent Video Generators
by: Huang, Chun-Hao Paul, et al.
Published: (2025)
by: Huang, Chun-Hao Paul, et al.
Published: (2025)
Memory-V2V: Memory-Augmented Video-to-Video Diffusion for Consistent Multi-Turn Editing
by: Lee, Dohun, et al.
Published: (2026)
by: Lee, Dohun, et al.
Published: (2026)
Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders
by: Lee, Dohun, et al.
Published: (2025)
by: Lee, Dohun, et al.
Published: (2025)
I2VControl-Camera: Precise Video Camera Control with Adjustable Motion Strength
by: Feng, Wanquan, et al.
Published: (2024)
by: Feng, Wanquan, et al.
Published: (2024)
Synergistic Global-space Camera and Human Reconstruction from Videos
by: Zhao, Yizhou, et al.
Published: (2024)
by: Zhao, Yizhou, et al.
Published: (2024)
SuperGaussian: Repurposing Video Models for 3D Super Resolution
by: Shen, Yuan, et al.
Published: (2024)
by: Shen, Yuan, et al.
Published: (2024)
COIN: Control-Inpainting Diffusion Prior for Human and Camera Motion Estimation
by: Li, Jiefeng, et al.
Published: (2024)
by: Li, Jiefeng, et al.
Published: (2024)
VideoHandles: Editing 3D Object Compositions in Videos Using Video Generative Priors
by: Koo, Juil, et al.
Published: (2025)
by: Koo, Juil, et al.
Published: (2025)
DGME-T: Directional Grid Motion Encoding for Transformer-Based Historical Camera Movement Classification
by: Lin, Tingyu, et al.
Published: (2025)
by: Lin, Tingyu, et al.
Published: (2025)
Geometry-Guided Camera Motion Understanding in VideoLLMs
by: Feng, Haoan, et al.
Published: (2026)
by: Feng, Haoan, et al.
Published: (2026)
Video Motion Transfer with Diffusion Transformers
by: Pondaven, Alexander, et al.
Published: (2024)
by: Pondaven, Alexander, et al.
Published: (2024)
Virtually Being: Customizing Camera-Controllable Video Diffusion Models with Multi-View Performance Captures
by: Xu, Yuancheng, et al.
Published: (2025)
by: Xu, Yuancheng, et al.
Published: (2025)
CamViG: Camera Aware Image-to-Video Generation with Multimodal Transformers
by: Marmon, Andrew, et al.
Published: (2024)
by: Marmon, Andrew, et al.
Published: (2024)
LightMotion: A Light and Tuning-free Method for Simulating Camera Motion in Video Generation
by: Song, Quanjian, et al.
Published: (2025)
by: Song, Quanjian, et al.
Published: (2025)
OmniCam: Unified Multimodal Video Generation via Camera Control
by: Yang, Xiaoda, et al.
Published: (2025)
by: Yang, Xiaoda, et al.
Published: (2025)
PipeFlow: Pipelined Processing and Motion-Aware Frame Selection for Long-Form Video Editing
by: Munir, Mustafa, et al.
Published: (2025)
by: Munir, Mustafa, et al.
Published: (2025)
Taming Hallucinations: Boosting MLLMs' Video Understanding via Counterfactual Video Generation
by: Huang, Zhe, et al.
Published: (2025)
by: Huang, Zhe, et al.
Published: (2025)
MotionFlow: Attention-Driven Motion Transfer in Video Diffusion Models
by: Meral, Tuna Han Salih, et al.
Published: (2024)
by: Meral, Tuna Han Salih, et al.
Published: (2024)
ActCam: Zero-Shot Joint Camera and 3D Motion Control for Video Generation
by: Khalifi, Omar El, et al.
Published: (2026)
by: Khalifi, Omar El, et al.
Published: (2026)
MOFA-Video: Controllable Image Animation via Generative Motion Field Adaptions in Frozen Image-to-Video Diffusion Model
by: Niu, Muyao, et al.
Published: (2024)
by: Niu, Muyao, et al.
Published: (2024)
EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance
by: Wang, Zun, et al.
Published: (2025)
by: Wang, Zun, et al.
Published: (2025)
Enabling Versatile Controls for Video Diffusion Models
by: Zhang, Xu, et al.
Published: (2025)
by: Zhang, Xu, et al.
Published: (2025)
Attend-Fusion: Efficient Audio-Visual Fusion for Video Classification
by: Awan, Mahrukh, et al.
Published: (2024)
by: Awan, Mahrukh, et al.
Published: (2024)
Ada-VE: Training-Free Consistent Video Editing Using Adaptive Motion Prior
by: Mahmud, Tanvir, et al.
Published: (2024)
by: Mahmud, Tanvir, et al.
Published: (2024)
SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer
by: Chen, Junsong, et al.
Published: (2025)
by: Chen, Junsong, et al.
Published: (2025)
MemCam: Memory-Augmented Camera Control for Consistent Video Generation
by: Gao, Xinhang, et al.
Published: (2026)
by: Gao, Xinhang, et al.
Published: (2026)
Human-Centric Video Anomaly Detection Through Spatio-Temporal Pose Tokenization and Transformer
by: Noghre, Ghazal Alinezhad, et al.
Published: (2024)
by: Noghre, Ghazal Alinezhad, et al.
Published: (2024)
TeleBoost: A Systematic Alignment Framework for High-Fidelity, Controllable, and Robust Video Generation
by: Liang, Yuanzhi, et al.
Published: (2026)
by: Liang, Yuanzhi, et al.
Published: (2026)
Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models
by: Liu, Huijie, et al.
Published: (2025)
by: Liu, Huijie, et al.
Published: (2025)
MotionShop: Zero-Shot Motion Transfer in Video Diffusion Models with Mixture of Score Guidance
by: Yesiltepe, Hidir, et al.
Published: (2024)
by: Yesiltepe, Hidir, et al.
Published: (2024)
MotionMatcher: Motion Customization of Text-to-Video Diffusion Models via Motion Feature Matching
by: Wu, Yen-Siang, et al.
Published: (2025)
by: Wu, Yen-Siang, et al.
Published: (2025)
Boximator: Generating Rich and Controllable Motions for Video Synthesis
by: Wang, Jiawei, et al.
Published: (2024)
by: Wang, Jiawei, et al.
Published: (2024)
LaMoGen: Laban Movement-Guided Diffusion for Text-to-Motion Generation
by: Kim, Heechang, et al.
Published: (2025)
by: Kim, Heechang, et al.
Published: (2025)
DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers
by: Wang, Lizhen, et al.
Published: (2025)
by: Wang, Lizhen, et al.
Published: (2025)
Towards Understanding Camera Motions in Any Video
by: Lin, Zhiqiu, et al.
Published: (2025)
by: Lin, Zhiqiu, et al.
Published: (2025)
Interpretable Motion-Attentive Maps: Spatio-Temporally Localizing Concepts in Video Diffusion Transformers
by: Jun, Youngjun, et al.
Published: (2026)
by: Jun, Youngjun, et al.
Published: (2026)
LAMP: Language-Assisted Motion Planning for Controllable Video Generation
by: Kizil, Muhammed Burak, et al.
Published: (2025)
by: Kizil, Muhammed Burak, et al.
Published: (2025)
Camera Movement Classification in Historical Footage: A Comparative Study of Deep Video Models
by: Lin, Tingyu, et al.
Published: (2025)
by: Lin, Tingyu, et al.
Published: (2025)
Similar Items
-
ViscoNet: Bridging and Harmonizing Visual and Textual Conditioning for ControlNet
by: Cheong, Soon Yau, et al.
Published: (2023) -
Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation
by: Jeong, Hyeonho, et al.
Published: (2024) -
JOG3R: Towards 3D-Consistent Video Generators
by: Huang, Chun-Hao Paul, et al.
Published: (2025) -
Memory-V2V: Memory-Augmented Video-to-Video Diffusion for Consistent Multi-Turn Editing
by: Lee, Dohun, et al.
Published: (2026) -
Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders
by: Lee, Dohun, et al.
Published: (2025)