MotionFlow:Learning Implicit Motion Flow for Complex Camera Trajectory Control in Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908558509998080 |
|---|---|
| author | Lei, Guojun Wang, Chi Wang, Yikai Li, Hong Song, Ying Xu, Weiwei |
| author_facet | Lei, Guojun Wang, Chi Wang, Yikai Li, Hong Song, Ying Xu, Weiwei |
| contents | Generating videos guided by camera trajectories poses significant challenges in achieving consistency and generalizability, particularly when both camera and object motions are present. Existing approaches often attempt to learn these motions separately, which may lead to confusion regarding the relative motion between the camera and the objects. To address this challenge, we propose a novel approach that integrates both camera and object motions by converting them into the motion of corresponding pixels. Utilizing a stable diffusion network, we effectively learn reference motion maps in relation to the specified camera trajectory. These maps, along with an extracted semantic object prior, are then fed into an image-to-video network to generate the desired video that can accurately follow the designated camera trajectory while maintaining consistent object motions. Extensive experiments verify that our model outperforms SOTA methods by a large margin. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_21119 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | MotionFlow:Learning Implicit Motion Flow for Complex Camera Trajectory Control in Video Generation Lei, Guojun Wang, Chi Wang, Yikai Li, Hong Song, Ying Xu, Weiwei Computer Vision and Pattern Recognition Generating videos guided by camera trajectories poses significant challenges in achieving consistency and generalizability, particularly when both camera and object motions are present. Existing approaches often attempt to learn these motions separately, which may lead to confusion regarding the relative motion between the camera and the objects. To address this challenge, we propose a novel approach that integrates both camera and object motions by converting them into the motion of corresponding pixels. Utilizing a stable diffusion network, we effectively learn reference motion maps in relation to the specified camera trajectory. These maps, along with an extracted semantic object prior, are then fed into an image-to-video network to generate the desired video that can accurately follow the designated camera trajectory while maintaining consistent object motions. Extensive experiments verify that our model outperforms SOTA methods by a large margin. |
| title | MotionFlow:Learning Implicit Motion Flow for Complex Camera Trajectory Control in Video Generation |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2509.21119 |