Generative Video Motion Editing with 3D Point Tracks

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lee, Yao-Chih, Zhang, Zhoutong, Huang, Jiahui, Wang, Jui-Hsien, Lee, Joon-Young, Huang, Jia-Bin, Shechtman, Eli, Li, Zhengqi
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918226124865536
author Lee, Yao-Chih
Zhang, Zhoutong
Huang, Jiahui
Wang, Jui-Hsien
Lee, Joon-Young
Huang, Jia-Bin
Shechtman, Eli
Li, Zhengqi
author_facet Lee, Yao-Chih
Zhang, Zhoutong
Huang, Jiahui
Wang, Jui-Hsien
Lee, Joon-Young
Huang, Jia-Bin
Shechtman, Eli
Li, Zhengqi
contents Camera and object motions are central to a video's narrative. However, precisely editing these captured motions remains a significant challenge, especially under complex object movements. Current motion-controlled image-to-video (I2V) approaches often lack full-scene context for consistent video editing, while video-to-video (V2V) methods provide viewpoint changes or basic object translation, but offer limited control over fine-grained object motion. We present a track-conditioned V2V framework that enables joint editing of camera and object motion. We achieve this by conditioning a video generation model on a source video and paired 3D point tracks representing source and target motions. These 3D tracks establish sparse correspondences that transfer rich context from the source video to new motions while preserving spatiotemporal coherence. Crucially, compared to 2D tracks, 3D tracks provide explicit depth cues, allowing the model to resolve depth order and handle occlusions for precise motion editing. Trained in two stages on synthetic and real data, our model supports diverse motion edits, including joint camera/object manipulation, motion transfer, and non-rigid deformation, unlocking new creative potential in video editing.
format Preprint
id arxiv_https___arxiv_org_abs_2512_02015
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Generative Video Motion Editing with 3D Point Tracks
Lee, Yao-Chih
Zhang, Zhoutong
Huang, Jiahui
Wang, Jui-Hsien
Lee, Joon-Young
Huang, Jia-Bin
Shechtman, Eli
Li, Zhengqi
Computer Vision and Pattern Recognition
Camera and object motions are central to a video's narrative. However, precisely editing these captured motions remains a significant challenge, especially under complex object movements. Current motion-controlled image-to-video (I2V) approaches often lack full-scene context for consistent video editing, while video-to-video (V2V) methods provide viewpoint changes or basic object translation, but offer limited control over fine-grained object motion. We present a track-conditioned V2V framework that enables joint editing of camera and object motion. We achieve this by conditioning a video generation model on a source video and paired 3D point tracks representing source and target motions. These 3D tracks establish sparse correspondences that transfer rich context from the source video to new motions while preserving spatiotemporal coherence. Crucially, compared to 2D tracks, 3D tracks provide explicit depth cues, allowing the model to resolve depth order and handle occlusions for precise motion editing. Trained in two stages on synthetic and real data, our model supports diverse motion edits, including joint camera/object manipulation, motion transfer, and non-rigid deformation, unlocking new creative potential in video editing.
title Generative Video Motion Editing with 3D Point Tracks
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.02015