DragEntity: Trajectory Guided Video Generation using Entity and Positional Relationships

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wan, Zhang, Tang, Sheng, Wei, Jiawei, Zhang, Ruize, Cao, Juan
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929541397610496
author Wan, Zhang
Tang, Sheng
Wei, Jiawei
Zhang, Ruize
Cao, Juan
author_facet Wan, Zhang
Tang, Sheng
Wei, Jiawei
Zhang, Ruize
Cao, Juan
contents In recent years, diffusion models have achieved tremendous success in the field of video generation, with controllable video generation receiving significant attention. However, existing control methods still face two limitations: Firstly, control conditions (such as depth maps, 3D Mesh) are difficult for ordinary users to obtain directly. Secondly, it's challenging to drive multiple objects through complex motions with multiple trajectories simultaneously. In this paper, we introduce DragEntity, a video generation model that utilizes entity representation for controlling the motion of multiple objects. Compared to previous methods, DragEntity offers two main advantages: 1) Our method is more user-friendly for interaction because it allows users to drag entities within the image rather than individual pixels. 2) We use entity representation to represent any object in the image, and multiple objects can maintain relative spatial relationships. Therefore, we allow multiple trajectories to control multiple objects in the image with different levels of complexity simultaneously. Our experiments validate the effectiveness of DragEntity, demonstrating its excellent performance in fine-grained control in video generation.
format Preprint
id arxiv_https___arxiv_org_abs_2410_10751
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DragEntity: Trajectory Guided Video Generation using Entity and Positional Relationships
Wan, Zhang
Tang, Sheng
Wei, Jiawei
Zhang, Ruize
Cao, Juan
Computer Vision and Pattern Recognition
In recent years, diffusion models have achieved tremendous success in the field of video generation, with controllable video generation receiving significant attention. However, existing control methods still face two limitations: Firstly, control conditions (such as depth maps, 3D Mesh) are difficult for ordinary users to obtain directly. Secondly, it's challenging to drive multiple objects through complex motions with multiple trajectories simultaneously. In this paper, we introduce DragEntity, a video generation model that utilizes entity representation for controlling the motion of multiple objects. Compared to previous methods, DragEntity offers two main advantages: 1) Our method is more user-friendly for interaction because it allows users to drag entities within the image rather than individual pixels. 2) We use entity representation to represent any object in the image, and multiple objects can maintain relative spatial relationships. Therefore, we allow multiple trajectories to control multiple objects in the image with different levels of complexity simultaneously. Our experiments validate the effectiveness of DragEntity, demonstrating its excellent performance in fine-grained control in video generation.
title DragEntity: Trajectory Guided Video Generation using Entity and Positional Relationships
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.10751