VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Tu, Yuanpeng, Luo, Hao, Chen, Xi, Ji, Sihui, Bai, Xiang, Zhao, Hengshuang
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916763662286848
author Tu, Yuanpeng
Luo, Hao
Chen, Xi
Ji, Sihui
Bai, Xiang
Zhao, Hengshuang
author_facet Tu, Yuanpeng
Luo, Hao
Chen, Xi
Ji, Sihui
Bai, Xiang
Zhao, Hengshuang
contents Despite significant advancements in video generation, inserting a given object into videos remains a challenging task. The difficulty lies in preserving the appearance details of the reference object and accurately modeling coherent motions at the same time. In this paper, we propose VideoAnydoor, a zero-shot video object insertion framework with high-fidelity detail preservation and precise motion control. Starting from a text-to-video model, we utilize an ID extractor to inject the global identity and leverage a box sequence to control the overall motion. To preserve the detailed appearance and meanwhile support fine-grained motion control, we design a pixel warper. It takes the reference image with arbitrary key-points and the corresponding key-point trajectories as inputs. It warps the pixel details according to the trajectories and fuses the warped features with the diffusion U-Net, thus improving detail preservation and supporting users in manipulating the motion trajectories. In addition, we propose a training strategy involving both videos and static images with a weighted loss to enhance insertion quality. VideoAnydoor demonstrates significant superiority over existing methods and naturally supports various downstream applications (e.g., talking head generation, video virtual try-on, multi-region editing) without task-specific fine-tuning.
format Preprint
id arxiv_https___arxiv_org_abs_2501_01427
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control
Tu, Yuanpeng
Luo, Hao
Chen, Xi
Ji, Sihui
Bai, Xiang
Zhao, Hengshuang
Computer Vision and Pattern Recognition
Despite significant advancements in video generation, inserting a given object into videos remains a challenging task. The difficulty lies in preserving the appearance details of the reference object and accurately modeling coherent motions at the same time. In this paper, we propose VideoAnydoor, a zero-shot video object insertion framework with high-fidelity detail preservation and precise motion control. Starting from a text-to-video model, we utilize an ID extractor to inject the global identity and leverage a box sequence to control the overall motion. To preserve the detailed appearance and meanwhile support fine-grained motion control, we design a pixel warper. It takes the reference image with arbitrary key-points and the corresponding key-point trajectories as inputs. It warps the pixel details according to the trajectories and fuses the warped features with the diffusion U-Net, thus improving detail preservation and supporting users in manipulating the motion trajectories. In addition, we propose a training strategy involving both videos and static images with a weighted loss to enhance insertion quality. VideoAnydoor demonstrates significant superiority over existing methods and naturally supports various downstream applications (e.g., talking head generation, video virtual try-on, multi-region editing) without task-specific fine-tuning.
title VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.01427