Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored Prompts

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Liang, Feng, Ma, Haoyu, He, Zecheng, Hou, Tingbo, Hou, Ji, Li, Kunpeng, Dai, Xiaoliang, Juefei-Xu, Felix, Azadi, Samaneh, Sinha, Animesh, Zhang, Peizhao, Vajda, Peter, Marculescu, Diana
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917919176261632
author Liang, Feng
Ma, Haoyu
He, Zecheng
Hou, Tingbo
Hou, Ji
Li, Kunpeng
Dai, Xiaoliang
Juefei-Xu, Felix
Azadi, Samaneh
Sinha, Animesh
Zhang, Peizhao
Vajda, Peter
Marculescu, Diana
author_facet Liang, Feng
Ma, Haoyu
He, Zecheng
Hou, Tingbo
Hou, Ji
Li, Kunpeng
Dai, Xiaoliang
Juefei-Xu, Felix
Azadi, Samaneh
Sinha, Animesh
Zhang, Peizhao
Vajda, Peter
Marculescu, Diana
contents Video personalization, which generates customized videos using reference images, has gained significant attention. However, prior methods typically focus on single-concept personalization, limiting broader applications that require multi-concept integration. Attempts to extend these models to multiple concepts often lead to identity blending, which results in composite characters with fused attributes from multiple sources. This challenge arises due to the lack of a mechanism to link each concept with its specific reference image. We address this with anchored prompts, which embed image anchors as unique tokens within text prompts, guiding accurate referencing during generation. Additionally, we introduce concept embeddings to encode the order of reference images. Our approach, Movie Weaver, seamlessly weaves multiple concepts-including face, body, and animal images-into one video, allowing flexible combinations in a single model. The evaluation shows that Movie Weaver outperforms existing methods for multi-concept video personalization in identity preservation and overall quality.
format Preprint
id arxiv_https___arxiv_org_abs_2502_07802
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored Prompts
Liang, Feng
Ma, Haoyu
He, Zecheng
Hou, Tingbo
Hou, Ji
Li, Kunpeng
Dai, Xiaoliang
Juefei-Xu, Felix
Azadi, Samaneh
Sinha, Animesh
Zhang, Peizhao
Vajda, Peter
Marculescu, Diana
Computer Vision and Pattern Recognition
Graphics
Machine Learning
Video personalization, which generates customized videos using reference images, has gained significant attention. However, prior methods typically focus on single-concept personalization, limiting broader applications that require multi-concept integration. Attempts to extend these models to multiple concepts often lead to identity blending, which results in composite characters with fused attributes from multiple sources. This challenge arises due to the lack of a mechanism to link each concept with its specific reference image. We address this with anchored prompts, which embed image anchors as unique tokens within text prompts, guiding accurate referencing during generation. Additionally, we introduce concept embeddings to encode the order of reference images. Our approach, Movie Weaver, seamlessly weaves multiple concepts-including face, body, and animal images-into one video, allowing flexible combinations in a single model. The evaluation shows that Movie Weaver outperforms existing methods for multi-concept video personalization in identity preservation and overall quality.
title Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored Prompts
topic Computer Vision and Pattern Recognition
Graphics
Machine Learning
url https://arxiv.org/abs/2502.07802