Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored Prompts
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866917919176261632 |
|---|---|
| author | Liang, Feng Ma, Haoyu He, Zecheng Hou, Tingbo Hou, Ji Li, Kunpeng Dai, Xiaoliang Juefei-Xu, Felix Azadi, Samaneh Sinha, Animesh Zhang, Peizhao Vajda, Peter Marculescu, Diana |
| author_facet | Liang, Feng Ma, Haoyu He, Zecheng Hou, Tingbo Hou, Ji Li, Kunpeng Dai, Xiaoliang Juefei-Xu, Felix Azadi, Samaneh Sinha, Animesh Zhang, Peizhao Vajda, Peter Marculescu, Diana |
| contents | Video personalization, which generates customized videos using reference images, has gained significant attention. However, prior methods typically focus on single-concept personalization, limiting broader applications that require multi-concept integration. Attempts to extend these models to multiple concepts often lead to identity blending, which results in composite characters with fused attributes from multiple sources. This challenge arises due to the lack of a mechanism to link each concept with its specific reference image. We address this with anchored prompts, which embed image anchors as unique tokens within text prompts, guiding accurate referencing during generation. Additionally, we introduce concept embeddings to encode the order of reference images. Our approach, Movie Weaver, seamlessly weaves multiple concepts-including face, body, and animal images-into one video, allowing flexible combinations in a single model. The evaluation shows that Movie Weaver outperforms existing methods for multi-concept video personalization in identity preservation and overall quality. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2502_07802 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored Prompts Liang, Feng Ma, Haoyu He, Zecheng Hou, Tingbo Hou, Ji Li, Kunpeng Dai, Xiaoliang Juefei-Xu, Felix Azadi, Samaneh Sinha, Animesh Zhang, Peizhao Vajda, Peter Marculescu, Diana Computer Vision and Pattern Recognition Graphics Machine Learning Video personalization, which generates customized videos using reference images, has gained significant attention. However, prior methods typically focus on single-concept personalization, limiting broader applications that require multi-concept integration. Attempts to extend these models to multiple concepts often lead to identity blending, which results in composite characters with fused attributes from multiple sources. This challenge arises due to the lack of a mechanism to link each concept with its specific reference image. We address this with anchored prompts, which embed image anchors as unique tokens within text prompts, guiding accurate referencing during generation. Additionally, we introduce concept embeddings to encode the order of reference images. Our approach, Movie Weaver, seamlessly weaves multiple concepts-including face, body, and animal images-into one video, allowing flexible combinations in a single model. The evaluation shows that Movie Weaver outperforms existing methods for multi-concept video personalization in identity preservation and overall quality. |
| title | Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored Prompts |
| topic | Computer Vision and Pattern Recognition Graphics Machine Learning |
| url | https://arxiv.org/abs/2502.07802 |