ConceptMaster: Multi-Concept Video Customization on Diffusion Transformer Models Without Test-Time Tuning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Huang, Yuzhou, Yuan, Ziyang, Liu, Quande, Wang, Qiulin, Wang, Xintao, Zhang, Ruimao, Wan, Pengfei, Zhang, Di, Gai, Kun
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908361206792192
author Huang, Yuzhou
Yuan, Ziyang
Liu, Quande
Wang, Qiulin
Wang, Xintao
Zhang, Ruimao
Wan, Pengfei
Zhang, Di
Gai, Kun
author_facet Huang, Yuzhou
Yuan, Ziyang
Liu, Quande
Wang, Qiulin
Wang, Xintao
Zhang, Ruimao
Wan, Pengfei
Zhang, Di
Gai, Kun
contents Text-to-video generation has made remarkable advancements through diffusion models. However, Multi-Concept Video Customization (MCVC) remains a significant challenge. We identify two key challenges for this task: 1) the identity decoupling issue, where directly adopting existing customization methods inevitably mix identity attributes when handling multiple concepts simultaneously, and 2) the scarcity of high-quality video-entity pairs, which is crucial for training a model that can well represent and decouple various customized concepts in video generation. To address these challenges, we introduce ConceptMaster, a novel framework that effectively addresses the identity decoupling issues while maintaining concept fidelity in video customization. Specifically, we propose to learn decoupled multi-concept embeddings and inject them into diffusion models in a standalone manner, which effectively guarantees the quality of customized videos with multiple identities, even for highly similar visual concepts. To overcome the scarcity of high-quality MCVC data, we establish a data construction pipeline, which enables collection of high-quality multi-concept video-entity data pairs across diverse scenarios. A multi-concept video evaluation set is further devised to comprehensively validate our method from three dimensions, including concept fidelity, identity decoupling ability, and video generation quality, across six different concept composition scenarios. Extensive experiments demonstrate that ConceptMaster significantly outperforms previous methods for video customization tasks, showing great potential to generate personalized and semantically accurate content for video diffusion models.
format Preprint
id arxiv_https___arxiv_org_abs_2501_04698
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ConceptMaster: Multi-Concept Video Customization on Diffusion Transformer Models Without Test-Time Tuning
Huang, Yuzhou
Yuan, Ziyang
Liu, Quande
Wang, Qiulin
Wang, Xintao
Zhang, Ruimao
Wan, Pengfei
Zhang, Di
Gai, Kun
Computer Vision and Pattern Recognition
Text-to-video generation has made remarkable advancements through diffusion models. However, Multi-Concept Video Customization (MCVC) remains a significant challenge. We identify two key challenges for this task: 1) the identity decoupling issue, where directly adopting existing customization methods inevitably mix identity attributes when handling multiple concepts simultaneously, and 2) the scarcity of high-quality video-entity pairs, which is crucial for training a model that can well represent and decouple various customized concepts in video generation. To address these challenges, we introduce ConceptMaster, a novel framework that effectively addresses the identity decoupling issues while maintaining concept fidelity in video customization. Specifically, we propose to learn decoupled multi-concept embeddings and inject them into diffusion models in a standalone manner, which effectively guarantees the quality of customized videos with multiple identities, even for highly similar visual concepts. To overcome the scarcity of high-quality MCVC data, we establish a data construction pipeline, which enables collection of high-quality multi-concept video-entity data pairs across diverse scenarios. A multi-concept video evaluation set is further devised to comprehensively validate our method from three dimensions, including concept fidelity, identity decoupling ability, and video generation quality, across six different concept composition scenarios. Extensive experiments demonstrate that ConceptMaster significantly outperforms previous methods for video customization tasks, showing great potential to generate personalized and semantically accurate content for video diffusion models.
title ConceptMaster: Multi-Concept Video Customization on Diffusion Transformer Models Without Test-Time Tuning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.04698