MC-LLaVA: Multi-Concept Personalized Vision-Language Model

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: An, Ruichuan, Yang, Sihan, Lu, Ming, Zhang, Renrui, Zeng, Kai, Luo, Yulin, Cao, Jiajun, Liang, Hao, Chen, Ying, She, Qi, Zhang, Shanghang, Zhang, Wentao
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910892042485760
author An, Ruichuan
Yang, Sihan
Lu, Ming
Zhang, Renrui
Zeng, Kai
Luo, Yulin
Cao, Jiajun
Liang, Hao
Chen, Ying
She, Qi
Zhang, Shanghang
Zhang, Wentao
author_facet An, Ruichuan
Yang, Sihan
Lu, Ming
Zhang, Renrui
Zeng, Kai
Luo, Yulin
Cao, Jiajun
Liang, Hao
Chen, Ying
She, Qi
Zhang, Shanghang
Zhang, Wentao
contents Current vision-language models (VLMs) show exceptional abilities across diverse tasks, such as visual question answering. To enhance user experience, recent studies investigate VLM personalization to understand user-provided concepts. However, they mainly focus on single-concept personalization, neglecting the existence and interplay of multiple concepts, which limits real-world applicability. This paper proposes the first multi-concept personalization paradigm, MC-LLaVA. Specifically, MC-LLaVA employs a multi-concept instruction tuning strategy, effectively integrating multiple concepts in a single training step. To reduce the costs related to joint training, we propose a personalized textual prompt that uses visual token information to initialize concept tokens. Additionally, we introduce a personalized visual prompt during inference, aggregating location confidence maps for enhanced recognition and grounding capabilities. To advance multi-concept personalization research, we further contribute a high-quality instruction tuning dataset. We carefully collect images with multiple characters and objects from movies and manually generate question-answer samples for multi-concept scenarios, featuring superior diversity. Comprehensive qualitative and quantitative experiments demonstrate that MC-LLaVA can achieve impressive multi-concept personalized responses, paving the way for VLMs to become better user-specific assistants. The code and dataset will be publicly available at https://github.com/arctanxarc/MC-LLaVA}.
format Preprint
id arxiv_https___arxiv_org_abs_2503_18854
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MC-LLaVA: Multi-Concept Personalized Vision-Language Model
An, Ruichuan
Yang, Sihan
Lu, Ming
Zhang, Renrui
Zeng, Kai
Luo, Yulin
Cao, Jiajun
Liang, Hao
Chen, Ying
She, Qi
Zhang, Shanghang
Zhang, Wentao
Computer Vision and Pattern Recognition
Artificial Intelligence
Current vision-language models (VLMs) show exceptional abilities across diverse tasks, such as visual question answering. To enhance user experience, recent studies investigate VLM personalization to understand user-provided concepts. However, they mainly focus on single-concept personalization, neglecting the existence and interplay of multiple concepts, which limits real-world applicability. This paper proposes the first multi-concept personalization paradigm, MC-LLaVA. Specifically, MC-LLaVA employs a multi-concept instruction tuning strategy, effectively integrating multiple concepts in a single training step. To reduce the costs related to joint training, we propose a personalized textual prompt that uses visual token information to initialize concept tokens. Additionally, we introduce a personalized visual prompt during inference, aggregating location confidence maps for enhanced recognition and grounding capabilities. To advance multi-concept personalization research, we further contribute a high-quality instruction tuning dataset. We carefully collect images with multiple characters and objects from movies and manually generate question-answer samples for multi-concept scenarios, featuring superior diversity. Comprehensive qualitative and quantitative experiments demonstrate that MC-LLaVA can achieve impressive multi-concept personalized responses, paving the way for VLMs to become better user-specific assistants. The code and dataset will be publicly available at https://github.com/arctanxarc/MC-LLaVA}.
title MC-LLaVA: Multi-Concept Personalized Vision-Language Model
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2503.18854