Personalized Visual Instruction Tuning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pi, Renjie, Zhang, Jianshu, Han, Tianyang, Zhang, Jipeng, Pan, Rui, Zhang, Tong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912065309900800
author Pi, Renjie
Zhang, Jianshu
Han, Tianyang
Zhang, Jipeng
Pan, Rui
Zhang, Tong
author_facet Pi, Renjie
Zhang, Jianshu
Han, Tianyang
Zhang, Jipeng
Pan, Rui
Zhang, Tong
contents Recent advancements in multimodal large language models (MLLMs) have demonstrated significant progress; however, these models exhibit a notable limitation, which we refer to as "face blindness". Specifically, they can engage in general conversations but fail to conduct personalized dialogues targeting at specific individuals. This deficiency hinders the application of MLLMs in personalized settings, such as tailored visual assistants on mobile devices, or domestic robots that need to recognize members of the family. In this paper, we introduce Personalized Visual Instruction Tuning (PVIT), a novel data curation and training framework designed to enable MLLMs to identify target individuals within an image and engage in personalized and coherent dialogues. Our approach involves the development of a sophisticated pipeline that autonomously generates training data containing personalized conversations. This pipeline leverages the capabilities of various visual experts, image generation models, and (multi-modal) large language models. To evaluate the personalized potential of MLLMs, we present a benchmark called P-Bench, which encompasses various question types with different levels of difficulty. The experiments demonstrate a substantial personalized performance enhancement after fine-tuning with our curated dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2410_07113
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Personalized Visual Instruction Tuning
Pi, Renjie
Zhang, Jianshu
Han, Tianyang
Zhang, Jipeng
Pan, Rui
Zhang, Tong
Computer Vision and Pattern Recognition
Recent advancements in multimodal large language models (MLLMs) have demonstrated significant progress; however, these models exhibit a notable limitation, which we refer to as "face blindness". Specifically, they can engage in general conversations but fail to conduct personalized dialogues targeting at specific individuals. This deficiency hinders the application of MLLMs in personalized settings, such as tailored visual assistants on mobile devices, or domestic robots that need to recognize members of the family. In this paper, we introduce Personalized Visual Instruction Tuning (PVIT), a novel data curation and training framework designed to enable MLLMs to identify target individuals within an image and engage in personalized and coherent dialogues. Our approach involves the development of a sophisticated pipeline that autonomously generates training data containing personalized conversations. This pipeline leverages the capabilities of various visual experts, image generation models, and (multi-modal) large language models. To evaluate the personalized potential of MLLMs, we present a benchmark called P-Bench, which encompasses various question types with different levels of difficulty. The experiments demonstrate a substantial personalized performance enhancement after fine-tuning with our curated dataset.
title Personalized Visual Instruction Tuning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.07113