PVChat: Personalized Video Chat with One-Shot Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Shi, Yufei, Yan, Weilong, Xu, Gang, Li, Yumeng, Chen, Yucheng, Li, Zhenxi, Yu, Fei Richard, Li, Ming, Yeo, Si Yong
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915447831527424
author Shi, Yufei
Yan, Weilong
Xu, Gang
Li, Yumeng
Chen, Yucheng
Li, Zhenxi
Yu, Fei Richard
Li, Ming
Yeo, Si Yong
author_facet Shi, Yufei
Yan, Weilong
Xu, Gang
Li, Yumeng
Chen, Yucheng
Li, Zhenxi
Yu, Fei Richard
Li, Ming
Yeo, Si Yong
contents Video large language models (ViLLMs) excel in general video understanding, e.g., recognizing activities like talking and eating, but struggle with identity-aware comprehension, such as "Wilson is receiving chemotherapy" or "Tom is discussing with Sarah", limiting their applicability in smart healthcare and smart home environments. To address this limitation, we propose a one-shot learning framework PVChat, the first personalized ViLLM that enables subject-aware question answering (QA) from a single video for each subject. Our approach optimizes a Mixture-of-Heads (MoH) enhanced ViLLM on a synthetically augmented video-QA dataset, leveraging a progressive image-to-video learning strategy. Specifically, we introduce an automated augmentation pipeline that synthesizes identity-preserving positive samples and retrieves hard negatives from existing video corpora, generating a diverse training dataset with four QA types: existence, appearance, action, and location inquiries. To enhance subject-specific learning, we propose a ReLU Routing MoH attention mechanism, alongside two novel objectives: (1) Smooth Proximity Regularization for progressive learning through exponential distance scaling and (2) Head Activation Enhancement for balanced attention routing. Finally, we adopt a two-stage training strategy, transitioning from image pre-training to video fine-tuning, enabling a gradual learning process from static attributes to dynamic representations. We evaluate PVChat on diverse datasets covering medical scenarios, TV series, anime, and real-world footage, demonstrating its superiority in personalized feature understanding after learning from a single video, compared to state-of-the-art ViLLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2503_17069
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PVChat: Personalized Video Chat with One-Shot Learning
Shi, Yufei
Yan, Weilong
Xu, Gang
Li, Yumeng
Chen, Yucheng
Li, Zhenxi
Yu, Fei Richard
Li, Ming
Yeo, Si Yong
Computer Vision and Pattern Recognition
Artificial Intelligence
Video large language models (ViLLMs) excel in general video understanding, e.g., recognizing activities like talking and eating, but struggle with identity-aware comprehension, such as "Wilson is receiving chemotherapy" or "Tom is discussing with Sarah", limiting their applicability in smart healthcare and smart home environments. To address this limitation, we propose a one-shot learning framework PVChat, the first personalized ViLLM that enables subject-aware question answering (QA) from a single video for each subject. Our approach optimizes a Mixture-of-Heads (MoH) enhanced ViLLM on a synthetically augmented video-QA dataset, leveraging a progressive image-to-video learning strategy. Specifically, we introduce an automated augmentation pipeline that synthesizes identity-preserving positive samples and retrieves hard negatives from existing video corpora, generating a diverse training dataset with four QA types: existence, appearance, action, and location inquiries. To enhance subject-specific learning, we propose a ReLU Routing MoH attention mechanism, alongside two novel objectives: (1) Smooth Proximity Regularization for progressive learning through exponential distance scaling and (2) Head Activation Enhancement for balanced attention routing. Finally, we adopt a two-stage training strategy, transitioning from image pre-training to video fine-tuning, enabling a gradual learning process from static attributes to dynamic representations. We evaluate PVChat on diverse datasets covering medical scenarios, TV series, anime, and real-world footage, demonstrating its superiority in personalized feature understanding after learning from a single video, compared to state-of-the-art ViLLMs.
title PVChat: Personalized Video Chat with One-Shot Learning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2503.17069