PEARL: Personalized Streaming Video Understanding Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zheng, Yuanhong, An, Ruichuan, Lin, Xiaopeng, Liu, Yuxing, Yang, Sihan, Zhang, Huanyu, Li, Haodong, Zhang, Qintong, Zhang, Renrui, Li, Guopeng, Zhang, Yifan, Li, Yuheng, Zhang, Wentao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918400410779648
author Zheng, Yuanhong
An, Ruichuan
Lin, Xiaopeng
Liu, Yuxing
Yang, Sihan
Zhang, Huanyu
Li, Haodong
Zhang, Qintong
Zhang, Renrui
Li, Guopeng
Zhang, Yifan
Li, Yuheng
Zhang, Wentao
author_facet Zheng, Yuanhong
An, Ruichuan
Lin, Xiaopeng
Liu, Yuxing
Yang, Sihan
Zhang, Huanyu
Li, Haodong
Zhang, Qintong
Zhang, Renrui
Li, Guopeng
Zhang, Yifan
Li, Yuheng
Zhang, Wentao
contents Human cognition of new concepts is inherently a streaming process: we continuously recognize new objects or identities and update our memories over time. However, current multimodal personalization methods are largely limited to static images or offline videos. This disconnects continuous visual input from instant real-world feedback, limiting their ability to provide the real-time, interactive personalized responses essential for future AI assistants. To bridge this gap, we first propose and formally define the novel task of Personalized Streaming Video Understanding (PSVU). To facilitate research in this new direction, we introduce PEARL-Bench, the first comprehensive benchmark designed specifically to evaluate this challenging setting. It evaluates a model's ability to respond to personalized concepts at exact timestamps under two modes: (1) Frame-level, focusing on a specific person or object in discrete frames, and (2) a novel Video-level, focusing on personalized actions unfolding across continuous frames. PEARL-Bench comprises 132 unique videos and 2,173 fine-grained annotations with precise timestamps. Concept diversity and annotation quality are strictly ensured through a combined pipeline of automated generation and human verification. To tackle this challenging new setting, we further propose PEARL, a plug-and-play, training-free strategy that serves as a strong baseline. Extensive evaluations across 8 offline and online models demonstrate that PEARL achieves state-of-the-art performance. Notably, it brings consistent PSVU improvements when applied to 3 distinct architectures, proving to be a highly effective and robust strategy. We hope this work advances vision-language model (VLM) personalization and inspires further research into streaming personalized AI assistants. Code is available at https://github.com/Yuanhong-Zheng/PEARL.
format Preprint
id arxiv_https___arxiv_org_abs_2603_20422
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PEARL: Personalized Streaming Video Understanding Model
Zheng, Yuanhong
An, Ruichuan
Lin, Xiaopeng
Liu, Yuxing
Yang, Sihan
Zhang, Huanyu
Li, Haodong
Zhang, Qintong
Zhang, Renrui
Li, Guopeng
Zhang, Yifan
Li, Yuheng
Zhang, Wentao
Computer Vision and Pattern Recognition
Artificial Intelligence
Information Retrieval
Human cognition of new concepts is inherently a streaming process: we continuously recognize new objects or identities and update our memories over time. However, current multimodal personalization methods are largely limited to static images or offline videos. This disconnects continuous visual input from instant real-world feedback, limiting their ability to provide the real-time, interactive personalized responses essential for future AI assistants. To bridge this gap, we first propose and formally define the novel task of Personalized Streaming Video Understanding (PSVU). To facilitate research in this new direction, we introduce PEARL-Bench, the first comprehensive benchmark designed specifically to evaluate this challenging setting. It evaluates a model's ability to respond to personalized concepts at exact timestamps under two modes: (1) Frame-level, focusing on a specific person or object in discrete frames, and (2) a novel Video-level, focusing on personalized actions unfolding across continuous frames. PEARL-Bench comprises 132 unique videos and 2,173 fine-grained annotations with precise timestamps. Concept diversity and annotation quality are strictly ensured through a combined pipeline of automated generation and human verification. To tackle this challenging new setting, we further propose PEARL, a plug-and-play, training-free strategy that serves as a strong baseline. Extensive evaluations across 8 offline and online models demonstrate that PEARL achieves state-of-the-art performance. Notably, it brings consistent PSVU improvements when applied to 3 distinct architectures, proving to be a highly effective and robust strategy. We hope this work advances vision-language model (VLM) personalization and inspires further research into streaming personalized AI assistants. Code is available at https://github.com/Yuanhong-Zheng/PEARL.
title PEARL: Personalized Streaming Video Understanding Model
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Information Retrieval
url https://arxiv.org/abs/2603.20422