Enhancing Subsequent Video Retrieval via Vision-Language Models (VLMs)

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Duan, Yicheng, Huang, Xi, Chen, Duo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912286820532224
author Duan, Yicheng
Huang, Xi
Chen, Duo
author_facet Duan, Yicheng
Huang, Xi
Chen, Duo
contents The rapid growth of video content demands efficient and precise retrieval systems. While vision-language models (VLMs) excel in representation learning, they often struggle with adaptive, time-sensitive video retrieval. This paper introduces a novel framework that combines vector similarity search with graph-based data structures. By leveraging VLM embeddings for initial retrieval and modeling contextual relationships among video segments, our approach enables adaptive query refinement and improves retrieval accuracy. Experiments demonstrate its precision, scalability, and robustness, offering an effective solution for interactive video retrieval in dynamic environments.
format Preprint
id arxiv_https___arxiv_org_abs_2503_17415
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Enhancing Subsequent Video Retrieval via Vision-Language Models (VLMs)
Duan, Yicheng
Huang, Xi
Chen, Duo
Computer Vision and Pattern Recognition
Artificial Intelligence
Information Retrieval
The rapid growth of video content demands efficient and precise retrieval systems. While vision-language models (VLMs) excel in representation learning, they often struggle with adaptive, time-sensitive video retrieval. This paper introduces a novel framework that combines vector similarity search with graph-based data structures. By leveraging VLM embeddings for initial retrieval and modeling contextual relationships among video segments, our approach enables adaptive query refinement and improves retrieval accuracy. Experiments demonstrate its precision, scalability, and robustness, offering an effective solution for interactive video retrieval in dynamic environments.
title Enhancing Subsequent Video Retrieval via Vision-Language Models (VLMs)
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Information Retrieval
url https://arxiv.org/abs/2503.17415