PREGEN: Uncovering Latent Thoughts in Composed Video Retrieval

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Serussi, Gabriele, Vainshtein, David, Kouchly, Jonathan, Di Castro, Dotan, Baskin, Chaim
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912834606071808
author Serussi, Gabriele
Vainshtein, David
Kouchly, Jonathan
Di Castro, Dotan
Baskin, Chaim
author_facet Serussi, Gabriele
Vainshtein, David
Kouchly, Jonathan
Di Castro, Dotan
Baskin, Chaim
contents Composed Video Retrieval (CoVR) aims to retrieve a video based on a query video and a modifying text. Current CoVR methods fail to fully exploit modern Vision-Language Models (VLMs), either using outdated architectures or requiring computationally expensive fine-tuning and slow caption generation. We introduce PREGEN (PRE GENeration extraction), an efficient and powerful CoVR framework that overcomes these limitations. Our approach uniquely pairs a frozen, pre-trained VLM with a lightweight encoding model, eliminating the need for any VLM fine-tuning. We feed the query video and modifying text into the VLM and extract the hidden state of the final token from each layer. A simple encoder is then trained on these pooled representations, creating a semantically rich and compact embedding for retrieval. PREGEN significantly advances the state of the art, surpassing all prior methods on standard CoVR benchmarks with substantial gains in Recall@1 of +27.23 and +69.59. Our method demonstrates robustness across different VLM backbones and exhibits strong zero-shot generalization to more complex textual modifications, highlighting its effectiveness and semantic capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2601_13797
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PREGEN: Uncovering Latent Thoughts in Composed Video Retrieval
Serussi, Gabriele
Vainshtein, David
Kouchly, Jonathan
Di Castro, Dotan
Baskin, Chaim
Computer Vision and Pattern Recognition
Composed Video Retrieval (CoVR) aims to retrieve a video based on a query video and a modifying text. Current CoVR methods fail to fully exploit modern Vision-Language Models (VLMs), either using outdated architectures or requiring computationally expensive fine-tuning and slow caption generation. We introduce PREGEN (PRE GENeration extraction), an efficient and powerful CoVR framework that overcomes these limitations. Our approach uniquely pairs a frozen, pre-trained VLM with a lightweight encoding model, eliminating the need for any VLM fine-tuning. We feed the query video and modifying text into the VLM and extract the hidden state of the final token from each layer. A simple encoder is then trained on these pooled representations, creating a semantically rich and compact embedding for retrieval. PREGEN significantly advances the state of the art, surpassing all prior methods on standard CoVR benchmarks with substantial gains in Recall@1 of +27.23 and +69.59. Our method demonstrates robustness across different VLM backbones and exhibits strong zero-shot generalization to more complex textual modifications, highlighting its effectiveness and semantic capabilities.
title PREGEN: Uncovering Latent Thoughts in Composed Video Retrieval
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.13797