RAPID: Retrieval-Augmented Parallel Inference Drafting for Text-Based Video Event Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nguyen, Long, Nguyen, Huy, Khuu, Bao, Luu, Huy, Le, Huy, Nguyen, Tuan, Quan, Tho
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915125382873088
author Nguyen, Long
Nguyen, Huy
Khuu, Bao
Luu, Huy
Le, Huy
Nguyen, Tuan
Quan, Tho
author_facet Nguyen, Long
Nguyen, Huy
Khuu, Bao
Luu, Huy
Le, Huy
Nguyen, Tuan
Quan, Tho
contents Retrieving events from videos using text queries has become increasingly challenging due to the rapid growth of multimedia content. Existing methods for text-based video event retrieval often focus heavily on object-level descriptions, overlooking the crucial role of contextual information. This limitation is especially apparent when queries lack sufficient context, such as missing location details or ambiguous background elements. To address these challenges, we propose a novel system called RAPID (Retrieval-Augmented Parallel Inference Drafting), which leverages advancements in Large Language Models (LLMs) and prompt-based learning to semantically correct and enrich user queries with relevant contextual information. These enriched queries are then processed through parallel retrieval, followed by an evaluation step to select the most relevant results based on their alignment with the original query. Through extensive experiments on our custom-developed dataset, we demonstrate that RAPID significantly outperforms traditional retrieval methods, particularly for contextually incomplete queries. Our system was validated for both speed and accuracy through participation in the Ho Chi Minh City AI Challenge 2024, where it successfully retrieved events from over 300 hours of video. Further evaluation comparing RAPID with the baseline proposed by the competition organizers demonstrated its superior effectiveness, highlighting the strength and robustness of our approach.
format Preprint
id arxiv_https___arxiv_org_abs_2501_16303
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RAPID: Retrieval-Augmented Parallel Inference Drafting for Text-Based Video Event Retrieval
Nguyen, Long
Nguyen, Huy
Khuu, Bao
Luu, Huy
Le, Huy
Nguyen, Tuan
Quan, Tho
Computation and Language
Information Retrieval
Retrieving events from videos using text queries has become increasingly challenging due to the rapid growth of multimedia content. Existing methods for text-based video event retrieval often focus heavily on object-level descriptions, overlooking the crucial role of contextual information. This limitation is especially apparent when queries lack sufficient context, such as missing location details or ambiguous background elements. To address these challenges, we propose a novel system called RAPID (Retrieval-Augmented Parallel Inference Drafting), which leverages advancements in Large Language Models (LLMs) and prompt-based learning to semantically correct and enrich user queries with relevant contextual information. These enriched queries are then processed through parallel retrieval, followed by an evaluation step to select the most relevant results based on their alignment with the original query. Through extensive experiments on our custom-developed dataset, we demonstrate that RAPID significantly outperforms traditional retrieval methods, particularly for contextually incomplete queries. Our system was validated for both speed and accuracy through participation in the Ho Chi Minh City AI Challenge 2024, where it successfully retrieved events from over 300 hours of video. Further evaluation comparing RAPID with the baseline proposed by the competition organizers demonstrated its superior effectiveness, highlighting the strength and robustness of our approach.
title RAPID: Retrieval-Augmented Parallel Inference Drafting for Text-Based Video Event Retrieval
topic Computation and Language
Information Retrieval
url https://arxiv.org/abs/2501.16303