Multimodal Lengthy Videos Retrieval Framework and Evaluation Metric

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Eltahir, Mohamed, Sarraj, Osamah, Bremoo, Mohammed, Khurd, Mohammed, Alfrihidi, Abdulrahman, Alshatiri, Taha, Almatrafi, Mohammad, Hussain, Tanveer
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913780328300544
author Eltahir, Mohamed
Sarraj, Osamah
Bremoo, Mohammed
Khurd, Mohammed
Alfrihidi, Abdulrahman
Alshatiri, Taha
Almatrafi, Mohammad
Hussain, Tanveer
author_facet Eltahir, Mohamed
Sarraj, Osamah
Bremoo, Mohammed
Khurd, Mohammed
Alfrihidi, Abdulrahman
Alshatiri, Taha
Almatrafi, Mohammad
Hussain, Tanveer
contents Precise video retrieval requires multi-modal correlations to handle unseen vocabulary and scenes, becoming more complex for lengthy videos where models must perform effectively without prior training on a specific dataset. We introduce a unified framework that combines a visual matching stream and an aural matching stream with a unique subtitles-based video segmentation approach. Additionally, the aural stream includes a complementary audio-based two-stage retrieval mechanism that enhances performance on long-duration videos. Considering the complex nature of retrieval from lengthy videos and its corresponding evaluation, we introduce a new retrieval evaluation method specifically designed for long-video retrieval to support further research. We conducted experiments on the YouCook2 benchmark, showing promising retrieval performance.
format Preprint
id arxiv_https___arxiv_org_abs_2504_04572
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multimodal Lengthy Videos Retrieval Framework and Evaluation Metric
Eltahir, Mohamed
Sarraj, Osamah
Bremoo, Mohammed
Khurd, Mohammed
Alfrihidi, Abdulrahman
Alshatiri, Taha
Almatrafi, Mohammad
Hussain, Tanveer
Computer Vision and Pattern Recognition
Precise video retrieval requires multi-modal correlations to handle unseen vocabulary and scenes, becoming more complex for lengthy videos where models must perform effectively without prior training on a specific dataset. We introduce a unified framework that combines a visual matching stream and an aural matching stream with a unique subtitles-based video segmentation approach. Additionally, the aural stream includes a complementary audio-based two-stage retrieval mechanism that enhances performance on long-duration videos. Considering the complex nature of retrieval from lengthy videos and its corresponding evaluation, we introduce a new retrieval evaluation method specifically designed for long-video retrieval to support further research. We conducted experiments on the YouCook2 benchmark, showing promising retrieval performance.
title Multimodal Lengthy Videos Retrieval Framework and Evaluation Metric
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.04572