Multimodal Lengthy Videos Retrieval Framework and Evaluation Metric

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Eltahir, Mohamed, Sarraj, Osamah, Bremoo, Mohammed, Khurd, Mohammed, Alfrihidi, Abdulrahman, Alshatiri, Taha, Almatrafi, Mohammad, Hussain, Tanveer
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913780328300544
author Eltahir, Mohamed
Sarraj, Osamah
Bremoo, Mohammed
Khurd, Mohammed
Alfrihidi, Abdulrahman
Alshatiri, Taha
Almatrafi, Mohammad
Hussain, Tanveer
author_facet Eltahir, Mohamed
Sarraj, Osamah
Bremoo, Mohammed
Khurd, Mohammed
Alfrihidi, Abdulrahman
Alshatiri, Taha
Almatrafi, Mohammad
Hussain, Tanveer
contents Precise video retrieval requires multi-modal correlations to handle unseen vocabulary and scenes, becoming more complex for lengthy videos where models must perform effectively without prior training on a specific dataset. We introduce a unified framework that combines a visual matching stream and an aural matching stream with a unique subtitles-based video segmentation approach. Additionally, the aural stream includes a complementary audio-based two-stage retrieval mechanism that enhances performance on long-duration videos. Considering the complex nature of retrieval from lengthy videos and its corresponding evaluation, we introduce a new retrieval evaluation method specifically designed for long-video retrieval to support further research. We conducted experiments on the YouCook2 benchmark, showing promising retrieval performance.
format Preprint
id arxiv_https___arxiv_org_abs_2504_04572
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multimodal Lengthy Videos Retrieval Framework and Evaluation Metric
Eltahir, Mohamed
Sarraj, Osamah
Bremoo, Mohammed
Khurd, Mohammed
Alfrihidi, Abdulrahman
Alshatiri, Taha
Almatrafi, Mohammad
Hussain, Tanveer
Computer Vision and Pattern Recognition
Precise video retrieval requires multi-modal correlations to handle unseen vocabulary and scenes, becoming more complex for lengthy videos where models must perform effectively without prior training on a specific dataset. We introduce a unified framework that combines a visual matching stream and an aural matching stream with a unique subtitles-based video segmentation approach. Additionally, the aural stream includes a complementary audio-based two-stage retrieval mechanism that enhances performance on long-duration videos. Considering the complex nature of retrieval from lengthy videos and its corresponding evaluation, we introduce a new retrieval evaluation method specifically designed for long-video retrieval to support further research. We conducted experiments on the YouCook2 benchmark, showing promising retrieval performance.
title Multimodal Lengthy Videos Retrieval Framework and Evaluation Metric
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.04572