Multimodal Lengthy Videos Retrieval Framework and Evaluation Metric
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866913780328300544 |
|---|---|
| author | Eltahir, Mohamed Sarraj, Osamah Bremoo, Mohammed Khurd, Mohammed Alfrihidi, Abdulrahman Alshatiri, Taha Almatrafi, Mohammad Hussain, Tanveer |
| author_facet | Eltahir, Mohamed Sarraj, Osamah Bremoo, Mohammed Khurd, Mohammed Alfrihidi, Abdulrahman Alshatiri, Taha Almatrafi, Mohammad Hussain, Tanveer |
| contents | Precise video retrieval requires multi-modal correlations to handle unseen vocabulary and scenes, becoming more complex for lengthy videos where models must perform effectively without prior training on a specific dataset. We introduce a unified framework that combines a visual matching stream and an aural matching stream with a unique subtitles-based video segmentation approach. Additionally, the aural stream includes a complementary audio-based two-stage retrieval mechanism that enhances performance on long-duration videos. Considering the complex nature of retrieval from lengthy videos and its corresponding evaluation, we introduce a new retrieval evaluation method specifically designed for long-video retrieval to support further research. We conducted experiments on the YouCook2 benchmark, showing promising retrieval performance. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2504_04572 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Multimodal Lengthy Videos Retrieval Framework and Evaluation Metric Eltahir, Mohamed Sarraj, Osamah Bremoo, Mohammed Khurd, Mohammed Alfrihidi, Abdulrahman Alshatiri, Taha Almatrafi, Mohammad Hussain, Tanveer Computer Vision and Pattern Recognition Precise video retrieval requires multi-modal correlations to handle unseen vocabulary and scenes, becoming more complex for lengthy videos where models must perform effectively without prior training on a specific dataset. We introduce a unified framework that combines a visual matching stream and an aural matching stream with a unique subtitles-based video segmentation approach. Additionally, the aural stream includes a complementary audio-based two-stage retrieval mechanism that enhances performance on long-duration videos. Considering the complex nature of retrieval from lengthy videos and its corresponding evaluation, we introduce a new retrieval evaluation method specifically designed for long-video retrieval to support further research. We conducted experiments on the YouCook2 benchmark, showing promising retrieval performance. |
| title | Multimodal Lengthy Videos Retrieval Framework and Evaluation Metric |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2504.04572 |