ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Khalil, Ahmad, Khalil, Mahmoud, Ngom, Alioune
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909585986551808
author Khalil, Ahmad
Khalil, Mahmoud
Ngom, Alioune
author_facet Khalil, Ahmad
Khalil, Mahmoud
Ngom, Alioune
contents In this paper, we introduce ResNetVLLM (ResNet Vision LLM), a novel cross-modal framework for zero-shot video understanding that integrates a ResNet-based visual encoder with a Large Language Model (LLM. ResNetVLLM addresses the challenges associated with zero-shot video models by avoiding reliance on pre-trained video understanding models and instead employing a non-pretrained ResNet to extract visual features. This design ensures the model learns visual and semantic representations within a unified architecture, enhancing its ability to generate accurate and contextually relevant textual descriptions from video inputs. Our experimental results demonstrate that ResNetVLLM achieves state-of-the-art performance in zero-shot video understanding (ZSVU) on several benchmarks, including MSRVTT-QA, MSVD-QA, TGIF-QA FrameQA, and ActivityNet-QA.
format Preprint
id arxiv_https___arxiv_org_abs_2504_14432
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task
Khalil, Ahmad
Khalil, Mahmoud
Ngom, Alioune
Computer Vision and Pattern Recognition
Artificial Intelligence
In this paper, we introduce ResNetVLLM (ResNet Vision LLM), a novel cross-modal framework for zero-shot video understanding that integrates a ResNet-based visual encoder with a Large Language Model (LLM. ResNetVLLM addresses the challenges associated with zero-shot video models by avoiding reliance on pre-trained video understanding models and instead employing a non-pretrained ResNet to extract visual features. This design ensures the model learns visual and semantic representations within a unified architecture, enhancing its ability to generate accurate and contextually relevant textual descriptions from video inputs. Our experimental results demonstrate that ResNetVLLM achieves state-of-the-art performance in zero-shot video understanding (ZSVU) on several benchmarks, including MSRVTT-QA, MSVD-QA, TGIF-QA FrameQA, and ActivityNet-QA.
title ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2504.14432