Mobile-VideoGPT: Fast and Accurate Model for Mobile Video Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shaker, Abdelrahman, Maaz, Muhammad, Gou, Chenhui, Rezatofighi, Hamid, Khan, Salman, Khan, Fahad Shahbaz
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914408991555584
author Shaker, Abdelrahman
Maaz, Muhammad
Gou, Chenhui
Rezatofighi, Hamid
Khan, Salman
Khan, Fahad Shahbaz
author_facet Shaker, Abdelrahman
Maaz, Muhammad
Gou, Chenhui
Rezatofighi, Hamid
Khan, Salman
Khan, Fahad Shahbaz
contents Video understanding models often struggle with high computational requirements, extensive parameter counts, and slow inference speed, making them inefficient for practical use. To tackle these challenges, we propose Mobile-VideoGPT, an efficient multimodal framework designed to operate with fewer than a billion parameters. Unlike traditional video large multimodal models (LMMs), Mobile-VideoGPT consists of lightweight dual visual encoders, efficient projectors, and a small language model (SLM), enabling real-time throughput. To further improve efficiency, we present an Attention-Based Frame Scoring mechanism to select the key-frames, along with an efficient token projector that prunes redundant visual tokens and preserves essential contextual cues. We evaluate our model across well-established six video understanding benchmarks (e.g., MVBench, EgoSchema, NextQA, and PercepTest). Our results show that Mobile-VideoGPT-0.5B can generate up to 46 tokens per second while outperforming existing state-of-the-art 0.5B-parameter models by 6 points on average with 40% fewer parameters and more than 2x higher throughput. Our code and models are publicly available at: https://github.com/Amshaker/Mobile-VideoGPT.
format Preprint
id arxiv_https___arxiv_org_abs_2503_21782
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mobile-VideoGPT: Fast and Accurate Model for Mobile Video Understanding
Shaker, Abdelrahman
Maaz, Muhammad
Gou, Chenhui
Rezatofighi, Hamid
Khan, Salman
Khan, Fahad Shahbaz
Computer Vision and Pattern Recognition
Video understanding models often struggle with high computational requirements, extensive parameter counts, and slow inference speed, making them inefficient for practical use. To tackle these challenges, we propose Mobile-VideoGPT, an efficient multimodal framework designed to operate with fewer than a billion parameters. Unlike traditional video large multimodal models (LMMs), Mobile-VideoGPT consists of lightweight dual visual encoders, efficient projectors, and a small language model (SLM), enabling real-time throughput. To further improve efficiency, we present an Attention-Based Frame Scoring mechanism to select the key-frames, along with an efficient token projector that prunes redundant visual tokens and preserves essential contextual cues. We evaluate our model across well-established six video understanding benchmarks (e.g., MVBench, EgoSchema, NextQA, and PercepTest). Our results show that Mobile-VideoGPT-0.5B can generate up to 46 tokens per second while outperforming existing state-of-the-art 0.5B-parameter models by 6 points on average with 40% fewer parameters and more than 2x higher throughput. Our code and models are publicly available at: https://github.com/Amshaker/Mobile-VideoGPT.
title Mobile-VideoGPT: Fast and Accurate Model for Mobile Video Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.21782