Small Vision-Language Models: A Survey on Compact Architectures and Techniques

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Patnaik, Nitesh, Nayak, Navdeep, Agrawal, Himani Bansal, Khamaru, Moinak Chinmoy, Bal, Gourav, Panda, Saishree Smaranika, Raj, Rishi, Meena, Vishal, Vadlamani, Kartheek
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910875023048704
author Patnaik, Nitesh
Nayak, Navdeep
Agrawal, Himani Bansal
Khamaru, Moinak Chinmoy
Bal, Gourav
Panda, Saishree Smaranika
Raj, Rishi
Meena, Vishal
Vadlamani, Kartheek
author_facet Patnaik, Nitesh
Nayak, Navdeep
Agrawal, Himani Bansal
Khamaru, Moinak Chinmoy
Bal, Gourav
Panda, Saishree Smaranika
Raj, Rishi
Meena, Vishal
Vadlamani, Kartheek
contents The emergence of small vision-language models (sVLMs) marks a critical advancement in multimodal AI, enabling efficient processing of visual and textual data in resource-constrained environments. This survey offers a comprehensive exploration of sVLM development, presenting a taxonomy of architectures - transformer-based, mamba-based, and hybrid - that highlight innovations in compact design and computational efficiency. Techniques such as knowledge distillation, lightweight attention mechanisms, and modality pre-fusion are discussed as enablers of high performance with reduced resource requirements. Through an in-depth analysis of models like TinyGPT-V, MiniGPT-4, and VL-Mamba, we identify trade-offs between accuracy, efficiency, and scalability. Persistent challenges, including data biases and generalization to complex tasks, are critically examined, with proposed pathways for addressing them. By consolidating advancements in sVLMs, this work underscores their transformative potential for accessible AI, setting a foundation for future research into efficient multimodal systems.
format Preprint
id arxiv_https___arxiv_org_abs_2503_10665
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Small Vision-Language Models: A Survey on Compact Architectures and Techniques
Patnaik, Nitesh
Nayak, Navdeep
Agrawal, Himani Bansal
Khamaru, Moinak Chinmoy
Bal, Gourav
Panda, Saishree Smaranika
Raj, Rishi
Meena, Vishal
Vadlamani, Kartheek
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
The emergence of small vision-language models (sVLMs) marks a critical advancement in multimodal AI, enabling efficient processing of visual and textual data in resource-constrained environments. This survey offers a comprehensive exploration of sVLM development, presenting a taxonomy of architectures - transformer-based, mamba-based, and hybrid - that highlight innovations in compact design and computational efficiency. Techniques such as knowledge distillation, lightweight attention mechanisms, and modality pre-fusion are discussed as enablers of high performance with reduced resource requirements. Through an in-depth analysis of models like TinyGPT-V, MiniGPT-4, and VL-Mamba, we identify trade-offs between accuracy, efficiency, and scalability. Persistent challenges, including data biases and generalization to complex tasks, are critically examined, with proposed pathways for addressing them. By consolidating advancements in sVLMs, this work underscores their transformative potential for accessible AI, setting a foundation for future research into efficient multimodal systems.
title Small Vision-Language Models: A Survey on Compact Architectures and Techniques
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2503.10665