Small Vision-Language Models: A Survey on Compact Architectures and Techniques
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910875023048704 |
|---|---|
| author | Patnaik, Nitesh Nayak, Navdeep Agrawal, Himani Bansal Khamaru, Moinak Chinmoy Bal, Gourav Panda, Saishree Smaranika Raj, Rishi Meena, Vishal Vadlamani, Kartheek |
| author_facet | Patnaik, Nitesh Nayak, Navdeep Agrawal, Himani Bansal Khamaru, Moinak Chinmoy Bal, Gourav Panda, Saishree Smaranika Raj, Rishi Meena, Vishal Vadlamani, Kartheek |
| contents | The emergence of small vision-language models (sVLMs) marks a critical advancement in multimodal AI, enabling efficient processing of visual and textual data in resource-constrained environments. This survey offers a comprehensive exploration of sVLM development, presenting a taxonomy of architectures - transformer-based, mamba-based, and hybrid - that highlight innovations in compact design and computational efficiency. Techniques such as knowledge distillation, lightweight attention mechanisms, and modality pre-fusion are discussed as enablers of high performance with reduced resource requirements. Through an in-depth analysis of models like TinyGPT-V, MiniGPT-4, and VL-Mamba, we identify trade-offs between accuracy, efficiency, and scalability. Persistent challenges, including data biases and generalization to complex tasks, are critically examined, with proposed pathways for addressing them. By consolidating advancements in sVLMs, this work underscores their transformative potential for accessible AI, setting a foundation for future research into efficient multimodal systems. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_10665 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Small Vision-Language Models: A Survey on Compact Architectures and Techniques Patnaik, Nitesh Nayak, Navdeep Agrawal, Himani Bansal Khamaru, Moinak Chinmoy Bal, Gourav Panda, Saishree Smaranika Raj, Rishi Meena, Vishal Vadlamani, Kartheek Computer Vision and Pattern Recognition Artificial Intelligence Computation and Language Machine Learning The emergence of small vision-language models (sVLMs) marks a critical advancement in multimodal AI, enabling efficient processing of visual and textual data in resource-constrained environments. This survey offers a comprehensive exploration of sVLM development, presenting a taxonomy of architectures - transformer-based, mamba-based, and hybrid - that highlight innovations in compact design and computational efficiency. Techniques such as knowledge distillation, lightweight attention mechanisms, and modality pre-fusion are discussed as enablers of high performance with reduced resource requirements. Through an in-depth analysis of models like TinyGPT-V, MiniGPT-4, and VL-Mamba, we identify trade-offs between accuracy, efficiency, and scalability. Persistent challenges, including data biases and generalization to complex tasks, are critically examined, with proposed pathways for addressing them. By consolidating advancements in sVLMs, this work underscores their transformative potential for accessible AI, setting a foundation for future research into efficient multimodal systems. |
| title | Small Vision-Language Models: A Survey on Compact Architectures and Techniques |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2503.10665 |