Saved in:
Bibliographic Details
Main Authors: Sharshar, Ahmed, Khan, Latif U., Ullah, Waseem, Guizani, Mohsen
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2502.07855
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912434878414848
author Sharshar, Ahmed
Khan, Latif U.
Ullah, Waseem
Guizani, Mohsen
author_facet Sharshar, Ahmed
Khan, Latif U.
Ullah, Waseem
Guizani, Mohsen
contents Vision Large Language Models (VLMs) combine visual understanding with natural language processing, enabling tasks like image captioning, visual question answering, and video analysis. While VLMs show impressive capabilities across domains such as autonomous vehicles, smart surveillance, and healthcare, their deployment on resource-constrained edge devices remains challenging due to processing power, memory, and energy limitations. This survey explores recent advancements in optimizing VLMs for edge environments, focusing on model compression techniques, including pruning, quantization, knowledge distillation, and specialized hardware solutions that enhance efficiency. We provide a detailed discussion of efficient training and fine-tuning methods, edge deployment challenges, and privacy considerations. Additionally, we discuss the diverse applications of lightweight VLMs across healthcare, environmental monitoring, and autonomous systems, illustrating their growing impact. By highlighting key design strategies, current challenges, and offering recommendations for future directions, this survey aims to inspire further research into the practical deployment of VLMs, ultimately making advanced AI accessible in resource-limited settings.
format Preprint
id arxiv_https___arxiv_org_abs_2502_07855
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Vision-Language Models for Edge Networks: A Comprehensive Survey
Sharshar, Ahmed
Khan, Latif U.
Ullah, Waseem
Guizani, Mohsen
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Vision Large Language Models (VLMs) combine visual understanding with natural language processing, enabling tasks like image captioning, visual question answering, and video analysis. While VLMs show impressive capabilities across domains such as autonomous vehicles, smart surveillance, and healthcare, their deployment on resource-constrained edge devices remains challenging due to processing power, memory, and energy limitations. This survey explores recent advancements in optimizing VLMs for edge environments, focusing on model compression techniques, including pruning, quantization, knowledge distillation, and specialized hardware solutions that enhance efficiency. We provide a detailed discussion of efficient training and fine-tuning methods, edge deployment challenges, and privacy considerations. Additionally, we discuss the diverse applications of lightweight VLMs across healthcare, environmental monitoring, and autonomous systems, illustrating their growing impact. By highlighting key design strategies, current challenges, and offering recommendations for future directions, this survey aims to inspire further research into the practical deployment of VLMs, ultimately making advanced AI accessible in resource-limited settings.
title Vision-Language Models for Edge Networks: A Comprehensive Survey
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2502.07855