Towards Efficient Large Vision-Language Models: A Comprehensive Survey on Inference Strategies

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Pathak, Surendra, Han, Bo
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911583689506816
author Pathak, Surendra
Han, Bo
author_facet Pathak, Surendra
Han, Bo
contents Although Large Vision Language Models (LVLMs) have demonstrated impressive multimodal reasoning capabilities, their scalability and deployment are constrained by massive computational requirements. In particular, the massive amount of visual tokens from high-resolution input data aggravates the situation due to the quadratic complexity of attention mechanisms. To address these issues, the research community has developed several optimization frameworks. This paper presents a comprehensive survey of the current state-of-the-art techniques for accelerating LVLM inference. We introduce a systematic taxonomy that categorizes existing optimization frameworks into four primary dimensions: visual token compression, memory management and serving, efficient architectural design, and advanced decoding strategies. Furthermore, we critically examine the limitations of these current methodologies and identify critical open problems to inspire future research directions in efficient multimodal systems.
format Preprint
id arxiv_https___arxiv_org_abs_2603_27960
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Towards Efficient Large Vision-Language Models: A Comprehensive Survey on Inference Strategies
Pathak, Surendra
Han, Bo
Machine Learning
Computation and Language
Computer Vision and Pattern Recognition
Although Large Vision Language Models (LVLMs) have demonstrated impressive multimodal reasoning capabilities, their scalability and deployment are constrained by massive computational requirements. In particular, the massive amount of visual tokens from high-resolution input data aggravates the situation due to the quadratic complexity of attention mechanisms. To address these issues, the research community has developed several optimization frameworks. This paper presents a comprehensive survey of the current state-of-the-art techniques for accelerating LVLM inference. We introduce a systematic taxonomy that categorizes existing optimization frameworks into four primary dimensions: visual token compression, memory management and serving, efficient architectural design, and advanced decoding strategies. Furthermore, we critically examine the limitations of these current methodologies and identify critical open problems to inspire future research directions in efficient multimodal systems.
title Towards Efficient Large Vision-Language Models: A Comprehensive Survey on Inference Strategies
topic Machine Learning
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.27960