Empirical Recipes for Efficient and Compact Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Jiabo, Li, Zhizhong, Sajadmanesh, Sina, Zhuang, Weiming, Lyu, Lingjuan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917351004307456
author Huang, Jiabo
Li, Zhizhong
Sajadmanesh, Sina
Zhuang, Weiming
Lyu, Lingjuan
author_facet Huang, Jiabo
Li, Zhizhong
Sajadmanesh, Sina
Zhuang, Weiming
Lyu, Lingjuan
contents Deploying vision-language models (VLMs) in resource-constrained settings demands low latency and high throughput, yet existing compact VLMs often fall short of the inference speedups their smaller parameter counts suggest. To explain this discrepancy, we conduct an empirical end-to-end efficiency analysis and systematically profile inference to identify the dominant bottlenecks. Based on these findings, we develop optimization recipes tailored to compact VLMs that substantially reduce latency while preserving accuracy. These techniques cut time to first token (TTFT) by 53% on InternVL3-2B and by 93% on SmolVLM-256M. Our recipes are broadly applicable across both VLM architectures and common serving frameworks, providing practical guidance for building efficient VLM systems. Beyond efficiency, we study how to extend compact VLMs with structured perception outputs and introduce the resulting model family, ArgusVLM. Across diverse benchmarks, ArgusVLM achieves strong performance while maintaining a compact and efficient design.
format Preprint
id arxiv_https___arxiv_org_abs_2603_16987
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Empirical Recipes for Efficient and Compact Vision-Language Models
Huang, Jiabo
Li, Zhizhong
Sajadmanesh, Sina
Zhuang, Weiming
Lyu, Lingjuan
Computer Vision and Pattern Recognition
Artificial Intelligence
Deploying vision-language models (VLMs) in resource-constrained settings demands low latency and high throughput, yet existing compact VLMs often fall short of the inference speedups their smaller parameter counts suggest. To explain this discrepancy, we conduct an empirical end-to-end efficiency analysis and systematically profile inference to identify the dominant bottlenecks. Based on these findings, we develop optimization recipes tailored to compact VLMs that substantially reduce latency while preserving accuracy. These techniques cut time to first token (TTFT) by 53% on InternVL3-2B and by 93% on SmolVLM-256M. Our recipes are broadly applicable across both VLM architectures and common serving frameworks, providing practical guidance for building efficient VLM systems. Beyond efficiency, we study how to extend compact VLMs with structured perception outputs and introduce the resulting model family, ArgusVLM. Across diverse benchmarks, ArgusVLM achieves strong performance while maintaining a compact and efficient design.
title Empirical Recipes for Efficient and Compact Vision-Language Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2603.16987