CARES: Context-Aware Resolution Selector for VLMs

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kimhi, Moshe, Shabtay, Nimrod, Giryes, Raja, Baskin, Chaim, Schwartz, Eli
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916069607735296
author Kimhi, Moshe
Shabtay, Nimrod
Giryes, Raja
Baskin, Chaim
Schwartz, Eli
author_facet Kimhi, Moshe
Shabtay, Nimrod
Giryes, Raja
Baskin, Chaim
Schwartz, Eli
contents Large vision-language models (VLMs) commonly process images at native or high resolution to remain effective across tasks. This inflates visual tokens ofter to 97-99% of total tokens, resulting in high compute and latency, even when low-resolution images would suffice. We introduce \emph{CARES}-a \textbf{C}ontext-\textbf{A}ware \textbf{R}esolution \textbf{S}elector, a lightweight preprocessing module that, given an image-query pair, predicts the \emph{minimal} sufficient input resolution. CARES uses a compact VLM (350M) to extract features and predict when a target pretrained VLM's response converges to its peak ability to answer correctly. Though trained as a discrete classifier over a set of optional resolutions, CARES interpolates continuous resolutions at inference for fine-grained control. Across five multimodal benchmarks spanning documents and natural images, as well as diverse target VLMs, CARES preserves task performance while reducing compute by up to 80%.
format Preprint
id arxiv_https___arxiv_org_abs_2510_19496
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CARES: Context-Aware Resolution Selector for VLMs
Kimhi, Moshe
Shabtay, Nimrod
Giryes, Raja
Baskin, Chaim
Schwartz, Eli
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Large vision-language models (VLMs) commonly process images at native or high resolution to remain effective across tasks. This inflates visual tokens ofter to 97-99% of total tokens, resulting in high compute and latency, even when low-resolution images would suffice. We introduce \emph{CARES}-a \textbf{C}ontext-\textbf{A}ware \textbf{R}esolution \textbf{S}elector, a lightweight preprocessing module that, given an image-query pair, predicts the \emph{minimal} sufficient input resolution. CARES uses a compact VLM (350M) to extract features and predict when a target pretrained VLM's response converges to its peak ability to answer correctly. Though trained as a discrete classifier over a set of optional resolutions, CARES interpolates continuous resolutions at inference for fine-grained control. Across five multimodal benchmarks spanning documents and natural images, as well as diverse target VLMs, CARES preserves task performance while reducing compute by up to 80%.
title CARES: Context-Aware Resolution Selector for VLMs
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2510.19496