CropVLM: Learning to Zoom for Fine-Grained Vision-Language Perception

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Carvalho, Miguel, Dias, Helder, Martins, Bruno
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913026585657344
author Carvalho, Miguel
Dias, Helder
Martins, Bruno
author_facet Carvalho, Miguel
Dias, Helder
Martins, Bruno
contents Vision-Language Models (VLMs) often struggle with tasks that require fine-grained image understanding, such as scene-text recognition or document analysis, due to perception limitations and visual fragmentation. To address these challenges, we introduce CropVLM as an external low-cost method for boosting performance, enabling VLMs to dynamically ''zoom in'' on relevant image regions, enhancing their ability to capture fine details. CropVLM is trained using reinforcement learning, without using human-labeled bounding boxes as a supervision signal, and without expensive synthetic evaluations. The model is trained once and can be paired with both open-source and proprietary VLMs to improve their performance. Our approach delivers significant improvements on tasks that require high-resolution image understanding, notably for benchmarks that are out-of-domain for the target VLM, without modifying or fine-tuning the VLM, thus avoiding catastrophic forgetting.
format Preprint
id arxiv_https___arxiv_org_abs_2511_19820
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CropVLM: Learning to Zoom for Fine-Grained Vision-Language Perception
Carvalho, Miguel
Dias, Helder
Martins, Bruno
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Vision-Language Models (VLMs) often struggle with tasks that require fine-grained image understanding, such as scene-text recognition or document analysis, due to perception limitations and visual fragmentation. To address these challenges, we introduce CropVLM as an external low-cost method for boosting performance, enabling VLMs to dynamically ''zoom in'' on relevant image regions, enhancing their ability to capture fine details. CropVLM is trained using reinforcement learning, without using human-labeled bounding boxes as a supervision signal, and without expensive synthetic evaluations. The model is trained once and can be paired with both open-source and proprietary VLMs to improve their performance. Our approach delivers significant improvements on tasks that require high-resolution image understanding, notably for benchmarks that are out-of-domain for the target VLM, without modifying or fine-tuning the VLM, thus avoiding catastrophic forgetting.
title CropVLM: Learning to Zoom for Fine-Grained Vision-Language Perception
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2511.19820