Efficient Architectures for High Resolution Vision-Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Carvalho, Miguel, Martins, Bruno
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915627009048576
author Carvalho, Miguel
Martins, Bruno
author_facet Carvalho, Miguel
Martins, Bruno
contents Vision-Language Models (VLMs) have recently experienced significant advancements. However, challenges persist in the accurate recognition of fine details within high resolution images, which limits performance in multiple tasks. This work introduces Pheye, a novel architecture that efficiently processes high-resolution images while training fewer parameters than similarly sized VLMs. Notably, Pheye achieves a high efficiency while maintaining strong performance, particularly in tasks that demand fine-grained image understanding and/or the handling of scene-text.
format Preprint
id arxiv_https___arxiv_org_abs_2501_02584
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Efficient Architectures for High Resolution Vision-Language Models
Carvalho, Miguel
Martins, Bruno
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Vision-Language Models (VLMs) have recently experienced significant advancements. However, challenges persist in the accurate recognition of fine details within high resolution images, which limits performance in multiple tasks. This work introduces Pheye, a novel architecture that efficiently processes high-resolution images while training fewer parameters than similarly sized VLMs. Notably, Pheye achieves a high efficiency while maintaining strong performance, particularly in tasks that demand fine-grained image understanding and/or the handling of scene-text.
title Efficient Architectures for High Resolution Vision-Language Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2501.02584