HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908415792513024 |
|---|---|
| author | SR, Nikitha Mathur, Aradhya Neeraj Menta, Tarun Ram Jain, Rishabh Sarkar, Mausoom |
| author_facet | SR, Nikitha Mathur, Aradhya Neeraj Menta, Tarun Ram Jain, Rishabh Sarkar, Mausoom |
| contents | The integration of high-resolution image features in modern multimodal large language models has demonstrated significant improvements in fine-grained visual understanding tasks, achieving high performance across multiple benchmarks. Since these features are obtained from large image encoders like ViT, they come with a significant increase in computational costs due to multiple calls to these encoders. In this work, we first develop an intuition for feature upsampling as a natural extension of high-resolution feature generation. Through extensive experiments and ablations, we demonstrate how a shallow feature enricher can achieve competitive results with tremendous reductions in training and inference time as well as computational cost, with upto 1.5x saving in FLOPs. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_17608 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs SR, Nikitha Mathur, Aradhya Neeraj Menta, Tarun Ram Jain, Rishabh Sarkar, Mausoom Computer Vision and Pattern Recognition The integration of high-resolution image features in modern multimodal large language models has demonstrated significant improvements in fine-grained visual understanding tasks, achieving high performance across multiple benchmarks. Since these features are obtained from large image encoders like ViT, they come with a significant increase in computational costs due to multiple calls to these encoders. In this work, we first develop an intuition for feature upsampling as a natural extension of high-resolution feature generation. Through extensive experiments and ablations, we demonstrate how a shallow feature enricher can achieve competitive results with tremendous reductions in training and inference time as well as computational cost, with upto 1.5x saving in FLOPs. |
| title | HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2506.17608 |