HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: SR, Nikitha, Mathur, Aradhya Neeraj, Menta, Tarun Ram, Jain, Rishabh, Sarkar, Mausoom
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908415792513024
author SR, Nikitha
Mathur, Aradhya Neeraj
Menta, Tarun Ram
Jain, Rishabh
Sarkar, Mausoom
author_facet SR, Nikitha
Mathur, Aradhya Neeraj
Menta, Tarun Ram
Jain, Rishabh
Sarkar, Mausoom
contents The integration of high-resolution image features in modern multimodal large language models has demonstrated significant improvements in fine-grained visual understanding tasks, achieving high performance across multiple benchmarks. Since these features are obtained from large image encoders like ViT, they come with a significant increase in computational costs due to multiple calls to these encoders. In this work, we first develop an intuition for feature upsampling as a natural extension of high-resolution feature generation. Through extensive experiments and ablations, we demonstrate how a shallow feature enricher can achieve competitive results with tremendous reductions in training and inference time as well as computational cost, with upto 1.5x saving in FLOPs.
format Preprint
id arxiv_https___arxiv_org_abs_2506_17608
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs
SR, Nikitha
Mathur, Aradhya Neeraj
Menta, Tarun Ram
Jain, Rishabh
Sarkar, Mausoom
Computer Vision and Pattern Recognition
The integration of high-resolution image features in modern multimodal large language models has demonstrated significant improvements in fine-grained visual understanding tasks, achieving high performance across multiple benchmarks. Since these features are obtained from large image encoders like ViT, they come with a significant increase in computational costs due to multiple calls to these encoders. In this work, we first develop an intuition for feature upsampling as a natural extension of high-resolution feature generation. Through extensive experiments and ablations, we demonstrate how a shallow feature enricher can achieve competitive results with tremendous reductions in training and inference time as well as computational cost, with upto 1.5x saving in FLOPs.
title HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.17608