Task-Aware Resolution Optimization for Visual Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Luo, Weiqing, Tan, Zhen, Li, Yifan, Zhao, Xinyu, Lee, Kwonjoon, Dariush, Behzad, Chen, Tianlong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915546049544192
author Luo, Weiqing
Tan, Zhen
Li, Yifan
Zhao, Xinyu
Lee, Kwonjoon
Dariush, Behzad
Chen, Tianlong
author_facet Luo, Weiqing
Tan, Zhen
Li, Yifan
Zhao, Xinyu
Lee, Kwonjoon
Dariush, Behzad
Chen, Tianlong
contents Real-world vision-language applications demand varying levels of perceptual granularity. However, most existing visual large language models (VLLMs), such as LLaVA, pre-assume a fixed resolution for downstream tasks, which leads to subpar performance. To address this problem, we first conduct a comprehensive and pioneering investigation into the resolution preferences of different vision-language tasks, revealing a correlation between resolution preferences with image complexity, and uncertainty variance of the VLLM at different image input resolutions. Building on this insight, we propose an empirical formula to determine the optimal resolution for a given vision-language task, combining these two factors. Second, based on rigorous experiments, we propose a novel parameter-efficient fine-tuning technique to extend the visual input resolution of pre-trained VLLMs to the identified optimal resolution. Extensive experiments on various vision-language tasks validate the effectiveness of our method.
format Preprint
id arxiv_https___arxiv_org_abs_2510_09822
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Task-Aware Resolution Optimization for Visual Large Language Models
Luo, Weiqing
Tan, Zhen
Li, Yifan
Zhao, Xinyu
Lee, Kwonjoon
Dariush, Behzad
Chen, Tianlong
Computer Vision and Pattern Recognition
Computation and Language
Real-world vision-language applications demand varying levels of perceptual granularity. However, most existing visual large language models (VLLMs), such as LLaVA, pre-assume a fixed resolution for downstream tasks, which leads to subpar performance. To address this problem, we first conduct a comprehensive and pioneering investigation into the resolution preferences of different vision-language tasks, revealing a correlation between resolution preferences with image complexity, and uncertainty variance of the VLLM at different image input resolutions. Building on this insight, we propose an empirical formula to determine the optimal resolution for a given vision-language task, combining these two factors. Second, based on rigorous experiments, we propose a novel parameter-efficient fine-tuning technique to extend the visual input resolution of pre-trained VLLMs to the identified optimal resolution. Extensive experiments on various vision-language tasks validate the effectiveness of our method.
title Task-Aware Resolution Optimization for Visual Large Language Models
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2510.09822