Zoomer: Adaptive Image Focus Optimization for Black-box MLLM

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qian, Jiaxu, Wang, Chendong, Yang, Yifan, Zhang, Chaoyun, Jiang, Huiqiang, Luo, Xufang, Kang, Yu, Lin, Qingwei, Zhang, Anlan, Jiang, Shiqi, Cao, Ting, Mao, Tianjun, Banerjee, Suman, Liu, Guyue, Rajmohan, Saravan, Zhang, Dongmei, Yang, Yuqing, Zhang, Qi, Qiu, Lili
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915700386299904
author Qian, Jiaxu
Wang, Chendong
Yang, Yifan
Zhang, Chaoyun
Jiang, Huiqiang
Luo, Xufang
Kang, Yu
Lin, Qingwei
Zhang, Anlan
Jiang, Shiqi
Cao, Ting
Mao, Tianjun
Banerjee, Suman
Liu, Guyue
Rajmohan, Saravan
Zhang, Dongmei
Yang, Yuqing
Zhang, Qi
Qiu, Lili
author_facet Qian, Jiaxu
Wang, Chendong
Yang, Yifan
Zhang, Chaoyun
Jiang, Huiqiang
Luo, Xufang
Kang, Yu
Lin, Qingwei
Zhang, Anlan
Jiang, Shiqi
Cao, Ting
Mao, Tianjun
Banerjee, Suman
Liu, Guyue
Rajmohan, Saravan
Zhang, Dongmei
Yang, Yuqing
Zhang, Qi
Qiu, Lili
contents Multimodal large language models (MLLMs) such as GPT-4o, Gemini Pro, and Claude 3.5 have enabled unified reasoning over text and visual inputs, yet they often hallucinate in real world scenarios especially when small objects or fine spatial context are involved. We pinpoint two core causes of this failure: the absence of region-adaptive attention and inflexible token budgets that force uniform downsampling, leading to critical information loss. To overcome these limitations, we introduce Zoomer, a visual prompting framework that delivers token-efficient, detail-preserving image representations for black-box MLLMs. Zoomer integrates (1) a prompt-aware emphasis module to highlight semantically relevant regions, (2) a spatial-preserving orchestration schema to maintain object relationships, and (3) a budget-aware strategy to adaptively allocate tokens between global context and local details. Extensive experiments on nine benchmarks and three commercial MLLMs demonstrate that Zoomer boosts accuracy by up to 27% while cutting image token usage by up to 67%. Our approach establishes a principled methodology for robust, resource-aware multimodal understanding in settings where model internals are inaccessible.
format Preprint
id arxiv_https___arxiv_org_abs_2505_00742
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Zoomer: Adaptive Image Focus Optimization for Black-box MLLM
Qian, Jiaxu
Wang, Chendong
Yang, Yifan
Zhang, Chaoyun
Jiang, Huiqiang
Luo, Xufang
Kang, Yu
Lin, Qingwei
Zhang, Anlan
Jiang, Shiqi
Cao, Ting
Mao, Tianjun
Banerjee, Suman
Liu, Guyue
Rajmohan, Saravan
Zhang, Dongmei
Yang, Yuqing
Zhang, Qi
Qiu, Lili
Computer Vision and Pattern Recognition
Artificial Intelligence
Image and Video Processing
Multimodal large language models (MLLMs) such as GPT-4o, Gemini Pro, and Claude 3.5 have enabled unified reasoning over text and visual inputs, yet they often hallucinate in real world scenarios especially when small objects or fine spatial context are involved. We pinpoint two core causes of this failure: the absence of region-adaptive attention and inflexible token budgets that force uniform downsampling, leading to critical information loss. To overcome these limitations, we introduce Zoomer, a visual prompting framework that delivers token-efficient, detail-preserving image representations for black-box MLLMs. Zoomer integrates (1) a prompt-aware emphasis module to highlight semantically relevant regions, (2) a spatial-preserving orchestration schema to maintain object relationships, and (3) a budget-aware strategy to adaptively allocate tokens between global context and local details. Extensive experiments on nine benchmarks and three commercial MLLMs demonstrate that Zoomer boosts accuracy by up to 27% while cutting image token usage by up to 67%. Our approach establishes a principled methodology for robust, resource-aware multimodal understanding in settings where model internals are inaccessible.
title Zoomer: Adaptive Image Focus Optimization for Black-box MLLM
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Image and Video Processing
url https://arxiv.org/abs/2505.00742