Zoomer: Adaptive Image Focus Optimization for Black-box MLLM
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915700386299904 |
|---|---|
| author | Qian, Jiaxu Wang, Chendong Yang, Yifan Zhang, Chaoyun Jiang, Huiqiang Luo, Xufang Kang, Yu Lin, Qingwei Zhang, Anlan Jiang, Shiqi Cao, Ting Mao, Tianjun Banerjee, Suman Liu, Guyue Rajmohan, Saravan Zhang, Dongmei Yang, Yuqing Zhang, Qi Qiu, Lili |
| author_facet | Qian, Jiaxu Wang, Chendong Yang, Yifan Zhang, Chaoyun Jiang, Huiqiang Luo, Xufang Kang, Yu Lin, Qingwei Zhang, Anlan Jiang, Shiqi Cao, Ting Mao, Tianjun Banerjee, Suman Liu, Guyue Rajmohan, Saravan Zhang, Dongmei Yang, Yuqing Zhang, Qi Qiu, Lili |
| contents | Multimodal large language models (MLLMs) such as GPT-4o, Gemini Pro, and Claude 3.5 have enabled unified reasoning over text and visual inputs, yet they often hallucinate in real world scenarios especially when small objects or fine spatial context are involved. We pinpoint two core causes of this failure: the absence of region-adaptive attention and inflexible token budgets that force uniform downsampling, leading to critical information loss. To overcome these limitations, we introduce Zoomer, a visual prompting framework that delivers token-efficient, detail-preserving image representations for black-box MLLMs. Zoomer integrates (1) a prompt-aware emphasis module to highlight semantically relevant regions, (2) a spatial-preserving orchestration schema to maintain object relationships, and (3) a budget-aware strategy to adaptively allocate tokens between global context and local details. Extensive experiments on nine benchmarks and three commercial MLLMs demonstrate that Zoomer boosts accuracy by up to 27% while cutting image token usage by up to 67%. Our approach establishes a principled methodology for robust, resource-aware multimodal understanding in settings where model internals are inaccessible. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_00742 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Zoomer: Adaptive Image Focus Optimization for Black-box MLLM Qian, Jiaxu Wang, Chendong Yang, Yifan Zhang, Chaoyun Jiang, Huiqiang Luo, Xufang Kang, Yu Lin, Qingwei Zhang, Anlan Jiang, Shiqi Cao, Ting Mao, Tianjun Banerjee, Suman Liu, Guyue Rajmohan, Saravan Zhang, Dongmei Yang, Yuqing Zhang, Qi Qiu, Lili Computer Vision and Pattern Recognition Artificial Intelligence Image and Video Processing Multimodal large language models (MLLMs) such as GPT-4o, Gemini Pro, and Claude 3.5 have enabled unified reasoning over text and visual inputs, yet they often hallucinate in real world scenarios especially when small objects or fine spatial context are involved. We pinpoint two core causes of this failure: the absence of region-adaptive attention and inflexible token budgets that force uniform downsampling, leading to critical information loss. To overcome these limitations, we introduce Zoomer, a visual prompting framework that delivers token-efficient, detail-preserving image representations for black-box MLLMs. Zoomer integrates (1) a prompt-aware emphasis module to highlight semantically relevant regions, (2) a spatial-preserving orchestration schema to maintain object relationships, and (3) a budget-aware strategy to adaptively allocate tokens between global context and local details. Extensive experiments on nine benchmarks and three commercial MLLMs demonstrate that Zoomer boosts accuracy by up to 27% while cutting image token usage by up to 67%. Our approach establishes a principled methodology for robust, resource-aware multimodal understanding in settings where model internals are inaccessible. |
| title | Zoomer: Adaptive Image Focus Optimization for Black-box MLLM |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Image and Video Processing |
| url | https://arxiv.org/abs/2505.00742 |