Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Haochen, Wang, Yuhao, Zhang, Tao, Zhou, Yikang, Li, Yanwei, Wang, Jiacong, Zheng, Jiani, Tian, Ye, Meng, Jiahao, Huang, Zilong, Mai, Guangcan, Wang, Anran, Tong, Yunhai, Wang, Zhuochen, Li, Xiangtai, Zhang, Zhaoxiang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915834762362880
author Wang, Haochen
Wang, Yuhao
Zhang, Tao
Zhou, Yikang
Li, Yanwei
Wang, Jiacong
Zheng, Jiani
Tian, Ye
Meng, Jiahao
Huang, Zilong
Mai, Guangcan
Wang, Anran
Tong, Yunhai
Wang, Zhuochen
Li, Xiangtai
Zhang, Zhaoxiang
author_facet Wang, Haochen
Wang, Yuhao
Zhang, Tao
Zhou, Yikang
Li, Yanwei
Wang, Jiacong
Zheng, Jiani
Tian, Ye
Meng, Jiahao
Huang, Zilong
Mai, Guangcan
Wang, Anran
Tong, Yunhai
Wang, Zhuochen
Li, Xiangtai
Zhang, Zhaoxiang
contents While Multimodal Large Language Models (MLLMs) excel at holistic understanding, they struggle in capturing the dense world with complex scenes, requiring fine-grained analysis of intricate details and object inter-relationships. Region-level MLLMs have been a promising step. However, previous attempts are generally optimized to understand given regions in isolation, neglecting crucial global contexts. To address this, we introduce Grasp Any Region (GAR) for comprehen- sive region-level visual understanding. Empowered by an effective RoI-aligned feature replay technique, GAR supports (1) precise perception by leveraging necessary global contexts, and (2) modeling interactions between multiple prompts. Together, it then naturally achieves (3) advanced compositional reasoning to answer specific free-form questions about any region, shifting the paradigm from passive description to active dialogue. Moreover, we construct GAR-Bench, which not only provides a more accurate evaluation of single-region comprehension, but also, more importantly, measures interactions and complex reasoning across multiple regions. Extensive experiments have demonstrated that GAR-1B not only maintains the state-of-the-art captioning capabilities, e.g., outperforming DAM-3B +4.5 on DLC-Bench, but also excels at modeling relationships between multiple prompts with advanced comprehension capabilities, even surpassing InternVL3-78B on GAR-Bench-VQA. More importantly, our zero-shot GAR-8B even outperforms in-domain VideoRefer-7B on VideoRefer-BenchQ, indicating its strong capabilities can be easily transferred to videos.
format Preprint
id arxiv_https___arxiv_org_abs_2510_18876
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs
Wang, Haochen
Wang, Yuhao
Zhang, Tao
Zhou, Yikang
Li, Yanwei
Wang, Jiacong
Zheng, Jiani
Tian, Ye
Meng, Jiahao
Huang, Zilong
Mai, Guangcan
Wang, Anran
Tong, Yunhai
Wang, Zhuochen
Li, Xiangtai
Zhang, Zhaoxiang
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
While Multimodal Large Language Models (MLLMs) excel at holistic understanding, they struggle in capturing the dense world with complex scenes, requiring fine-grained analysis of intricate details and object inter-relationships. Region-level MLLMs have been a promising step. However, previous attempts are generally optimized to understand given regions in isolation, neglecting crucial global contexts. To address this, we introduce Grasp Any Region (GAR) for comprehen- sive region-level visual understanding. Empowered by an effective RoI-aligned feature replay technique, GAR supports (1) precise perception by leveraging necessary global contexts, and (2) modeling interactions between multiple prompts. Together, it then naturally achieves (3) advanced compositional reasoning to answer specific free-form questions about any region, shifting the paradigm from passive description to active dialogue. Moreover, we construct GAR-Bench, which not only provides a more accurate evaluation of single-region comprehension, but also, more importantly, measures interactions and complex reasoning across multiple regions. Extensive experiments have demonstrated that GAR-1B not only maintains the state-of-the-art captioning capabilities, e.g., outperforming DAM-3B +4.5 on DLC-Bench, but also excels at modeling relationships between multiple prompts with advanced comprehension capabilities, even surpassing InternVL3-78B on GAR-Bench-VQA. More importantly, our zero-shot GAR-8B even outperforms in-domain VideoRefer-7B on VideoRefer-BenchQ, indicating its strong capabilities can be easily transferred to videos.
title Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2510.18876