RegionGPT: Towards Region Understanding Vision Language Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guo, Qiushan, De Mello, Shalini, Yin, Hongxu, Byeon, Wonmin, Cheung, Ka Chun, Yu, Yizhou, Luo, Ping, Liu, Sifei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909127408615424
author Guo, Qiushan
De Mello, Shalini
Yin, Hongxu
Byeon, Wonmin
Cheung, Ka Chun
Yu, Yizhou
Luo, Ping
Liu, Sifei
author_facet Guo, Qiushan
De Mello, Shalini
Yin, Hongxu
Byeon, Wonmin
Cheung, Ka Chun
Yu, Yizhou
Luo, Ping
Liu, Sifei
contents Vision language models (VLMs) have experienced rapid advancements through the integration of large language models (LLMs) with image-text pairs, yet they struggle with detailed regional visual understanding due to limited spatial awareness of the vision encoder, and the use of coarse-grained training data that lacks detailed, region-specific captions. To address this, we introduce RegionGPT (short as RGPT), a novel framework designed for complex region-level captioning and understanding. RGPT enhances the spatial awareness of regional representation with simple yet effective modifications to existing visual encoders in VLMs. We further improve performance on tasks requiring a specific output scope by integrating task-guided instruction prompts during both training and inference phases, while maintaining the model's versatility for general-purpose tasks. Additionally, we develop an automated region caption data generation pipeline, enriching the training set with detailed region-level captions. We demonstrate that a universal RGPT model can be effectively applied and significantly enhancing performance across a range of region-level tasks, including but not limited to complex region descriptions, reasoning, object classification, and referring expressions comprehension.
format Preprint
id arxiv_https___arxiv_org_abs_2403_02330
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle RegionGPT: Towards Region Understanding Vision Language Model
Guo, Qiushan
De Mello, Shalini
Yin, Hongxu
Byeon, Wonmin
Cheung, Ka Chun
Yu, Yizhou
Luo, Ping
Liu, Sifei
Computer Vision and Pattern Recognition
Vision language models (VLMs) have experienced rapid advancements through the integration of large language models (LLMs) with image-text pairs, yet they struggle with detailed regional visual understanding due to limited spatial awareness of the vision encoder, and the use of coarse-grained training data that lacks detailed, region-specific captions. To address this, we introduce RegionGPT (short as RGPT), a novel framework designed for complex region-level captioning and understanding. RGPT enhances the spatial awareness of regional representation with simple yet effective modifications to existing visual encoders in VLMs. We further improve performance on tasks requiring a specific output scope by integrating task-guided instruction prompts during both training and inference phases, while maintaining the model's versatility for general-purpose tasks. Additionally, we develop an automated region caption data generation pipeline, enriching the training set with detailed region-level captions. We demonstrate that a universal RGPT model can be effectively applied and significantly enhancing performance across a range of region-level tasks, including but not limited to complex region descriptions, reasoning, object classification, and referring expressions comprehension.
title RegionGPT: Towards Region Understanding Vision Language Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.02330