VHM: Versatile and Honest Vision Language Model for Remote Sensing Image Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pang, Chao, Weng, Xingxing, Wu, Jiang, Li, Jiayu, Liu, Yi, Sun, Jiaxing, Li, Weijia, Wang, Shuai, Feng, Litong, Xia, Gui-Song, He, Conghui
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913617436213248
author Pang, Chao
Weng, Xingxing
Wu, Jiang
Li, Jiayu
Liu, Yi
Sun, Jiaxing
Li, Weijia
Wang, Shuai
Feng, Litong
Xia, Gui-Song
He, Conghui
author_facet Pang, Chao
Weng, Xingxing
Wu, Jiang
Li, Jiayu
Liu, Yi
Sun, Jiaxing
Li, Weijia
Wang, Shuai
Feng, Litong
Xia, Gui-Song
He, Conghui
contents This paper develops a Versatile and Honest vision language Model (VHM) for remote sensing image analysis. VHM is built on a large-scale remote sensing image-text dataset with rich-content captions (VersaD), and an honest instruction dataset comprising both factual and deceptive questions (HnstD). Unlike prevailing remote sensing image-text datasets, in which image captions focus on a few prominent objects and their relationships, VersaD captions provide detailed information about image properties, object attributes, and the overall scene. This comprehensive captioning enables VHM to thoroughly understand remote sensing images and perform diverse remote sensing tasks. Moreover, different from existing remote sensing instruction datasets that only include factual questions, HnstD contains additional deceptive questions stemming from the non-existence of objects. This feature prevents VHM from producing affirmative answers to nonsense queries, thereby ensuring its honesty. In our experiments, VHM significantly outperforms various vision language models on common tasks of scene classification, visual question answering, and visual grounding. Additionally, VHM achieves competent performance on several unexplored tasks, such as building vectorizing, multi-label classification and honest question answering. We will release the code, data and model weights at https://github.com/opendatalab/VHM .
format Preprint
id arxiv_https___arxiv_org_abs_2403_20213
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle VHM: Versatile and Honest Vision Language Model for Remote Sensing Image Analysis
Pang, Chao
Weng, Xingxing
Wu, Jiang
Li, Jiayu
Liu, Yi
Sun, Jiaxing
Li, Weijia
Wang, Shuai
Feng, Litong
Xia, Gui-Song
He, Conghui
Computer Vision and Pattern Recognition
This paper develops a Versatile and Honest vision language Model (VHM) for remote sensing image analysis. VHM is built on a large-scale remote sensing image-text dataset with rich-content captions (VersaD), and an honest instruction dataset comprising both factual and deceptive questions (HnstD). Unlike prevailing remote sensing image-text datasets, in which image captions focus on a few prominent objects and their relationships, VersaD captions provide detailed information about image properties, object attributes, and the overall scene. This comprehensive captioning enables VHM to thoroughly understand remote sensing images and perform diverse remote sensing tasks. Moreover, different from existing remote sensing instruction datasets that only include factual questions, HnstD contains additional deceptive questions stemming from the non-existence of objects. This feature prevents VHM from producing affirmative answers to nonsense queries, thereby ensuring its honesty. In our experiments, VHM significantly outperforms various vision language models on common tasks of scene classification, visual question answering, and visual grounding. Additionally, VHM achieves competent performance on several unexplored tasks, such as building vectorizing, multi-label classification and honest question answering. We will release the code, data and model weights at https://github.com/opendatalab/VHM .
title VHM: Versatile and Honest Vision Language Model for Remote Sensing Image Analysis
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.20213