Vision-Language Models in Remote Sensing: Current Progress and Future Trends

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Xiang, Wen, Congcong, Hu, Yuan, Yuan, Zhenghang, Zhu, Xiao Xiang
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917627753922560
author Li, Xiang
Wen, Congcong
Hu, Yuan
Yuan, Zhenghang
Zhu, Xiao Xiang
author_facet Li, Xiang
Wen, Congcong
Hu, Yuan
Yuan, Zhenghang
Zhu, Xiao Xiang
contents The remarkable achievements of ChatGPT and GPT-4 have sparked a wave of interest and research in the field of large language models for Artificial General Intelligence (AGI). These models provide intelligent solutions close to human thinking, enabling us to use general artificial intelligence to solve problems in various applications. However, in remote sensing (RS), the scientific literature on the implementation of AGI remains relatively scant. Existing AI-related research in remote sensing primarily focuses on visual understanding tasks while neglecting the semantic understanding of the objects and their relationships. This is where vision-language models excel, as they enable reasoning about images and their associated textual descriptions, allowing for a deeper understanding of the underlying semantics. Vision-language models can go beyond visual recognition of RS images, model semantic relationships, and generate natural language descriptions of the image. This makes them better suited for tasks requiring visual and textual understanding, such as image captioning, and visual question answering. This paper provides a comprehensive review of the research on vision-language models in remote sensing, summarizing the latest progress, highlighting challenges, and identifying potential research opportunities.
format Preprint
id arxiv_https___arxiv_org_abs_2305_05726
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Vision-Language Models in Remote Sensing: Current Progress and Future Trends
Li, Xiang
Wen, Congcong
Hu, Yuan
Yuan, Zhenghang
Zhu, Xiao Xiang
Computer Vision and Pattern Recognition
Artificial Intelligence
The remarkable achievements of ChatGPT and GPT-4 have sparked a wave of interest and research in the field of large language models for Artificial General Intelligence (AGI). These models provide intelligent solutions close to human thinking, enabling us to use general artificial intelligence to solve problems in various applications. However, in remote sensing (RS), the scientific literature on the implementation of AGI remains relatively scant. Existing AI-related research in remote sensing primarily focuses on visual understanding tasks while neglecting the semantic understanding of the objects and their relationships. This is where vision-language models excel, as they enable reasoning about images and their associated textual descriptions, allowing for a deeper understanding of the underlying semantics. Vision-language models can go beyond visual recognition of RS images, model semantic relationships, and generate natural language descriptions of the image. This makes them better suited for tasks requiring visual and textual understanding, such as image captioning, and visual question answering. This paper provides a comprehensive review of the research on vision-language models in remote sensing, summarizing the latest progress, highlighting challenges, and identifying potential research opportunities.
title Vision-Language Models in Remote Sensing: Current Progress and Future Trends
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2305.05726