Locate Anything on Earth: Advancing Open-Vocabulary Object Detection for Remote Sensing Community

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pan, Jiancheng, Liu, Yanxing, Fu, Yuqian, Ma, Muyuan, Li, Jiahao, Paudel, Danda Pani, Van Gool, Luc, Huang, Xiaomeng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916644794662912
author Pan, Jiancheng
Liu, Yanxing
Fu, Yuqian
Ma, Muyuan
Li, Jiahao
Paudel, Danda Pani
Van Gool, Luc
Huang, Xiaomeng
author_facet Pan, Jiancheng
Liu, Yanxing
Fu, Yuqian
Ma, Muyuan
Li, Jiahao
Paudel, Danda Pani
Van Gool, Luc
Huang, Xiaomeng
contents Object detection, particularly open-vocabulary object detection, plays a crucial role in Earth sciences, such as environmental monitoring, natural disaster assessment, and land-use planning. However, existing open-vocabulary detectors, primarily trained on natural-world images, struggle to generalize to remote sensing images due to a significant data domain gap. Thus, this paper aims to advance the development of open-vocabulary object detection in remote sensing community. To achieve this, we first reformulate the task as Locate Anything on Earth (LAE) with the goal of detecting any novel concepts on Earth. We then developed the LAE-Label Engine which collects, auto-annotates, and unifies up to 10 remote sensing datasets creating the LAE-1M - the first large-scale remote sensing object detection dataset with broad category coverage. Using the LAE-1M, we further propose and train the novel LAE-DINO Model, the first open-vocabulary foundation object detector for the LAE task, featuring Dynamic Vocabulary Construction (DVC) and Visual-Guided Text Prompt Learning (VisGT) modules. DVC dynamically constructs vocabulary for each training batch, while VisGT maps visual features to semantic space, enhancing text features. We comprehensively conduct experiments on established remote sensing benchmark DIOR, DOTAv2.0, as well as our newly introduced 80-class LAE-80C benchmark. Results demonstrate the advantages of the LAE-1M dataset and the effectiveness of the LAE-DINO method.
format Preprint
id arxiv_https___arxiv_org_abs_2408_09110
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Locate Anything on Earth: Advancing Open-Vocabulary Object Detection for Remote Sensing Community
Pan, Jiancheng
Liu, Yanxing
Fu, Yuqian
Ma, Muyuan
Li, Jiahao
Paudel, Danda Pani
Van Gool, Luc
Huang, Xiaomeng
Computer Vision and Pattern Recognition
Object detection, particularly open-vocabulary object detection, plays a crucial role in Earth sciences, such as environmental monitoring, natural disaster assessment, and land-use planning. However, existing open-vocabulary detectors, primarily trained on natural-world images, struggle to generalize to remote sensing images due to a significant data domain gap. Thus, this paper aims to advance the development of open-vocabulary object detection in remote sensing community. To achieve this, we first reformulate the task as Locate Anything on Earth (LAE) with the goal of detecting any novel concepts on Earth. We then developed the LAE-Label Engine which collects, auto-annotates, and unifies up to 10 remote sensing datasets creating the LAE-1M - the first large-scale remote sensing object detection dataset with broad category coverage. Using the LAE-1M, we further propose and train the novel LAE-DINO Model, the first open-vocabulary foundation object detector for the LAE task, featuring Dynamic Vocabulary Construction (DVC) and Visual-Guided Text Prompt Learning (VisGT) modules. DVC dynamically constructs vocabulary for each training batch, while VisGT maps visual features to semantic space, enhancing text features. We comprehensively conduct experiments on established remote sensing benchmark DIOR, DOTAv2.0, as well as our newly introduced 80-class LAE-80C benchmark. Results demonstrate the advantages of the LAE-1M dataset and the effectiveness of the LAE-DINO method.
title Locate Anything on Earth: Advancing Open-Vocabulary Object Detection for Remote Sensing Community
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2408.09110