Popeye: A Unified Visual-Language Model for Multi-Source Ship Detection from Remote Sensing Imagery

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Wei, Cai, Miaoxin, Zhang, Tong, Lei, Guoqiang, Zhuang, Yin, Mao, Xuerui
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914833316708352
author Zhang, Wei
Cai, Miaoxin
Zhang, Tong
Lei, Guoqiang
Zhuang, Yin
Mao, Xuerui
author_facet Zhang, Wei
Cai, Miaoxin
Zhang, Tong
Lei, Guoqiang
Zhuang, Yin
Mao, Xuerui
contents Ship detection needs to identify ship locations from remote sensing (RS) scenes. Due to different imaging payloads, various appearances of ships, and complicated background interference from the bird's eye view, it is difficult to set up a unified paradigm for achieving multi-source ship detection. To address this challenge, in this article, leveraging the large language models (LLMs)'s powerful generalization ability, a unified visual-language model called Popeye is proposed for multi-source ship detection from RS imagery. Specifically, to bridge the interpretation gap between the multi-source images for ship detection, a novel unified labeling paradigm is designed to integrate different visual modalities and the various ship detection ways, i.e., horizontal bounding box (HBB) and oriented bounding box (OBB). Subsequently, the hybrid experts encoder is designed to refine multi-scale visual features, thereby enhancing visual perception. Then, a visual-language alignment method is developed for Popeye to enhance interactive comprehension ability between visual and language content. Furthermore, an instruction adaption mechanism is proposed for transferring the pre-trained visual-language knowledge from the nature scene into the RS domain for multi-source ship detection. In addition, the segment anything model (SAM) is also seamlessly integrated into the proposed Popeye to achieve pixel-level ship segmentation without additional training costs. Finally, extensive experiments are conducted on the newly constructed ship instruction dataset named MMShip, and the results indicate that the proposed Popeye outperforms current specialist, open-vocabulary, and other visual-language models for zero-shot multi-source ship detection.
format Preprint
id arxiv_https___arxiv_org_abs_2403_03790
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Popeye: A Unified Visual-Language Model for Multi-Source Ship Detection from Remote Sensing Imagery
Zhang, Wei
Cai, Miaoxin
Zhang, Tong
Lei, Guoqiang
Zhuang, Yin
Mao, Xuerui
Computer Vision and Pattern Recognition
Ship detection needs to identify ship locations from remote sensing (RS) scenes. Due to different imaging payloads, various appearances of ships, and complicated background interference from the bird's eye view, it is difficult to set up a unified paradigm for achieving multi-source ship detection. To address this challenge, in this article, leveraging the large language models (LLMs)'s powerful generalization ability, a unified visual-language model called Popeye is proposed for multi-source ship detection from RS imagery. Specifically, to bridge the interpretation gap between the multi-source images for ship detection, a novel unified labeling paradigm is designed to integrate different visual modalities and the various ship detection ways, i.e., horizontal bounding box (HBB) and oriented bounding box (OBB). Subsequently, the hybrid experts encoder is designed to refine multi-scale visual features, thereby enhancing visual perception. Then, a visual-language alignment method is developed for Popeye to enhance interactive comprehension ability between visual and language content. Furthermore, an instruction adaption mechanism is proposed for transferring the pre-trained visual-language knowledge from the nature scene into the RS domain for multi-source ship detection. In addition, the segment anything model (SAM) is also seamlessly integrated into the proposed Popeye to achieve pixel-level ship segmentation without additional training costs. Finally, extensive experiments are conducted on the newly constructed ship instruction dataset named MMShip, and the results indicate that the proposed Popeye outperforms current specialist, open-vocabulary, and other visual-language models for zero-shot multi-source ship detection.
title Popeye: A Unified Visual-Language Model for Multi-Source Ship Detection from Remote Sensing Imagery
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.03790