Dual-Perspective United Transformer for Object Segmentation in Optical Remote Sensing Images

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sun, Yanguang, Yan, Jiexi, Qian, Jianjun, Xu, Chunyan, Yang, Jian, Luo, Lei
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916812557385728
author Sun, Yanguang
Yan, Jiexi
Qian, Jianjun
Xu, Chunyan
Yang, Jian
Luo, Lei
author_facet Sun, Yanguang
Yan, Jiexi
Qian, Jianjun
Xu, Chunyan
Yang, Jian
Luo, Lei
contents Automatically segmenting objects from optical remote sensing images (ORSIs) is an important task. Most existing models are primarily based on either convolutional or Transformer features, each offering distinct advantages. Exploiting both advantages is valuable research, but it presents several challenges, including the heterogeneity between the two types of features, high complexity, and large parameters of the model. However, these issues are often overlooked in existing the ORSIs methods, causing sub-optimal segmentation. For that, we propose a novel Dual-Perspective United Transformer (DPU-Former) with a unique structure designed to simultaneously integrate long-range dependencies and spatial details. In particular, we design the global-local mixed attention, which captures diverse information through two perspectives and introduces a Fourier-space merging strategy to obviate deviations for efficient fusion. Furthermore, we present a gated linear feed-forward network to increase the expressive ability. Additionally, we construct a DPU-Former decoder to aggregate and strength features at different layers. Consequently, the DPU-Former model outperforms the state-of-the-art methods on multiple datasets. Code: https://github.com/CSYSI/DPU-Former.
format Preprint
id arxiv_https___arxiv_org_abs_2506_21866
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dual-Perspective United Transformer for Object Segmentation in Optical Remote Sensing Images
Sun, Yanguang
Yan, Jiexi
Qian, Jianjun
Xu, Chunyan
Yang, Jian
Luo, Lei
Computer Vision and Pattern Recognition
Automatically segmenting objects from optical remote sensing images (ORSIs) is an important task. Most existing models are primarily based on either convolutional or Transformer features, each offering distinct advantages. Exploiting both advantages is valuable research, but it presents several challenges, including the heterogeneity between the two types of features, high complexity, and large parameters of the model. However, these issues are often overlooked in existing the ORSIs methods, causing sub-optimal segmentation. For that, we propose a novel Dual-Perspective United Transformer (DPU-Former) with a unique structure designed to simultaneously integrate long-range dependencies and spatial details. In particular, we design the global-local mixed attention, which captures diverse information through two perspectives and introduces a Fourier-space merging strategy to obviate deviations for efficient fusion. Furthermore, we present a gated linear feed-forward network to increase the expressive ability. Additionally, we construct a DPU-Former decoder to aggregate and strength features at different layers. Consequently, the DPU-Former model outperforms the state-of-the-art methods on multiple datasets. Code: https://github.com/CSYSI/DPU-Former.
title Dual-Perspective United Transformer for Object Segmentation in Optical Remote Sensing Images
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.21866